Video diffusion models are just overqualified depth estimators! Deterministic... single-pass depth estimation based on WanV2.1. - SOTA 5.5 AbsRel on ScanNet - data-efficient than baselines; - no temporal flicker + infinite-length estimation w/ zero scale drift.show more

Wildminder
49,391 次观看 • 4 个月前
Depth Any Video with Scalable Synthetic Data AI physicists... and chemists continue to make strides in depth estimation from video. Check out this new paper featuring some impressive examples. See the thread for more details (unfortunately no code yet). Abstract: Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge through two key innovations. First, we develop a scalable synthetic data pipeline, capturing real-time video depth data from diverse game environments, yielding 40,000 video clips of 5-second duration, each with precise depth annotations. Second, we leverage the powerful priors of generative video diffusion models to handle real-world videos effectively, integrating advanced techniques such as rotary position encoding and flow matching to further enhance flexibility and efficiency. Unlike previous models, which are limited to fixed-length video sequences, our approach introduces a novel mixed-duration training strategy that handles videos of varying lengths and performs robustly across different frame rates 0 - even on single frames. At inference, we propose a depth interpolation method that enables our model to infer high-resolution video depth across sequences of up to 150 frames. Our model outperforms all previous generative depth models in terms of spatial accuracy and temporal consistency.show more

MrNeRF
27,428 次观看 • 1 年前
Wonderland: Navigating 3D Scenes from a Single Image Contributions:... • First, we introduce a representation for controllable 3D generation by leveraging the generative priors from camera-guided video diffusion models. Unlike image models, video diffusion models are trained on extensive video datasets. This enables them to capture comprehensive spatial relationships within scenes across multiple views and embed a form of "3D awareness" in their latent space, which allows us to maintain 3D consistency in novel view synthesis. • Second, to achieve controllable novel view generation, we empower video models with precise control over specified camera motions. We introduce a novel dual-branch conditioning mechanism that effectively incorporates desired diverse camera trajectories into the video diffusion model. This enables expansion of a single image into a multi-view consistent capture of a 3D scene with precise pose control. • Third, to achieve efficient 3D reconstruction, we directly transform video latents into 3DGS. We propose a novel latent-based large reconstruction model (LaLRM) that lifts video latents to 3D in a feed-forward manner. With this design, during inference, our model directly predicts 3DGS from a single input image, effectively aligning the generation and reconstruction tasks—and bridging image space and 3D space—through the video latent space. Compared with reconstructing scenes from images, the video latent space offers a 256× spatial-temporal reduction while retaining essential and consistent 3D structural details. Such a high degree of compression is crucial, as it allows the LaLRM to handle a wider range of 3D scenes within the reconstruction framework, with the same memory constraints.show more

MrNeRF
52,801 次观看 • 1 年前
🚨 JUST IN: THIS FREE TOOL JUST REPLACED FOUR... AI IMAGE AND VIDEO SUBSCRIPTIONS AT ONCE. Midjourney. Krea. Higgsfield. Openart. One repo. 200+ models. Zero dollars a month. Here is what it actually does. It is a full image and video studio that runs in your browser or as a desktop app. Text to image, image to image, text to video, image to video, lip sync, cinema mode with real camera controls. All of it. 4,500 people already starred this. What you get for free: → 50+ image models including Flux, Midjourney v7, Ideogram, GPT-4o, Seedream → 60+ video models including Kling, Sora, Veo, Runway, Wan, Hailuo → lip sync studio with 9 dedicated models. upload a portrait and audio and it talks → cinema studio with real camera controls. lens, focal length, aperture, film stock → feed up to 14 reference images into one generation → self-hosted. your data never leaves your machine The crazy part is there is also a hosted version that needs zero setup. Just open the link and start generating. Now the math. Midjourney Standard: $30/month Krea AI Pro: $30/month Higgsfield Plus: $49/month Openart AI: $15/month That is $124 a month. $1,488 a year. This repo does everything all four do. With more models than any of them. For free. Forever. No subscription. No vendor lock-in. MIT licensed. Download it in one click on Mac or Windows. Someone should have told me about this sooner. I feel like an idiot. ( save this )show more

Kanika
14,749 次观看 • 3 个月前
🚨 Anthropic committed up to 1M TPU chips for... Claude. Openai is leasing TPUs for chatgpt inference. Here's How kernels work on TPUs (deep dive 2/6 by emi) pallas is Google's answer to kernel writing. a python kernel SDK built on JAX. still very experimental (jax.experimental.pallas). on TPU it compiles through mosaic; on GPU it lowers to triton. if you know CUDA, the syntax will feel familiar but the execution model is completely different. in CUDA, grid=(4,4) launches 16 blocks running simultaneously across SMs. in pallas, those 16 iterations run one after another in lexicographic order. no threads. no warps. no blocks. no occupancy tuning. a TPU is a sequential machine with a very wide vector register — more like a CPU than a GPU. performance comes from width: a 128x128 systolic array doing matmul and an 8x128 SIMD vector unit doing everything else. maximum parallelism on chip: 2, one per TensorCore in megacore mode. three concepts replace CUDA's thread/block/grid hierarchy. Refs are mutable memory references. because execution is sequential, each iteration safely accumulates without atomics. in CUDA you'd need atomics or a separate reduction pass. the memory model is also very different from NVIDIA's. zero hardware caches. VMEM is 32-128 MiB of software-managed scratchpad — 500-1000x larger than GPU shared memory per SM. all data must be explicitly DMA'd from HBM to VMEM before any computation touches it. four levels: HBM → VMEM → VREGs → MXU/VPU, plus SMEM for scalar control data. every byte of data movement is your responsibility. this is like CUDA shared memory except it's 500x bigger and there's no cache fallback. pipelining is mandatory. without double-buffering HBM→VMEM transfers, the MXU just stalls waiting for data. this is the single most important optimization on TPU. and because grid execution is sequential and deterministic, consecutive iterations that need the same input block skip the redundant HBM transfer automatically, impossible on GPU where block execution order is undefined. the compilation pipeline is unlike anything in this series: python → jaxpr → stableHLO → XLA HLO (71+ optimization passes) → LLO (78+ passes) → 322-bit VLIW bundles. the compiler packs instructions for scalar, vector, matrix, and DMA units into a single 322-bit word. everything in that bundle executes in parallel, with no runtime scheduling.show more

wafer
33,134 次观看 • 13 天前
Kling 3.0 on Yapper Prompt: FORMAT: 15s / 8... shots / editorial contact sheet SUBJECT: Young woman with long blue-white gradient hair, dark leather armor with scale-textured panels, mounted on or posed beside a colossal dragon. ENVIRONMENT: Volcanic ridge at golden hour, soft directional sunlight through ash haze, no artificial light, no flash. MOOD: High-fashion creature editorial, every frame a magazine spread candidate. MUSIC: Minimal ambient bass pulse, slow and deliberate. COLOR LOGIC: Naturalistic Film Print Emulation STYLE: Fashion editorial, beauty portraiture LOGIC RULE: Each shot reads as a still editorial frame with near-frozen poses. Natural light only, shaped by ash diffusion and golden-hour direction. Rider and dragon share the frame as co-subjects. No flash, no strobe, no artificial fill. SHOT SEQUENCE: SHOT 1: Full-length, 50mm / Rider stands in profile against dragon's folded wing, one hand on hip, blue-white hair backlit by golden haze, dragon scales fill background as texture wall / SFX: soft wind SHOT 2: Hard cut. Medium close-up, 85mm / Rider faces camera, chin slightly lifted, dragon's jaw rests just behind her shoulder, shallow depth of field melts scales into bokeh / SFX: silence SHOT 3: Hard cut. Detail insert, 100mm macro / Rider's leather gauntlet resting on dragon's neck ridge, golden light raking across both textures, skin and scale side by side / SFX: faint ember crackle SHOT 4: Hard cut. Wide, 35mm / Rider seated sidesaddle on dragon's back, legs crossed, hair swept by updraft, ash particles drifting like golden confetti / SFX: low wind hum SHOT 5: Hard cut. Over-the-shoulder from dragon's head, 40mm / Rider looks back over her shoulder toward camera, face half-lit by warm side light, dragon horn frames the top of shot / SFX: deep breath SHOT 6: Hard cut. Low angle, 24mm / Rider standing on dragon's foreleg, full body, arms relaxed at sides, dragon wing spread behind her as a dark canopy, golden rim light on hair edges / SFX: membrane stretch SHOT 7: Hard cut. Extreme close-up, 135mm / Rider's eye and cheekbone, golden-hour catchlight in iris, single strand of blue-white hair across face, dragon scale texture reflected in pupil / SFX: heartbeat SHOT 8: Hard cut. Ultra-wide, 14mm / Rider walks away from camera along dragon's spine toward the head, dragon lifts chin to sky, fire by mouth both silhouetted against amber sunset / SFX: low brass swellshow more

Zara
19,755 次观看 • 3 个月前
🚨 BREAKING — one of the strongest OpenClaw setups... on Polymarket just went public. A trader reportedly started with ~$100–200 and scaled it to ~$3.7M. No insider access. No political connections. Just a developer running his own automation built with OpenClaw. Profile → Copytrade → I went through the framework myself. What surprised me: There’s no huge infrastructure. No complex quant stack. No giant data pipelines. Just clean logic and disciplined automation. After about 8 hours analyzing it, the strategy breaks down into three parts. 1) “Free money” via NO positions The bot targets outcomes with near-zero probability. Instead of chasing big wins, it accumulates a massive number of small high-probability NO trades. Not speculation — systematic probability harvesting. 2) Logical arbitrage Sometimes Outcome A logically implies Outcome B, but markets don’t adjust instantly. The bot detects these inconsistencies and enters before repricing happens. By the time the headline reaches traders, the window is already closed. 3) Retail-driven markets Sports and political markets are dominated by retail flow and emotional reactions. Prices overshoot, spreads widen, and inefficiencies appear constantly. The bot sits in those gaps and clips small edges repeatedly. Scale is the edge. 4,192 trades executed. Individually small. Together they compounded into roughly ~$3.7M profit. Largest single win: $1,464,152. The equity curve is almost vertical. It’s not about predicting events. It’s about exploiting structural inefficiencies faster than the crowd.show more

Discover
186,472 次观看 • 4 个月前
SPACEX’S STARLINK As a B787 pilot, it pains me... when posts about aviation are this wrong on every detail & every conclusion. Starlink is the best option for airliners. Amazon is a distant second. ♦️ BEST-IN-CLASS Starlink Aviation already uses proven flat phased-array antennas — no gimbals, no moving parts. Amazon’s is the same tech, not different or better. ♦️ TWO > ONE The speed claims are way off. Starlink’s real-world performance on airlines already beats Amazon’s unproven promises. There is no 250 Mbps cap. And one larger antenna isn’t automatically superior — its wider profile creates more aerodynamic drag than Starlink’s two smaller inline antennas. Two antennas also give dispatch reliability: if one fails, Wi-Fi still works. Starlink installs are famously quick and reliable. ♦️ GLOBAL SERVICE Major airlines are global operations, I cross 80 time zones every month — the size of the constellation is the differentiator. Coverage dead zones mean I don’t get real-time updated turbulence plots or full-storm radar maps in the flight deck. Saying the antenna is THE bottleneck doesn’t make it true. A Wi-Fi antenna with no signal from satellites is just useless dead weight. Check out the gaps on the Amazon constellation below. ♦️ MARKETING HYPE The AWS private interconnect is a marketing bullet, but nothing an airliner actually needs. Compute for the plane sits on the plane. Operational data exchange is heavily regulated and runs on dedicated satcom datalinks (ACARS and CPDLC). Looping AWS in adds zero safety benefit and simply creates another potential hacker entry point. As for analytics — airlines already excel at that on the ground where it belongs. ♦️ LOCKED-IN? Long-term lock-in to a single cloud provider is a disadvantage to many. Delta clearly got a discount on the AWS today to bundle in the promise of Wi-Fi tomorrow. But what happens once the introductory discounts disappear and Delta gave up all leverage? Starlink is delivering today at global scale. Amazon is still selling PowerPoint slides. Facts matter in aviation. Videos - Left: Starlink satellites Right: Amazon satellitesshow more

Amy
83,362 次观看 • 3 个月前
This week is already so hot. 🔥 Massive release... from Decart : Lucy 2.0 a World Editing Model running at 1080p, 30FPS in realtime. This is truly exciting, the era of real-time generative reality is here. We are moving from watching AI video to living inside AI video. A breakthrough model capable of transforming the visual world in real-time. Moving beyond offline rendering, Lucy 2.0 delivers high-fidelity 1080p video generation with near-zero latency. Lucy 2.0 literally "redraws" the entire world pixel-by-pixel, while you are watching it. e.g. If you want to be an anime character, it doesn't just put a mask on you. It turns your skin into anime skin, your hair into anime hair, and the lighting in your room into anime lighting. Lucy 2.0 is also trained to stop the generated video from slowly falling apart over time, so the same stream can run much longer without faces and details drifting. So why is this a "Massive Deal"? Traditional AI video-generation model takes a prompt, you wait 10–20 minutes, and the computer "bakes" a video for you. You couldn't touch it or change it while it was happening. But Lucy 2.0 works like a mirror. It happens in real-time (30 frames per second). There is no waiting. You move your hand, the AI character moves its hand instantly. The craziest part isn't the visuals; it's the physics. Usually, AI hallucinations are glitchy—hands merge into faces, walls melt. Lucy 2.0 understands how the world works without being told. It knows that if you take off a helmet, there is hair underneath. It knows that if you splash water, droplets fly. It learned "physics" just by watching millions of videos. The physical behavior you see emerges from learned visual dynamics, not from engineered geometry or explicit physics engines. Their official technical report explicitly states that the model does not use traditional 3D engines, depth maps, or wireframes. It is a "pure diffusion model."show more

Rohan Paul
12,761 次观看 • 5 个月前
Hermes Agent + Higgsfield Marketing Studio = AI UGC... Content Factory I built a fully automated system inside Higgsfield that repurposes, localizes, and launches winning TikTok Shop content across hundreds of creator-style accounts. It's so effective it feels like running Facebook ads in 2008. No actors. No products in hand. No ghost creators. Just viral TikTok Shop sales - 24/7. The results speak louder than any pitch: • CPMs as low as $0.10 • 550+ cinematic, product-ready ads per day from a single prompt • 100 hooks tested in the time it used to take to test 10 • $100/mo replacing a $50k+ creative budget Here's the full pipeline - all native inside Higgsfield Marketing Studio: > Hermes Agent analyzes your product, scrapes Meta Ads + TikTok Ads, identifies winning content, and localizes every angle to your brand. > Seedance 2.0 turns data into AI UGC ads - captions, pacing, hooks, your website showcase, all auto-edited inside Higgsfield Marketing Studio. > AI UGC personas are spun up with realistic faces, voices, and personalities - cloned voiceovers in seconds. > Our phone farm pushes every finished video straight to TikTok Shop, daily, on autopilot. >No setup. No switching between five tools. Everything lives inside Higgsfield Marketing Studio. Here's how it actually runs: Hermes Agent researches the niche, scrapes winning TikTok Shop videos, and rebuilds them with fresh hooks, angles, and UGC visuals tailored to your brand. Agents create and post daily to affiliate accounts - fully automated. Then we activate the MPS (Multi-Platform Swarm): once a concept wins on TikTok Shop, Higgsfield deploys hundreds of AI Agents to flood the niche with variations that all drive back to our shot. Most brands are still paying $300–$500 per video. Testing 10 hooks costs $5,000 and takes three weeks. With this system, we test 100 hooks in the same timeframe - and the winners scale automatically. TikTok doesn't reward the best video. It rewards the brand that shows up the most - with content that converts. The brands automating content at scale will be the biggest winners of 2026.show more

Noah Frydberg | Tiktok Shop For Brands
27,037 次观看 • 3 个月前
Researchers made KMeans 200x faster. And the new technique... also beats approaches like cuML and FAISS. Flash-KMeans is an IO-aware implementation of exact KMeans that redesigns the algorithm around modern GPU bottlenecks. By attacking the memory bottlenecks directly, Flash-KMeans achieves: - 33x speedup over cuML - 200x speedup over FAISS This speedup comes from how it moves through GPU memory. Standard KMeans runs in two steps, and both are bottlenecked by reads and writes to GPU memory: 1) The first step matches every point to its nearest centroid. Standard KMeans computes the full point-to-centroid distance matrix, writes it out to GPU memory, then reads it back to find each nearest centroid. That write-then-read round trip is the bottleneck. Flash-KMeans combines the distance calculation with the nearest-centroid step, so the result is computed on-chip and the full matrix is never written out. 2) The second step recomputes each centroid by averaging the points assigned to it. Standard KMeans has thousands of threads writing into the same centroid slots at once, so they stall waiting for their turn. Flash-KMeans sorts points by cluster first, turning scattered writes into sequential reductions that read and write memory in one efficient pass. Using these two optimizations at the million-scale, Flash-KMeans completes a standard KMeans iteration in a few milliseconds. The video below depicts this in action. Several reasons why this is important: KMeans has always been an offline primitive. Something you run once to preprocess data and move on. These speedups make the approach viable in several runtime-critical systems. ↳ Vector indices like FAISS use KMeans to build search indices. Faster KMeans means you can re-index dynamically as data changes. ↳ LLM quantization methods need KMeans to find optimal weight codebooks, per layer, repeatedly. What takes hours could now take minutes. ↳ MoE models need fast token routing at inference time. Flash-KMeans makes it viable to run this inside the inference loop, not just in preprocessing. I have shared the paper in the replies. That said, memory is the real constraint Flash-KMeans solves, and the problem is not just limited to clustering. The vectors a RAG system stores after indexing create similar bottlenecks. I wrote a detailed walkthrough recently on cutting this vector memory by 32x with binary quantization, querying 36M+ vectors in a few milliseconds. Read it below.show more

Avi Chawla
89,234 次观看 • 1 个月前
Proud to announce the in-depth collaboration between Kingnet and... Alibaba Cloud in AI Gaming. Alibaba Cloud provides world-leading cloud computing, big data, and AI services, with disclosed revenue exceeding $15 billion in 2024, which is one of the most renowned global server providers. When two superpowers collide, the game changes. 🌊AI Gaming R&D By integrating Qwen 's LLM and Alibaba Cloud 's PAI platform (including PAI-iTAG, PAI-Designer, PAI-DSW, PAI-DLC, and PAI-EAS), Kingnet has emerged as one of the gaming industry's pioneers in AIGC-powered content generation and AI rendering. Together, we are accelerating the realization of no-code game development. 🌊GPU Computing Resources Alibaba Cloud delivers GPU-accelerated elastic computing services with exceptional processing power, supporting diverse workloads including deep learning, scientific computing, graphics visualization, and video processing - providing robust GPU computing capabilities for KingnetAI's demanding requirements. 🌊Cloud Service Optimization Cloud server deployment has become the mainstream choice for small and mid-sized game studios in global operations. Leveraging Alibaba Cloud server advantages, we will develop and deploy more cloud-native games to meet user demands. The disruptive innovation we're bringing to the industry: 🔸Minute-scale game asset production replaces traditional week/month-long cycles 🔸Single-digit dollar development costs VS traditional four-figure entry thresholds 🔸AI-powered NPCs with behavioral engines deliver dynamic player interactions, breaking static story constraints, etc. 🔜Kingnet AI V2 is approaching launch. The Agent system and game generation engine will be officially deployed across 3 chains: 🔹Leveraging Solana high throughput and low gas fee , Solana has consistently been a developer favorite, latest product will be deployed on Solana - with users paying $SOL for on-demand asset creation fees. 🔹Another key partner is BNB Chain ,We are actively participating in both the #BNBAIHack and the latest MVB 10. Powered by BNB Chain long-standing support for AI innovation. Kingnet V2 and NFT drop will be deployed on BNB Chain, providing developers and the community with comprehensive game-generation tools and support. 🔹As an early strategic partner of Kingnet, TON 💎 @TONEastAsia was one of the earliest chain to connect Web2 and Web3, Kingnet V2 will be deployed on TON, providing TON game developers with low-cost, high-efficiency asset generation, and supporting users to use $TON as an asset generation cost. The Future of AI Gaming is coming.show more

Kingnet AI
149,774 次观看 • 1 年前
my 8 GB VRAM gaming laptop is absolutely going... to hate me for this. but I still did it. ran a 31b dense model (Gemma 4 31b Q4) with only 8 GB VRAM last week I ran Gemma 4 26B A4B a mixture of experts model on my RTX 4060 and hit 25–28 tokens/sec using llama.cpp's new MTP support. smooth. snappy. but MoE has a secret: it only activates 4B parameters per token despite having 26B total. that's why it flies. so the real question started haunting me. what if I throw a full, no tricks, every parameter fires on every token, 31B DENSE model at the same machine? # Hardware: GPU: NVIDIA RTX 4060, 8 GB VRAM RAM: 16 GB CPU: Intel Core i7 H Laptop. Gaming. Modest. The model: gemma-4-31B-it-qat-UD-Q4_K_XL.gguf (model's unsloth huggingface link in the comments) This is Google DeepMind's flagship dense model in the Gemma 4 family that can run on single consumer GPU. It packs a hybrid attention architecture, supports up to 256K context natively, and is QAT (Quantization Aware Training) optimized, meaning it retains far more quality than standard post training quants at the same bit depth. This is NOT the MoE. This is 31 BILLION dense parameters, every single one of them loaded. # the flags I used: -m gemma-4-31B-it-qat-UD-Q4_K_XL.gguf -cnv --spec-type draft-mtp --spec-draft-model mtp-gemma-4-31B-it.gguf --spec-draft-n-max 8 --spec-draft-p-min 0.6 -c 6000 -v Multi Token Prediction (MTP) is still active here. Separate draft GGUF required, same as the 26B setup. # Results: → Decode: ~3 tokens/sec → Prefill: ~2 tokens/sec → Context: 6000 tokens → Hardware crying quietly in the corner: yes so is 3 tps actually usable? For real time back and forth chat? Not ideal. You're not having a fluid conversation at 3 tps. but slow ≠ useless. And this is where it gets genuinely interesting. think about how senior devs actually work in a real team. But when something is architectural, deeply complex, or needs serious reasoning? they walk down the hall and escalate to the senior. That's exactly the local AI agent architecture this unlocks: → Fast orchestrator model (Gemma 4 26B MoE at 25+ tps) handles routing, simple queries, tool calls, memory. The junior dev. → Gemma 4 31B dense is the senior, called only when the fast model genuinely hits a wall. Hard multi step reasoning. Complex code generation. Deep architectural decisions. The agentic loop stays fast. Only the hard hops touch the 31B. That's a legitimate production grade local AI architecture on a budget hardware. (requires 2 8gb gpus) other workflows where 3 tps is completely fine: - overnight batch jobs. summarize documents, extract structured data, review code. Fire it off. Sleep. wake up to results. - One shot deep reasoning - Silent code audit loops, you write and test, the 31B reviews diffs and flags issues in the background between your sprints - Any workflow where output quality > output speed A few weeks ago, nobody was running a 30B+ dense model on a single consumer GPU with 8 GB VRAM. At all. Now we're doing it on an Intel i7-H gaming laptop with a NVIDIA RTX 4060, thanks to llama.cpp + QAT quants + MTP speculative drafting. Google DeepMind said the Gemma 4 31B targets "consumer GPUs and workstations." They were not exaggerating. The hardware bar to run serious frontier class models locally keeps dropping. the tools are here. the models are here. you just have to be willing to abuse your laptop a little. what workflows would you actually run on a local 3 tps 31B dense model? genuinely curious. drop it below.show more

Alok
63,545 次观看 • 1 个月前
Would you underestimate her just because she wears a... school uniform? GPT Image 2 + Seedance 2.0 on Sjolt Try Canvas: prompt Character Identity Lock (Highest Priority): Use the exact same young East Asian woman from the provided reference character sheet. Preserve 100% identical facial features, face shape, eye shape, nose, lips, skin tone, hairstyle, hair color, proportions, and overall identity throughout the entire video. Do not redesign, reinterpret, or substitute the character. She must remain instantly recognizable as the same person from the reference image. She has shoulder-length wavy silver-gray hair with subtle blue undertones, bright expressive eyes, fair skin, and a confident slight smile that naturally transitions into a focused, determined combat expression. She wears the identical navy blue Korean high school uniform blazer over a gray sweater vest, white collared shirt, striped tie, and matching school skirt from the reference character sheet. Video Prompt: A cinematic, hyper-realistic action sequence inside a chaotic South Korean high school classroom. The classroom is filled with overturned desks, scattered chairs, flying notebooks, broken pencils, and papers drifting through the air. Bright natural daylight streams through large classroom windows, creating realistic highlights, soft shadows, and cinematic contrast. The young female student moves with incredible speed, confidence, and precision as she expertly defends herself against multiple aggressive male students wearing matching Korean school uniforms. Every movement is fluid, athletic, and grounded in realistic martial arts choreography. The camera remains highly dynamic, featuring cinematic handheld tracking shots, fast push-ins, orbit shots, dramatic slow-motion moments, whip pans, low-angle hero shots, and close-up impact shots. Capture rapid combinations of punches, clean high kicks, evasive footwork, parries, elbow strikes, blocks, and throws. Desks slide across the floor, chairs topple over, and dust particles catch the sunlight, emphasizing the intensity of the action. Maintain a high shutter-speed action-photography aesthetic with crisp motion detail, subtle motion blur only during extremely fast movements, physically accurate body mechanics, realistic cloth simulation, natural hair physics, authentic facial expressions, and believable impact reactions. Keep the camera frequently returning to sharp close-ups of her face to reinforce character continuity and emotional intensity. Her silver-gray hair flows naturally with every movement while her determined eyes remain locked on her opponents. Photorealistic cinematic quality, 4K HDR, ultra-detailed skin textures, realistic lighting, volumetric daylight, physically based rendering, shallow depth of field during close-ups, blockbuster Korean action film aesthetic, empowering heroine energy, consistent facial identity throughout every frame, no face drift, no character variation, no animation-style exaggeration.show more

Sharon Riley
26,184 次观看 • 3 天前
Stanford researchers did it again. They just built the... agent-native version of Git. When an agent works on a longer task, the run builds up a lot of state. This includes files edited/created, a dev server, a database, installed packages, KV cache, etc. Say the agent is at step 10 and makes a mistake, maybe it misreads a traceback and rewrites a file that was actually fine. The tests start failing, and the run goes off track, although everything through step eight was correct. By default, the agent just tries to fix it, which creates more edits and tool calls. This burns more tokens and grows the context. The other options are a person stepping in to redirect it or restarting the whole run from step one. That's wasteful, because it pays for every model/tool call again and re-prefills the context. Moreover, since an agent's run is non-deterministic, it doesn't reproduce the same early steps anyway. The reason it's hard to just jump back exactly to a previous correct step and resume from there is that the trajectory is only a message log. It records what the agent said and which tools it called, but not the live state underneath. That state includes things like memory, open file handles, child processes, installed packages, /tmp, and KV cache. None of that is in the log. Git can version the files, but it doesn't snapshot the running process or the KV cache. Checking out step eight moves the files back, but the process is still sitting in step-ten memory with a cold cache. Shepherd is a runtime layer by Stanford that records the run as a trace of typed events rather than a flat log. Each agent-environment interaction becomes a commit, similar to Git, but it tracks the live run. Its commit includes the agent process and the filesystem together, copy-on-write, so a branch carries the actual state and not just the files. Going back to a previous step is then a single call that forks from that commit and continues from the exact state. The copy-on-write fork is roughly five times faster than docker commit, and because the prompt prefix through step eight is unchanged, the KV cache is reused over 95% on replay, so early steps aren't reprocessed again. Once the run can be forked, a meta-agent can sit on top and operate it. It watches the trace and reverts as soon as it looks wrong, before the bad write is committed. In practice, it's just Python calling fork, replay, and revert on the trace, rather than a separate control plane wired into the harness. Not everything is reversible though. Files and sandbox changes undo themselves, but a database write has no automatic undo, so it needs a matching undo step set up in advance. Something external, like a sent email or a real charge, can't be undone, so the supervisor's job there is to catch it before it fires. They tested this on a few public benchmarks. On CooperBench, where two agents work on the same codebase, adding a live supervisor took the pair-coding pass rate from 28.8% to 54.7%. It's still early and labeled alpha. The benefit mostly shows up when a run gets branched a lot over a heavy sandbox state, which is exactly where restarting wastes the most tokens and time. If Git was made to make file changes reversible, Shepherd is trying to do the same thing for a live agent run. Shepherd Repo: (don't forget to star it ⭐ ) That said, Shepherd reverts a bad step inside a run. The harness around it, the prompts, tools, and checks the supervisor relies on, still drifts across runs as models and dependencies change. Akshay wrote about making that harness repair itself, where a failing trace gets diagnosed, the fix is verified against the exact input that failed, and the failure is locked as a regression test so it can't recur. Read it below.show more

Avi Chawla
438,580 次观看 • 17 天前
This looks so neat and clean Created by using... GPT Image 2 + Seedance 2.0 on TapNow Prompt Ultra-luxury cinematic fashion construction film. STRICTLY follow all 12 storyboard panels sequentially without skipping, merging, shortening, reordering, or improvising any panel. Every shot must transition smoothly into the next in exact numerical order from Panel 1 through Panel 12. Total runtime exactly 155 seconds. Maintain absolute continuity in lighting, material behavior, camera language, scale progression, and object identity throughout the entire film. Only ONE shoe exists during the entire video — never show a pair under any circumstance. Visual format: 65mm IMAX film aesthetic, macro-capable Panavision anamorphic lens, ultra-sharp macro detail, shallow cinematic depth of field, fine editorial film grain throughout, subtle horizontal anamorphic lens flares ONLY from thread highlights and suede edge speculars. Infinity cyclorama studio environment with seamless pale grey-to-white gradient background where the floor curves upward into the wall with no visible horizon line. Warm-neutral soft key light from above camera-left creating one consistent clean shadow falling back-right in every shot. Lighting inspired by Loewe / Hermès luxury editorial campaigns — soft but directional, warm and premium, never harsh, never blue, never high contrast, never crushed blacks. Stable cinematography only. No jitter, no AI morphing, no flicker, no ghosting. The entire narrative is a meditative material-transformation journey where a single deep crimson liquid droplet gradually evolves into a handcrafted suede slingback heel through couture construction and hidden comfort engineering. Every material interaction obeys realistic physics. Diegetic sound only — no music, no score. PANEL 1 — THE VOID (00:00–00:10) Wide locked-off cinematic shot of an empty infinity cyclorama studio with pale grey-to-white seamless gradient background. Atmospheric dust floating slowly in warm studio light. Exposure breathing subtly. Silence and soft room tone dominate. After several seconds, one tiny deep crimson droplet slowly falls into frame from above in near weightless slow motion, rotating slightly while catching warm highlights. The droplet lands center frame on the polished cyclorama floor with realistic liquid surface tension physics. It briefly holds spherical form before gently flattening outward. One soft shadow falls back-right. Sound: low room tone, elegant bell-like tick on impact, delicate reverb decay. PANEL 2 — SURFACE TENSION (00:10–00:22) Cut to macro floor-level close-up. Camera slowly orbits around the flattened crimson liquid pool. Surface tension creates organic rounded edges and subtle thickness variation. Warm key light reflects softly like satin lacquer. Tiny ripples propagate outward and gradually settle. Floating dust visible in the beam of light. The liquid slowly begins thickening microscopically as if memory is forming beneath the surface. Edges transition from glossy wetness toward velvety softness. Sound: soft viscous liquid movement, faint atmospheric resonance. PANEL 3 — THE FIRST FIBERS (00:22–00:36) Extreme macro push-in across the crimson surface. Thousands of microscopic suede nap fibers begin extruding upward organically in rhythmic waves. Individual fibers catch warm highlights differently depending on density and angle. The transformation moves naturally across the surface like wind through grass. Matte suede texture gradually replaces liquid gloss entirely. Camera glides slowly across the newly forming velvet landscape. Sound: microscopic fiber rustling, delicate textile friction. PANEL 4 — MATERIAL MEMORY (00:36–00:50) Macro tracking shot across fully formed crimson suede terrain. Every velvet strand visible in extreme detail. Invisible pressure waves beneath the suede subtly shift the nap direction, creating tonal changes across the material surface. Warm light rolls gently over the velvet texture. The material gradually lifts upward from the floor, beginning to imply the sculptural shape of a pointed shoe toe. Sound: soft velvet brushing, low warm resonance. PANEL 5 — THE THREAD ARRIVES (00:50–01:06) Extreme-extreme macro couture construction sequence. One single crimson thread enters frame illuminated by warm directional light. FPV-style camera flies beside the advancing thread as it stitches into the suede surface. Thread fibers visibly twist under tension. The thread enters the suede, disappears beneath the surface, emerges again at the next stitch peak, pulls taut, and repeats rhythmically. Camera passes over every stitch mountain in sequence. Each stitch slightly reshapes surrounding suede. Sound: couture stitching ticks, thread tension tightening, soft textile compression. PANEL 6 — CONSTRUCTION OF FORM (01:06–01:20) Transition from macro to medium scale. Behind the advancing stitch line, the shoe body begins solidifying into recognizable structure. Pointed toe box forms first, then elegant sidewalls rise upward with sculptural precision. The slingback silhouette slowly emerges from previously flat suede terrain. Camera performs a slow floating cinematic arc around the forming structure. Matte crimson suede texture remains perfectly consistent. Sound: restrained structural textile movement, distant stitching continuation. PANEL 7 — THE INTERIOR (01:20–01:34) Macro cutaway revealing the inside of the forming shoe. Cream-colored lambskin lining flows organically into place beneath the crimson suede shell. The lining smooths itself naturally against elegant internal curves. Strong visual contrast between warm cream lambskin and deep crimson suede under soft editorial lighting. Materials appear tactile and luxurious. Sound: soft leather settling, delicate tactile friction. PANEL 8 — ENGINEERED COMFORT (01:34–01:48) Macro cross-sectional engineering sequence. Internal comfort layers assemble one by one with realistic material physics. Dense foam base layer forms first with subtle porous texture. Softer memory foam settles gently above and compresses naturally under its own weight. Cream lambskin velvet seals the upper layer. The completed cushion compresses once slowly and rebounds naturally, demonstrating softness and resilience. Camera glides across microscopic velvet interior texture. Sound: soft pneumatic settling, muted resonance, cushion compression. PANEL 9 — THE HEEL SCULPTURE (01:48–02:02) Macro cinematic focus on the kitten heel gradually forming from the same crimson suede structure beneath the shoe body. Elegant curvature emerges slowly and becomes refined and balanced. Camera tracks upward along the heel contour toward the slingback strap. Warm highlights skim softly across suede edges. Sound: low structural resonance, faint textile shaping. PANEL 10 — FINAL REFINEMENT (02:02–02:18) Medium macro editorial beauty detailing. Camera slowly explores completed craftsmanship: stitch consistency, velvet nap direction, edge finishing, cream lambskin softness, seamless slingback geometry. Tiny dust particles drift through warm studio light. The shoe settles microscopically as if materials are naturally relaxing into final form. Sound: quiet room tone, soft textile creaks, near silence. PANEL 11 — HERO EMERGENCE (02:18–02:32) Slow cinematic pullback from macro detail into full product reveal. The fully completed deep crimson suede slingback heel stands alone centered on the infinity cyclorama floor. Same warm-neutral key light from above camera-left and same single clean shadow falling back-right for continuity with Panel 1. Camera movement extremely slow and controlled. Every material reads clearly: deep crimson suede upper, cream lambskin interior, sculpted kitten heel, elegant slingback strap. Sound: deep low resonant tone gradually fading into room tone. PANEL 12 — THE FINAL HOLD (02:32–02:35) Locked-off front three-quarter hero composition. The single crimson suede slingback heel remains perfectly still at center frame on the seamless infinity cyclorama. Velvet nap subtly catches warm light with natural tonal variation. Silence gradually overtakes the room tone. Hold on pure luxury restraint. No text. No logo. No end card. Global Constraints: Only ONE shoe visible at all times. Never show a pair. No hardware, buckles, laces, or metallic elements. No logos, subtitles, typography, branding, or watermarks. No lighting changes between shots. Maintain consistent deep crimson suede color. Preserve ultra-sharp macro clarity and stable cinematography. No morphing artifacts, no flicker, no surreal deformation. Realistic liquid surface tension. Realistic suede microfiber extrusion. Visible thread twist and accurate stitch tension. Natural foam compression and rebound physics. Consistent matte velvet nap behavior. Diegetic sound only — liquid, textile, stitching, room resonance. No music. No soundtrack.show more

Aaliya
33,224 次观看 • 2 个月前
British Writer Pens The Best Description Of Trump I’ve... Read “Why do some British people not like Donald Trump? A few things spring to mind. Trump lacks certain qualities which the British traditionally esteem. For instance, he has no class, no charm, no coolness, no credibility, no compassion, no wit, no warmth, no wisdom, no subtlety, no sensitivity, no self-awareness, no humility, no honour and no grace – all qualities, funnily enough, with which his predecessor Mr. Obama was generously blessed. So for us, the stark contrast does rather throw Trump’s limitations into embarrassingly sharp relief. Plus, we like a laugh. And while Trump may be laughable, he has never once said anything wry, witty or even faintly amusing – not once, ever. I don’t say that rhetorically, I mean it quite literally: not once, not ever. And that fact is particularly disturbing to the British sensibility – for us, to lack humour is almost inhuman. But with Trump, it’s a fact. He doesn’t even seem to understand what a joke is – his idea of a joke is a crass comment, an illiterate insult, a casual act of cruelty. Trump is a troll. And like all trolls, he is never funny and he never laughs; he only crows or jeers. And scarily, he doesn’t just talk in crude, witless insults – he actually thinks in them. His mind is a simple bot-like algorithm of petty prejudices and knee-jerk nastiness. There is never any under-layer of irony, complexity, nuance or depth. It’s all surface. Some Americans might see this as refreshingly upfront. Well, we don’t. We see it as having no inner world, no soul. And in Britain we traditionally side with David, not Goliath. All our heroes are plucky underdogs: Robin Hood, Dick Whittington, Oliver Twist. Trump is neither plucky, nor an underdog. He is the exact opposite of that. He’s not even a spoiled rich-boy, or a greedy fat-cat. He’s more a fat white slug. A Jabba the Hutt of privilege. And worse, he is that most unforgivable of all things to the British: a bully. That is, except when he is among bullies; then he suddenly transforms into a snivelling sidekick instead. There are unspoken rules to this stuff – the Queensberry rules of basic decency – and he breaks them all. He punches downwards – which a gentleman should, would, could never do – and every blow he aims is below the belt. He particularly likes to kick the vulnerable or voiceless – and he kicks them when they are down. So the fact that a significant minority – perhaps a third – of Americans look at what he does, listen to what he says, and then think ‘Yeah, he seems like my kind of guy’ is a matter of some confusion and no little distress to British people, given that: • Americans are supposed to be nicer than us, and mostly are. • You don’t need a particularly keen eye for detail to spot a few flaws in the man. This last point is what especially confuses and dismays British people, and many other people too; his faults seem pretty bloody hard to miss. After all, it’s impossible to read a single tweet, or hear him speak a sentence or two, without staring deep into the abyss. He turns being artless into an art form; he is a Picasso of pettiness; a Shakespeare of shit. His faults are fractal: even his flaws have flaws, and so on ad infinitum. God knows there have always been stupid people in the world, and plenty of nasty people too. But rarely has stupidity been so nasty, or nastiness so stupid. He makes Nixon look trustworthy and George W look smart. In fact, if Frankenstein decided to make a monster assembled entirely from human flaws – he would make a Trump. And a remorseful Doctor Frankenstein would clutch out big clumpfuls of hair and scream in anguish: ‘My God… what… have… I… created?' If being a twat was a TV show, Trump would be the boxed set.” -Nate Whiteshow more

Republicans against Trump
3,007,220 次观看 • 2 年前
This is my "feel the AGI" moment: I used... GPT-5.6 Sol to train my own autocorrect model that outperforms GPT-5.6 Sol (wtf??) I have no ML background. I have no idea what I'm doing. I just kept pushing Sol until it spat out a SOTA model. And I spent $0. The motivation: Years of talking to AI have made me terrible at typing. Rather than fix my skill issue, I decided to throw more AI at it. My idea was: instead of autocorrect that interrupts my flow, I want to type fast with mistakes and have AI clean it up after. I wanted the smallest local model possible, for speed, for battery life, for science! So I decided to train my own. Inspired by Andrej Karpathy’s autoresearch, I ran Codex /goal with this setup: pick an experiment, try it, record the results to a doc, throw it out if it fails, and plan the next experiment without repeating failures. I gave a few examples that had to pass, tight latency targets, and let it run. Sol did some amazing things. First, it scanned benchmarks and shortlisted base models: Qwen 3.5, Gemma 4, Liquid LFM 2.5. It found a dataset on HuggingFace for typed text. Then it built a simulator for fingers striking a Mac keyboard, modeling the physical layout with a Gaussian distribution around each key. It simulated striking the wrong key, wrong order, fat-fingering, etc. With the models + data + simulator, it fine-tuned using MLX right on my MacBook. It had a working prototype within an hour! But accuracy was pretty poor. — Problem 1: Tokenization Sol read papers, ran tests, and identified that the tokenizer was the bottleneck. Tokenization makes typos hard for the model to see, so it memorizes mappings instead of using its language priors. Sol tried ByT5, Google’s tokenizer-free byte-level LLM. This made a big improvement, but the model is old and lacked the knowledge needed to reach Sol performance. Sol dug deeper and realized a tokenizer-free model isn’t needed; instead, it used T5Gemma, an encoder-decoder model. This can understand the input deeply before producing output, and furthermore, Sol could post-train the encoder to improve performance. This gave a much higher ceiling. — Problem 2: Loss function Now the model was correcting some typos perfectly, but ignoring most. Sol realized that standard cross-entropy loss was teaching the model to avoid edits, because the vast majority of characters in the training data were left unmodified. The fix was wild: Sol wrote a custom loss function that byte-aligns the source and target strings, uses a dynamic programming algorithm to compute the minimum edits between the two, then weights correct edits much higher than copies. After a lot of tuning, this dramatically improved accuracy. — Problem 3: Autoregression One failure mode remained: if the model made a mistake, it couldn’t backtrack. It could only predict the next token. Teaching it to “think” like a reasoning model would solve this, but would be far too slow. Sol found a beautiful solution: instead of greedily predicting the next token, beam search over all possibilities. This parallelizes the exploration instead of one linear chain-of-thought. At the end, choose the path with highest cumulative log probability. This worked great, but made the experience worse, since the user wouldn’t see progress until the whole search was done. To fix this, Sol made a clever observation: after each search step, the longest common prefix among surviving branches is guaranteed to appear in the final result, so it can be displayed immediately. As the search progresses, weaker paths are dropped and the prefix grows, so the user sees continuous progress. Sol built all this as a custom MLX pipeline that does the parallel decoding on the MacBook GPU, with just ~40ms TTFT. It’s crazy fast and entirely local. — Final eval (error reduction rate, higher is better): - Apple autocorrect: 49.66% - GPT-5.6 Luna: 82.47% - GPT-5.6 Terra: 87.64% - GPT-5.6 Sol: 90.56% - Our model (1.7B): 91.02% Final cost: - 1 quota reset (thanks Tibo) - $0 (And yes, I verified there's no cheating. In fact, we test words scrubbed from the training data to prove the model isn’t memorizing) There were a ton more details and tangents I could write about: contrastive learning, GRPO, DPO, dynamic masking, and more. Sol is a fascinating and creative model. It blew my mind so many times. Don’t let a lack of experience stop you: Sol makes AI experiments accessible to anyone!show more

Anshu
177,177 次观看 • 7 天前
Seedance 2.0 on Higgsfield AI delivers some of the... best camera motion, consistency, and creative control I've seen in AI video generation. Full Open-sourced prompts & assets below: SCENE CONTEXT A bright summer afternoon on the coastal road: the young man drives the mint scooter down toward the sea with the young woman riding behind him, arms around his waist — an easy, happy ride past the railway crossing along the water. ACTIVE REFERENCES >> — young woman, 20 years old, 165 cm tall, slender, straight dark brown hair with side-swept bangs pinned by a small black clip, freckles across her cheeks and nose. 100% matches the reference. >> — young man, 22 years old, 178 cm tall, lean, sun-tanned, messy dark hair under a tan baseball cap worn backwards. 100% matches the reference. >> — vehicle: vintage mint-green scooter with a brown leather saddle, chrome mirrors and silver wheels. 100% matches the reference. >> — location: coastal road curving downhill past a railway crossing with yellow-and-black crossbuck signs, utility poles and wires, stone embankment walls, an orange convex traffic mirror on a pole, the open sea with white-capped waves behind. LOCATION MAP The road from >> curves downhill through the midground toward the railway crossing, the sea filling the background beyond it. Stone embankments rise on both sides, the orange convex mirror stands on the right shoulder in the near foreground, utility poles line the curve. Their path: down the curve, past the crossing, along the water toward screen-left. Primary light: bright seaside daylight, sun high, wind off the sea. FIRST FRAME AND SPATIAL BLOCKING The first visible frame already contains >> rolling down the curve with both riders aboard — >> driving, hands on the grips, >> seated close behind him, arms wrapped around his waist, her head just above his shoulder, a full head shorter than him. No empty establishing frame, no delayed reveal. He drives in every segment; she is always the passenger. FORMAT MODE Controlled four-segment multi-shot sequence: one INSERT CUT and two HARD CUTS. Real-time motion at an easy unhurried scooter pace. Every segment is shot handheld — no static shot anywhere in the sequence. OPTICS LENS LOCK SEGMENT 1 = 47° diagonal field of view, standard normal lens character, camera 12 to 15 meters at the roadside, the scooter and both riders full in frame with the crossing and sea behind, straight lines rectilinear, no fisheye. Soft vintage lens rendering: gentle edge softness, mild halation in the bright sky and sea glare, even brightness across the whole frame — no vignette, corners stay as bright as the center. This rendering applies to every segment. LENS LOCK SEGMENT 2 = 29° diagonal field of view, short telephoto character, camera 3 to 4 meters tracking alongside from a following vehicle, close two-shot of their faces and shoulders, the sea streaming soft behind them. LENS LOCK SEGMENT 3 = 29°, camera 1.5 to 2 meters, tight insert on her hands clasped at his stomach, the mint body and brown saddle below, road surface blurring past. LENS LOCK SEGMENT 4 = 47°, camera 10 to 12 meters behind the orange convex mirror on the right shoulder, the mirror large in the near foreground reflecting the road, the real scooter passing through the frame and receding along the sea. No drift mid-segment. CAMERA Handheld in every segment with no exceptions — a real operator at the roadside and in a following vehicle: the frame breathes with shoulder sway and soft micro-tremor visible in every second, small late reframes chasing the scooter and easing back; the tracking shot carries gentle road vibration on top of the hand movement; the insert trembles slightly more; the mirror wide breathes slower but never freezes. No tripod stillness, no gimbal smoothness, no stabilization anywhere. On top, the footage behaves like an old film print running through a projector: constant subtle gate weave, faint exposure flicker, occasional tiny dust specks and hairline scratches, image soft and slightly diffused like an aged 16mm print — never sharp, never digitally clean, no vignette or darkened corners at any moment. ACTION TIMING 0.0s to 3.5s — Roadside wide: the mint scooter putters down the curve at an easy pace, leaning gently with the bend; >> relaxed at the grips, >> pressed close behind him, her hair and skirt hem streaming in the sea wind; they pass the yellow-and-black crossing signs with the white-capped sea glittering beyond. 3.5s HARD CUT 3.5s to 6.5s — Tracking close two-shot: she rests her chin almost on his shoulder and says something teasing into his ear — lips moving without audible words; he barks a laugh, shaking his head, cap holding snug; she grins wide against the wind, bangs whipping, eyes squinting happily. 6.5s INSERT CUT 6.5s to 8.5s — Tight insert: her hands clasped over his stomach, fingers laced, giving a little squeeze as the scooter sways through a bend; the mint body flexes light reflections, the road surface streams underneath in soft blur. 8.5s HARD CUT 8.5s to 12.0s — Wide past the orange convex mirror: the tiny reflection of the scooter slides across the round mirror in the foreground a beat before the real scooter enters and crosses the frame, unhurried, the two of them small against the vast bright sea; she tips her head back and laughs into the wind as they recede along the coast; the engine putter fades. PHYSICS The scooter carries real combined weight: soft suspension compression over road seams, a gentle lean into each bend with both bodies tilting as one, slight throttle sway she counterbalances by gripping tighter; engine vibration trembles through their sleeves; wind at riding speed streams her hair, his tee and her skirt hem backward continuously with fabric flutter; the sea wind adds gusts; the convex mirror reflection tracks their motion with true optics. LIGHTING Bright seaside daylight only — no artificial light. Aged film print look: the sky and the glittering sea bloom into a soft white-gold haze with visible halation rings, gentle glow hanging in the air, creamy highlights rolling off softly. Faded pastel grade of an old print: lifted milky blacks, warm ivory and honey tones over softened sea blues, the mint scooter body reading as a gentle washed pastel green, the orange mirror and yellow-black signs as warm muted accents — never oversaturated; slightly yellowed whites, low contrast, colors gently washed as if the print has aged for twenty years, heavy visible film grain crawling in every frame, delicate haze. Exposure stays natural across the frame — no added vignette, no darkened edges or corners. The whole image reads as an old 2000s Japanese film discovered on a dusty reel. No crisp modern digital look, no cool color cast. AUDIO SFX only, with the worn texture of an old optical soundtrack — slightly muffled, faint constant hiss: the soft putter of the small scooter engine rising and fading with the throttle, wind buffeting past, waves breaking below the road, gull cries, her bright laugh snatched by the wind, the faint tick of the engine at the far end. No music, no intelligible spoken words, no captions, no score. POSITIVE LOCKS Identities lock 100% to >> and >> in every segment — same outfits as their references throughout, her natural 165 cm and his 178 cm with true relative proportions, his cap staying backwards and snug at riding speed in every shot. >> drives in every segment; >> rides pillion with her arms around his waist from first frame to last, hands unclasping never. The scooter stays 100% >> — mint-green body, brown saddle, chrome mirrors — in every shot. Road geography stays consistent with >> across all cuts: downhill curve, crossing signs, embankments, orange mirror on the right shoulder, sea always beyond the road; travel direction constant toward screen-left. Handheld breathing holds in every single segment, and the aged-film texture holds identically throughout: grain, gate weave, flicker, dust, halation and faded grade never weaken, frame stays vignette-free with even corner brightness; no segment turns rigid, stabilized, sharp or digitally clean. Only these two people appear; the road stays otherwise empty, no cars, no train.show more

WasifAI
14,191 次观看 • 5 天前
I would like to explain the latest batch of... viral videos I'm working on to the bemused brainrot-curious reader who is not familiar with "the culture". Why are these characters, mixed with this song, going viral? It's all about connecting infinite referential mirrors. What makes this video interesting are not its individual parts but the signifier links it draws. Let's look at the individual parts: ONE: The song is a Brazilian funk or "pancadão" song called MC Lan e MC WM - Sua Amiga Vou Pegar, these days part of what's broadly referred as Brazilian phonk or just phonk (not to be confused with the original phonk, a Memphis-derived genre from the early 2010s built around chopped Three 6 Mafia samples, cowbells and lo-fi tape hiss and etc. The Brazilian version comes an entirely different lineage and got its name adapted from “funk” to “phonk” exclusively because the names sounded similar. It has a similarly menacing posture but swaps the rap cadence for funk's 4/4 with kicks on 1 and 3 rhythm and a much heavier, distorted 808 synth sound). Phonk is often used for its exaggerated reverb feeling bass lines to signify power, style or simply "aura", which you can take as a shorthand for poise, coolness, being de-bon-air and a general detached positive feeling of high status. Aura. Because most users cannot understand the Portuguese lyrics (which are often quite vulgar and sexual), the singing takes the characteristic of a chant, something to be appreciated entirely for its sound, texture and gravitas. The vocals are just another instrument where you can appreciate the menace and swagger of the delivery directly without the cognitive friction of meaning. Non-Portuguese-speaking audiences are not missing anything they were supposed to get, they get “the vibe” that matters, which is not lyrical. These songs are often paired with (male) characters that are taken to display these traits like American Psycho's Patrick Bateman (yes, yes I know that’s the opposite of what you should feel about the character), Peaky Blinder's Thomas Shelby and a menagerie of anime characters like Satoru Gojo (Jujutsu Kaisen), Yujiro Hanma (Baki) and Goku and, really, any male character that is just a little bit cool. TWO: The man in the suit is a minor Family Guy character called Tom Tucker. The reference comes from a scene where Meg sees him walking through her school and says "It's Tom Tucker from the news!” We then cut to her POV, where he is walking in slow motion with soft romantic music swelling and birds chirping, the whole love-at-first-sight trope. Then a camera crew member off-screen yells "hurry up Mr. Tucker," and we get to see he is not walking in slow motion because Meg is infatuated, he is just walking that slowly in real life. Only the music and the birds were in her head. The gag is built on the viewer recognizing the romantic-slow-motion trope, briefly accepting it as the scene's reality, and then being shown that we (and Meg) projected the trope onto what is actually just a man walking very slowly. HA! The original gag is already about projection: a neutral image (slow walk) being assigned an external meaning (romance) by a viewer's pattern-recognition. This is what makes the edit-culture appropriation work so well. The clip got stripped of its context, paired with phonk and text overlays (AURA or “Me and the boys going to detention”), and retroactively assigned a new meaning, only this time it’s the cinematic nonchalant walk, the slow deliberate gait that signifies a man who knows he's the most important thing in the frame (ta la any 1980s Schwazerneggerian action movie hero walking away from an explosion without looking back, every yakuza boss entering a room, every western gunslinger approaching the duel). The edit is ostensibly projecting a trope onto a neutral image. The first projection was romance; the second projection is aura. Family Guy clips and gifs are easy to access and repost, which makes it a readily available and easy to use building block. The show has, through sheer volume of output and over two decades of YouTube and cable TV saturation, become a kind of public-domain visual library, a default vocabulary that any editor can pull from knowing the audience will recognize the source without having to be told, and we can just keep loading meaning onto it. THREE: The character in the background is Tom, from Tom and Jerry, doing a pose made famous by an iShowSpeed fan who encountered him during a livestream. By quickly and correctly identifying Speed by his full legal name ("Darren Jason Watkins Jr"), she showcased herself to be a true fan, which he responded to with his characteristic exaggerated reactions. The pose the girl hit, with the knowing look to the camera, produced a perfect “aura moment” complete commitment, zero irony, the unshakeable conviction that what she was doing was the coolest possible thing to do. As a result, the clip then got endlessly edited with "aura 🥶🥶🥶" captions to canonize it. Aura, in this lexicon, is not granted by the universe; it is summoned by the person's own belief that they have it and by displaying the correct attitude. Tom is also dressed as the previously mentioned Thomas Shelby from Peaky Blinders, which is itself a double signifier. The name match (“Thomas”, get it?) and the suit-and-flat-cap costume turn the cartoon cat into a stand-in for the perhaps most used "high-aura" male character of the past decade, the brooding gangster patriarch whose every cigarette drag has been set to phonk, cinematic scores and electronic music a thousand times over. On top of that, he is made entirely out of chrome, a popular trope of asking ChatGPT (one of the few AI tools people have easy and broad access to) to render things out of very high quality materials to indicate "rarity" or "status" like diamonds, platinum and etc. A sign that itself descends from a longer lineage of in-game cosmetic rarity tiers (League of Legends, MMOs, various skin economy freemium game, the Fortnite battle pass, the Pokémon shiny, dacha games and etc) where material finish is the visual shorthand of value. So "chrome" or "platinum" Tom on top of all previous signifiers signals a “maximized” or “maxxd” version. The image is suppose to invoke the superlative highest possible tier, rarest-drop, legendary-rarity version of aura, the way a kid in a playground would describe their dad as not just strong but the strongest in the world. FOUR: Finally, the background black hole calls back to the original Tom image, where he is surrounded by the universe itself, having ascended. The character has transcended the diegetic frame of his own cartoon and now exists at a cosmological scale, with the black hole standing in for the kind of unmotivated, vibes-based "cosmic" imagery that has become the default background for any video trying to signify that something Big is happening (the same visual motif that has powered comic book characters, anime transformations, video game power ups and anything wants to feel grandiose or “epic” without specifying what about). The black hole means significance in the abstract. At this point I think you understand the mechanism at play here. None of these references resolve to a stable meaning on their own. Tom Tucker is “cool” only in the very short context in which his image served as a substrate; he was convenient footage to pair with a song, and the absurdity of doing an "aura edit" on such a minor, strange character scene makes it all funnier and easier to share. Tom-the-cat is doing the aura pose > the aura pose comes from the iShowSpeed girl > the iShowSpeed girl was cool because she correctly played her part in an established bit of a large streamer with the correct timing and theatrical flair > the bit was cool because it was a shared convention unified by a popular central streamer figure > the convention existed because phonk edits had already trained this exact scenario to be read as confidence-plus-detachment as aura > the chrome finish points to AI image generation quirks > the AI image generation style can be mapped to gaming visual rarity shorthands; the gaming rarity tiers point to a much older logic of precious-metal-as-status. Each step on the referential chain is propped by the one behind it, and the one behind it is propped up by the one behind that, so on and so forth. There is no natural endpoint, the entire structure functions more akin to a network than a linked list. If you stop at any single point and ask "but why is particular signifier cool or funny or interesting”, the answer is always "because of the thing behind it.” It’s hyper-citation, Here, what matters is the structure of the whole rather than the content. This is structure is what I mean by infinite referential mirrors. The rate at which a concept is referencing, remixing and calling back to another is what’s interesting. In other words, It’s the velocity that matters. The chain of recognitions, each "I get that reference," and the cumulative effect of getting six references stacked on top of each other a short span of time gives you the feeling that you are participating in something dense and alive, because it allows you to recognize the shared meme ecosystem of the platform that you are participating in, even if only a glimpse of it. You are inside the culture rather than outside it. The brainrot-curious reader who watches this video and feels nothing, has “failed” to understand the joke because they are outside the hall of mirrors I am describing. You can only get the magic if you step in and start counting the reflections: the song, the suit, the cat, the chrome, the black hole, the transitions the video uses. You are looking at connected parts of this network of symbols and at the speed at which one image hands you off to the next. The entire thirteen-second clip is functioning as a single compressed referential payload that decompresses in the viewer's head into a small private essay exactly like this one. The video allows you to recognize yourself as someone capable of decoding it, and that recognition is the reward. That’s why media like this goes viral.show more

Pleometric
68,776 次观看 • 2 个月前
🚨OPERATIONAL UPDATE: ISRAEL U.S. WAR WITH THE ISLAMIC REPUBLIC... - Reporting Window: 3/14 to 3/16 • The war widened further into the Gulf economy, with drone incidents near Dubai Airport, disruption at Fujairah energy infrastructure (also in the UAE), and continued pressure around the Strait of Hormuz • Israel continued deep strike waves inside Iran while expanding ground operations against Hezbollah in southern Lebanon • Iran continued missile salvos toward Israel while activating proxy pressure fronts across Iraq • The U.S. began pushing for a multinational Hormuz security coalition while calibrating escalation to avoid a global oil shock The last 48 hours showed that the war is no longer just about missile exchanges between Israel and Iran. It is increasingly a contest over the regional system itself: energy flows, maritime routes, proxy networks, and command infrastructure. While Israel continues to degrade Iran’s military capabilities, Iran is trying to widen the battlefield economically and geographically. 📽️VIDEO 1: U.S. strikes on Iran’s Kharg Island oil export hub. 📽️VIDEO 2: Smoke rising from oil infrastructure in Fujairah after drone debris caused a fire. *⃣ GULF FRONT: THE ECONOMIC WAR DEEPENED This remained the most strategically important development. Following earlier strikes around Kharg Island and threats to the Strait of Hormuz, pressure on Gulf infrastructure continued. Drone incidents and debris related fires disrupted operations near Fujairah’s energy infrastructure, one of the world’s largest bunkering hubs in the UAE. Shortly afterward, a drone related incident near Dubai International Airport ignited a fuel tank and temporarily disrupted flight operations before authorities contained the fire and resumed traffic. These incidents show the war repeatedly touching civilian energy and logistics infrastructure in the UAE, not just military facilities. Even limited disruptions matter in this region. The Strait of Hormuz handles roughly 20 percent of global oil supply, and repeated incidents have pushed oil prices above $100 during the week. Iran does not need to fully close Hormuz to achieve strategic leverage. Persistent disruption alone can force insurance spikes, rerouting of shipping, and higher global energy prices. *⃣ KHARG ISLAND: IRAN’S ECONOMIC JUGULAR IS NOW UNDER DIRECT PRESSURE One of the most consequential developments in the war involves Kharg Island, Iran’s primary oil export terminal. Historically, roughly 85 to 90 percent of Iran’s crude exports pass through Kharg, making it the single most important node in the country’s energy economy. Recent U.S. strikes targeted military infrastructure associated with IRGC naval operations near the island, particularly facilities linked to mine laying capability and coastal missile systems. These strikes appear connected to Washington’s warning that Iran must not deploy naval mines in the Strait of Hormuz. There are also scattered reports of secondary explosions and possible infrastructure damage near the port, though there is no credible confirmation that the export terminal itself has been destroyed. That distinction is important. Destroying Kharg outright would cripple Iran’s oil exports overnight and likely trigger a massive oil price spike. Instead, the current targeting pattern appears designed to threaten Iran’s economic lifeline without fully collapsing it, maintaining pressure while avoiding the most extreme global economic consequences. *⃣ IRAN: STRIKES INSIDE THE CORE CONTINUED Inside Iran, Israeli strike activity remained intense. Over the last 48 hours, strikes targeted command infrastructure, missile launch systems, air defense networks, and military production sites across Tehran and other strategic locations. The campaign also hit Mehrabad Airport, where Israeli officials reported destroying aircraft associated with Iran’s leadership. Across central and western Iran, the strike map remains broad. Tehran, Karaj, and several other military zones have continued to appear in overnight strike reporting. This suggests the campaign is still focused on systematically degrading Iran’s military capacity, particularly missile infrastructure and command networks. Rather than shifting toward a narrow endgame phase, the strikes indicate a continued effort to keep Iran’s launch capabilities suppressed. *⃣ LEBANON: THE NORTHERN FRONT IS EXPANDING The Lebanon front also escalated further. Israeli forces expanded ground operations in southern Lebanon and reportedly encircled Khiyam, pushing westward toward the Litani River. This represents a larger ground posture than earlier border operations and indicates Israel is attempting to shape the battlefield against Hezbollah rather than simply retaliating against rocket launches. At the same time, Hezbollah continued firing rockets and drones toward northern Israel, maintaining pressure on the northern front even after suffering extensive infrastructure losses earlier in the war. Israel has continued heavy strikes on Hezbollah infrastructure in Lebanon, including operational sites and logistical facilities tied to the group’s missile network. The northern theater now appears to be entering a phase of attrition and positional pressure, rather than the limited cross border exchanges that characterized earlier weeks. *⃣ IRAQ: PROXY PRESSURE ON THE UNITED STATES CONTINUES Iran aligned militias continued attacks on U.S. positions across Iraq. Over the past several days these groups have launched drones and rockets against American bases and diplomatic infrastructure, including a missile strike on the helipad area of the U.S. embassy compound in Baghdad. These attacks serve two purposes: ➡️First, they impose direct costs on U.S. operations in the region. ➡️Second, they force the United States to divert resources toward base defense and interception missions. Even when damage is limited, the attacks expand the battlefield and complicate the operational environment for U.S. forces. *⃣ IRANIAN MISSILE ATTACKS CONTINUE Iran continued launching missiles toward Israel during this period. Several salvos targeted southern Israel and the Negev region, triggering repeated air raid alerts. Despite continued launches, the overall military impact of these attacks appears limited. Israeli air defense systems including Iron Dome, David’s Sling, and U.S. systems such as THAAD have maintained high interception rates. Iran is still able to launch missiles and drones, but the sustained strikes on launchers, command centers, and production facilities appear to be reducing the scale of its barrages compared to the opening days of the war. *⃣ WASHINGTON: TRUMP SIGNALS A LONGER STRATEGIC GAME The political messaging from Washington over the past 48 hours has also clarified the broader strategic direction. Publicly, Trump has suggested the war could end soon and that most major Iranian targets have already been struck. Operationally, however, the signals point toward preparation for a longer campaign. Washington has begun pushing for a multinational coalition to secure the Strait of Hormuz, urging countries that depend on Gulf energy flows to participate in maritime security operations. At the same time, the United States has been careful not to push escalation to the point of triggering a global energy shock. This balancing act helps explain why certain targets, such as Kharg Island’s export terminal, have been threatened but not completely destroyed. The strategic posture appears to be: ➡️sustain pressure on Iran’s military capability ➡️protect global energy flows ➡️avoid triggering a catastrophic oil price spike Israel’s priorities are somewhat different. Israel is focused primarily on maximizing military degradation of Iran and Hezbollah, even if that increases regional escalation risks. The dynamic between Washington’s economic caution and Israel’s military pressure is likely to shape the next phase of the conflict. *⃣ WHAT CHANGED IN THE LAST 48 HOURS Three developments stand out: ➡️First, the Gulf economic front is becoming central to the war. Drone incidents near Dubai and Fujairah show that the conflict is now directly touching regional infrastructure and global energy flows. ➡️Second, Israel continues to widen the battlefield rather than narrow it. Deep strikes inside Iran and expanded operations in Lebanon suggest the campaign is still in a degradation phase. ➡️Third, the United States is beginning to shift toward coalition management of the conflict, particularly around Hormuz, while trying to prevent the war from triggering a global energy crisis. In short, the war is evolving from a direct military confrontation into a broader struggle over regional stability, energy flows, and long term strategic balance in the Middle East.show more

Inside_Israel_Intel
460,417 次观看 • 4 个月前