Loading video...

Video Failed to Load

Go Home

Twitch streamers are cooked. FlowAct-R1 by ByteDance: - Interactive humanoid video generation; - streams infinite-length 480p 25fps, ~1.5s TTFF; - one-shot full-body control; - based on Seedance; - distillation cuts inference to just 3 NFEs.

33,598 views • 6 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

In just one week, Binh Pham and I trained a full-body Unitree G1. Here's a recap: 1. Secured a Unitree G1 humanoid through a LinkedIn post 2. Deployed TWIST2 full-body teleoperation pipelines 3. Adapted TWIST2 for Zed stereo camera & collected full-body teleoperation samples (carried by Binh Pham ) 4. Adapted & fine-tuned NVIDIA Gr00T N1.5 VLA on the TWIST2 public datasets, which I fine-tuned on an 8xNVIDIA H100 Cluster. We picked Gr00T N1.5 as it was trained with Unitree G1 embodiment data. 5. Adapted the TWIST2 codebase to stream in the actions from Gr00T via ZMQ using a co-located NVIDIA H100 for ~200ms inference latency 6. Tested the model in sim, then deployed to the real-world Unitree G1. We streamed a training sample observation to the VLA (as we didn't want to break robot in case real observations were OOD) We were the first team in the world to deploy the full TWIST2 data collection pipeline to the unitree g1 :) Much more work ahead though, which I'll work on as a side-project over the next months: 1. Exploring the various types of 'world models': video backbones, dynamics models, v-jepa-2 models. I believe these will generalize better & train much more data-efficiently than VLM backbones 2. Speeding up inference - I believe low-latency robotics inference will be a big challenge. There are many works in video diffusion which I'd like to test (e.g. SageAttention, SparseAttention, Drifting Models). Perhaps also writing custom CUDA kernels. 3. Economics of inference scaling :) What will be the compute demands as we scale inference up to millions of humanoids? Will it run on edge or on distributed 'co-located' inference clusters? These are questions I'd like to answer. Adapted TWIST2 codebase: Adapted Gr00T-N1.5 codebase: The ETH Robotics Club are doing a cool GTC Golden ticket competition with NVIDIA , so this is my submission :) The DGX Spark compute will get me a long way with initial prototyping & especially working on inference optimization for next-gen Blackwell GPUs #NVIDIAGTC #GOLDENTICKET #ETHRC

Arnie Ramesh

14,815 views • 6 months ago

Had a couple of run ins with the biggest ghoster in EUW Challenger recently. In just a few days I saw him ghost Raider, Agurin, Lil TommyG, Sinerias and me and I didn't even look that hard. Most people are aware of it as ghosters are fairly rare and obvious in high elo. This problem is exacerbated by the fact that this is apparently a pro player for Bushido Wildcats. BW CINGENEM#KINGO or Ronaldo is the biggest ghoster I've seen in recent years. Possible name change incoming. Apparently he streams on Kick and always has multiple browsers watching twitch, never types so he doesn't get listed as an active viewer within the chat list. Every lobby he will scan multiple streams to see which streamers are out of game/in queue(q times in high elo go up to 20 minutes) then he will ask for a ban of one while he himself bans another person from the lobby. This is just unfathomable to me. High elo players, streamers and pro players all have a vested interest in league of legends success. Nobody wants a toxic environment where everyone is afraid to broadcast their gameplay. To act like this as all 3 is completely parasitic and it relies on the good will and passion for the game of everybody else. When encountered this behavior should be shunned and shamed. You are welcome to verify and check his vods based on opgg list of games vs various streamers. I am just tired of going into games and having to question whether there is foul play be it on my side or theirs. League of legends shines best when everybody plays fair but we all compete with each other. Friendly atmosphere where we all have the same goals. Improve, climb, have fun and entertain the audiences.

Azer Dugalić

119,445 views • 1 year ago

RIP Arcads 🤯 I built a Claude skill that vibe-codes UGC ads on demand. One product + one prompt = the AI creator, the script, the scene-by-scene shot list, and finished video. All inside Claude. Perfect for DTC brands and agencies who can't afford to keep paying $500-$1,500 per UGC video and waiting 2 weeks for revisions. If you're briefing creators, mailing PR boxes, waiting on first cuts, then asking for re-shoots because the hook didn't land... This skill eliminates the entire loop: → Tell Claude the product, ad angle, and length → Skill writes the GPT Image 2.0 prompt to generate the AI creator from scratch → Skill writes every scene prompt, dialogue line, and delivery direction → Pipe it into Seedance 2.0 with character + product + voice locked → Speed up + caption in CapCut → Ship the ad in 20 minutes No more paying $11 per video on Arcads. What you get: → Perfect character consistency across every scene → Voice consistency that holds clip-to-clip (small CapCut hack inside the tutorial) → Real product fidelity using your actual product photo as a reference → Multi-scene day-in-life, testimonial, and action-shot formats out of the box Built 100% with a Claude skill + Seedance 2.0. I recorded a full step-by-step tutorial showing the exact workflow + the 3 finished AI UGC ads I made for Rhode, AG1, and Barebells. Want to see the full breakdown? > Like this post > Comment "UGC" And I'll send it over (must be following so I can DM)

Mike Futia

34,310 views • 3 months ago

-- What holds it together -- ✂️PAPER CUTS.- When the image is the end ⤵️ What you bring to life is the journey towards the image: how the pieces come to be as they are. The movement is the idea, and your image in Seedance is the final frame. ---------------------------------------------------- What happens to the object? 1⃣ By the time the object TRANFORMS into something else. It takes place in a continuous shot, a single camera that zooms in slowly and never cuts away. A cut would break the transformation; the hypnotic effect lies precisely in the fact that it doesn’t flicker. Template.- [@.RE IMAGE] is the LAST frame. [GLOBAL] [Technique + light + backdrop + focus]. [Real materials]. One continuous hypnotic transformation — a single take, no cuts. The same [OBJECT A] does not appear in pieces; it [VERB: melts / folds / blooms / dissolves] into [OBJECT B]. Slow surreal dreamlike drift, one unbroken slow push, slight stop-motion shimmer, no snapping. 0-1s: [state A, intact and recognizable]. 1-2.5s: [the change begins — "the same material begins to..."]. 2.5-3.5s: [the change at full — "...becomes..."]. 3.5-5s: [it settles]. The push eases to rest. Locked, exact match to the last frame. [LOGIC RULE] one continuous same-lens push, never cutting. [A] morphs into [B], [details that must not deform] stay legible, no warping. hypnotic drift. SFX: [a sound that also transforms]. no music, normal speed. 2⃣ When one thing LEADS to another, we need each step to be a link in the chain of causality. Template.- [@.RE IMAGE] is the LAST frame. [GLOBAL] [Technique + light + backdrop + focus]. [Real materials]. Hypnotic stop-motion paper cadence, slight frame-step, brisk causal montage, ~1.5s per shot, no naturalistic motion, no slow-mo. [cut] [shot + camera] / [carries on from previous cut] / [ACTION: what it DOES, not how it looks]. SFX: [beat-anchored hit]. [cut] ... (link 2 — chains from link 1) [cut] ... (link 3) [cut] ... (link 4) [cut] ... ease back to reveal / [link 5]. Locked, exact match to the last frame. SFX: ... [LOGIC RULE].- [materials], [what must NOT happen], no warping, [text legible if any]. ~1.5s per shot, no slow-mo. no music.

AlexandrIA

41,320 views • 1 month ago

Claymotion ads are crushing it on Meta right now. Built a free claude skill to make them 👇 If you've been scrolling Meta lately, you've seen them — stop-motion clay characters, tactile textures, weirdly satisfying to watch. CTRs are 2-3x the feed average. Almost nobody is running them. The problem: they look impossible to make unless you have a studio. They're not. You just need the right prompts. So I packaged the prompt system as a Claude Code skill. It's free. Here's what it does: Paste your product URL. Out comes a full claymotion ad plan: 1/ Shot-by-shot storyboard 5-7 shots with the narrative arc. Setup → product reveal → payoff → CTA. 2/ Image prompt per shot Exact prompt you paste into Midjourney, Nano Banana, or any image gen. Camera angle, lighting, clay texture specs, character details — dialed in for consistency across shots. 3/ Video prompt per shot The animation prompt you paste into Kling, Veo, Seedance, or Sora. Motion direction, pacing, transitions — so the shots actually flow. 4/ VO script per shot Voiceover copy written for rhythm. Timed to the shot length. Hook, body, CTA — all on brand. 5/ Music + sfx direction Tone notes for the track. Specific sfx cues per shot (squish, pop, whoosh) You take the outputs. Paste them into your image + video generators. Stitch the shots. Record the VO. A full claymotion ad in under an hour, at the cost of a few API credits. Instead of $3,000 and 3 weeks with an animation studio. Why claymotion works right now: → Pattern break — nothing else in the feed looks like it → Tactile feel — clay reads as "real" even when AI-generated → High dwell time — people watch the whole thing → Cheap to test — 5-10 variations per product is now feasible Comment "Clay" and I'll send you: → The Claude Code skill (free) → A starter prompt pack → 3 example storyboards so you can see the output (must be connected)

Ahad Shams

16,855 views • 4 months ago

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

128,678 views • 1 month ago

f*ck it. i'm LEAKING my entire Seedance 2.5 AI UGC system i cracked the formula for generations that look like real UGC filmed on a phone.. no camera, no actor, no set video below has been generated 100% by Seedance 2.5 + Claude... here's the FULL system: 1. open Seedance 2.5 and upload 3 references: Image your character, image2 the environment, David Carbon the style frame 2. assign every file a role in the prompt: "Image is the main character, image2 is the location, match David Carbon for the look".. never leave a file unassigned or the model guesses 3. write a timestamped storyboard instead of one messy paragraph: "0-8s: wide shot, she picks up the product. 9-16s: close-up, reaction. 17-30s: talks to camera".. the model follows it beat by beat like a shooting script 4. kill the AI look with one line: "retain real micro pores and skin texture, natural imperfections, no beauty retouching, no commercial polish" 5. add global rules at the end: "no subtitles, no fast cuts, same person in every frame" one generation. 30 seconds of footage. consistent character, zero drift now how to turn it into money, fast: > faceless youtube: stack scenes into 8-10 min docs. history, true crime, luxury pay $8-15 RPM, channels clear $5-20k/mo > UGC ads for brands: DTC brands pay creators $300-2,000 per video. yours cost minutes > sell the service: run this exact system for local businesses as a monthly retainer and the timing is unfair: Seedance 2.5 went live on Higgsfield today with 33 days of unlimited generations at zero credit cost on your plan a full month of unlimited practice and library-building before anyone else catches on most people will scroll past this don't be most people next i guess to make it automated P.S. thanks Higgsfield for sponsoring this post

Ronin

14,935 views • 6 days ago

$AMD $5 Trillion MC Is Inevitable Long Term👑 This thread will focus more on Inference! 2026 EPYC "Venice" $TSM 2nm to save Large GW Scale Inference by 40% more than Prior Turin gen. Context: EPYC Turin achieves ~$0.001 per million tokens for batch inference vs $0.02-$0.12/ million tokens as I wrote the thread below. Venice is going to lower cost down to $0.0005-$0.0006/Million Tokens. OpenAI spent roughly $20B on Inference and Training, where 80-90% of that was for Inference per Analysts. AKA Renting Compute is Expensive AF! In this thread, I want to focus on why most analysts and investors are underestimating the role EPYC "Venice" and future Gen on overall Data center revenue. And $TSM ramping up 2nm supply early is a confirmation that AMD will be a major buyer long term. I will also link the thread the Gap between AMD Analysts & Reality and 2nm Ramp Thread so you have more comprehensive view of what I'm writing here. Before I go into detail this is my 2026 Projection: AI GPUs: $35-$50B EPYC Data Center: $15B-$17B Client Segment: $12-$13B Gaming: $6B Embedded: $4B-$5B Total Revenue $70-$100B Non-GAAP net income $18B-$25B Non-GAAP EPS $10.97-$15.40 Foward P/E 55x-70x= $603-$1,078 AMD's Analysts are projecting $0 Revenue for MI450 and sluggish EPYC Growth. Meaning, all analysts are either full of 💩 or Sexist, you decide! Analysts are also projecting 0% growth on AMD "Secret Weapon" Chip as $MSFT said we are at significant Windows refresh and upgrade cycle. Do you think TSMC would allocate more 2nm supply to $AMD at $0 MI450 revenue and sluggish EPYC? 1. EPYC is going to be the leader in lowest Inference! Current Turin cost saving is 95% vs $NVDA or 98-99% on Inference cost when you factor in renting Inference compute from Amazon Web Services, Microsoft Azure, or $NVDA Neocloud pets. TSMC claimed: 10-15% higher performance at iso-power, 25-30% lower power at iso-speed, and ~15% higher transistor density compared to 3nm. This reduces operational expenses (energy, cooling) while increasing throughput per chip. EPYC Turin achieves ~$0.001 per million tokens for batch inference (via vLLM on models like Llama 3 70B), driven by high core counts and low hardware costs. EPYC Venice offers ~1.7x overall performance and up to 70% more compute capability per core, with up to 256 cores (512 threads). Enhanced vector/AI instructions and open-source firmware (openSIL) optimize for inference workloads. AMD Incorporates AI Engines (now part of AMD's XDNA) for on-chip acceleration, improving efficiency for low-latency and edge inference. This reduces reliance on discrete GPUs, lowering system complexity and TCO. Venice SKUs are projected at $3,000-$15,000 ($5,000 for 256-core flagship), far below NVIDIA Rubin ($50,000-$90,000) or AMD's own MI450 GPUs ($40,000-$50,000). High memory bandwidth (up to 1.6 TB/s) supports efficient batch inference. Venice is designed exactly for Large customers that want to lower Inference Cost and MI450 Helios is for Customers that want Training at lowest TCO, TDP as well as lower Upfront 1GW scale(Full build $35-$40B vs $NVDA $55B-$80B). 2. Real World Example: OpenAI's 2025 inference spend reached ~$20B, escalating to even higher total compute rental (mostly inference) amid token volume growth(from video generating). By 2026, with usage doubling (consistent with industry trends: token demand grows 2-5x YoY), assume OpenAI processes ~1,800 billion million-tokens annually $NVDA Blackwell at $0.02-$0.12 is $36B(most optimized) Rubin is projected to be at $0.01/million tokens or $18B annual Inference Cost vs $AMD Venice $0.0005/million tokens or $0.9B annual Inference Cost => Massive saving for OpenAI or anyone that are paying 80-90% Annual Bill for Inference compute. In short, it is unsustainable to pay this much rent vs owning for all current AI players for the medium to long term. Rubin excels in low-latency decode (if Groq integration from $20B deal in 2027-2028), but Venice dominates batch (80% of inference by 2030). Actual savings depend on deployment scale (OpenAI's 6GW AMD plans), electricity rates, and software maturity. If Rubin only hits $0.03, savings swell to $53.1B vs. $17.1B. 3. Will running Inference on Venice and future Gen slow down response generation in 2026 and beyond? Human perception of "fast enough" for chat, agents, search augmentation, summarization, coding assistance is roughly Meaning, EPYC may generate $100B a year on data center revenue, Hence $MSFT $AMZN $META $GOOGL OpenAI xAI and 42+ Countries are leaning AMD for Inference, because the cost saving is MASSIVE! 4. Regular users (you, me, people using ChatGPT, Claude, Gemini, Grok, Perplexity...) are extremely unlikely to notice any slowdown and in many cases might even experience slightly faster or more consistent response times if the industry heavily shifts toward AMD EPYC for inference. What actually happens when companies save massively on inference? When OpenAI , Anthropic , Gemini , Grok Meta .... save billions on the batch/enterprise/RAG layer using EPYC Venice, they typically do one or more of these things with the savings, none of which make your chat slower but enhancing their bottom line(Profit) ~Keep prices the same → make more profit ~Lower subscription prices / increase free tier limits ~Train bigger & better models more frequently ~Offer longer context windows ~Add more reasoning steps / tool calls / agents per query ~Improve multimodal capabilities ~Build more data centers / reduce throttling during peaks In practice the consumer experience usually gets better, not worse, when inference becomes dramatically cheaper. Prime example is $META leaning AMD heavily or currently AMD largest customer. or Grok 2 to Grok 3 heavily used AMD for Inference saving. And most Grok Users reported Groke responses snappier, not slower. 5. What does this mean for potential Revenue? Noted that TSMC is massively ramping 2nm supply for $AMD both MI450 and EPYC. EPYC Conservative projection: FY2025: $10.5B(best Est) FY2026: $16B FY2027: $29B FY2028: $49B FY2029: $75B FY2030: $100B Large customers: $META OpenAI $MSFT $AMZN $GOOGL xAI (Apple?) Smaller customer: $DELL $HPE $SMCI and 42+ other countries. The roadmap to $5 Trillion is very much inevitable as Inference Cost from Renting or owning $NVDA are too high, but $NVDA will still dominate Training market share, where MI families are likely to take 15-20% market share, but the TAM is also expanding Rapidly. Most Institutions are projecting $2-$3Trillion TAM by 2030. $NVDA said $4 Trillion. Dr. Lisa Su said $1 Trillion+ by 2030. So you decide on how much TAM. If you enjoy this kind of analysis, Slap the Like/Repost and Bookmark to please the X Algo as it is Free.99! If you want to support my work further, consider subscribe to see more in-depth analysis! Alright, that is it. Not Financial Advice!

Mike

102,223 views • 7 months ago