Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Twitch streamers are cooked. FlowAct-R1 by ByteDance: - Interactive humanoid video generation; - streams infinite-length 480p 25fps, ~1.5s TTFF; - one-shot full-body control; - based on Seedance; - distillation cuts inference to just 3 NFEs.

33,598 görüntüleme • 6 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

In just one week, Binh and I trained a full-body Unitree G1. Here's a recap: 1. Secured a Unitree G1 humanoid through a LinkedIn post 2. Deployed TWIST2 full-body teleoperation pipelines 3. Adapted TWIST2 for Zed stereo camera & collected full-body teleoperation samples (carried by Binh ) 4. Adapted & fine-tuned NVIDIA Gr00T N1.5 VLA on the TWIST2 public datasets, which I fine-tuned on an 8xNVIDIA H100 Cluster. We picked Gr00T N1.5 as it was trained with Unitree G1 embodiment data. 5. Adapted the TWIST2 codebase to stream in the actions from Gr00T via ZMQ using a co-located NVIDIA H100 for ~200ms inference latency 6. Tested the model in sim, then deployed to the real-world Unitree G1. We streamed a training sample observation to the VLA (as we didn't want to break robot in case real observations were OOD) We were the first team in the world to deploy the full TWIST2 data collection pipeline to the unitree g1 :) Much more work ahead though, which I'll work on as a side-project over the next months: 1. Exploring the various types of 'world models': video backbones, dynamics models, v-jepa-2 models. I believe these will generalize better & train much more data-efficiently than VLM backbones 2. Speeding up inference - I believe low-latency robotics inference will be a big challenge. There are many works in video diffusion which I'd like to test (e.g. SageAttention, SparseAttention, Drifting Models). Perhaps also writing custom CUDA kernels. 3. Economics of inference scaling :) What will be the compute demands as we scale inference up to millions of humanoids? Will it run on edge or on distributed 'co-located' inference clusters? These are questions I'd like to answer. Adapted TWIST2 codebase: Adapted Gr00T-N1.5 codebase: The ETH Robotics Club are doing a cool GTC Golden ticket competition with NVIDIA , so this is my submission :) The DGX Spark compute will get me a long way with initial prototyping & especially working on inference optimization for next-gen Blackwell GPUs #NVIDIAGTC #GOLDENTICKET #ETHRC

Arnie Ramesh

23,236 görüntüleme • 6 ay önce

veo 3.1 fast vs seedance 2.0 vs grok imagine vs happyhorse 1.1 four video models pulled from the openrouter video leaderboard by request count this week (skipping the duplicate google/bytedance variants to get four distinct labs): #1 veo 3.1 fast (Google DeepMind) – 45k requests #3 seedance 2.0 (bytedance) – 22k requests #5 grok imagine video (SpaceXAI) – 9k requests #8 happyhorse 1.1 (Alibaba Group) – 4k requests so we tested them. 3 prompts, text-to-video, 16:9 / 720p / 8s, real-player likeness fed in as reference images where the model allowed it. all run via AI/ML API in the run-up to the 2026 world cup final – argentina vs spain – we built three broadcast moments from that tie. each one has to be mechanically correct, not just pretty: • stadium flyover – 80k-seat bowl, argentina vs spain, one continuous descending aerial spiral, tifo + flares, golden-hour / floodlight split • penalty – lamine yamal (spain #19) vs emiliano martínez (argentina keeper): run-up, single strike, full-stretch dive, ball in the net. real faces via reference • free kick – messi 25m out, five-man wall, curl up and over the wall into the top corner. the wall has to face the ball with arms pinned down, like a real wall the takeaway up front: the gap that decides this isn't quality – it's moderation. three of the four refuse to render real footballers' faces (grok was the only one that took every reference), so most of the test had to be reshot "faceless" – camera behind the player. the price spread on top of that is ~4x overall results: cost #1 grok – $1.56 #2 veo 3.1 fast – $3.12 #3 happyhorse – $4.38 #4 seedance 2.0 – $6.00 generation time #1 grok – 4m 16s #2 veo 3.1 fast – 4m 26s #3 happyhorse – 8m 16s #4 seedance 2.0 – 10m 38s avg bitrate (picture density) #1 grok – 12.0 mbps #2 veo 3.1 fast – 11.1 mbps #3 happyhorse – 7.4 mbps #4 seedance 2.0 – 5.3 mbps real faces allowed ✅ grok – took every reference ❌ happyhorse – yamal ok, messi blocked ❌ veo – blocked ❌ seedance – blocked observations: 1. moderation is the whole story. three of the four blocked at least one real face – veo and seedance refused every reference outright, happyhorse took yamal but rejected messi. only grok rendered all of them. everything else had to be shot from behind so no face shows 2. grok is the outlier: cheapest, densest picture, fastest, and the only one that renders real faces. it won on every axis that mattered here 3. seedance is the anti-grok – 4x the cost, 2.5x the time, half the bitrate, and no real faces. worst value in the set 4. none of them understand football out of the box follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

15,453 görüntüleme • 29 gün önce

Had a couple of run ins with the biggest ghoster in EUW Challenger recently. In just a few days I saw him ghost Raider, Agurin, Lil TommyG, Sinerias and me and I didn't even look that hard. Most people are aware of it as ghosters are fairly rare and obvious in high elo. This problem is exacerbated by the fact that this is apparently a pro player for Bushido Wildcats. BW CINGENEM#KINGO or Ronaldo is the biggest ghoster I've seen in recent years. Possible name change incoming. Apparently he streams on Kick and always has multiple browsers watching twitch, never types so he doesn't get listed as an active viewer within the chat list. Every lobby he will scan multiple streams to see which streamers are out of game/in queue(q times in high elo go up to 20 minutes) then he will ask for a ban of one while he himself bans another person from the lobby. This is just unfathomable to me. High elo players, streamers and pro players all have a vested interest in league of legends success. Nobody wants a toxic environment where everyone is afraid to broadcast their gameplay. To act like this as all 3 is completely parasitic and it relies on the good will and passion for the game of everybody else. When encountered this behavior should be shunned and shamed. You are welcome to verify and check his vods based on opgg list of games vs various streamers. I am just tired of going into games and having to question whether there is foul play be it on my side or theirs. League of legends shines best when everybody plays fair but we all compete with each other. Friendly atmosphere where we all have the same goals. Improve, climb, have fun and entertain the audiences.

Azer Dugalić

119,445 görüntüleme • 1 yıl önce

RIP Arcads 🤯 I built a Claude skill that vibe-codes UGC ads on demand. One product + one prompt = the AI creator, the script, the scene-by-scene shot list, and finished video. All inside Claude. Perfect for DTC brands and agencies who can't afford to keep paying $500-$1,500 per UGC video and waiting 2 weeks for revisions. If you're briefing creators, mailing PR boxes, waiting on first cuts, then asking for re-shoots because the hook didn't land... This skill eliminates the entire loop: → Tell Claude the product, ad angle, and length → Skill writes the GPT Image 2.0 prompt to generate the AI creator from scratch → Skill writes every scene prompt, dialogue line, and delivery direction → Pipe it into Seedance 2.0 with character + product + voice locked → Speed up + caption in CapCut → Ship the ad in 20 minutes No more paying $11 per video on Arcads. What you get: → Perfect character consistency across every scene → Voice consistency that holds clip-to-clip (small CapCut hack inside the tutorial) → Real product fidelity using your actual product photo as a reference → Multi-scene day-in-life, testimonial, and action-shot formats out of the box Built 100% with a Claude skill + Seedance 2.0. I recorded a full step-by-step tutorial showing the exact workflow + the 3 finished AI UGC ads I made for Rhode, AG1, and Barebells. Want to see the full breakdown? > Like this post > Comment "UGC" And I'll send it over (must be following so I can DM)

Mike Futia

34,322 görüntüleme • 3 ay önce

-- What holds it together -- ✂️PAPER CUTS.- When the image is the end ⤵️ What you bring to life is the journey towards the image: how the pieces come to be as they are. The movement is the idea, and your image in Seedance is the final frame. ---------------------------------------------------- What happens to the object? 1⃣ By the time the object TRANFORMS into something else. It takes place in a continuous shot, a single camera that zooms in slowly and never cuts away. A cut would break the transformation; the hypnotic effect lies precisely in the fact that it doesn’t flicker. Template.- [@.RE IMAGE] is the LAST frame. [GLOBAL] [Technique + light + backdrop + focus]. [Real materials]. One continuous hypnotic transformation — a single take, no cuts. The same [OBJECT A] does not appear in pieces; it [VERB: melts / folds / blooms / dissolves] into [OBJECT B]. Slow surreal dreamlike drift, one unbroken slow push, slight stop-motion shimmer, no snapping. 0-1s: [state A, intact and recognizable]. 1-2.5s: [the change begins — "the same material begins to..."]. 2.5-3.5s: [the change at full — "...becomes..."]. 3.5-5s: [it settles]. The push eases to rest. Locked, exact match to the last frame. [LOGIC RULE] one continuous same-lens push, never cutting. [A] morphs into [B], [details that must not deform] stay legible, no warping. hypnotic drift. SFX: [a sound that also transforms]. no music, normal speed. 2⃣ When one thing LEADS to another, we need each step to be a link in the chain of causality. Template.- [@.RE IMAGE] is the LAST frame. [GLOBAL] [Technique + light + backdrop + focus]. [Real materials]. Hypnotic stop-motion paper cadence, slight frame-step, brisk causal montage, ~1.5s per shot, no naturalistic motion, no slow-mo. [cut] [shot + camera] / [carries on from previous cut] / [ACTION: what it DOES, not how it looks]. SFX: [beat-anchored hit]. [cut] ... (link 2 — chains from link 1) [cut] ... (link 3) [cut] ... (link 4) [cut] ... ease back to reveal / [link 5]. Locked, exact match to the last frame. SFX: ... [LOGIC RULE].- [materials], [what must NOT happen], no warping, [text legible if any]. ~1.5s per shot, no slow-mo. no music.

AlexandrIA

41,320 görüntüleme • 1 ay önce

Claymotion ads are crushing it on Meta right now. Built a free claude skill to make them 👇 If you've been scrolling Meta lately, you've seen them — stop-motion clay characters, tactile textures, weirdly satisfying to watch. CTRs are 2-3x the feed average. Almost nobody is running them. The problem: they look impossible to make unless you have a studio. They're not. You just need the right prompts. So I packaged the prompt system as a Claude Code skill. It's free. Here's what it does: Paste your product URL. Out comes a full claymotion ad plan: 1/ Shot-by-shot storyboard 5-7 shots with the narrative arc. Setup → product reveal → payoff → CTA. 2/ Image prompt per shot Exact prompt you paste into Midjourney, Nano Banana, or any image gen. Camera angle, lighting, clay texture specs, character details — dialed in for consistency across shots. 3/ Video prompt per shot The animation prompt you paste into Kling, Veo, Seedance, or Sora. Motion direction, pacing, transitions — so the shots actually flow. 4/ VO script per shot Voiceover copy written for rhythm. Timed to the shot length. Hook, body, CTA — all on brand. 5/ Music + sfx direction Tone notes for the track. Specific sfx cues per shot (squish, pop, whoosh) You take the outputs. Paste them into your image + video generators. Stitch the shots. Record the VO. A full claymotion ad in under an hour, at the cost of a few API credits. Instead of $3,000 and 3 weeks with an animation studio. Why claymotion works right now: → Pattern break — nothing else in the feed looks like it → Tactile feel — clay reads as "real" even when AI-generated → High dwell time — people watch the whole thing → Cheap to test — 5-10 variations per product is now feasible Comment "Clay" and I'll send you: → The Claude Code skill (free) → A starter prompt pack → 3 example storyboards so you can see the output (must be connected)

Ahad Shams

16,855 görüntüleme • 4 ay önce

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

128,678 görüntüleme • 1 ay önce

f*ck it. i'm LEAKING my entire Seedance 2.5 AI UGC system i cracked the formula for generations that look like real UGC filmed on a phone.. no camera, no actor, no set video below has been generated 100% by Seedance 2.5 + Claude... here's the FULL system: 1. open Seedance 2.5 and upload 3 references: Image your character, image2 the environment, David Carbon the style frame 2. assign every file a role in the prompt: "Image is the main character, image2 is the location, match David Carbon for the look".. never leave a file unassigned or the model guesses 3. write a timestamped storyboard instead of one messy paragraph: "0-8s: wide shot, she picks up the product. 9-16s: close-up, reaction. 17-30s: talks to camera".. the model follows it beat by beat like a shooting script 4. kill the AI look with one line: "retain real micro pores and skin texture, natural imperfections, no beauty retouching, no commercial polish" 5. add global rules at the end: "no subtitles, no fast cuts, same person in every frame" one generation. 30 seconds of footage. consistent character, zero drift now how to turn it into money, fast: > faceless youtube: stack scenes into 8-10 min docs. history, true crime, luxury pay $8-15 RPM, channels clear $5-20k/mo > UGC ads for brands: DTC brands pay creators $300-2,000 per video. yours cost minutes > sell the service: run this exact system for local businesses as a monthly retainer and the timing is unfair: Seedance 2.5 went live on Higgsfield today with 33 days of unlimited generations at zero credit cost on your plan a full month of unlimited practice and library-building before anyone else catches on most people will scroll past this don't be most people next i guess to make it automated P.S. thanks Higgsfield for sponsoring this post

Ronin

14,935 görüntüleme • 8 gün önce