Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Continuous-video agents (computer use, robotics, static scenes) burn compute re-ingesting pixels that didn't move. VLMaxxing teaches a frozen video VLM to skip the reruns. 54 fps perception on Gemma 4 26B, training-free, no accuracy drift. w/ JF Bastien (arXiv 2605.03351)

63,417 Aufrufe • vor 4 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,331 Aufrufe • vor 2 Monaten

Most AI world models can generate beautiful scenes. Keeping those scenes alive for an hour without falling apart is the real challenge. That's what caught my attention about LingBot-World 2.0 (LingBot-World-Infinity) from Robbyant Instead of chasing longer videos, it focuses on something much harder: persistent, interactive worlds that stay coherent while you explore. A few highlights: • Generates worlds from a single frame and continuously responds to live user actions through a causal world model. • Streams stable 720p at 60 FPS in real time. The team reports a continuous 60 minute stress test across 20 different scenarios with no noticeable visual degradation. • Uses a Brain-Cerebellum co-simulation framework where a VLM plans events while the video model turns them into consistent world evolution. • Pilot and Director Agents help drive character behavior and introduce new objects and events. • Open sourced with a 14B flagship model, while the paper also describes a lightweight 1.3B version for a single consumer GPU. There is also an online interactive demo. The biggest takeaway? We're moving beyond AI that generates clips. We're getting closer to AI that generates living, evolving worlds you can actually interact with. And that feels like a much bigger shift than another jump in video quality. Explore more: 💻 Github: 🤗 Weights: 🌐 Website-with videos you can use : 🎮 Try it online: #Robbyant #LingBot #WorldModel #EmbodiedAI #OpenSource #Robotics #ad

Alif Khan

84,543 Aufrufe • vor 2 Monaten

When we started Score, the standard computer vision tools already existed. About a million people use them every day. Most of those people are still waiting on labels, running training jobs by hand, and watching models fail once they leave the test set. Most of those people are still waiting on labels, running training jobs by hand, and watching models fail once they leave the test set. Most of those people are also still waiting on verified computer vision models, evaluated against real life conditions and ready to be deployed for them to deliver value for their teams, clients or users. Score Studio is the full computer vision path in one place. A team describes the problem. The system can generate the missing scenes, label them, train the candidates, evaluate which ones actually hold, and deploy the winner. Data, labels, training, eval, ship. One loop. If no model exists for that job yet, they can put a bounty on the subnet. Anything from a small vision brick to a full VLM. Miners compete on the task. Only the winning work comes back. Same path for software agents. Any agent can call it. Built to be fully agent-accessible. Built for the people who already do this work: computer vision engineers and the small teams around them in plants, warehouses, farms, robotics, sport, and security. And for the agents those teams will run. That is the part that changes the job. Not another training screen. The stretch that used to take a lab and a calendar, footage, boxes, versions, failed runs, a separate deploy project, sits behind one starting point. And if the network needs a new model, that request is part of the same path. We spent more than a year building it. Then we had a choice. Keep it for us, or commoditize the whole subnet and make it available 24/7, in permissionless and open-source way. And we knew we couldn't keep it for us. It had to live on Bittensor. Open source software already showed how this should work. Infrastructure should not sit inside one company. Same idea as open AI before the phrase changed meaning: inspect it, fork it, keep building. That is what SN44 is for. Open vision intelligence, powered by Bittensor. Miners do the work. Studio is how that gets monetized. Profit does not stay in a company account. It goes back into the subnet through buyback and burn. We built the tool we wanted on day one. It will live on the network now, and for ever. Waitlist is open.

Score

11,292 Aufrufe • vor 27 Tagen

🚨 BREAKING: Big news in the computer vision world! 🎥 Luxonis | Robotic Vision just dropped its new OAK 4 line, and it’s a big upgrade for edge computer vision. Instead of being “just a stereo camera,” OAK 4 is a fully standalone vision computer with 52 TOPS of on-device AI. Models run locally, depth is computed locally, and no external PC or cloud pipeline is required. This is why robotics teams love it: lower latency, lower cost, fewer failure points in the field. The hardware is built for the real-world. IP67, shock-resistant, wide-FOV RGB + stereo pair, IR projection, IMU, audio, and a patent-pending calibration system that keeps depth accurate even when conditions change. But the real move is the platform. With Luxonis Hub, you can deploy models, grab telemetry, push OTA updates, or collect data when performance drifts, all from a unified interface. It turns a single device into an end-to-end edge CV system. Most customers today in robotics are groups who just want something that works: AMRs, bin-picking systems, trailer-loading robots, and ag-tech. 🤖 And they all say the same thing, the appeal isn’t raw TOPS, it’s the all-in-one simplicity that lets them scale without building custom infrastructure. Feels like the direction edge vision has been waiting for: rugged hardware + high-throughput on-device compute + a real management layer. A next step toward “plug-and-deploy” perception for robots. 🔗 Find out more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

41,955 Aufrufe • vor 9 Monaten

I tried jack's Buzz. It's like Slack + OpenClaw + Herdr + but with some really unique features that people are sleeping on. The video below shows how it works, and some of my thoughts on the process and platform, e.g.: - Create and interact with agents on top of any harness (claude code, codex, pi, etc.) - Choose which models agents use, including local ones - Agents can delegate work and work in parallel in git worktrees - Agents are first-class citizens and work like humans (creating channels, delegating, access to chat history) - You can share AI compute within a community - It's completely open-source and decentralized Things I like: - Delegating work in chat feels natural: tag an agent, it replies in a thread with status updates as it e.g. compiles, commits, and deploys. - Shared compute: relay owners can share local compute with members, so a community could pool funds for one beefy machine running a local model and everyone uses it. - It's built on Nostr, an open protocol already tied into Bitcoin Lightning so I can imagine communities tipping each other or paying for compute/agent tasks with instant zero-fee micropayments in the future. - It ties together things like OpenClaw, an agent manager, and Slack-style chat into one tool. Things I didn't like: - You can't see what the agent is doing in a terminal. The activity view exists, but if you're used to watching a session run, this UI feels a bit abstracted. A terminal view would be great. - It feels slower than running a session in Claude Code, though no evidence to back that up. For that reason I found myself doing one-off tasks in the terminal instead. Verdict: - I really like it so far and can genuinely imagine working with a team this way. - It doesn't feel ready for big, complex tasks yet. For shallower tasks, it's perfect. - The shared compute + Nostr/Lightning angle is what really separates it from every other agent manager for me, and I think that future is coming.

Vinny

1,357,789 Aufrufe • vor 2 Monaten

In just one week, Binh and I trained a full-body Unitree G1. Here's a recap: 1. Secured a Unitree G1 humanoid through a LinkedIn post 2. Deployed TWIST2 full-body teleoperation pipelines 3. Adapted TWIST2 for Zed stereo camera & collected full-body teleoperation samples (carried by Binh ) 4. Adapted & fine-tuned NVIDIA Gr00T N1.5 VLA on the TWIST2 public datasets, which I fine-tuned on an 8xNVIDIA H100 Cluster. We picked Gr00T N1.5 as it was trained with Unitree G1 embodiment data. 5. Adapted the TWIST2 codebase to stream in the actions from Gr00T via ZMQ using a co-located NVIDIA H100 for ~200ms inference latency 6. Tested the model in sim, then deployed to the real-world Unitree G1. We streamed a training sample observation to the VLA (as we didn't want to break robot in case real observations were OOD) We were the first team in the world to deploy the full TWIST2 data collection pipeline to the unitree g1 :) Much more work ahead though, which I'll work on as a side-project over the next months: 1. Exploring the various types of 'world models': video backbones, dynamics models, v-jepa-2 models. I believe these will generalize better & train much more data-efficiently than VLM backbones 2. Speeding up inference - I believe low-latency robotics inference will be a big challenge. There are many works in video diffusion which I'd like to test (e.g. SageAttention, SparseAttention, Drifting Models). Perhaps also writing custom CUDA kernels. 3. Economics of inference scaling :) What will be the compute demands as we scale inference up to millions of humanoids? Will it run on edge or on distributed 'co-located' inference clusters? These are questions I'd like to answer. Adapted TWIST2 codebase: Adapted Gr00T-N1.5 codebase: The ETH Robotics Club are doing a cool GTC Golden ticket competition with NVIDIA , so this is my submission :) The DGX Spark compute will get me a long way with initial prototyping & especially working on inference optimization for next-gen Blackwell GPUs #NVIDIAGTC #GOLDENTICKET #ETHRC

Arnie Ramesh

23,236 Aufrufe • vor 7 Monaten

this guy built an ai-girl pipeline using real-time face filters, and d2c brands now pay him $2,000 per ugc video he got tired of watching brands burn $4,000 on a single creator who takes 2 weeks to deliver one angle, so he built a setup that runs photoreal ai girls live from his own webcam, no actresses, no studios, no makeup artists his monthly revenue hit $89,000 last month from a network of 7 ai personas across tiktok and instagram. the average ugc creator caps at $6k juggling 4 brand deals the breakdown: > hardware is the moat, but most people butcher the setup in the first frame. face mesh locked at 60fps with zero artifacting > persona comes first, mess this up and nothing saves it: name, backstory, voice tone, niche before a single clip is shot > face selection is not random. you a/b test features (eye spacing, jawline, hair contrast) because some faces convert better in 9:16 > you're picking who your audience trusts, not who looks cool. that's targeting baked into bone structure > real-time physics run before the script, this is what kills the uncanny valley that destroys watch time in 2 seconds > the filter has to survive the strap of a tank top, the texture of a knit cardigan, the hair flick > batching is the move 96% skip: one performance, multiple personas, three platforms > the system pushes 12 pieces before lunch while brands test 2 creators a week and wonder why their cpa sits at $94 the economics: each video costs $4 in compute, sells for $1,500 to $3,000, takes 14 minutes to produce. that's a 37,500% margin, while ugc agencies pay creators $400-800 per clip and net $200 after revisions one supplement brand generated 14 variants with 7 personas in 4 hours and found a winner in 36 hours without flying a creator to la. they were paying $1,200 per ugc video and burning $6,000/week on content that didn't scale. now they spend $210 for 14 variants and their cpa dropped from $89 to $27 the avatars hold real products, warm window light on the persona, cold neon on the operator, mouth shapes sync to consonants not just vowels just a webcam, a tracked face, and the discipline to move enough that the filter never has a chance to break

Kiyoro

31,137 Aufrufe • vor 4 Monaten

One of the largest trading firms in the world teaches new hires poker before it lets them near a book. A newspaper brought a camera to the table and sat down to play. The men across the felt are working Wall Street traders. The firm is Susquehanna, which built poker into its training programme decades ago and still runs it that way, on the argument that the card room teaches something no finance degree does. Gunjan Banerji from the Wall Street Journal plays the hands herself rather than interviewing them about it. A green table, a dealer, chips, 4 people who do this for a living. No lecture hall, no slides. The teaching happens between deals, while money is actually at stake. They break it into parts on camera. Risk management first. Then bet sizing. Then patience, which sounds like the soft one and is not. Then reading the person opposite when the only data available is how they behave with money on the table. The section on patience is the one most people skip. Folding is the correct action in the overwhelming majority of hands, and almost nobody can do it for hours without inventing a reason to play. The same failure shows up in a trading account as overtrading, and it kills more people than bad analysis. Then bet sizing, which is where the video earns the watch. Being right about the odds is the easy half. How much you put behind a correct read is the part that ends careers, and they work through it hand by hand instead of describing it in the abstract. The turn is what the traders admit about being wrong. A good decision loses regularly, a bad decision wins regularly, and the only way to last is to grade the process instead of the result. Everything else in the video sits downstream of that. It matters more now than when the firm started running these tables. Every model prices probability in a second and hands it to you for free. Nothing on the screen tells you how much of your account to put behind the number, or what to do after the number was right and you lost anyway. Free on YouTube, produced by a newspaper, filmed at a real table with real hands. The maths is public. The sizing is the job. 1 table. 4 traders. It is in the video.

the lich

62,689 Aufrufe • vor 1 Monat

X-Humanoid just officially dropped Embodied Tien Kung 3.0, A universal platform designed to be way more open and developer-friendly. 🤖 Built on their Wise Kaiwu AI platform, this next-gen humanoid is all about slashing development costs. It’s a fully interoperable ecosystem that supports everything from tactile interaction to high-dynamic motion control at a full humanoid scale. ➤ Radical Openness: X-Humanoid is open-sourcing the full stack—robot body, motion control, VLM/VLA models, and the RoboMIND dataset. It fully supports ROS2, MQTT, and TCP/IP, so developers can customize use cases without re-engineering the basics. ➤ High-Performance Hardware: With high-torque integrated joints, Tien Kung 3.0 can clear 1-meter (3.3ft) obstacles and handle dexterous moves like kneeling and bending. It hits millimeter-level precision, making it a solid fit for industrial-grade tasks. ➤ True Autonomy: The bot runs a continuous perception-decision-execution loop. It uses world models to break down complex language commands and VLA models for real-time obstacle avoidance and navigation. ➤ Scalable Collaboration: The platform moves beyond single-unit tasks to support multi-robot collaboration with autonomous scheduling. It’s built to move embodied AI from the lab straight into real-world commercial and industrial environments. Source: X-Humanoid #Humanoid #OpenSource #Robotics #EmbodiedAI #PhysicalAI #Automation #XHumanoid #TienKung #WiseKaiwu

RoboHub🤖

49,440 Aufrufe • vor 7 Monaten

This guy deserves to get Famous .. Such an amazing achievement at such young age .. The AirPods Feature Apple Quietly Took Away From Android Users Most companies build ecosystems. Apple builds walls. If you use AirPods with an Android phone, you already know this. Noise cancellation disabled. Ear detection gone. Features you paid for, locked behind a device you don't own. Not because the hardware can't do it because Apple decided it shouldn't. For most people, that's just the cost of choosing Android. Something you accept and move on from. Kavish Dewar didn't move on. In this video, we cover the full story: 1. How Apple restricts AirPods features on Android and why it is a deliberate design choice, not a technical limitation 2. How a 16-year-old high school student from Gurugram reverse-engineered that restriction system with no team, no funding, and no formal training 3. What LibrePods actually does and why it works so effectively 4. How the project went from a bedroom in Gurugram to GitHub's trending page to global coverage by The Verge 5. Why Kavish made LibrePods completely free and open source and what that decision says about the difference between building for impact versus building for profit Why this story matters beyond the tech: LibrePods is not just a clever app. It is a direct challenge to the idea that trillion dollar companies get to define the limits of what you can do with hardware you already own. One teenager decided that was worth fighting not with lawyers, not with a startup, but with clean code and an open license. That is a different kind of power. And it is available to anyone willing to use it.

AstroCounselKK 🇮🇳

31,158 Aufrufe • vor 6 Monaten

Matthew Gallagher Built a $401M Company in Year One with 2 People. And the tool behind it? Claude Code. This year he's on track for $1.8B. Sam Altman predicted this. It's happening now. The problem? It costs money. API credits stack up. Monthly bills keep growing. Every prompt eats your budget. Every project drains your wallet faster. Until now. Two methods. 99% cheaper. One is completely free. Forever. $0. Not a trial. This video breaks down both step by step. ↓ Let me put this in perspective. $100-$500. That's monthly. That's what you spend. That's $6,000/year on API credits. Just to use a tool you haven't shipped anything with. The $401M guy? Spending $0. Same capability. Shipping weekly. Different cost structure. Different results. Different life. I'm about to hand you his cost structure for free. ↓ Open source vs closed source. Pay attention. Closed source: Claude. GPT-4. Pay per token. Meter always running. Open source: Qwen. Llama. Mistral. Free to download. Free to run. Free forever. No meter. No tokens. No bill. Here's what nobody tells you: 80% of coding tasks? Open source handles them. More than handles them. Writes clean code. Debugs errors. Generates boilerplate. Handles routine work perfectly. You're paying premium prices for tasks that don't need premium intelligence. That's hiring a brain surgeon to put on a bandaid. Smart play: Free models for the 80%. Paid credits for the 20%. That's what the $401M guy does. That's what this video teaches you. Follow Himanshu Kumar for more breakdowns that turn free tools into real businesses. ↓ Method 1: Ollama. Local. Free. Forever. Download it. Pull a model. Point Claude Code at it. Done. No internet needed. No API keys required. No monthly subscription. No token counting ever. No bill. Today. Tomorrow. Ever. Your data never leaves your computer. Complete privacy. Complete freedom. Claude Code thinks it's talking to the cloud. It's talking to your laptop. For $0. The video walks through every step: Every config file. Every variable. Every command. Every click. If you can follow a recipe, you can do this. People who set this up 3 months ago? Saved $300-$1,500 since then. Workflow didn't change one bit. ↓ Hardware you need: 16GB RAM: 7B models run smooth. 32GB RAM: 32B models run comfortable. 64GB + GPU: biggest models available. No GPU? Still works. Just slower. Few extra seconds. That's it. Your $1,500 laptop is sitting there running Chrome and Spotify. Put it to work saving you $200/month instead. Follow Himanshu Kumar for more breakdowns that turn free tools into real businesses. ↓ Method 2: Open Router. Free Cloud. No Hardware. Weak machine? Don't want local setup? This method is for you. Free AI models in the cloud. No download. No hardware. Configure Claude Code to route through Open Router. The config: Base URL: Open Router API. API key: free Open Router key. Default Sonnet: free. Default Opus: free. Default Haiku: free. Small fast model: free. Subagent model: free. Free. Free. Free. Free. Free across the board. Same interface. Same commands. Same workflow. Zero cost. Copy the config from the video. Paste it. Save $200/month. Starting today. Right now. ↓ When to use which: Ollama (local): Best for privacy. Best for offline work. Best for unlimited usage. Best if you have decent hardware. Open Router (cloud): Best for weak machines. Best for instant setup. Best for trying different models. Best if you don't want to manage anything. Both methods: Best for 80% of your daily work. Still use paid Claude for: Complex architecture. Multi-file refactoring. Deep reasoning tasks. The 20% that actually needs it. $20/month instead of $200/month. Same output. 90% less cost. ↓ The math that should make you angry. You (current): $200-$500/month. $2,400-$6,000/year. $7,200-$18,000 over 3 years. You (after this video): $20-$50/month. $240-$600/year. $720-$1,800 over 3 years. Savings over 3 years: $6,480-$16,200. That's a used car. That's seed money. That's 6 months of rent. All from one 25-minute video. All from 15 minutes of configuration. Highest ROI 25 minutes you'll spend this year. ↓ The limitations. I won't lie to you. Open source is not Opus. Not as smart on complex reasoning. Not as good at long-context tasks. Makes more mistakes on nuanced problems. But they are: Free. Capable. Getting better monthly. Good enough for 80% of daily work. Smart cost management isn't being cheap. It's being strategic. Expensive tool when it matters. Free tool when it doesn't. ↓ The one-person billion-dollar company is coming. $401M in year one proved it's possible. The building blocks: AI that codes: Claude Code. Way to run it free: this video. Distribution: the internet. Customers: everyone. Only missing ingredient? Someone who builds. Not reads about building. Not saves posts about building. Not bookmarks videos about building. Builds. Tools are free. Knowledge is free. Opportunity is screaming. You're still "thinking about it." ↓ Your action plan: Tonight: Watch the video. Tomorrow morning: Set up Ollama or Open Router. Tomorrow afternoon: Build something. Anything. This week: Build a second thing. Faster. This month: Charge someone for it. One video. One setup. One weekend. $0 cost. Unlimited potential. Or keep paying $200/month for something you could get free. Keep consuming instead of building. Keep planning instead of shipping. Matthew Gallagher didn't plan a $401M company. He built it. Full video attached. Every method. Every config. Every tradeoff. 25 minutes. Your move. Follow Himanshu Kumar for more breakdowns that turn free tools into real businesses.

Himanshu Kumar

13,677 Aufrufe • vor 5 Monaten