Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Humans can see in high-res, high-FPS in real-time. Why can't VLMs? Introducing AutoGaze: ViTs/VLMs "gaze" only at key video regions! Up to 4-100x token savings, 19x speedup, and enables scaling to 4K-res 1K-frame videos. 📄 🌐 🤗 (1/n)🧵

162,275 Aufrufe • vor 5 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

In my previous post about the #Nioh3 demo, I mentioned that when the game isn’t consistently holding 60 FPS or 120 FPS, several issues can show up. Here’s a clearer breakdown of what's going on. The main reason is that the game speed and physics are tied to the frame rate. When you enable the 120 FPS cap, the engine start overreacting and pushes CPU usage very high, even if the game is still running at around 60 FPS (as shown in the first example). In the second clip, you can see the uneven, jittery camera movement I mentioned before. Notice how much smoother 60 FPS and 120 FPS look compared to the unlocked side. Also, camera movement at 60 FPS is slightly faster than at 120 FPS. In the third clip, player movement is slower when the frame rate is unlocked. This doesn’t happen all the time, but it shows up often enough to be noticeable. Because of how the Katana Engine behaves, the game is clearly designed around 60 FPS. Running at 120 FPS is possible, but it’s only recommended if your system can maintain that target almost all the time, which isn’t easy to achieve. There’s also an alternative workaround where you select the 60 (locked) option and enable Frame Generation (DLSS or FSR 3), as shown in the last clip. The downside is that DLSS Frame Generation tends to show the same stuttery look as when the frame rate isn’t holding a fixed target, likely due to Reflex keeping the frame rate slightly below target. FSR Frame Generation, on the other hand, looks much smoother and works better here.

BenchmarKing

17,820 Aufrufe • vor 7 Monaten

Two of the people most responsible for scaling the transformer are now betting on a next act. Jerry Tworek ran the Reasoning 🍓 team at OpenAI. rohan anil was a pre-training lead on Gemini after years at Google Brain and Anthropic. They just started to find what comes next. Their core argument: (1) models are trained in the lab but deployed in the real world and can't keep learning once they leave; (2) AI research is done by humans today but models will be able to explore and uncover new advances more rapidly and systematically (controversial but timely w this week's petition). The conversation covers: — why Jerry expected AGI in 2025 and what changed his mind — the two kinds of learning from experience, and why RL only captures one — the computational depth problem baked into today's architectures — why the biggest labs can't afford to look for a transformer replacement — the kernel competition where humans + $100K of coding agents found a 60x speedup no frontier model comes close to — a definition of AGI you can actually test: a model that improves itself with no human in the loop 00:00 Introduction 01:46 Appreciating Transformers 02:44 Scaling Hits Limits 04:54 Why Architecture Matters 05:32 RL Reality Check 07:32 Test Time Learning 09:52 Economics Of Scaling 12:47 Why Start A Company 14:24 Rohan On Transformers 19:11 Computational Depth Problem 20:32 When Transformers Top Out 23:22 Beyond Reinforcement Learning 26:41 Optimization And Efficiency 34:24 Building An Automated Lab 39:45 Kernel Automation Roadmap

Sonya Huang 🐥

232,559 Aufrufe • vor 1 Monat

We believe we’re the first robotics company to demonstrate a robot peeling an apple with dual dexterous human-like hands. This breakthrough closes a key gap in robotics, achieving bimanual, contact-rich manipulation and moving far beyond the limits of simple grippers. 🧵↓ Today’s AI models (VLMs) are excellent at perception but struggle with action. Controlling high-degree-of-freedom hands for tasks like this is incredibly complex, and precise finger-level teleoperation is nearly impossible for humans. Our first step was a shared-autonomy system: rather than controlling every finger, the operator triggers pre-learned skills like a “rotate apple or tennis ball” primitive via a keyboard press or pedal. This makes scalable data collection and RL training possible. How does the AI manage this? We created "MoDE-VLA" (Mixture of Dexterous Experts). It fuses vision, language, force, and touch data by using a team of specialist "experts," making control in high-dimensional spaces stable and effective. The combination of these two innovations allows for seamless, contact-rich manipulation. The human provides high-level guidance, and the robot executes the complex in-hand coordination required. This work paves the way for robots that can safely handle delicate tasks in human environments. Want the full technical details? 📄 Read the full research paper: Visit us at NVIDIA GTC Booth #1838, Hall 3 to learn more! #Robotics #AI #DexterousManipulation #VLA #NVIDIAGTC Nancy Villicaña NVIDIA GTC

Sharpa

20,429 Aufrufe • vor 5 Monaten