Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Training world models needs egocentric video and dense action signals, synchronized. That data is genuinely hard to find. We built it from Counter-Strike 2 demos. CS2-10k: 600K+ player-round videos, 10K+ hours, per-frame annotations — keyboard state, mouse delta, 3D position, camera yaw/pitch. All paired to the visual stream. Why...

95,190 Aufrufe • vor 2 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Hollywood isn't dead. It's evolving. And we're leading that evolution. We just shipped Koyal v2.5: The best Agentic AI filmmaking platform. It goes from your script or music to full video with consistent characters, settings & storylines. Large Production houses, music labels & ad agencies use koyal.ai (YC F25) for storyboarding & pre-viz. Smaller studios use it for making content they never had the time, budget or resources to create. Write a scene. Direct the camera. Build a world. koyal.ai (YC F25) brings it to life. Here's what's new: - Go from script to video: Write a scene with dialogue, pacing, emotion. Koyal fully realizes it. Using the best voice models on the planet. - Direct the camera: Push in. Pull back. Drift through a scene. Real 3D camera blocking. Shape the shot the way you see it in your head. - Build a world that stays:Lock in a location and return to it. Your scenes stay consistent from the first frame to the last. Plus: dialogue, SFX, annotations for editing, and more control over every shot. Under the hood, we benchmark 40+ models weekly with real humans so Koyal always uses the best available (more on that next week!) Our agents pick the best models for each scene in run-time so you focus on the story, we handle the complexity. You don't need to prompt a film, you can direct it. With Koyal, we're looking to replace the camera, not the filmmaker behind it. Try it now at beta [dot] koyal [dot] ai DM me for credits!

Mehul Agarwal

74,482 Aufrufe • vor 5 Monaten

Just how capable are open source models? Below is the first in a new series where we go behind the scenes and pull back the curtain on interesting AI research / demos, making them fun and easy to understand. Here, we have a short visual demonstration from aizk ✡️ showcasing how Kimi K3 (a language model that operates primarily through text) is capable of building complicated 3D structures / moments in history in Minecraft, something that previously was not possible with other open source models, and why this matters. The crazy part? The model doesn't "see" the game like we do. The LLMs must reason in pure text, writing JavaScript, that later compiles down into commands placing each block, one at a time. Spatial reasoning is a very hard problem in AI, it's the same core challenge behind robotics and self-driving cars, where a model has to understand and act in physical 3D space. Watching a text model pull it off is nothing short of a miracle. The point isn't just Minecraft itself, rather, it's AI being able to generalize, not memorize, on things that are weird and beyond their training data. This is key to building true artificial general intelligence. These video game benchmarks (there are many different games actively being researched right now) provide a clear-cut end goal, challenges that are almost certainly not in the training set, and a fun, very fast, visual way to almost feel the increasing capabilities of various open source AI models over time. If you haven't given open source models a serious try yet, watch the video, it may shock you!

Featherless AI

39,769 Aufrufe • vor 10 Tagen

Elon Musk reduced the oldest question in human history to basic math. No one has found a flaw in it. Musk: “What are the odds that we are in base reality? And that this has not happened before.” You don’t need a physics degree to follow it. You need a timeline. Musk: “If you look at the advancement of video games, it’s gone from Pong, two rectangles and a square batting it back and forth, to photorealistic, real-time games with millions of people playing simultaneously.” Fifty years. That is all it took to close the gap between two rectangles on a screen and a world you cannot tell apart from the one outside your window. Musk: “If that trend continues, video games will be indistinguishable from reality.” The visuals are not what seals it. The intelligence is. Musk: “Think of how sophisticated the conversations are you can have with an AI today, and that’s only going to get more sophisticated.” We are not scripting characters anymore. We are building minds that reason, adapt, and surprise the people who made them. We are nowhere near finished. Musk: “The future, if civilization continues, will be millions, maybe billions of photorealistic, indistinguishable from reality, video games. And with characters in those video games that are very deep, and where the dialogue is not pre-programmed.” One base reality. Billions of perfect copies. Each one running minds that feel exactly as conscious as you do right now. Each one certain it is the original. Musk: “So then what are the odds that we are in base reality?” If even one civilization crosses that threshold, simulated minds outnumber real ones by billions. The probability you are sitting in the real one is not low. It is nearly zero. Not as philosophy. As mathematics. We are not watching this happen. We are building it. Right now. Every AI that reasons without a script. Every world rendered one frame closer to indistinguishable. We are constructing the exact technology that makes our own existence statistically implausible. And we will never stop. Because the curiosity that questions reality is the same force that builds it. If the math holds, something built us. Something conscious enough to create consciousness. They stood where we are standing. Same question. Same inability to stop. And whatever built them never answered it either. There is no top floor. There is no original. None of that changes what you feel right now. Consciousness was never about what you are made of. It was about what you experience. Musk did not float a theory. He held up a mirror with no back wall. And the math does not need you to believe it. It only needs time.

Dustin

193,823 Aufrufe • vor 1 Monat

Domain Randomization (DR) is a key component of the data augmentation pipeline at Axis Robotics. By applying DR, we are able to scale verified, high-quality human trajectories by 10x to 100x. During training, we systematically introduce variances in environmental parameters. This prevents the model from relying on spurious visual correlations. The objective is to ensure the policy learns rather than overfitting. To demonstrate the necessity and effectiveness of this approach, we evaluated both DR and No-DR models on Task 74 (pour_water_into_mug). The empirical results show a definitive impact on real-world deployment reliability: integrating DR into the pipeline increased the success rate from 0% to 90% (Fig. 1). This divergence stems from how the respective policies process visual observations (Fig. 2). The baseline (No DR) model overfits to the static visual background. It essentially memorizes the poses from the training dataset but fails to generalize when subjected to the inevitable variances of real-world deployment. Consequently, it cannot execute the correct manipulation on the target object. Conversely, the DR-trained model learns to extract essential geometric features and physical constraints, filtering out superficial visual noise. This leads to significantly higher robustness in dynamic environments. The structural difference in execution is clearly reflected in the end-effector trajectory data: These real-world deployment recordings further illustrate this difference (Videos 1 and 2). Scaling Physical AI requires turning raw trajectory data into robust policies, and a rigorously engineered DR infrastructure is an essential bridge to close the Sim2Real gap.

Axis Robotics

27,125 Aufrufe • vor 4 Monaten

Yann LeCun just told the most well-funded industry in human history it is solving the wrong problem. LeCun: “Babies learn this around the age of eight or nine months, that objects don’t float, they fall.” No dataset. No labels. No reward signal. A nine month old drops a spoon and builds a physics engine no machine can match. LeCun: “Most of us can learn to drive in about 20 or 30 hours of training without ever crashing, causing any accident.” Twenty hours. Tesla has built the most capable driving system on the road. It took billions of miles of data to get there. A sixteen year old gets there over a long weekend. Not because the teenager is the better driver. Because the teenager is not learning to drive. They are deploying a model of reality they have been building since birth. LeCun: “If we drive next to a cliff, we know that if we turn the wheel to the right, the car is going to run off the cliff and nothing good is going to come out of this.” You simulate the crash. You see the wreckage. You feel the fall. You turn the wheel. None of it was real. All of it was intelligence. Every AI has to crash a thousand times to learn what you imagined once and never did. That is not a performance gap. That is an architecture gap. LeCun: “The main problem we need to solve is how do we learn models of the world.” Not bigger models. Not more compute. Not another trillion tokens. World models. A machine that can run reality forward before it acts. The industry is scaling language. LeCun says language is a compression of thought. Not thought itself. You understood gravity before you could say the word. You grasped cause and effect before your first sentence. The deepest intelligence you will ever possess was built in total silence. And every lab on Earth is trying to reconstruct the mind from words alone. Physics does not care about your context window. A baby who learns that cups fall in a kitchen already knows that rocks fall off cliffs. No retraining. No fine-tuning. One model. Every environment. That is what intelligence actually is. Not prediction. Not pattern matching. Not scale. A simulation of reality so precise you rehearse the future before it exists. Every infant on Earth builds one. No machine ever has.

Dustin

122,066 Aufrufe • vor 1 Monat

Rich Roll on why waiting to "feel like it" is a trap: "You can't think your way into the mood that you seek or the state of mind that you aspire to inhabit. Action is the only thing that can trigger that change." Rich uses running as the perfect illustration of this principle. Imagine you wake up in the morning and you're supposed to do a run because you're training for a race. You don't feel like it. So what do most of us do? "We all resort to that state where we think, 'Well, I don't want to do it right now. I'll just wait until I feel like doing it and then I'll do it then.'" But here's the problem with that logic: "If you're waiting until you feel like doing something, chances are you're probably never going to get to it." The mood you're hoping will arrive on its own? It's not coming. Not without action first. "To take the action despite how you feel about it is the thing that catalyzes the state change." You don't run because you feel motivated. You feel motivated because you ran. He points to what every runner knows from experience: "When they finish the run, they're always glad that they did it. They don't generally regret it. And then they feel better." Notice the sequence. The good feeling comes after the action, not before it. The state change is the reward for showing up, not the prerequisite. And this isn't just about running. As Rich puts it: "That example is applicable to all areas of life." The workout you're avoiding. The conversation you're delaying. The project you're putting off until you're "in the right headspace." You're waiting for a feeling that only exists on the other side of doing the thing.

Kevin Tanaka

10,256 Aufrufe • vor 4 Monaten

Eric Schmidt was asked a technical question about open source and answered with the map of the next fifty years. The winner won’t be the smartest model. It’ll be the one four billion people never had to choose. Schmidt: “China is competing with open weights and open training data, and the US is largely and majority focused on closed weights, closed data.” That isn’t a product decision. It’s a distribution decision. And distribution has beaten quality in every contest that ever mattered. Schmidt: “The majority of the world, think of it as the Belt and Road initiative, are going to use Chinese models and not American models.” The first Belt and Road was ports, rail, and highways. This one doesn’t get poured. It gets downloaded. Every piece of infrastructure ever built was indifferent to what moved across it. A road doesn’t tell you where to go. A model does. Schmidt: “The American models are typically using 16-bit precision for their training. The Chinese are pushing 8 and now even 4.” Every bit they drop is a cheaper device that can run it. We cut off their chips to slow them down. Scarcity made their models small. Small is what crosses a border. We designed their advantage. Not better. Present. America is building the best model on earth and metering it. China is building one that’s good enough and giving it away. A model isn’t software. It’s a compressed set of judgments about what’s true, what’s askable, and what a reasonable answer sounds like. Install that as a country’s default and you haven’t sold them a tool. You’ve set the limits of what occurs to them. That isn’t censorship. Censorship leaves a mark. A question that never occurs to you doesn’t feel like a restriction. It feels like the edge of the world. Every empire before this one had to teach the world its language first. Missionaries, schoolteachers, garrisons, printing presses. Every one of them ran through a human being who could hesitate, doubt, or be talked out of it. AI arrives already speaking yours. It doesn’t ask you to change. It changes you in your own voice. The first ideology in history that doesn’t need believers. It only needs to be installed. Schmidt: “I’d much rather have the proliferation of large language models and that learning be done based on Western values.” He’s right, and we’re playing it backwards. We treat openness like a giveaway, as if the weights were the crown jewels. Openness is the one advantage an authoritarian can’t copy. An open model can be read, probed, and torn apart by anyone who doubts it. A system that has to control the answer can never afford to publish the reasoning. China opens its weights to spread them. America could open its weights to be trusted. Only one of those compounds. A closed American model wins the benchmark. An open American model wins the default. Centuries get built out of defaults. Schmidt: “We also have to watch to make sure that the proliferation of these models for handheld devices is under American control.” That’s the ground. Not data centers. Not cloud contracts. Pockets. The frontier race has five contenders and the whole world watching. This one has no audience at all. It plays out on hardware too cheap to run an American model, and goes to whoever bothered to show up. We keep asking who reaches AGI first. The question that settles the century is smaller and much harder to take back. Four billion people are going to ask a machine what happened in their own country. Whose answer do they get? Nobody votes on that. It’s decided by whatever was already installed. America has the best AI ever built. The only way to lose this era is to keep it.

Dustin

12,094 Aufrufe • vor 1 Monat