Loading video...

Video Failed to Load

Go Home

GEN-1 plays the 🐚 shell game, trained on just 1 hr of robot data. It also generalizes to unseen objects, like berkay 's car keys. Physical AI models should be capable of benchmark tasks like this one. It's interesting for the all the reasons Rhoda AI calls out --...

860,208 views • 3 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

For generative AI to become an interesting art tool, we need much more control over the output. The slot-machine-like nature of pure text-to-image leaves too much to chance. Using the "Real-time Latent Consistency Model" that I'm using in the example here, is the first time I truly got a glimpse of a future where we'll be able to use our artistic skills and sensibility, to get control over AI image gen. Systems like these will never be able to match the quality or originality of a skilled artist, it won't surprise us in the same way an artist can. Things are a mess in terms of the training data these models are based on, and the questions about copyright concerns and about a time when everything will look the same are very valid. At some point capabilities like these will be embedded in photoshop, and anyone will be able to generate a pretty picture. But to create interesting designs, to tell original stories and to surprise us, we need creatives and artists with something on their mind. We'll be able to create immersive worlds, by making brush-strokes and sculpt marks, without needing to worry about all the dials, plugins, wires of our 3d and 2d tools today. I love to sculpt, I love to draw, and I love to explore new mediums and new ways to create. The Gen AI tools we have today are far from perfect, and things need to be steered in a better direction. For that we need artists to help point the way. Gen AI isn't going away.. it's too powerful and has the potential to allow us to tell stories like never before. Like all other big technological shifts, tech like this will come at a cost, but it will also open up new opportunities and empower a new generation of storytellers. I might be naive, but I believe that human ingenuity and creativity will persevere in this new world ♥️ #art #ai

Martin Nebelong

1,660,244 views • 2 years ago

Coinbase CEO Explains “Reverse Prompting” and the Rise of the AI CEO Brian Armstrong: “One of the big pushes we made in the last year was we got our own internal hosted AI model that was connected to all of our data sources, right?” “So it's like every Slack message, every Google doc, Salesforce data, Confluence, you know.” “So now the data is all aggregated and I've started to ask it really… it's not just like prompting it, ‘Hey, can you write this kind of memo for me,’ or something.” “I'm asking these AI agents now, ‘As CEO, what should I be aware of in the company that I might not be aware of?’ And it'll tell me, ‘Did you know that there's actually disagreement on this team about the strategy?’ And I was like, actually, I didn't know that.” “This is like reverse prompting. So instead of telling the AI agent what you want it to do, you ask it what you should be thinking more about.” @jason: “It's a mentor. It's a coach.” Brian: “Yeah. Like, what could make me a better CEO? And it's like, ‘Well, I looked at how you spent your time in the last quarter and here's how you said that you wanted to spend it, but you actually spent 32% of your time on this instead of 20%.’” “I've asked it other questions like, ‘What's the thing that I changed my mind on the most over the last year?’ Things like that.” “It'll prompt you with information you should be thinking about instead of the other way around.” Thanks to our partner for making this happen!: Our episode is sponsored by the New York Stock Exchange - a modern marketplace and exchange for building the future. It all happens at the NYSE 🏛.

The All-In Podcast

80,524 views • 6 months ago

This is one-shot assembly: you show examples of what to build, and the robot just does it. (see original post: To share more on how this works, the robot is controlled in real time by a neural network that takes in video pixels and outputs 100Hz actions. The video below is part of the raw input passed directly into the model. I also like this view (at 1x speed) because it shows more of the (I think very cool) subtle moments of dexterity near the fingertips 👌 One-shot assembly seemed like a dream even just a year ago — it's not easy. It requires both the high-level reasoning of "what to build" (recognizing the geometry of the structures presented by the human), and the low-level visuomotor control of "how to build it" (purposefully re-orienting individual pieces and nudging them together in place). While possible to manually engineer a complex system for this (e.g. w/ hierarchical control, or explicit state representations), we were curious if our own Foundation model could do it all end-to-end with just some post-training data. Surprisingly, it just worked. Nothing about the recipe is substantially different than any other demo we’ve run in the past, and we’re excited about its implications on model capabilities: • On contextual reasoning, these models can (i) attend to task-related pixels in the peripheral view of the video inputs, and (ii) retain this knowledge in-context while ignoring irrelevant background. This is useful for generalizing to a wide range of real workflows: e.g. paying attention to what’s coming down the conveyor line, or glancing at the instructions displayed on a nearby monitor. • On dexterity, these models can produce contact-rich "commonsense" behaviors that can be difficult to pre-program or write language instructions for e.g. rolling a brick slightly to align its studs against the bottom of another, re-grasping to get a better grip or to move out of the way before a forceful press, or gently pushing the corners of a brick against the mat to rotate it in hand and stand it up vertically (i.e. extrinsic dexterity). These aspects work together to form a capability that resembles fast adaptation — a hallmark of intelligence, relevant for real use cases. This has also expanded my own perspective on what's possible with robot learning, using a recipe that's repeatable for many more skills. This milestone stands on top of the solid technical foundations we’ve built here at Generalist: hardcore controls & hardware, all in-house built models, and a data engine that "just works." We're a small group of hyper-focused engineers, and hands-down the highest talent-density team I’ve ever worked with. We're accelerating and scaling aggressively towards unlocking next-generation robot intelligence. Building Legos is just one example, and it's clear to me that we're headed towards a future where robots can do just about anything we want them to. Its coming, and we're going to make it happen.

Andy Zeng

49,443 views • 10 months ago