Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

I’m back, with a little something… GPC 1, a new general-purpose classifier enabling: bounding boxes, poses, coordinates, direct mechanical control and so much more, all in milliseconds. Post-trained on millions of datapoints and thousands of unique problem classes learned directly from you guys posting on X, GPC-1 is a...

43,276 görüntüleme • 15 gün önce •via X (Twitter)

39 Yorum

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

i wish i can give everyone a hug but here’s the next best thing 🤗

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

GPC-1 uses a new dual training technique, jointly training the Qwen base model and add on adapters. As it is, its great, but if you want to customize more, the model was literally trained to be really easy to fine tune for anything 🚀

Yohei profil fotoğrafı
Yohei15 gün önce

this is awesome! thx for sharing. i did something similar but went the opposite direction of using a small vlm to run locally, though no boxes or anything.

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

Love it!

rose🌲🌲 profil fotoğrafı
rose🌲🌲15 gün önce

mmmm yes, this won't be used for robotics and anything that contributes to society but used to trade sports prediction books with cameras within stadiums

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

humanity at its best 😂

Max profil fotoğrafı
Max15 gün önce

you cooked with this one. since the Qwen backbone supports video, curious if you tried GPC-1 on videos?

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

not yet but good idea!

Stephane profil fotoğrafı
Stephane15 gün önce

@yoheinakajima Very nice. Congrats. Now I’m curious as to if I can integrate this and how well it will work for the avatar portion of my twin.

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

@yoheinakajima Tons of medical data is being training on now

Stephane profil fotoğrafı
Stephane15 gün önce

@yoheinakajima What’s the medical data that you’re training on and what’s the goal with that training?

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

@yoheinakajima goal rn is to optimize hospital scheduling for surgery based on patients (but training data is broad, including data on how to control lab equipment)

Brjan | AI Builder profil fotoğrafı
Brjan | AI Builder15 gün önce

if it really operates in milliseconds, that could seriously change automation workflows

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

It rly does and it rly will, already deploying it in a few labs rn

Jasper Croome profil fotoğrafı
Jasper Croome15 gün önce

Super cool, classifiers are having the moment they deserve! Did you try this out on maps or geospatial stuff at all?

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

Good idea! Will include it in v1.1

Jasper Croome profil fotoğrafı
Jasper Croome15 gün önce

Happy to give specific use cases / feedback if & when you get into it

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

DM me anytime, I’m always into it

ζ Pedram ζ profil fotoğrafı
ζ Pedram ζ15 gün önce

Cool, wasm port?

John profil fotoğrafı
John15 gün önce

Awesome video Harsha

Wayne Workman profil fotoğrafı
Wayne Workman15 gün önce

This is so cool. There's not enough time in the day to try out all the new models people are creating. How much vram does this use?

Marc profil fotoğrafı
Marc15 gün önce

Congrats Harsha! We’re currently in stealth mode building a humanoid robot 🤖 and would love to learn more about GPC-1. Do you see it being useful for robot control, like watching a human do a task and have a robot replicate the pose and how would you advise using it for this?

varick lim profil fotoğrafı
varick lim15 gün önce

woah this was exactly what i hoped jev could do

Kiran Jd profil fotoğrafı
Kiran Jd15 gün önce

@yoheinakajima Will this work for multiple objects in video?

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

@yoheinakajima Yup, just describe the object in the key (i.e person_to_the_left_x_axis, not the most elegant solution rn but working on it)

Kiran Jd profil fotoğrafı
Kiran Jd15 gün önce

@yoheinakajima I could see so may things being useful in sport tracking. I’m working on analysing a football training drill from coach’s perspective. I would love to try it

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

@yoheinakajima lmk how it performs for you!

Harsh profil fotoğrafı
Harsh15 gün önce

This is so cool! Whats the latency like?

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

Depends on what hardware you use, under 100ms on a h100

girardi's mind profil fotoğrafı
girardi's mind15 gün önce

I'm not sure for robotics because those arms pass through each other, but looks okay for less critical setups

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

ironing those issues out now

Alakaham profil fotoğrafı
Alakaham15 gün önce

Damn...just an unrelated question...how do you make these prod demo videos?

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

one shotted with Gemini and Astra

Mathew Chan profil fotoğrafı
Mathew Chan15 gün önce

How does this do on OCR tasks on docs?

Harsha Gundala profil fotoğrafı
Harsha Gundala15 gün önce

Trained on many OCR examples, lmk how it performs at your use case

robot 2.0 profil fotoğrafı
robot 2.015 gün önce

@CDerinbogaz

Aakash Gnanakumar profil fotoğrafı
Aakash Gnanakumar15 gün önce

Bro is too smart!!

Robin | Poker x AI profil fotoğrafı
Robin | Poker x AI15 gün önce

Impressive that you got continuous numeric output in ms, could be a game changer for real time robotics pipelines

pulsatingGenius profil fotoğrafı
pulsatingGenius15 gün önce

Yolo with natural language

Benzer Videolar

Geoffrey Hinton says a big language model runs on about 1% of your brain's connections and still ends up knowing more than you: "So in your brain, you have a hundred trillion connections, roughly speaking. Okay. That's a lot. And you only live for about two billion seconds. That's not much." "If you compare how many seconds you live for, with how many connections you've got, you have a whole lot more connections than experiences." "Now with these neural nets, it's sort of the other way round. They only have of the order of a trillion connections. So like 1% of your connections, even in a big language model, many of them fewer, but they get thousands of times more experience than you." "So the big language models are solving the problem with not many connections, only a trillion. How do I make use of a huge amount of experience?" "And back propagation is really, really good at packing huge amounts of knowledge into not many connections." "But that's not the problem we're solving. We've got huge numbers of connections, not much experience. We need to sort of extract the most we can from each experience." Two to three billion seconds is the whole budget. Everything you know, you learned inside it. So evolution built you to squeeze a lot out of very little. Hinton's point is that a language model has the opposite problem and the opposite fix, and backprop turned out to be extremely good at that fix. Worth noticing what this predicts about failure. A system running on 1% of your wiring and thousands of times your experience is not going to fail the way you do. You fail from having seen too few examples. It fails from compressing too many into too little, and the compression is where the errors get made. That is a strange thing to be deploying into hospitals and courts with no way to inspect it. We test these systems by asking them questions, which tells you what came out. Nobody can yet look at a trillion connections and say what got packed in. - Geoffrey Hinton, Nobel laureate and Turing Award winner, on StarTalk (StarTalk) with Neil deGrasse Tyson.

Karl Mehta

662,652 görüntüleme • 2 ay önce

David Friedberg Explains Why AI Will Break Every Economic Model: Will Output Surpass Consumptive Capacity? I think fundamentally, if you're driving productivity with AI, you're driving leverage on human time and leverage on capital. The question is, how quickly can you drive that up? And that's a function of how much consumption there is, how much capacity there is for consumption. So if your earnings are the same, but things are getting more expensive, you're not happy. If your earnings go up by 10% and things stay the same price, you got 10% more than you had last year, you're gonna be happy. I just think all humans are driven by this need to consume more each year than they did last year. So I think for me, that's the lower limit on consumptive capacity in the world. The question that we're now facing, which we've never faced in human history before, is there an upper limit on consumptive capacity? Because AI creates such a profound shift in productivity and in leverage that normally you would say, “Hey, when we get a new tool or we get new leverage in a system, we build a new technology, we can make more with less.” Therefore, everyone gets access to more things for the same price or the cost of things that they consume come down by a certain price. But there may be a situation now where the ability to make stuff exceeds the capacity to consume stuff, and that is something that I don't think we’ve faced before. And I think that's where a lot of the models start to break. Just general economic models, just general productivity models and general social models. And this goes to the point about, like, what is everyone going to do? In the same way that I think we've argued that maybe SaaS was a transitory business phenomenon that existed between the foundation of the internet and the era of AI, it may be the case that knowledge work in general is also a transitory phenomenon that only existed between the foundation of the computer, or computing tools, and the existence of AI, generally speaking. And if all of that goes away very quickly, and all of those people can be redistributed and recast into doing other higher level, more creative things, their productivity goes up by 100x, is there really a consumer on the other end of all of that productivity? Is there really enough consumptive capacity? And I think that's the profound question that we all face.

The All-In Podcast

81,924 görüntüleme • 7 ay önce

I asked Garry Tan how to use meta prompting to get better at AI: "My partners at YC Jared Friedman and Pete Koomen showed me how to do this. You can take almost anything that you do all the time and just drop it into a context window. And then say, “Here’s a bunch of inputs and outputs." And maybe you also add a bunch of notes. And then you tell it, “Write me a prompt that can act as an agent that takes this input and makes this output over here.” You can do this for almost any type of knowledge work. And you can even introspect. "What are things you notice that I did to convert this from the input to the output?”. And then you can just start using the prompt. Initially, it’s going to suck. Because it’s just not that smart yet. But what’s funny is now, I also use it to Iterate my writing. You can be very direct, "I would never say that", "Don’t say it like this", or "Oh, you used the long word there, use the short word". Just speak to it conversationally. And then when you're happy with the output, you can use that new output to make a new prompt. "Based on this conversation, give me a better initial prompt that incorporates all the things we talked about." And you can do this with literally everything. And in theory, there’s so much it applies to that people do day-to-day. You could use it for tweets. You could use it for editing podcasts. You can use it for pretty much everything. I have a folder of prompts that I use all the time. My YouTube prompt is on v27 or something. I'll go through this process with all the different max models. I'll use GPT 5.2 Pro. I’ll use Grok. I'll use Claude. Then, I’ll take all the outputs from all the models and put them into Claude and say "Here’s my prompt, here’s the output from four LLMs, including yourself. Rate each response and tell me what the pros and cons of each approach are." And I usually say "give it to me in numbered form". And then you can agree with one, disagree with two, tell it three is this or that. And then after that, you say given all of this, synthesize it."

The Peel

51,632 görüntüleme • 7 ay önce