Loading video...

Video Failed to Load

Go Home

I’m back, with a little something… GPC 1, a new general-purpose classifier enabling: bounding boxes, poses, coordinates, direct mechanical control and so much more, all in milliseconds. Post-trained on millions of datapoints and thousands of unique problem classes learned directly from you guys posting on X, GPC-1 is a...

43,276 views • 15 days ago •via X (Twitter)

39 Comments

Harsha Gundala's profile picture
Harsha Gundala15 days ago

i wish i can give everyone a hug but here’s the next best thing 🤗

Harsha Gundala's profile picture
Harsha Gundala15 days ago

GPC-1 uses a new dual training technique, jointly training the Qwen base model and add on adapters. As it is, its great, but if you want to customize more, the model was literally trained to be really easy to fine tune for anything 🚀

Yohei's profile picture
Yohei15 days ago

this is awesome! thx for sharing. i did something similar but went the opposite direction of using a small vlm to run locally, though no boxes or anything.

Harsha Gundala's profile picture
Harsha Gundala15 days ago

Love it!

rose🌲🌲's profile picture
rose🌲🌲15 days ago

mmmm yes, this won't be used for robotics and anything that contributes to society but used to trade sports prediction books with cameras within stadiums

Harsha Gundala's profile picture
Harsha Gundala15 days ago

humanity at its best 😂

Max's profile picture
Max15 days ago

you cooked with this one. since the Qwen backbone supports video, curious if you tried GPC-1 on videos?

Harsha Gundala's profile picture
Harsha Gundala15 days ago

not yet but good idea!

Stephane's profile picture
Stephane15 days ago

@yoheinakajima Very nice. Congrats. Now I’m curious as to if I can integrate this and how well it will work for the avatar portion of my twin.

Harsha Gundala's profile picture
Harsha Gundala15 days ago

@yoheinakajima Tons of medical data is being training on now

Stephane's profile picture
Stephane15 days ago

@yoheinakajima What’s the medical data that you’re training on and what’s the goal with that training?

Harsha Gundala's profile picture
Harsha Gundala15 days ago

@yoheinakajima goal rn is to optimize hospital scheduling for surgery based on patients (but training data is broad, including data on how to control lab equipment)

Brjan | AI Builder's profile picture
Brjan | AI Builder15 days ago

if it really operates in milliseconds, that could seriously change automation workflows

Harsha Gundala's profile picture
Harsha Gundala15 days ago

It rly does and it rly will, already deploying it in a few labs rn

Jasper Croome's profile picture
Jasper Croome15 days ago

Super cool, classifiers are having the moment they deserve! Did you try this out on maps or geospatial stuff at all?

Harsha Gundala's profile picture
Harsha Gundala15 days ago

Good idea! Will include it in v1.1

Jasper Croome's profile picture
Jasper Croome15 days ago

Happy to give specific use cases / feedback if & when you get into it

Harsha Gundala's profile picture
Harsha Gundala15 days ago

DM me anytime, I’m always into it

ζ Pedram ζ's profile picture
ζ Pedram ζ15 days ago

Cool, wasm port?

John's profile picture
John15 days ago

Awesome video Harsha

Wayne Workman's profile picture
Wayne Workman15 days ago

This is so cool. There's not enough time in the day to try out all the new models people are creating. How much vram does this use?

Marc's profile picture
Marc15 days ago

Congrats Harsha! We’re currently in stealth mode building a humanoid robot 🤖 and would love to learn more about GPC-1. Do you see it being useful for robot control, like watching a human do a task and have a robot replicate the pose and how would you advise using it for this?

varick lim's profile picture
varick lim15 days ago

woah this was exactly what i hoped jev could do

Kiran Jd's profile picture
Kiran Jd15 days ago

@yoheinakajima Will this work for multiple objects in video?

Harsha Gundala's profile picture
Harsha Gundala15 days ago

@yoheinakajima Yup, just describe the object in the key (i.e person_to_the_left_x_axis, not the most elegant solution rn but working on it)

Kiran Jd's profile picture
Kiran Jd15 days ago

@yoheinakajima I could see so may things being useful in sport tracking. I’m working on analysing a football training drill from coach’s perspective. I would love to try it

Harsha Gundala's profile picture
Harsha Gundala15 days ago

@yoheinakajima lmk how it performs for you!

Harsh's profile picture
Harsh15 days ago

This is so cool! Whats the latency like?

Harsha Gundala's profile picture
Harsha Gundala15 days ago

Depends on what hardware you use, under 100ms on a h100

girardi's mind's profile picture
girardi's mind15 days ago

I'm not sure for robotics because those arms pass through each other, but looks okay for less critical setups

Harsha Gundala's profile picture
Harsha Gundala15 days ago

ironing those issues out now

Alakaham's profile picture
Alakaham15 days ago

Damn...just an unrelated question...how do you make these prod demo videos?

Harsha Gundala's profile picture
Harsha Gundala15 days ago

one shotted with Gemini and Astra

Mathew Chan's profile picture
Mathew Chan15 days ago

How does this do on OCR tasks on docs?

Harsha Gundala's profile picture
Harsha Gundala15 days ago

Trained on many OCR examples, lmk how it performs at your use case

robot 2.0's profile picture
robot 2.015 days ago

@CDerinbogaz

Aakash Gnanakumar's profile picture
Aakash Gnanakumar15 days ago

Bro is too smart!!

Robin | Poker x AI's profile picture
Robin | Poker x AI15 days ago

Impressive that you got continuous numeric output in ms, could be a game changer for real time robotics pipelines

pulsatingGenius's profile picture
pulsatingGenius15 days ago

Yolo with natural language

Related Videos

Geoffrey Hinton says a big language model runs on about 1% of your brain's connections and still ends up knowing more than you: "So in your brain, you have a hundred trillion connections, roughly speaking. Okay. That's a lot. And you only live for about two billion seconds. That's not much." "If you compare how many seconds you live for, with how many connections you've got, you have a whole lot more connections than experiences." "Now with these neural nets, it's sort of the other way round. They only have of the order of a trillion connections. So like 1% of your connections, even in a big language model, many of them fewer, but they get thousands of times more experience than you." "So the big language models are solving the problem with not many connections, only a trillion. How do I make use of a huge amount of experience?" "And back propagation is really, really good at packing huge amounts of knowledge into not many connections." "But that's not the problem we're solving. We've got huge numbers of connections, not much experience. We need to sort of extract the most we can from each experience." Two to three billion seconds is the whole budget. Everything you know, you learned inside it. So evolution built you to squeeze a lot out of very little. Hinton's point is that a language model has the opposite problem and the opposite fix, and backprop turned out to be extremely good at that fix. Worth noticing what this predicts about failure. A system running on 1% of your wiring and thousands of times your experience is not going to fail the way you do. You fail from having seen too few examples. It fails from compressing too many into too little, and the compression is where the errors get made. That is a strange thing to be deploying into hospitals and courts with no way to inspect it. We test these systems by asking them questions, which tells you what came out. Nobody can yet look at a trillion connections and say what got packed in. - Geoffrey Hinton, Nobel laureate and Turing Award winner, on StarTalk (StarTalk) with Neil deGrasse Tyson.

Karl Mehta

662,652 views • 2 months ago

David Friedberg Explains Why AI Will Break Every Economic Model: Will Output Surpass Consumptive Capacity? I think fundamentally, if you're driving productivity with AI, you're driving leverage on human time and leverage on capital. The question is, how quickly can you drive that up? And that's a function of how much consumption there is, how much capacity there is for consumption. So if your earnings are the same, but things are getting more expensive, you're not happy. If your earnings go up by 10% and things stay the same price, you got 10% more than you had last year, you're gonna be happy. I just think all humans are driven by this need to consume more each year than they did last year. So I think for me, that's the lower limit on consumptive capacity in the world. The question that we're now facing, which we've never faced in human history before, is there an upper limit on consumptive capacity? Because AI creates such a profound shift in productivity and in leverage that normally you would say, “Hey, when we get a new tool or we get new leverage in a system, we build a new technology, we can make more with less.” Therefore, everyone gets access to more things for the same price or the cost of things that they consume come down by a certain price. But there may be a situation now where the ability to make stuff exceeds the capacity to consume stuff, and that is something that I don't think we’ve faced before. And I think that's where a lot of the models start to break. Just general economic models, just general productivity models and general social models. And this goes to the point about, like, what is everyone going to do? In the same way that I think we've argued that maybe SaaS was a transitory business phenomenon that existed between the foundation of the internet and the era of AI, it may be the case that knowledge work in general is also a transitory phenomenon that only existed between the foundation of the computer, or computing tools, and the existence of AI, generally speaking. And if all of that goes away very quickly, and all of those people can be redistributed and recast into doing other higher level, more creative things, their productivity goes up by 100x, is there really a consumer on the other end of all of that productivity? Is there really enough consumptive capacity? And I think that's the profound question that we all face.

The All-In Podcast

81,924 views • 7 months ago

I asked Garry Tan how to use meta prompting to get better at AI: "My partners at YC Jared Friedman and Pete Koomen showed me how to do this. You can take almost anything that you do all the time and just drop it into a context window. And then say, “Here’s a bunch of inputs and outputs." And maybe you also add a bunch of notes. And then you tell it, “Write me a prompt that can act as an agent that takes this input and makes this output over here.” You can do this for almost any type of knowledge work. And you can even introspect. "What are things you notice that I did to convert this from the input to the output?”. And then you can just start using the prompt. Initially, it’s going to suck. Because it’s just not that smart yet. But what’s funny is now, I also use it to Iterate my writing. You can be very direct, "I would never say that", "Don’t say it like this", or "Oh, you used the long word there, use the short word". Just speak to it conversationally. And then when you're happy with the output, you can use that new output to make a new prompt. "Based on this conversation, give me a better initial prompt that incorporates all the things we talked about." And you can do this with literally everything. And in theory, there’s so much it applies to that people do day-to-day. You could use it for tweets. You could use it for editing podcasts. You can use it for pretty much everything. I have a folder of prompts that I use all the time. My YouTube prompt is on v27 or something. I'll go through this process with all the different max models. I'll use GPT 5.2 Pro. I’ll use Grok. I'll use Claude. Then, I’ll take all the outputs from all the models and put them into Claude and say "Here’s my prompt, here’s the output from four LLMs, including yourself. Rate each response and tell me what the pros and cons of each approach are." And I usually say "give it to me in numbered form". And then you can agree with one, disagree with two, tell it three is this or that. And then after that, you say given all of this, synthesize it."

The Peel

51,632 views • 7 months ago