Загрузка видео...

Не удалось загрузить видео

На главную

I’m back, with a little something… GPC 1, a new general-purpose classifier enabling: bounding boxes, poses, coordinates, direct mechanical control and so much more, all in milliseconds. Post-trained on millions of datapoints and thousands of unique problem classes learned directly from you guys posting on X, GPC-1 is a...

43,276 просмотров • 15 дней назад •via X (Twitter)

Комментарии: 39

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

i wish i can give everyone a hug but here’s the next best thing 🤗

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

GPC-1 uses a new dual training technique, jointly training the Qwen base model and add on adapters. As it is, its great, but if you want to customize more, the model was literally trained to be really easy to fine tune for anything 🚀

Фото профиля Yohei
Yohei15 дней назад

this is awesome! thx for sharing. i did something similar but went the opposite direction of using a small vlm to run locally, though no boxes or anything.

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

Love it!

Фото профиля rose🌲🌲
rose🌲🌲15 дней назад

mmmm yes, this won't be used for robotics and anything that contributes to society but used to trade sports prediction books with cameras within stadiums

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

humanity at its best 😂

Фото профиля Max
Max15 дней назад

you cooked with this one. since the Qwen backbone supports video, curious if you tried GPC-1 on videos?

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

not yet but good idea!

Фото профиля Stephane
Stephane15 дней назад

@yoheinakajima Very nice. Congrats. Now I’m curious as to if I can integrate this and how well it will work for the avatar portion of my twin.

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

@yoheinakajima Tons of medical data is being training on now

Фото профиля Stephane
Stephane15 дней назад

@yoheinakajima What’s the medical data that you’re training on and what’s the goal with that training?

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

@yoheinakajima goal rn is to optimize hospital scheduling for surgery based on patients (but training data is broad, including data on how to control lab equipment)

Фото профиля Brjan | AI Builder
Brjan | AI Builder15 дней назад

if it really operates in milliseconds, that could seriously change automation workflows

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

It rly does and it rly will, already deploying it in a few labs rn

Фото профиля Jasper Croome
Jasper Croome15 дней назад

Super cool, classifiers are having the moment they deserve! Did you try this out on maps or geospatial stuff at all?

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

Good idea! Will include it in v1.1

Фото профиля Jasper Croome
Jasper Croome15 дней назад

Happy to give specific use cases / feedback if & when you get into it

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

DM me anytime, I’m always into it

Фото профиля ζ Pedram ζ
ζ Pedram ζ15 дней назад

Cool, wasm port?

Фото профиля John
John15 дней назад

Awesome video Harsha

Фото профиля Wayne Workman
Wayne Workman15 дней назад

This is so cool. There's not enough time in the day to try out all the new models people are creating. How much vram does this use?

Фото профиля Marc
Marc15 дней назад

Congrats Harsha! We’re currently in stealth mode building a humanoid robot 🤖 and would love to learn more about GPC-1. Do you see it being useful for robot control, like watching a human do a task and have a robot replicate the pose and how would you advise using it for this?

Фото профиля varick lim
varick lim15 дней назад

woah this was exactly what i hoped jev could do

Фото профиля Kiran Jd
Kiran Jd15 дней назад

@yoheinakajima Will this work for multiple objects in video?

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

@yoheinakajima Yup, just describe the object in the key (i.e person_to_the_left_x_axis, not the most elegant solution rn but working on it)

Фото профиля Kiran Jd
Kiran Jd15 дней назад

@yoheinakajima I could see so may things being useful in sport tracking. I’m working on analysing a football training drill from coach’s perspective. I would love to try it

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

@yoheinakajima lmk how it performs for you!

Фото профиля Harsh
Harsh15 дней назад

This is so cool! Whats the latency like?

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

Depends on what hardware you use, under 100ms on a h100

Фото профиля girardi's mind
girardi's mind15 дней назад

I'm not sure for robotics because those arms pass through each other, but looks okay for less critical setups

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

ironing those issues out now

Фото профиля Alakaham
Alakaham15 дней назад

Damn...just an unrelated question...how do you make these prod demo videos?

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

one shotted with Gemini and Astra

Фото профиля Mathew Chan
Mathew Chan15 дней назад

How does this do on OCR tasks on docs?

Фото профиля Harsha Gundala
Harsha Gundala15 дней назад

Trained on many OCR examples, lmk how it performs at your use case

Фото профиля robot 2.0
robot 2.015 дней назад

@CDerinbogaz

Фото профиля Aakash Gnanakumar
Aakash Gnanakumar15 дней назад

Bro is too smart!!

Фото профиля Robin | Poker x AI
Robin | Poker x AI15 дней назад

Impressive that you got continuous numeric output in ms, could be a game changer for real time robotics pipelines

Фото профиля pulsatingGenius
pulsatingGenius15 дней назад

Yolo with natural language

Похожие видео

Geoffrey Hinton says a big language model runs on about 1% of your brain's connections and still ends up knowing more than you: "So in your brain, you have a hundred trillion connections, roughly speaking. Okay. That's a lot. And you only live for about two billion seconds. That's not much." "If you compare how many seconds you live for, with how many connections you've got, you have a whole lot more connections than experiences." "Now with these neural nets, it's sort of the other way round. They only have of the order of a trillion connections. So like 1% of your connections, even in a big language model, many of them fewer, but they get thousands of times more experience than you." "So the big language models are solving the problem with not many connections, only a trillion. How do I make use of a huge amount of experience?" "And back propagation is really, really good at packing huge amounts of knowledge into not many connections." "But that's not the problem we're solving. We've got huge numbers of connections, not much experience. We need to sort of extract the most we can from each experience." Two to three billion seconds is the whole budget. Everything you know, you learned inside it. So evolution built you to squeeze a lot out of very little. Hinton's point is that a language model has the opposite problem and the opposite fix, and backprop turned out to be extremely good at that fix. Worth noticing what this predicts about failure. A system running on 1% of your wiring and thousands of times your experience is not going to fail the way you do. You fail from having seen too few examples. It fails from compressing too many into too little, and the compression is where the errors get made. That is a strange thing to be deploying into hospitals and courts with no way to inspect it. We test these systems by asking them questions, which tells you what came out. Nobody can yet look at a trillion connections and say what got packed in. - Geoffrey Hinton, Nobel laureate and Turing Award winner, on StarTalk (StarTalk) with Neil deGrasse Tyson.

Karl Mehta

662,652 просмотров • 2 месяцев назад

David Friedberg Explains Why AI Will Break Every Economic Model: Will Output Surpass Consumptive Capacity? I think fundamentally, if you're driving productivity with AI, you're driving leverage on human time and leverage on capital. The question is, how quickly can you drive that up? And that's a function of how much consumption there is, how much capacity there is for consumption. So if your earnings are the same, but things are getting more expensive, you're not happy. If your earnings go up by 10% and things stay the same price, you got 10% more than you had last year, you're gonna be happy. I just think all humans are driven by this need to consume more each year than they did last year. So I think for me, that's the lower limit on consumptive capacity in the world. The question that we're now facing, which we've never faced in human history before, is there an upper limit on consumptive capacity? Because AI creates such a profound shift in productivity and in leverage that normally you would say, “Hey, when we get a new tool or we get new leverage in a system, we build a new technology, we can make more with less.” Therefore, everyone gets access to more things for the same price or the cost of things that they consume come down by a certain price. But there may be a situation now where the ability to make stuff exceeds the capacity to consume stuff, and that is something that I don't think we’ve faced before. And I think that's where a lot of the models start to break. Just general economic models, just general productivity models and general social models. And this goes to the point about, like, what is everyone going to do? In the same way that I think we've argued that maybe SaaS was a transitory business phenomenon that existed between the foundation of the internet and the era of AI, it may be the case that knowledge work in general is also a transitory phenomenon that only existed between the foundation of the computer, or computing tools, and the existence of AI, generally speaking. And if all of that goes away very quickly, and all of those people can be redistributed and recast into doing other higher level, more creative things, their productivity goes up by 100x, is there really a consumer on the other end of all of that productivity? Is there really enough consumptive capacity? And I think that's the profound question that we all face.

The All-In Podcast

81,924 просмотров • 7 месяцев назад

I asked Garry Tan how to use meta prompting to get better at AI: "My partners at YC Jared Friedman and Pete Koomen showed me how to do this. You can take almost anything that you do all the time and just drop it into a context window. And then say, “Here’s a bunch of inputs and outputs." And maybe you also add a bunch of notes. And then you tell it, “Write me a prompt that can act as an agent that takes this input and makes this output over here.” You can do this for almost any type of knowledge work. And you can even introspect. "What are things you notice that I did to convert this from the input to the output?”. And then you can just start using the prompt. Initially, it’s going to suck. Because it’s just not that smart yet. But what’s funny is now, I also use it to Iterate my writing. You can be very direct, "I would never say that", "Don’t say it like this", or "Oh, you used the long word there, use the short word". Just speak to it conversationally. And then when you're happy with the output, you can use that new output to make a new prompt. "Based on this conversation, give me a better initial prompt that incorporates all the things we talked about." And you can do this with literally everything. And in theory, there’s so much it applies to that people do day-to-day. You could use it for tweets. You could use it for editing podcasts. You can use it for pretty much everything. I have a folder of prompts that I use all the time. My YouTube prompt is on v27 or something. I'll go through this process with all the different max models. I'll use GPT 5.2 Pro. I’ll use Grok. I'll use Claude. Then, I’ll take all the outputs from all the models and put them into Claude and say "Here’s my prompt, here’s the output from four LLMs, including yourself. Rate each response and tell me what the pros and cons of each approach are." And I usually say "give it to me in numbered form". And then you can agree with one, disagree with two, tell it three is this or that. And then after that, you say given all of this, synthesize it."

The Peel

51,632 просмотров • 7 месяцев назад