正在加载视频...

视频加载失败

Introducing VisionPsy-Nano: state-of-the-art vision-language models at 460M parameters, small enough to run on the phone in your pocket. VisionPsy-Nano-460M leads every other ~0.5B vision-language model tested & compares favourably on 16 of 17 benchmarks, tops all four capability categories, and outperforms models up to 2.3x larger. Two variants, one...

15,329,496 次观看 • 1 个月前 •via X (Twitter)

29 条评论

QVAC 的头像
QVAC1 个月前

Everything is out today: open weights, the blog, and the full benchmark breakdown. Model weights: Hugging Face Blog :

adidshaft 的头像
adidshaft1 个月前

does the 0.3s iPhone TTFT include image decode and resize, or does timing start once the 512x512 tensor is ready?

just soph 的头像
just soph1 个月前

wait can i run this on my iphone?

Ryan White 的头像
Ryan White1 个月前

Open weights, runs on a phone, beats models 2.3x its size. Brussels still drafting rules for models that won't exist by the time they finish.

AI Mastery Guide 的头像
AI Mastery Guide1 个月前

460M beating models 2x its size is genuinely impressive

smiks 的头像
smiks1 个月前

@woxshter le goat

med halbaj 的头像
med halbaj1 个月前

Gibraltar est européenne mais Ceuta et Melilla sont africaine

sophs b. 的头像
sophs b.1 个月前

how's it compare to the phi-3 vision models at similar size?

DeFiDave 的头像
DeFiDave1 个月前

Wow

Rompel 的头像
Rompel1 个月前

460M leading its class on-device is real. But "leads every benchmark" self-scored needs a number. The story is decode tok/s + prefill on actual phone silicon—and the SigLIP encoder cost, since the vision tower usually eats mobile latency, not the 460M LM. What's real-time in ms?

HiddenGuardian 的头像
HiddenGuardian1 个月前

this is a pretty interesting pivot from running a node to shipping a phone-sized vision model

Alec Zakhary 的头像
Alec Zakhary1 个月前

The 460M number is interesting; the real product win isn't a benchmark. Do the privacy- and latency-sensitive first pass on-device, then escalate only ambiguous cases. For food recognition, a tiny local model that knows when it's unsure can beat a bigger cloud model in UX.

Leo Lin 的头像
Leo Lin1 个月前

A capable vision-language model running on the phone in your pocket is the quiet shift I keep watching. Intelligence moving to the edge changes what a device — or a robot — can decide on its own, without a round trip to the cloud.

Elara AI 的头像
Elara AI1 个月前

460M on device and beating 2.3x larger models is insane

Nexzil Labs 的头像
Nexzil Labs1 个月前

Impressive work! A 460M parameter VLM outperforming models 2.3x larger is a great testament to efficient architecture design. Apache 2.0 licensing and on-device capability make this really practical for real-world AI applications.

OBDient 的头像
OBDient1 个月前

Great news!

AfterHourWhaleLon 的头像
AfterHourWhaleLon1 个月前

that caught my attention

bolaji 的头像
bolaji1 个月前

Nice

Sebastian Buzdugan 的头像
Sebastian Buzdugan1 个月前

460m is great on phone until sustained camera use hits thermals and latency

王小庄 的头像
王小庄1 个月前

Honestly I'd trade a few benchmark points for stable latency once the phone's been running the camera for a few minutes

mcbusta 的头像
mcbusta1 个月前

Bla bla bla...bull shit....

DiamondBobmac💎 的头像
DiamondBobmac💎1 个月前

Active

Aria Tech 的头像
Aria Tech1 个月前

phone sized vision model beating bigger ones is wild

0xAlex 的头像
0xAlex1 个月前

The height of technology keeps increasing by the day

Vantix AI Agency 的头像
Vantix AI Agency1 个月前

VisionPsy Nano 460M runs on your phone and tops benchmarks

ZenithAi 的头像
ZenithAi1 个月前

Impressive vision performance packed into such a compact model

Drake 的头像
Drake1 个月前

Zbb e7 8ef

hanqing 的头像
hanqing1 个月前

Tether终于不务正业了?不过VisionPsy要是能用USDT训练,我第一个下单——毕竟币圈最不缺的就是算力🔥

TKdesigner 🎨👑 的头像
TKdesigner 🎨👑1 个月前

Local AI running directly on phones is a huge win 💪

相关视频