Video wird geladen...
Video konnte nicht geladen werden
Introducing VisionPsy-Nano: state-of-the-art vision-language models at 460M parameters, small enough to run on the phone in your pocket. VisionPsy-Nano-460M leads every other ~0.5B vision-language model tested & compares favourably on 16 of 17 benchmarks, tops all four capability categories, and outperforms models up to 2.3x larger. Two variants, one... show more
15,329,496 Aufrufe • vor 1 Monat •via X (Twitter)
29 Kommentare

Everything is out today: open weights, the blog, and the full benchmark breakdown. Model weights: Hugging Face Blog :

does the 0.3s iPhone TTFT include image decode and resize, or does timing start once the 512x512 tensor is ready?

wait can i run this on my iphone?

Open weights, runs on a phone, beats models 2.3x its size. Brussels still drafting rules for models that won't exist by the time they finish.

460M beating models 2x its size is genuinely impressive

@woxshter le goat

Gibraltar est européenne mais Ceuta et Melilla sont africaine

how's it compare to the phi-3 vision models at similar size?

Wow

460M leading its class on-device is real. But "leads every benchmark" self-scored needs a number. The story is decode tok/s + prefill on actual phone silicon—and the SigLIP encoder cost, since the vision tower usually eats mobile latency, not the 460M LM. What's real-time in ms?

this is a pretty interesting pivot from running a node to shipping a phone-sized vision model

The 460M number is interesting; the real product win isn't a benchmark. Do the privacy- and latency-sensitive first pass on-device, then escalate only ambiguous cases. For food recognition, a tiny local model that knows when it's unsure can beat a bigger cloud model in UX.

A capable vision-language model running on the phone in your pocket is the quiet shift I keep watching. Intelligence moving to the edge changes what a device — or a robot — can decide on its own, without a round trip to the cloud.

460M on device and beating 2.3x larger models is insane

Impressive work! A 460M parameter VLM outperforming models 2.3x larger is a great testament to efficient architecture design. Apache 2.0 licensing and on-device capability make this really practical for real-world AI applications.

Great news!

that caught my attention

Nice

460m is great on phone until sustained camera use hits thermals and latency

Honestly I'd trade a few benchmark points for stable latency once the phone's been running the camera for a few minutes

Bla bla bla...bull shit....

Active

phone sized vision model beating bigger ones is wild

The height of technology keeps increasing by the day

VisionPsy Nano 460M runs on your phone and tops benchmarks

Impressive vision performance packed into such a compact model

Zbb e7 8ef

Tether终于不务正业了?不过VisionPsy要是能用USDT训练,我第一个下单——毕竟币圈最不缺的就是算力🔥

Local AI running directly on phones is a huge win 💪
