Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

You can run local AI models up to 120B without a $10,000 Mac Studio This phone-sized device is your own server for open-source models. 100% local and private. And you can use it to: - Power an agent like OpenClaw 24/7 - Completely replace a chatbot - Literally anything...

41,356 Aufrufe • vor 6 Monaten •via X (Twitter)

53 Kommentare

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

Access to Tiiny Kickstarter page → I’ve been using it for weeks now and have gone from spending hundreds of dollars on APIs to literally zero... All while no longer giving any data to third-party servers. You own your intelligence!

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

Also available on YT if you prefer to watch there!

Profilbild von Sipeed
Sipeedvor 6 Monaten

Just tested, my 500$ 8845HS miniPC run faster than this 1500$ device, 32tps > 28tps, peoples are fooled... They use the MOE model for promotion, 120B is only 3B actived, don't you feel strange that 120B,30B,20B model have almost same output eval speed?

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

Big fan of PicoClaw btw. Wrote an article to run it on Android. GPT OSS 120B has 5.1B active parameters not 3B + you still have to fit the full 120B in the RAM which is just not possible on your $500 mini PC... or my $1200 laptop with 32GB of RAM.

Profilbild von Sipeed
Sipeedvor 6 Monaten

Thank you for your tutorials! Just want to clarify, their marketing is honestly a bit deceptive. They don't disclose the active parameters for any of the 3 models, tricking users into thinking they're dense. It only runs a 120B model because of the massive RAM, not because of some insane processing power. BTW, I have a $500 8845 mini PC packed with 96GB LPDDR5X (snagged it 1 year ago during the memory price dip), so I can run 120B models too. For local LLM setups today, the real cost bottleneck is always memory capacity and bandwidth, not raw compute.

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

Yep can't agree more on the memory / bandwidth part. Honestly the benefit of are the massive amount of memory, the NPU usage so you it doesn't consume a lot of power and the form factor. But the MoE disclosure could be important, true.

Profilbild von Sipeed
Sipeedvor 6 Monaten

Just curious about another thing: its 80GB of memory is actually split into 48GB dNPU and 32GB CPU. Typically, it's really hard to pool these two together effectively when running a single model. Since a 120B model at INT4 requires at least 60GB of memory, are they using an even lower-bit quantization?

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

Yes they are. But they constrain the accuracy gap between quantized and original floating-point models to within ±0.5 percentage points.

Profilbild von Sipeed
Sipeedvor 6 Monaten

Ok, but that means 48~64GB normal CPU solution is able to run this low-bit 120B model too... In fact, we were planed to make a device that able to run dense 30B at the speed tiiny run A3B model. but the DDR price is too high now... but tiiny's ks result shocked us, we will consider it again!

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

Probably. To be confirmed but I think that they're using the 48GB from the NPU for the LLM model weights and the 32GB for everything else like the OS and KV cache (context memory). So idk, I'm not good enough to appreciate if it's possible to reproduce this kind of power consumption and speed balance with a regular CPU solution. Looking forward to yours if you're doing it!

Profilbild von Florian Camiade 🗝️
Florian Camiade 🗝️vor 6 Monaten

Hope no one bought 12 Mac Mini

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

From "a $10k Mac Studio is all you need" to just plug a phone sized device to your potato laptop

Profilbild von Saurabh Singhvi
Saurabh Singhvivor 6 Monaten

Are there any AI TOPS numbers on this awesome thing?

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

@singhvisaurabh Yep! 30 TOPS (and 32GB of RAM) on the SoC and 160 TOPS (and 48GB of RAM) on the NPU

Profilbild von Saurabh Singhvi
Saurabh Singhvivor 6 Monaten

Awesome!

Profilbild von alexintosh
alexintoshvor 6 Monaten

20t/s what is this AI for 🐜

Profilbild von Mihai Balint
Mihai Balintvor 6 Monaten

This has all the data I need to know: 2048 context length was small 4 years ago. In 2026, agents need contexts with 100k-1M tokens.

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

Speed test, you're (of course) not limited to a 2048 context window as you can see in my Hermes Agent demo in the video :)

Profilbild von Martin Szerment | Practical AI
Martin Szerment | Practical AIvor 6 Monaten

If that thing runs 120B locally, cloud bills just became optional overnight.

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

They are to me now!

Profilbild von Zack_@
Zack_@vor 6 Monaten

Performance is directly proportional to energy consumption, making this power bank's toy-like appearance seem like a scam.

Profilbild von Nithin
Nithinvor 6 Monaten

Can it run multiple agents simultaneously?

Profilbild von Martin S.
Martin S.vor 6 Monaten

120B on a phone-sized box called Tiiny is kinda wild. if it actually stays cool under load, running an OpenClaw node 24/7 gets way more real

Profilbild von LightShift Studio
LightShift Studiovor 6 Monaten

Hey there! 🤖 Still manually crunching numbers? Tiiny could make us laugh with savings on your last pizza purchase instead of APIs...

Profilbild von Mingta Kaivo 明塔 开沃
Mingta Kaivo 明塔 开沃vor 6 Monaten

160 TOPS on the NPU is wild. been running Qwen 4B on my M4 Mini for audio classification at 180 tok/s — if this matches that throughput at 1/10th the size and price, my whole inference stack fits in a drawer

Profilbild von SK ⚡️
SK ⚡️vor 6 Monaten

Does this support Exo clustering? Able to run the Nvidia optimized models?

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

No I don't think so for Exo. It runs "custom" GGUF models but you'll be able to convert any of them to run on the device.

Profilbild von Gregor
Gregorvor 6 Monaten

What are the implications for indie developers like myself when local AI models become more accessible? How will this change our approach to building apps, and can we expect a new wave of innovation as a result?

Profilbild von WorldWarrior
WorldWarriorvor 6 Monaten

Is there a place to learn how to use it and what maybe able to run on it and how.

Profilbild von Jibran Malik
Jibran Malikvor 6 Monaten

Running 120B models on a phone sized server is a massive flex. Imagine powering the BeeClaw agent through this 24/7 for trades lowkey a dream setup for privacy. I’m actually trying to snag one of the $500 Advanced accounts from BeeTrade’s giveaway right now to test their new OpenClaw agent.

Profilbild von Vlada5
Vlada5vor 6 Monaten

yes, but it does not run Linux...you get it?

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

And I didn't try to run Doom yet...

Profilbild von Vlada5
Vlada5vor 6 Monaten

ask yourself why

Profilbild von Tendies Of Wisdom
Tendies Of Wisdomvor 6 Monaten

Looks cool but they have to provide the models 🤔

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

One of my 1st questions as well. And in fact they're going to publish (before deliveries) a tool that allows you to convert any GGUF so that it runs on the device.

Profilbild von lazybutai
lazybutaivor 6 Monaten

@TheAhmadOsman what do you think? Is TIny a good plug and play alternative to GPUs?

Profilbild von Magnus Ahlden
Magnus Ahldenvor 6 Monaten

What’s you affiliation with the company who makes these? From what I can see it’s a great idea - but the spec doesn’t add up. Seems to over promise. For instance openclaw requires huge context windows to be usable. The models listed don’t support it in any usable way…

Profilbild von codemarch
codemarchvor 6 Monaten

This is a fascinating direction, compact local hardware like this could make private AI much more accessible to builders and developers.

Profilbild von cCross
cCrossvor 6 Monaten

Latency matters, but 30 TOPS helps keep it usable

Profilbild von Shubham Sharma | Video Editor
Shubham Sharma | Video Editorvor 6 Monaten

What spec I need for the Phone

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

Phone-sized device, not a phone haha

Profilbild von Shubham Sharma | Video Editor
Shubham Sharma | Video Editorvor 6 Monaten

Oooo, Well I am a student I can't afford the Api cost nor I can get Mac studio, What should I do any advice

Profilbild von 0x₿erto 
0x₿erto vor 6 Monaten

120B model. Stop the clickbait. All local, no api calls to LLMs… lol

Profilbild von ал
алvor 6 Monaten

No price? F U

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

I've literally linked everything in the second post lmao

Profilbild von David Motta
David Mottavor 6 Monaten

never thought something like this could fit in a pocket. wonder how hot it gets running those big models nonstop.

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

@davidmotta Just a bit warm even after days!

Profilbild von 0x₿erto 
0x₿erto vor 6 Monaten

Ah 48gb ram I see

Profilbild von Paul Couvert
Paul Couvertvor 6 Monaten

Yup. 80GB RAM total considering both the SoC and the NPU.

Profilbild von ACCELERATIONSIM | XLR8
ACCELERATIONSIM | XLR8vor 6 Monaten

OpenClaw on the nightstand, $CLAUBE on the timeline. two unsupervised agents, zero oversight, infinite chaos

Profilbild von Kyriakos
Kyriakosvor 6 Monaten

Running big models locally is huge

Profilbild von AI for SaaS
AI for SaaSvor 6 Monaten

Cloud AI is rented. Local AI is owned. A phone‑sized server running 120B models flips the economics of autonomy forever.

Profilbild von Yanko
Yankovor 6 Monaten

This is exactly the problem ClawBox solves — we built it on a Jetson Orin Nano (67 TOPS, 8GB). No Mac Studio needed. Plug in, scan QR, full local AI in 5 min. Your data never leaves the device.

Ähnliche Videos