Loading video...

Video Failed to Load

Go Home

You can run local AI models up to 120B without a $10,000 Mac Studio This phone-sized device is your own server for open-source models. 100% local and private. And you can use it to: - Power an agent like OpenClaw 24/7 - Completely replace a chatbot - Literally anything...

41,356 views • 6 months ago •via X (Twitter)

53 Comments

Paul Couvert's profile picture
Paul Couvert6 months ago

Access to Tiiny Kickstarter page → I’ve been using it for weeks now and have gone from spending hundreds of dollars on APIs to literally zero... All while no longer giving any data to third-party servers. You own your intelligence!

Paul Couvert's profile picture
Paul Couvert6 months ago

Also available on YT if you prefer to watch there!

Sipeed's profile picture
Sipeed6 months ago

Just tested, my 500$ 8845HS miniPC run faster than this 1500$ device, 32tps > 28tps, peoples are fooled... They use the MOE model for promotion, 120B is only 3B actived, don't you feel strange that 120B,30B,20B model have almost same output eval speed?

Paul Couvert's profile picture
Paul Couvert6 months ago

Big fan of PicoClaw btw. Wrote an article to run it on Android. GPT OSS 120B has 5.1B active parameters not 3B + you still have to fit the full 120B in the RAM which is just not possible on your $500 mini PC... or my $1200 laptop with 32GB of RAM.

Sipeed's profile picture
Sipeed6 months ago

Thank you for your tutorials! Just want to clarify, their marketing is honestly a bit deceptive. They don't disclose the active parameters for any of the 3 models, tricking users into thinking they're dense. It only runs a 120B model because of the massive RAM, not because of some insane processing power. BTW, I have a $500 8845 mini PC packed with 96GB LPDDR5X (snagged it 1 year ago during the memory price dip), so I can run 120B models too. For local LLM setups today, the real cost bottleneck is always memory capacity and bandwidth, not raw compute.

Paul Couvert's profile picture
Paul Couvert6 months ago

Yep can't agree more on the memory / bandwidth part. Honestly the benefit of are the massive amount of memory, the NPU usage so you it doesn't consume a lot of power and the form factor. But the MoE disclosure could be important, true.

Sipeed's profile picture
Sipeed6 months ago

Just curious about another thing: its 80GB of memory is actually split into 48GB dNPU and 32GB CPU. Typically, it's really hard to pool these two together effectively when running a single model. Since a 120B model at INT4 requires at least 60GB of memory, are they using an even lower-bit quantization?

Paul Couvert's profile picture
Paul Couvert6 months ago

Yes they are. But they constrain the accuracy gap between quantized and original floating-point models to within ±0.5 percentage points.

Sipeed's profile picture
Sipeed6 months ago

Ok, but that means 48~64GB normal CPU solution is able to run this low-bit 120B model too... In fact, we were planed to make a device that able to run dense 30B at the speed tiiny run A3B model. but the DDR price is too high now... but tiiny's ks result shocked us, we will consider it again!

Paul Couvert's profile picture
Paul Couvert6 months ago

Probably. To be confirmed but I think that they're using the 48GB from the NPU for the LLM model weights and the 32GB for everything else like the OS and KV cache (context memory). So idk, I'm not good enough to appreciate if it's possible to reproduce this kind of power consumption and speed balance with a regular CPU solution. Looking forward to yours if you're doing it!

Florian Camiade 🗝️'s profile picture
Florian Camiade 🗝️6 months ago

Hope no one bought 12 Mac Mini

Paul Couvert's profile picture
Paul Couvert6 months ago

From "a $10k Mac Studio is all you need" to just plug a phone sized device to your potato laptop

Saurabh Singhvi's profile picture
Saurabh Singhvi6 months ago

Are there any AI TOPS numbers on this awesome thing?

Paul Couvert's profile picture
Paul Couvert6 months ago

@singhvisaurabh Yep! 30 TOPS (and 32GB of RAM) on the SoC and 160 TOPS (and 48GB of RAM) on the NPU

Saurabh Singhvi's profile picture
Saurabh Singhvi6 months ago

Awesome!

alexintosh's profile picture
alexintosh6 months ago

20t/s what is this AI for 🐜

Mihai Balint's profile picture
Mihai Balint6 months ago

This has all the data I need to know: 2048 context length was small 4 years ago. In 2026, agents need contexts with 100k-1M tokens.

Paul Couvert's profile picture
Paul Couvert6 months ago

Speed test, you're (of course) not limited to a 2048 context window as you can see in my Hermes Agent demo in the video :)

Martin Szerment | Practical AI's profile picture
Martin Szerment | Practical AI6 months ago

If that thing runs 120B locally, cloud bills just became optional overnight.

Paul Couvert's profile picture
Paul Couvert6 months ago

They are to me now!

Zack_@'s profile picture
Zack_@6 months ago

Performance is directly proportional to energy consumption, making this power bank's toy-like appearance seem like a scam.

Nithin's profile picture
Nithin6 months ago

Can it run multiple agents simultaneously?

Martin S.'s profile picture
Martin S.6 months ago

120B on a phone-sized box called Tiiny is kinda wild. if it actually stays cool under load, running an OpenClaw node 24/7 gets way more real

LightShift Studio's profile picture
LightShift Studio6 months ago

Hey there! 🤖 Still manually crunching numbers? Tiiny could make us laugh with savings on your last pizza purchase instead of APIs...

Mingta Kaivo 明塔 开沃's profile picture
Mingta Kaivo 明塔 开沃6 months ago

160 TOPS on the NPU is wild. been running Qwen 4B on my M4 Mini for audio classification at 180 tok/s — if this matches that throughput at 1/10th the size and price, my whole inference stack fits in a drawer

SK ⚡️'s profile picture
SK ⚡️6 months ago

Does this support Exo clustering? Able to run the Nvidia optimized models?

Paul Couvert's profile picture
Paul Couvert6 months ago

No I don't think so for Exo. It runs "custom" GGUF models but you'll be able to convert any of them to run on the device.

Gregor's profile picture
Gregor6 months ago

What are the implications for indie developers like myself when local AI models become more accessible? How will this change our approach to building apps, and can we expect a new wave of innovation as a result?

WorldWarrior's profile picture
WorldWarrior6 months ago

Is there a place to learn how to use it and what maybe able to run on it and how.

Jibran Malik's profile picture
Jibran Malik6 months ago

Running 120B models on a phone sized server is a massive flex. Imagine powering the BeeClaw agent through this 24/7 for trades lowkey a dream setup for privacy. I’m actually trying to snag one of the $500 Advanced accounts from BeeTrade’s giveaway right now to test their new OpenClaw agent.

Vlada5's profile picture
Vlada56 months ago

yes, but it does not run Linux...you get it?

Paul Couvert's profile picture
Paul Couvert6 months ago

And I didn't try to run Doom yet...

Vlada5's profile picture
Vlada56 months ago

ask yourself why

Tendies Of Wisdom's profile picture
Tendies Of Wisdom6 months ago

Looks cool but they have to provide the models 🤔

Paul Couvert's profile picture
Paul Couvert6 months ago

One of my 1st questions as well. And in fact they're going to publish (before deliveries) a tool that allows you to convert any GGUF so that it runs on the device.

lazybutai's profile picture
lazybutai6 months ago

@TheAhmadOsman what do you think? Is TIny a good plug and play alternative to GPUs?

Magnus Ahlden's profile picture
Magnus Ahlden6 months ago

What’s you affiliation with the company who makes these? From what I can see it’s a great idea - but the spec doesn’t add up. Seems to over promise. For instance openclaw requires huge context windows to be usable. The models listed don’t support it in any usable way…

codemarch's profile picture
codemarch6 months ago

This is a fascinating direction, compact local hardware like this could make private AI much more accessible to builders and developers.

cCross's profile picture
cCross6 months ago

Latency matters, but 30 TOPS helps keep it usable

Shubham Sharma | Video Editor's profile picture
Shubham Sharma | Video Editor6 months ago

What spec I need for the Phone

Paul Couvert's profile picture
Paul Couvert6 months ago

Phone-sized device, not a phone haha

Shubham Sharma | Video Editor's profile picture
Shubham Sharma | Video Editor6 months ago

Oooo, Well I am a student I can't afford the Api cost nor I can get Mac studio, What should I do any advice

0x₿erto 's profile picture
0x₿erto 6 months ago

120B model. Stop the clickbait. All local, no api calls to LLMs… lol

ал's profile picture
ал6 months ago

No price? F U

Paul Couvert's profile picture
Paul Couvert6 months ago

I've literally linked everything in the second post lmao

David Motta's profile picture
David Motta6 months ago

never thought something like this could fit in a pocket. wonder how hot it gets running those big models nonstop.

Paul Couvert's profile picture
Paul Couvert6 months ago

@davidmotta Just a bit warm even after days!

0x₿erto 's profile picture
0x₿erto 6 months ago

Ah 48gb ram I see

Paul Couvert's profile picture
Paul Couvert6 months ago

Yup. 80GB RAM total considering both the SoC and the NPU.

ACCELERATIONSIM | XLR8's profile picture
ACCELERATIONSIM | XLR86 months ago

OpenClaw on the nightstand, $CLAUBE on the timeline. two unsupervised agents, zero oversight, infinite chaos

Kyriakos's profile picture
Kyriakos6 months ago

Running big models locally is huge

AI for SaaS's profile picture
AI for SaaS6 months ago

Cloud AI is rented. Local AI is owned. A phone‑sized server running 120B models flips the economics of autonomy forever.

Yanko's profile picture
Yanko6 months ago

This is exactly the problem ClawBox solves — we built it on a Jetson Orin Nano (67 TOPS, 8GB). No Mac Studio needed. Plug in, scan QR, full local AI in 5 min. Your data never leaves the device.

Related Videos