Video wird geladen...
Video konnte nicht geladen werden
You can run local AI models up to 120B without a $10,000 Mac Studio This phone-sized device is your own server for open-source models. 100% local and private. And you can use it to: - Power an agent like OpenClaw 24/7 - Completely replace a chatbot - Literally anything... show more
41,356 Aufrufe • vor 6 Monaten •via X (Twitter)
53 Kommentare

Access to Tiiny Kickstarter page → I’ve been using it for weeks now and have gone from spending hundreds of dollars on APIs to literally zero... All while no longer giving any data to third-party servers. You own your intelligence!

Also available on YT if you prefer to watch there!

Just tested, my 500$ 8845HS miniPC run faster than this 1500$ device, 32tps > 28tps, peoples are fooled... They use the MOE model for promotion, 120B is only 3B actived, don't you feel strange that 120B,30B,20B model have almost same output eval speed?

Big fan of PicoClaw btw. Wrote an article to run it on Android. GPT OSS 120B has 5.1B active parameters not 3B + you still have to fit the full 120B in the RAM which is just not possible on your $500 mini PC... or my $1200 laptop with 32GB of RAM.

Thank you for your tutorials! Just want to clarify, their marketing is honestly a bit deceptive. They don't disclose the active parameters for any of the 3 models, tricking users into thinking they're dense. It only runs a 120B model because of the massive RAM, not because of some insane processing power. BTW, I have a $500 8845 mini PC packed with 96GB LPDDR5X (snagged it 1 year ago during the memory price dip), so I can run 120B models too. For local LLM setups today, the real cost bottleneck is always memory capacity and bandwidth, not raw compute.

Yep can't agree more on the memory / bandwidth part. Honestly the benefit of are the massive amount of memory, the NPU usage so you it doesn't consume a lot of power and the form factor. But the MoE disclosure could be important, true.

Just curious about another thing: its 80GB of memory is actually split into 48GB dNPU and 32GB CPU. Typically, it's really hard to pool these two together effectively when running a single model. Since a 120B model at INT4 requires at least 60GB of memory, are they using an even lower-bit quantization?

Yes they are. But they constrain the accuracy gap between quantized and original floating-point models to within ±0.5 percentage points.

Ok, but that means 48~64GB normal CPU solution is able to run this low-bit 120B model too... In fact, we were planed to make a device that able to run dense 30B at the speed tiiny run A3B model. but the DDR price is too high now... but tiiny's ks result shocked us, we will consider it again!

Probably. To be confirmed but I think that they're using the 48GB from the NPU for the LLM model weights and the 32GB for everything else like the OS and KV cache (context memory). So idk, I'm not good enough to appreciate if it's possible to reproduce this kind of power consumption and speed balance with a regular CPU solution. Looking forward to yours if you're doing it!

Hope no one bought 12 Mac Mini

From "a $10k Mac Studio is all you need" to just plug a phone sized device to your potato laptop

Are there any AI TOPS numbers on this awesome thing?

@singhvisaurabh Yep! 30 TOPS (and 32GB of RAM) on the SoC and 160 TOPS (and 48GB of RAM) on the NPU

Awesome!

20t/s what is this AI for 🐜

This has all the data I need to know: 2048 context length was small 4 years ago. In 2026, agents need contexts with 100k-1M tokens.

Speed test, you're (of course) not limited to a 2048 context window as you can see in my Hermes Agent demo in the video :)

If that thing runs 120B locally, cloud bills just became optional overnight.

They are to me now!

Performance is directly proportional to energy consumption, making this power bank's toy-like appearance seem like a scam.

Can it run multiple agents simultaneously?

120B on a phone-sized box called Tiiny is kinda wild. if it actually stays cool under load, running an OpenClaw node 24/7 gets way more real

Hey there! 🤖 Still manually crunching numbers? Tiiny could make us laugh with savings on your last pizza purchase instead of APIs...

160 TOPS on the NPU is wild. been running Qwen 4B on my M4 Mini for audio classification at 180 tok/s — if this matches that throughput at 1/10th the size and price, my whole inference stack fits in a drawer

Does this support Exo clustering? Able to run the Nvidia optimized models?

No I don't think so for Exo. It runs "custom" GGUF models but you'll be able to convert any of them to run on the device.

What are the implications for indie developers like myself when local AI models become more accessible? How will this change our approach to building apps, and can we expect a new wave of innovation as a result?

Is there a place to learn how to use it and what maybe able to run on it and how.

Running 120B models on a phone sized server is a massive flex. Imagine powering the BeeClaw agent through this 24/7 for trades lowkey a dream setup for privacy. I’m actually trying to snag one of the $500 Advanced accounts from BeeTrade’s giveaway right now to test their new OpenClaw agent.

yes, but it does not run Linux...you get it?

And I didn't try to run Doom yet...

ask yourself why

Looks cool but they have to provide the models 🤔

One of my 1st questions as well. And in fact they're going to publish (before deliveries) a tool that allows you to convert any GGUF so that it runs on the device.

@TheAhmadOsman what do you think? Is TIny a good plug and play alternative to GPUs?

What’s you affiliation with the company who makes these? From what I can see it’s a great idea - but the spec doesn’t add up. Seems to over promise. For instance openclaw requires huge context windows to be usable. The models listed don’t support it in any usable way…

This is a fascinating direction, compact local hardware like this could make private AI much more accessible to builders and developers.

Latency matters, but 30 TOPS helps keep it usable

What spec I need for the Phone

Phone-sized device, not a phone haha

Oooo, Well I am a student I can't afford the Api cost nor I can get Mac studio, What should I do any advice

120B model. Stop the clickbait. All local, no api calls to LLMs… lol

No price? F U

I've literally linked everything in the second post lmao

never thought something like this could fit in a pocket. wonder how hot it gets running those big models nonstop.

@davidmotta Just a bit warm even after days!

Ah 48gb ram I see

Yup. 80GB RAM total considering both the SoC and the NPU.

OpenClaw on the nightstand, $CLAUBE on the timeline. two unsupervised agents, zero oversight, infinite chaos

Running big models locally is huge

Cloud AI is rented. Local AI is owned. A phone‑sized server running 120B models flips the economics of autonomy forever.

This is exactly the problem ClawBox solves — we built it on a Jetson Orin Nano (67 TOPS, 8GB). No Mac Studio needed. Plug in, scan QR, full local AI in 5 min. Your data never leaves the device.

