
Cactus Compute
@cactuscompute • 3,055 subscribers
Automation Foundation Models For Tiny Devices
Videos

We release Needle 3: A Sliceable 8-29MB automation foundation model that can match DeepSeek V4 Flash. One set of weights, every depth from 2 to 20 layers a model of its own, 25-121M parameters at CQ2-bit, built on our Simple Attention Networks and running locally at up to 4k tokens/sec decode speed on a Raspberry Pi 5. Needle does not chat. Every turn is a function call: give it the tools your app exposes and it picks the right ones and fills every argument from what the user said, or hand it a schema and it returns a typed record. Ask for something no tool covers and you get an empty list, not a guess. That trade is lets 121M parameters trained on 360B tokens of structured data beat models 10x their size on mobile tool calls and match 2-3x bigger models on structured JSON extraction. It runs on mobiles, wearables, smart home devices, small robots and microcontrollers, with prebuilt engines for macOS, Linux, Windows, Android, iOS, watchOS, tvOS, the browser and WASI hosts. Try it in your browser:
Cactus Compute401,091 просмотров • 5 дней назад
Больше нет контента для загрузки