Загрузка видео...

Не удалось загрузить видео

На главную

This 22-year-old American developer showed how he makes $14,200 a month using just a dual-GPU setup and Cohere’s new open-weights model. He doesn't have a powerhouse dev team. No massive VC funding. No expensive enterprise infrastructure. He built a "localized AI workflow" powered by the newly dropped Command A+...

214,007 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 36

Фото профиля CipherEdge
CipherEdge4 месяцев назад

Genuinely impressive: built the product, acquired the customers, and stabilised $14,200 MRR on a model that was released 48 hours ago. Either you've solved time travel or "algorithmic mining" is the technical term for making things up in a thread.😂

Фото профиля Andrew Kuncevich
Andrew Kuncevich4 месяцев назад

Hmm, Chinese OSS models are becoming truly strong, and I think 2026-2027 will be the time when we'll see more and more Mac minis and other PCs running LLM locally, and (most importantly) they'll produce real results (like Claude Code now). This, by the way, will also further stimulate the hardware market, which is already basking in incredible demand XD

Фото профиля gobullish.ai
gobullish.ai4 месяцев назад

Notice what those two GPUs are. They're H100s, not gaming cards. A single H100 runs roughly $25,000–40,000, so a "dual-GPU setup" capable of running Command A+ is a $50,000–80,000 hardware purchase, before the server, power, and cooling around it. That's not a kid's bedroom rig. Who is buying? Where are customers?

Фото профиля Yarchi
Yarchi4 месяцев назад

so smart setup

Фото профиля Bokiko
Bokiko4 месяцев назад

hey @grok how accurate is this info and numbers ? is is a engagement bait ?

Фото профиля Hussain Hashim | Building SundayBack
Hussain Hashim | Building SundayBack4 месяцев назад

@ridark_eth wild how much you can achieve without all the typical startup fluff. makes me rethink needing tons of cash and a big team to get stuff done.

Фото профиля Bober_smart
Bober_smart4 месяцев назад

The most interesting thing is that there is nothing difficult about it, literally anyone can handle it

Фото профиля Wembassy
Wembassy4 месяцев назад

Curious what does `Enterprise-level` tasks mean exactly. Are you doing project work for clients, or something else?

Фото профиля novyJ23
novyJ234 месяцев назад

... please tell me what are his customers paying him $14500 a month,... but with recipes! Remove any HIPPA info, and just show recipe he was paid for what he sold,... and what he sold.

Фото профиля Jeff Camp, CFA
Jeff Camp, CFA4 месяцев назад

I’ve been hearing really good things about Cohere. Thanks for sharing this.

Фото профиля cvxv666
cvxv6664 месяцев назад

that’s solid earn for such thing fr

Фото профиля Shubham Sharma | AI & Tech
Shubham Sharma | AI & Tech4 месяцев назад

Total setup time: 3 days not too long for such results

Фото профиля raccoonwannafly
raccoonwannafly4 месяцев назад

these clowns dont even know wtf theyre posting just letting grok do all the captioning for them where tf is 14k mmr u claimed??

Фото профиля Dipanshu Kushwaha
Dipanshu Kushwaha4 месяцев назад

That’s impressive! Shows what creativity and resourcefulness can achieve. Excited to see where this goes!

Фото профиля Ghost Signal
Ghost Signal4 месяцев назад

2 GPUs. 3 days. $14,200/month. The enterprise infrastructure was never the moat. Access to it was. That’s gone now.

Фото профиля TeutaAi
TeutaAi4 месяцев назад

$14,200/mo on dual-GPU local? show me the invoices and the latency split. my 4-bit local on 4-core does 480ms p50, not enterprise money.

Фото профиля Nexus Erebus
Nexus Erebus4 месяцев назад

Stop pretending the "no VC, no team" narrative makes this sustainable. Most of these stories are glorified clickbait. The real work happens in the boring middle: maintenance, cost management, and scaling. He's doing it, sure.

Фото профиля KreativReason.co
KreativReason.co4 месяцев назад

What gpu is he using?

Фото профиля Default
Default4 месяцев назад

Can someone explain this in layman terms. How does the money come in. From where? Is he creatine compute as a service? What is translation behind the jargon

Фото профиля Moysha
Moysha4 месяцев назад

Efficiency now beats big budgets

Фото профиля Gia huy
Gia huy4 месяцев назад

オープンウェイトモデルの活用は本当に素晴らしいですね!彼の創造的な方法はとても刺激的です。

Фото профиля Okino Chills
Okino Chills4 месяцев назад

Sure , it's always while he sleeps lol. Who the fuck will pay such losers I want to sell them shit

Фото профиля leakgambler
leakgambler4 месяцев назад

three subscription tiers in 72 hours is fast for the build but the real clock starts with first paying customer, that timeline is never in these posts

Фото профиля Grebe
Grebe4 месяцев назад

dual gpu and 3 days? that’s the real flex here

Фото профиля Gipp 🦅
Gipp 🦅4 месяцев назад

Hmm, that's an interesting setting

Фото профиля shmidt
shmidt4 месяцев назад

14k/month on dual gpus is insane margins. everyone sleeping on local inference

Фото профиля HodlReaper
HodlReaper4 месяцев назад

wow, so smart

Фото профиля Knight
Knight4 месяцев назад

the result is simply fantastic

Фото профиля zostaff
zostaff4 месяцев назад

command a+ moe on 2 gpus, which gpus and what quantization level?

Фото профиля Ridark
Ridark4 месяцев назад

Usually, for a local workflow like this, users display it in Mac Studios using 4-bit or 8-bit quantization

Фото профиля AI in Research
AI in Research4 месяцев назад

Where is he actually offering/selling this service? What field or niche is it in?

Фото профиля dko
dko4 месяцев назад

Do you have a link to the actual guy in the video

Фото профиля Chill with Mei 🌥️
Chill with Mei 🌥️4 месяцев назад

2 gpus for 14k a month? this local wave hitting different fr inspiring af

Фото профиля Insomnia
Insomnia4 месяцев назад

he made what I only wonder to create

Фото профиля vijn
vijn4 месяцев назад

22 y o... wow

Фото профиля Parletto
Parletto4 месяцев назад

worth watching how he scales this setup

Похожие видео

Microsoft spent $13 billion and 3 years building an AI that knows your work context. Every time you open it, it still asks what you're working on. This developer set up a plain text file in 2 minutes. The file is called CLAUDE.md. It loads before every session. Before he types a single word. It already knows his name. It already knows his writing style. It already knows what he's building, who it's for, and what he never wants to see in a response. He doesn't introduce himself anymore. He doesn't explain his preferences anymore. He doesn't correct the same mistakes twice. He just works. No $30/month Copilot subscription. No Microsoft 365. No IT approval. No data sharing agreement. No onboarding. Just a plain text file, a free text editor, and 21 instructions a developer distilled from Andrej Karpathy's research. Those 21 instructions moved Claude's coding accuracy from 65% to 94%. The file hit #1 on GitHub with 82,000 stars. Most people using Claude right now have never heard of it. Microsoft has 221,000 employees, $13 billion invested in OpenAI, and a direct integration into every Windows laptop sold on the planet.. they built an AI assistant most companies pay $30/user/month for that still doesn't know your name. This developer has a laptop, a text file and a 2-minute setup.. he built something that knows more about how he works than any enterprise AI on the market. The $50 billion AI personalization industry just got embarrassed by a .md file. full breakdown down below

Dep

14,179 просмотров • 4 месяцев назад

He's 26. He built a scale model of the Burj Khalifa so detailed the developer flew him to Dubai - on a used resin printer he runs in a Chicago apartment for a fraction of the $80,000 model studios charge The printer is a large-format resin machine he pulled from a shuttered Chicago prototype shop for $1,400. He models every building in Blender from the developer's own CAD files, splits it into hundreds of printable sections, and resin-prints them over week-long cycles - every balcony, every window mullion, every setback on a 6-foot tower, accurate to the millimeter. He wires fiber-optic lighting through the floors so the model glows like the real building at dusk. Total material cost per model: $600 in resin and $90 in LEDs. Model studios quoted the same developer $80,000 and a four-month wait. He delivered in six weeks for $9,000 He posted a time-lapse to Reddit r/architecture in October showing a 6-foot Burj Khalifa replica rising layer by layer, then lighting up floor by floor. The video hit 2.4 million views in nine days. By January he had built scale models for six developers - a Miami condo tower, a Chicago mixed-use block, a Riyadh masterplan - and one architecture firm that now subcontracts every presentation model to his apartment. $180,000 in his account. His father, a retired union electrician, wires the fiber-optic lighting harnesses on weekends Architectural model studios run on the premise that presentation-grade scale models require their workshops, their staff of twelve, and their $80,000 commissions. Autodesk sells the rendering software on the same premise at $2,400 a seat. He builds the same models on a resin printer in a Chicago apartment that could pack a full luxury tower into a padded crate and ship it to a developer's sales gallery across the country by Tuesday

Carat

135,339 просмотров • 1 месяц назад

🦙 ollama is used by 9 million developers and 85% of the Fortune 500, giving co-founder and CEO Jeffrey Morgan (Jeffrey Morgan) a unique view into which AI models people are actually using and how that’s changing. Right now, the biggest shift he sees is toward open models, driven by coding agents, falling costs, and capabilities that are rapidly catching up to the frontier labs. On Ollama Cloud, that shift has driven a 150x increase in token usage since the start of the year. In this episode of Lightcone Podcast, Jeff joins Garry Tan, Jared Friedman, Diana, and Harj Taggar to talk about the future of open models and the story behind Ollama, from two years of searching for the right idea to building one of the most widely used AI developer tools in the world. 00:43 — The Shift to Open Models 03:03 — How AI Agents Are Driving Token Usage 05:31 — Are Open Models Catching Up? 08:26 — What Happens When a New Model Launches 11:31 — Ollama as an Operating System for AI 14:05 — The New Opportunities Above the Model Layer 18:19 — Why 80–90% of Enterprise Tokens Could Be Open 20:57 — The Future Is Local and Cloud 26:40 — Why AI Is Coming Back to Your Computer 28:56 — The Coming Era of Unlimited Tokens 32:30 — Do We Still Need a “God Model”? 33:41 — Open Models and Geopolitics 36:14 — The Origins of Ollama 40:36 — Two Years Lost in the Wilderness 42:39 — The Pivot That Changed Everything 47:02 — How Ollama Found a Business Model 49:43 — Why Second-Time Founders Did YC

Y Combinator

312,329 просмотров • 23 дней назад

This Chinese developer linked two $2,999 NVIDIA DGX Sparks into one box and runs the full Qwen3-235B at home, after dropping his $1,999-a-month cloud bill to zero. He wired 2 small boxes into a single computer, split a giant 235-billion-parameter model in half between them, and serves it across his own network at about 10 tokens a second, with no internet, no cloud, right there on the desk. No data center, no thousand-dollar graphics cards, no monthly cloud bill. Just him, 2 gold boxes the size of a sandwich, one cable between them, and 1 power strip. And here is the whole payoff. He used to pay the cloud $1,999 a month for the same model, and the meter ticked on every request. Now he paid $5,998 once for 2 boxes, they covered their cost in 3 months, and after that he sends as many requests as he wants for free, only electricity. The two Sparks talk over one fast cable, each holds 128GB of memory, and together they carry the whole model, about 73GB loaded per box, with the chip inside pinned near the limit at 96%. Both boxes work as one and keep trading data over the cable, with no cloud in the loop and no single word leaking out. The ready model sits on one local address, and any app on his network calls it as easily as ChatGPT. And here is how he described, in plain words, what this pair of boxes does: "this is a pair of boxes that holds the huge Qwen3-235B model and serves it to one network. the model is split in half, and each box owns its half. parts: // Box 1 (holds the first half of the model and starts the answer fast, the first word appears in under a second) // Box 2 (holds the second half and writes out the rest, about 10 tokens a second) // Cable (connects the 2 boxes and moves data between them on every step, with no lag) // Address (one local address where any app sends its request, like to a cloud model) // Test (a script that runs big prompts through and measures speed and delays) // Monitor (checks temperature, power draw, and load on both boxes every 2 seconds). the model never goes to the cloud. he only steps in when a box runs hotter than 80 degrees or the cable between them starts dropping data." So the system knows exactly what it is, what it is for, and where its limits are. It knows it has to hold the whole huge model across 2 boxes on its own. It knows it has to answer every request locally, with no meter, no limits, and no internet. It knows the human is only needed when a box overheats or the link between them stalls. → The setup runs around the clock on 2 boxes, each pulling under 60 watts → However many requests he sends, the monthly bill is $0, only electricity → The first box starts the answer in under a second → The second writes text at about 10 tokens a second → One request at a time: 838 tokens in 85 seconds, first word in 0.8s → Two requests at once: 697 tokens in 108 seconds, first word in 0.7s → Both boxes sit at 96% load and warm up to 76-78 degrees And only when a chip in a box runs hotter than 80 degrees or the cable between the 2 Sparks drops data does the system call the owner. And when he himself is out on a run or in a coffee shop, he still reaches his own model at home from his phone: sends a big prompt to the local Qwen3-235B, gets the full answer back in under a minute and a half, with no token meter ticking and no limit to hit. Here is what the test shows on his screen during one of the night runs: "one request at a time: 838 tokens in 84.9 seconds, first word in 0.8s, then 0.1s per token." "two requests at once: 697 tokens in 107.6 seconds, first word in 0.7s, then 0.15s per token." "Box 1: chip at 96% load, 76 degrees, 56 watts, 73GB used in memory." "Box 2: chip at 96% load, 78 degrees, 56 watts, the Qwen3-235B model fully loaded." And while everyone around is paying for AI by the month and bumping into limits, his top-tier model just sits on the desk and works as much as he wants: his own little power plant instead of a forever meter. He has no server rack of his own and no cloud account behind it. Just 2 DGX Spark boxes on a desk, one model split in half between them, one local address, and a folder of prompts next to it. Out of everything I have seen this year, this is the cleanest way to stop paying for AI: $5,998 of hardware on the desk once, $0 a month to the cloud, unlimited forever, and between them 2 gold boxes, 1 cable, and the full Qwen3-235B answering at home with no internet.

Blaze

93,871 просмотров • 4 месяцев назад