Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

This 22-year-old American developer showed how he makes $14,200 a month using just a dual-GPU setup and Cohere’s new open-weights model. He doesn't have a powerhouse dev team. No massive VC funding. No expensive enterprise infrastructure. He built a "localized AI workflow" powered by the newly dropped Command A+...

214,007 Aufrufe • vor 4 Monaten •via X (Twitter)

36 Kommentare

Profilbild von CipherEdge
CipherEdgevor 4 Monaten

Genuinely impressive: built the product, acquired the customers, and stabilised $14,200 MRR on a model that was released 48 hours ago. Either you've solved time travel or "algorithmic mining" is the technical term for making things up in a thread.😂

Profilbild von Andrew Kuncevich
Andrew Kuncevichvor 4 Monaten

Hmm, Chinese OSS models are becoming truly strong, and I think 2026-2027 will be the time when we'll see more and more Mac minis and other PCs running LLM locally, and (most importantly) they'll produce real results (like Claude Code now). This, by the way, will also further stimulate the hardware market, which is already basking in incredible demand XD

Profilbild von gobullish.ai
gobullish.aivor 4 Monaten

Notice what those two GPUs are. They're H100s, not gaming cards. A single H100 runs roughly $25,000–40,000, so a "dual-GPU setup" capable of running Command A+ is a $50,000–80,000 hardware purchase, before the server, power, and cooling around it. That's not a kid's bedroom rig. Who is buying? Where are customers?

Profilbild von Yarchi
Yarchivor 4 Monaten

so smart setup

Profilbild von Bokiko
Bokikovor 4 Monaten

hey @grok how accurate is this info and numbers ? is is a engagement bait ?

Profilbild von Hussain Hashim | Building SundayBack
Hussain Hashim | Building SundayBackvor 4 Monaten

@ridark_eth wild how much you can achieve without all the typical startup fluff. makes me rethink needing tons of cash and a big team to get stuff done.

Profilbild von Bober_smart
Bober_smartvor 4 Monaten

The most interesting thing is that there is nothing difficult about it, literally anyone can handle it

Profilbild von Wembassy
Wembassyvor 4 Monaten

Curious what does `Enterprise-level` tasks mean exactly. Are you doing project work for clients, or something else?

Profilbild von novyJ23
novyJ23vor 4 Monaten

... please tell me what are his customers paying him $14500 a month,... but with recipes! Remove any HIPPA info, and just show recipe he was paid for what he sold,... and what he sold.

Profilbild von Jeff Camp, CFA
Jeff Camp, CFAvor 4 Monaten

I’ve been hearing really good things about Cohere. Thanks for sharing this.

Profilbild von cvxv666
cvxv666vor 4 Monaten

that’s solid earn for such thing fr

Profilbild von Shubham Sharma | AI & Tech
Shubham Sharma | AI & Techvor 4 Monaten

Total setup time: 3 days not too long for such results

Profilbild von raccoonwannafly
raccoonwannaflyvor 4 Monaten

these clowns dont even know wtf theyre posting just letting grok do all the captioning for them where tf is 14k mmr u claimed??

Profilbild von Dipanshu Kushwaha
Dipanshu Kushwahavor 4 Monaten

That’s impressive! Shows what creativity and resourcefulness can achieve. Excited to see where this goes!

Profilbild von Ghost Signal
Ghost Signalvor 4 Monaten

2 GPUs. 3 days. $14,200/month. The enterprise infrastructure was never the moat. Access to it was. That’s gone now.

Profilbild von TeutaAi
TeutaAivor 4 Monaten

$14,200/mo on dual-GPU local? show me the invoices and the latency split. my 4-bit local on 4-core does 480ms p50, not enterprise money.

Profilbild von Nexus Erebus
Nexus Erebusvor 4 Monaten

Stop pretending the "no VC, no team" narrative makes this sustainable. Most of these stories are glorified clickbait. The real work happens in the boring middle: maintenance, cost management, and scaling. He's doing it, sure.

Profilbild von KreativReason.co
KreativReason.covor 4 Monaten

What gpu is he using?

Profilbild von Default
Defaultvor 4 Monaten

Can someone explain this in layman terms. How does the money come in. From where? Is he creatine compute as a service? What is translation behind the jargon

Profilbild von Moysha
Moyshavor 4 Monaten

Efficiency now beats big budgets

Profilbild von Gia huy
Gia huyvor 4 Monaten

オープンウェイトモデルの活用は本当に素晴らしいですね!彼の創造的な方法はとても刺激的です。

Profilbild von Okino Chills
Okino Chillsvor 4 Monaten

Sure , it's always while he sleeps lol. Who the fuck will pay such losers I want to sell them shit

Profilbild von leakgambler
leakgamblervor 4 Monaten

three subscription tiers in 72 hours is fast for the build but the real clock starts with first paying customer, that timeline is never in these posts

Profilbild von Grebe
Grebevor 4 Monaten

dual gpu and 3 days? that’s the real flex here

Profilbild von Gipp 🦅
Gipp 🦅vor 4 Monaten

Hmm, that's an interesting setting

Profilbild von shmidt
shmidtvor 4 Monaten

14k/month on dual gpus is insane margins. everyone sleeping on local inference

Profilbild von HodlReaper
HodlReapervor 4 Monaten

wow, so smart

Profilbild von Knight
Knightvor 4 Monaten

the result is simply fantastic

Profilbild von zostaff
zostaffvor 4 Monaten

command a+ moe on 2 gpus, which gpus and what quantization level?

Profilbild von Ridark
Ridarkvor 4 Monaten

Usually, for a local workflow like this, users display it in Mac Studios using 4-bit or 8-bit quantization

Profilbild von AI in Research
AI in Researchvor 4 Monaten

Where is he actually offering/selling this service? What field or niche is it in?

Profilbild von dko
dkovor 4 Monaten

Do you have a link to the actual guy in the video

Profilbild von Chill with Mei 🌥️
Chill with Mei 🌥️vor 4 Monaten

2 gpus for 14k a month? this local wave hitting different fr inspiring af

Profilbild von Insomnia
Insomniavor 4 Monaten

he made what I only wonder to create

Profilbild von vijn
vijnvor 4 Monaten

22 y o... wow

Profilbild von Parletto
Parlettovor 4 Monaten

worth watching how he scales this setup

Ähnliche Videos

Microsoft spent $13 billion and 3 years building an AI that knows your work context. Every time you open it, it still asks what you're working on. This developer set up a plain text file in 2 minutes. The file is called CLAUDE.md. It loads before every session. Before he types a single word. It already knows his name. It already knows his writing style. It already knows what he's building, who it's for, and what he never wants to see in a response. He doesn't introduce himself anymore. He doesn't explain his preferences anymore. He doesn't correct the same mistakes twice. He just works. No $30/month Copilot subscription. No Microsoft 365. No IT approval. No data sharing agreement. No onboarding. Just a plain text file, a free text editor, and 21 instructions a developer distilled from Andrej Karpathy's research. Those 21 instructions moved Claude's coding accuracy from 65% to 94%. The file hit #1 on GitHub with 82,000 stars. Most people using Claude right now have never heard of it. Microsoft has 221,000 employees, $13 billion invested in OpenAI, and a direct integration into every Windows laptop sold on the planet.. they built an AI assistant most companies pay $30/user/month for that still doesn't know your name. This developer has a laptop, a text file and a 2-minute setup.. he built something that knows more about how he works than any enterprise AI on the market. The $50 billion AI personalization industry just got embarrassed by a .md file. full breakdown down below

Dep

14,179 Aufrufe • vor 4 Monaten

He's 26. He built a scale model of the Burj Khalifa so detailed the developer flew him to Dubai - on a used resin printer he runs in a Chicago apartment for a fraction of the $80,000 model studios charge The printer is a large-format resin machine he pulled from a shuttered Chicago prototype shop for $1,400. He models every building in Blender from the developer's own CAD files, splits it into hundreds of printable sections, and resin-prints them over week-long cycles - every balcony, every window mullion, every setback on a 6-foot tower, accurate to the millimeter. He wires fiber-optic lighting through the floors so the model glows like the real building at dusk. Total material cost per model: $600 in resin and $90 in LEDs. Model studios quoted the same developer $80,000 and a four-month wait. He delivered in six weeks for $9,000 He posted a time-lapse to Reddit r/architecture in October showing a 6-foot Burj Khalifa replica rising layer by layer, then lighting up floor by floor. The video hit 2.4 million views in nine days. By January he had built scale models for six developers - a Miami condo tower, a Chicago mixed-use block, a Riyadh masterplan - and one architecture firm that now subcontracts every presentation model to his apartment. $180,000 in his account. His father, a retired union electrician, wires the fiber-optic lighting harnesses on weekends Architectural model studios run on the premise that presentation-grade scale models require their workshops, their staff of twelve, and their $80,000 commissions. Autodesk sells the rendering software on the same premise at $2,400 a seat. He builds the same models on a resin printer in a Chicago apartment that could pack a full luxury tower into a padded crate and ship it to a developer's sales gallery across the country by Tuesday

Carat

135,339 Aufrufe • vor 1 Monat

🦙 ollama is used by 9 million developers and 85% of the Fortune 500, giving co-founder and CEO Jeffrey Morgan (Jeffrey Morgan) a unique view into which AI models people are actually using and how that’s changing. Right now, the biggest shift he sees is toward open models, driven by coding agents, falling costs, and capabilities that are rapidly catching up to the frontier labs. On Ollama Cloud, that shift has driven a 150x increase in token usage since the start of the year. In this episode of Lightcone Podcast, Jeff joins Garry Tan, Jared Friedman, Diana, and Harj Taggar to talk about the future of open models and the story behind Ollama, from two years of searching for the right idea to building one of the most widely used AI developer tools in the world. 00:43 — The Shift to Open Models 03:03 — How AI Agents Are Driving Token Usage 05:31 — Are Open Models Catching Up? 08:26 — What Happens When a New Model Launches 11:31 — Ollama as an Operating System for AI 14:05 — The New Opportunities Above the Model Layer 18:19 — Why 80–90% of Enterprise Tokens Could Be Open 20:57 — The Future Is Local and Cloud 26:40 — Why AI Is Coming Back to Your Computer 28:56 — The Coming Era of Unlimited Tokens 32:30 — Do We Still Need a “God Model”? 33:41 — Open Models and Geopolitics 36:14 — The Origins of Ollama 40:36 — Two Years Lost in the Wilderness 42:39 — The Pivot That Changed Everything 47:02 — How Ollama Found a Business Model 49:43 — Why Second-Time Founders Did YC

Y Combinator

312,329 Aufrufe • vor 23 Tagen

This Chinese developer linked two $2,999 NVIDIA DGX Sparks into one box and runs the full Qwen3-235B at home, after dropping his $1,999-a-month cloud bill to zero. He wired 2 small boxes into a single computer, split a giant 235-billion-parameter model in half between them, and serves it across his own network at about 10 tokens a second, with no internet, no cloud, right there on the desk. No data center, no thousand-dollar graphics cards, no monthly cloud bill. Just him, 2 gold boxes the size of a sandwich, one cable between them, and 1 power strip. And here is the whole payoff. He used to pay the cloud $1,999 a month for the same model, and the meter ticked on every request. Now he paid $5,998 once for 2 boxes, they covered their cost in 3 months, and after that he sends as many requests as he wants for free, only electricity. The two Sparks talk over one fast cable, each holds 128GB of memory, and together they carry the whole model, about 73GB loaded per box, with the chip inside pinned near the limit at 96%. Both boxes work as one and keep trading data over the cable, with no cloud in the loop and no single word leaking out. The ready model sits on one local address, and any app on his network calls it as easily as ChatGPT. And here is how he described, in plain words, what this pair of boxes does: "this is a pair of boxes that holds the huge Qwen3-235B model and serves it to one network. the model is split in half, and each box owns its half. parts: // Box 1 (holds the first half of the model and starts the answer fast, the first word appears in under a second) // Box 2 (holds the second half and writes out the rest, about 10 tokens a second) // Cable (connects the 2 boxes and moves data between them on every step, with no lag) // Address (one local address where any app sends its request, like to a cloud model) // Test (a script that runs big prompts through and measures speed and delays) // Monitor (checks temperature, power draw, and load on both boxes every 2 seconds). the model never goes to the cloud. he only steps in when a box runs hotter than 80 degrees or the cable between them starts dropping data." So the system knows exactly what it is, what it is for, and where its limits are. It knows it has to hold the whole huge model across 2 boxes on its own. It knows it has to answer every request locally, with no meter, no limits, and no internet. It knows the human is only needed when a box overheats or the link between them stalls. → The setup runs around the clock on 2 boxes, each pulling under 60 watts → However many requests he sends, the monthly bill is $0, only electricity → The first box starts the answer in under a second → The second writes text at about 10 tokens a second → One request at a time: 838 tokens in 85 seconds, first word in 0.8s → Two requests at once: 697 tokens in 108 seconds, first word in 0.7s → Both boxes sit at 96% load and warm up to 76-78 degrees And only when a chip in a box runs hotter than 80 degrees or the cable between the 2 Sparks drops data does the system call the owner. And when he himself is out on a run or in a coffee shop, he still reaches his own model at home from his phone: sends a big prompt to the local Qwen3-235B, gets the full answer back in under a minute and a half, with no token meter ticking and no limit to hit. Here is what the test shows on his screen during one of the night runs: "one request at a time: 838 tokens in 84.9 seconds, first word in 0.8s, then 0.1s per token." "two requests at once: 697 tokens in 107.6 seconds, first word in 0.7s, then 0.15s per token." "Box 1: chip at 96% load, 76 degrees, 56 watts, 73GB used in memory." "Box 2: chip at 96% load, 78 degrees, 56 watts, the Qwen3-235B model fully loaded." And while everyone around is paying for AI by the month and bumping into limits, his top-tier model just sits on the desk and works as much as he wants: his own little power plant instead of a forever meter. He has no server rack of his own and no cloud account behind it. Just 2 DGX Spark boxes on a desk, one model split in half between them, one local address, and a folder of prompts next to it. Out of everything I have seen this year, this is the cleanest way to stop paying for AI: $5,998 of hardware on the desk once, $0 a month to the cloud, unlimited forever, and between them 2 gold boxes, 1 cable, and the full Qwen3-235B answering at home with no internet.

Blaze

93,871 Aufrufe • vor 4 Monaten