Загрузка видео...

Не удалось загрузить видео

На главную

6 dead office PCs at $50 each run a Kubernetes cluster that AWS would bill $10,000 a year for. Most home labs die before they boot. Someone specs a Xeon. Prices 128GB of RAM. Opens 40 tabs. Buys nothing. The build that actually runs is uglier than that. HP...

28,629 просмотров • 2 дней назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

FOUR DIFFERENT VENDORS ARE NOW SHIPPING GB10 MINI PCs WITH 128GB UNIFIED MEMORY, AND ONE MICROTIK CRS 804 SWITCH CAN CONNECT UP TO EIGHT OF THEM INTO A 1 TERABYTE LOCAL AI CLUSTER 00:00 he points at the MikroTik CRS 804, "you need some kind of switch that'll handle QSFP56 ports like these", the interconnect that makes the whole cluster possible the GB10 ecosystem is no longer just Nvidia. Dell Pro Max GB10, ASUS Ascent GX10, and MSI Edge Expert all ship the same Grace Blackwell Superchip with 128GB of coherent memory. same silicon, different cases, same 200 gigabit ports on the back the CRS 804 is what connects them at prosumer prices. four 400 gigabit QSFP56 ports on one 1U chassis, breakout cables that split each port into two 200 gigabit lanes. one switch drives eight GB10 units in parallel do the math. eight nodes at 128GB each equals 1024GB of pooled unified memory across the cluster. run vLLM, shard a frontier model across all eight, and inference happens locally on hardware that fits in half a rack the real limiter revealed in the stress test was never throttling. it was interconnect topology, exactly the layer this switch fixes at a fraction of enterprise switch pricing $400 a month for combined chatgpt pro and claude code max hits $4,800 a year per developer. a small team of five running through this cluster pays back inside eight months and never expires the article covers the buying ladder for a single desk. this post is proof of the cluster ladder that starts where the desk one ends save this before the GB10 lineup grows past four vendors and prosumer cluster switches move upmarket

NO1ennn

59,725 просмотров • 1 месяц назад

One wallet on Polymarket is literally named "PBot-6" Every single position: 5-minute crypto Up-or-Down. It enters in the last 90 seconds before resolution. Every time. That's not a trader. That's a machine reading Chainlink's lag. I found 6 more just like it. Columbia published the proof last year. 25% of Polymarket volume is wash trading. 14% of wallets are coordinated. Everyone retweeted it. Nobody did anything with it. I did. 86 million trades. One Claude prompt. "build a graph where nodes = wallets, edges = pairs that traded together 10+ times" 412,000 wallet pairs. 23 cluster candidates. 4 filters later: 7 survivors. What I found: Cluster 1. Buy 1-8¢ entries on thin Saturday books while BTC moves on Binance. +3,400% PnL. That's not prediction. That's coordination. Cluster 2. PBot-6 and friends. Enter 90 seconds before resolution. Aligned with Chainlink lag. 71% of entries within 45 seconds of each other. Cluster 3. Buy 0.1¢ entries on "Seoul Mayoral 2026" months out. Entry 0.1¢ → current 50¢. +47,000% PnL. They're not predicting winners. They're farming optionality at sub-penny prices. Cluster 6. $35,000 positions days before token launches. Markets resolve their way every time. This one reads like insider allocation lists. Three rules to fade them: -> Cluster pumps a market. Wait 8 min. Fade the 50% retracement. -> 2+ cluster wallets exit what you're holding. Dump in 30 seconds. No debate. -> Cluster 4 is >40% of last-hour volume. The entire order book is fiction. Skip. 31 days. $300 in. 142 trades. 71% win rate. $16,300 out. Sharpe 2.6. Copytrade here: Stack cost: $25/month. Claude API + $5 VPS. Nothing paywalled. Repos: The clusters are still running right now. PBot-6 opened 3 positions this weekend. The graph doesn't lie. 99.9% will say this is cap. The 0.01% will run the pipeline.

Trackmind

20,082 просмотров • 3 месяцев назад

One guy built this app in a month and now it makes him more than $1,000,000 a year. No team. No investors. No marketing department. One developer. One month. One clever idea. The app is a camera for events. In its first month it got 100,000 downloads and he did not spend a single dollar on ads. Here is how he did it because the most interesting part is the growth itself. He did not bolt marketing onto the product. He made using the product the marketing itself. After that things kick in that almost nobody figures out. 1. You cannot use the app alone. For it to work the host has to pull every guest into it. Each install drags in dozens more right away. 2. One wedding is not one user but a whole crowd at once. 200 people scan one code in an evening and install the app. No ad brings that many for the same money and here the money is zero. 3. The guest becomes the host. He liked it at someone else's wedding and a month later he throws his own event and brings his own people. The loop spins itself and for free. 4. It does not look like an ad. To the guest it is a gift not some app forced on him. So they install it gladly and all of them do. 5. It all runs on emotion. A wedding. Memories. Shared shots. People film it and show their own people and a new wave comes in. Now let us count the money plain and honest. The subscription runs from 2 to 50 dollars. Say only every 20th person pays. That is 5,000 people out of 100,000. The average check a modest 20 dollars. 5,000 times 20 is 100,000 dollars a month. More than 3,000 a day. More than $1,000,000 a year. And all of this is one guy in a month without a single dollar on ads. He did not win on budget and not on a team. He won by sewing distribution into the very use of the product. You can lift almost any product this way. Could you build something like this on your own or is it just luck?

Blaze

10,703 просмотров • 1 месяц назад

Persi Diaconis walked into a lecture at the University of Washington, held up a coin, and told a room of physicists that Richard Feynman was fooled by it his entire life. He was right. In 2007, Diaconis proved a coin flip is not 50/50. It lands on the side it started on about 51 percent of the time. The bias comes from the physics of rotation under gravity. Every physicist since Newton had assumed 50/50 without ever testing it. Every trading model built on that assumption is running on the same lie. The lecture was on Feynman's book "The Meaning of it All." Diaconis quoted the most famous line in it: "the first principle is that you must not fool yourself, and you are the easiest person to fool." Then he pointed out that Feynman himself was fooled by every coin he ever flipped. Feynman's own rule would have killed the 50/50 assumption on day one. The market is the same setup at scale. Every model that assumes independent 50/50 outcomes at the base layer is built on a physical impossibility. Order flow, positioning, forced flows, expiries all leave biases larger than 1 percent. Your gut cannot see them. The math already knows they are there. Diaconis's rule: before you trust a random process, check it. Actually check it. Not with a simulation. With a proof or an experiment. The coin is where you start. The chart is where the same rule pays. The only random thing about markets is how thoroughly you refuse to check them.

veles

18,798 просмотров • 1 месяц назад

Today, December 6, marks the anniversary of the demolition of the Babri Masjid. As a child, whenever the issue came up, I used to wonder why we needed either a temple or a mosque at that spot at all—one large school and one large hospital would have been far more useful for father, a typical Leftist, always dismissed it as political drama manufactured by the Congress party. My mother, firmly on the Right, insisted there was nothing wrong in reclaiming the site for a temple and that a mosque could easily be built elsewhere, outside the sacred precinct or even outside the city.But once I began reading the actual history—how a temple marking the birthplace of Lord Ram had been destroyed in the very city of his birth, and a new structure raised over it—something shifted inside me. I tried to imagine the same thing happening to the holiest sites of another faith: the shrine of Imam Ali razed, Karbala desecrated, or a new building erected over the ruins of Mecca. The anger and grief would be overwhelming. For Hindus, the pain of what happened in Ayodhya runs even deeper.Hindus have every right to a grand, magnificent Shri Ram Mandir at that janmabhoomi—just as the faithful of every religion cherish and protect their most sacred places, whether Karbala, Mecca, Jerusalem, Amritsar, or on this December 6, I extend my heartfelt congratulations to every Hindu for the restoration of the Bhavya Ram Mandir, for reclaiming and reviving what was unjustly taken centuries ago. Jai Siya Ram. Maula Ali Madad

Farzana Ali Mazari 🇮🇳

74,247 просмотров • 8 месяцев назад

UC Berkeley just open-sourced FreeToken. (2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU. let me explain how: all three models mentioned above are Mixture-of-Experts, and that is what FreeToken takes advantage of. each layer holds hundreds of separate experts plus a small router that picks a few of them per token. Qwen3.6-35B activates roughly 3B of its 35B parameters per token. DeepSeek-V4-Flash picks 6 of 256 experts per layer, so 13B of its 284B run at a time. so compute was never the bottleneck. the weights a single step touches fit comfortably on a consumer GPU. every expert the router might pick still has to exist somewhere. they sit in system RAM, and the GPU keeps a cache of the ones the model has been using recently. so everything comes down to what happens when the router picks an expert that is not on the GPU. there are two ways to serve that miss: 1. copy it over PCIe and run it on the GPU 2. run it on the CPU, where it already lives both read from the same system memory, so they compete for one pool of bandwidth instead of adding to each other. existing engines pick one option and freeze it when the model loads. but routing changes on every token, so a fixed choice misses most of what the model asks for. FreeToken measures both bandwidths on your machine and splits each step's misses between the two paths in proportion. the GPU and CPU results then merge exactly, with no approximation. two machines with the same GPU can end up wanting opposite strategies, which I did not expect. a 5090 in a gaming desktop should push nearly everything over PCIe, while an 8GB laptop is better off computing most misses on the CPU. none of that is readable off a spec sheet, so the engine profiles it once per machine. the second half of the design is about agents. coding agents constantly rewrite their own history, and every edit normally forces thousands of tokens back through prefill. FreeToken saves its checkpoints at the exact boundaries agent frameworks cut on, so it only reprocesses the new part. its slowest first token stays under 44 seconds, while llama.cpp peaks at 232 and KTransformers at 946. it serves the OpenAI and Anthropic APIs under Apache 2.0, so Claude Code and Codex can point at it directly. releasing weights publicly decides who can download a model, not who can afford to run one. frontier open models keep shipping, and running them still assumes a rented cluster. meanwhile there are over a hundred million consumer machines with discrete GPUs sitting mostly idle. closing that gap was never a hardware problem, and work like this is what turns open weights into something you can actually use. paper: repo: almost every idea in this post, from why memory bandwidth decides the outcome to why moving weights costs more than computing on them, comes straight out of how a GPU is built. I wrote a detailed primer on that. the article is quoted below.

Akshay 🚀

318,151 просмотров • 3 дней назад

One guy keeps a farm of Mac minis on his desk and says each $600 box brings him $2,000 a month while he sleeps. AND THE HARDWARE ACTUALLY WORKS. But the number is not even the interesting part. The broken part is HOW: his AI no longer sits in a chat window. It sees the screen, moves the mouse itself, types and clicks the interface like a human at a computer. That is it. While most people still run AI in a chat and ask it for text, he sat Claude down right at the computer and put it to work with its hands. He automated not a single task but the workplace itself. How it actually works: on every Mac mini Claude runs with computer use turned on and the official Claude API docs spell it out: screenshot capture, mouse control, keyboard input, desktop automation. The agent opens the browser and the apps itself and runs the boring routine on a schedule: pulls leads, fills the CRM, checks orders, runs QA on the site. One box, one quiet worker that does not sleep and does not ask for a salary. His math is simple: a Mac mini is $600 once, Claude Max is $200 a month, and a live white-collar worker on the same routine costs a business $4,000 and up. So he rents out each node to a client as an AI worker for about $2,000 and 6 Mac minis come out to around $12,000 a month with costs a bit over $1,000 on subscriptions. But the $12,000 is his projection not a revenue dashboard: the video has no client, no task log, no working automation at all. The real asset here is not the stack of hardware but the one repeatable process the agent actually closes. Because a Mac mini on its own earns nothing. The money shows up exactly where the boring browser routine used to be done by hand for a salary and now you can hand it to an agent for the price of a subscription. Computer use is still in beta, almost nobody builds a service on it and the demand for cheap GUI routine is huge. The window is open for literally the next few months. Most people will watch this, laugh at the "$600 AI worker" and close it. And the ones who actually put an agent on one boring task and grind it into a repeat will ride this wave while it is still empty. Would you sit an AI right at your own computer on the boring routine or are you still clicking through it by hand?

Sorven

12,259 просмотров • 2 месяцев назад

Aravind Srinivas just described a future most founders are pretending they are ready for. One person. One machine. A company that runs itself. Srinivas: “Buy a Mac mini, set up a Perplexity personal computer, and run their business on that.” Not a side project. Not a pitch deck. A real business with real revenue while the founder is not in the building. AI runs the ads. Handles SEO. Integrates Stripe. Ships features. Answers customers. All of it executing without a single employee. Srinivas: “Have this all working while you can be sipping wine in Napa.” But before he sold the dream he killed the one most people are already chasing. Srinivas: “Everybody talks about this one-person one-billion-dollar company. It’s not truly moving the GDP by one billion. It’s not truly creating new value.” One researcher collecting a billion in equity does not grow an economy. It rearranges numbers between balance sheets. Nothing gets built. No customer gets served. That is not value creation. That is valuation creation. Srinivas wants no part of it. What he described is the opposite. The person driving Uber between shifts who has the idea but not the payroll. Not the engineering. Not the marketing. Not the support staff. That person gets a machine that replaces all of it. Hundreds of thousands in revenue. Millions. Generated by autonomous systems doing the work that used to require ten employees and a burn rate. Not paper wealth. Not valuation theater. Output that moves through an economy and touches real customers. That is what moves GDP. Not one person worth a billion dollars. A million people each building something worth a million. That math rewrites a country. Then Srinivas said the part that separates him from every hype merchant in the room. Srinivas: “Everybody thinks AI is already there. It’s not there yet. Someone has to do that hard work.” The vision is real. The infrastructure is not. The agents are not autonomous. The integrations are not seamless. The plumbing is not finished. Someone has to wire the APIs. Connect the billing. Build the bridge between what a founder wants and what a machine can deliver. That work is not a keynote. It is not a tweet thread. It is engineering that nobody wants to do and everybody will depend on. Whoever finishes it first does not just build a product. They hand every ambitious person on Earth a company they can run alone. The corporations that need five hundred people to do what one founder with the right infrastructure could do are not efficient. They are exposed. And the person building the thing that exposes them just told you exactly what it looks like. He also told you it is not going to build itself.

Dustin

64,593 просмотров • 5 месяцев назад

your agent loop needs 8 exits. most people ship only one. (explained with triggers) 1) goal met → an evaluator scores the output against a rubric, and the run stops on a pass. → fires when the work is measurably done, not when the model says it is done. 2) turn cap → a hard ceiling on iterations, counted and enforced by the harness, not the prompt. → fires on the task it was never going to finish, before you pay to find that out. 3) budget cap → a limit on tokens or dollars, whichever one runs out first. → fires mid-run, which is exactly why it is the exit that saves you the 3am bill. 4) wall clock → a deadline on elapsed time, independent of how much progress was made. → fires when the run collides with a deploy window or the start of business hours. 5) no progress → hash the state every turn and compare it against the last few. → fires when three turns in a row change nothing. busy is not the same as moving. 6) human interrupt → an approval gate before risky steps, plus a kill switch that lives outside the loop. → fires whenever you decide, and it is the one exit the model cannot argue with. 7) error threshold → a counter of consecutive failures that resets on any success. → fires at n in a row, so it halts instead of retrying into the same wall all night. 8) external event → a webhook or a poll on whatever the task was actually about. → fires when the PR merged or the ticket closed and the work stopped mattering. a loop with one exit hangs. a loop with eight is a system. write the exits before you write the prompt.

Hanako

181,713 просмотров • 1 месяц назад