正在加载视频...

视频加载失败

China is making Dario Amodei's AI slowdown proposal worthless. The thing doing it is a 510 GB file that anyone can download for free. Two days before that essay went out, DeepSeek shipped a model called V4.1 Flash. The weights went straight onto Hugging Face under an MIT license,...

400,709 次观看 • 6 天前 •via X (Twitter)

12 条评论

The AI Therapist 的头像
The AI Therapist6 天前

DeepSeek dropped a 510GB file. The inference cost just crashed through the floor. Everyone is watching training. Nobody’s looking at their bill to run it tomorrow. open weights win when they ship for free and everyone else bills per token.

ClassicMonk 的头像
ClassicMonk6 天前

If you slowdown then you are out of game. Even just staying in game alone is not enough , you have to be first, no prizes for coming 2nd or last.

Marek Kapelinski 的头像
Marek Kapelinski6 天前

So true.

Milena Harito 的头像
Milena Harito6 天前

DeepSeek model V4.1 Flash in open acess on on Hugging Face .. are we in the road toward the commoditization of the models ?

First Sauce Labs 的头像
First Sauce Labs6 天前

Yeah, the "slowdown" idea didn't really account for the actual pace of open source releases. That ship was always going to sail.

Brian 的头像
Brian6 天前

So morons think anyone should be able to download tech that can hack any system and teach you to make bioweapons that could kill billions? Unbelievable!

Tyty 的头像
Tyty6 天前

I keep thinking about this. If Jacob Coxon never spoke publicly about the dangers of AI… would Sam, Elon, Dario and everyone else still be talking about slowing AI down? Did they already know the risks? Or did something actually change? I don’t know. But it feels like a question worth asking.

Steven Online 的头像
Steven Online6 天前

When the biggest AI companies agree to slow everyone down, ask who benefits. Show us the evidence, accept independent scrutiny, and protect open models and affordable access. “Safety” must not become a convenient excuse to pull up the ladder behind you.

kaboski 的头像
kaboski6 天前

Quando olho para o debate global sobre o futuro da Inteligência Artificial (IA), vejo que estamos em uma encruzilhada retórica perigosa. De um lado, manifestos e depoimentos de pesquisadores de segurança de grandes laboratórios, como OpenAI e Anthropic, pintam um cenário apocalíptico, alertando sobre "perda de controle irreversível" e riscos de "extinção humana". Do outro lado — onde eu me posiciono —, ergue-se uma defesa robusta baseada em fatos, benefícios palpáveis e na urgência de resolver as crises reais do presente. Para mim, uma análise racional e detalhada revela que esse alarmismo existencial carece de fundamentos técnicos sólidos. Acredito que desacelerar a IA seria o verdadeiro erro histórico contra o progresso humano. meu primeiro passo é expor as falhas estruturais e os vieses que noto nesses discursos catastróficos. Cientistas da computação renomados que sigo e estudo, como Yann LeCun (Cientista-Chefe de IA da Meta) e Andrew Ng (fundador do Google Brain), apontam exatamente as mesmas lacunas que eu percebo. Muitos desses profissionais que alertam sobre o "fim do mundo" pertencem a uma área muito específica chamada "Alinhamento de IA". Eu vejo aí um claro conflito de interesse intelectual: o financiamento, a relevância de mercado e a atenção midiática desse nicho dependem diretamente de eles convencerem o mundo de que o problema que tentam resolver é uma ameaça iminente. Não concordo com a ideia de que uma inteligência superior desenvolverá, automaticamente, um desejo de dominação ou tirania. Como bem define LeCun, o desejo de poder é uma característica biológica de primatas, fruto da evolução e da competição por recursos — não é uma propriedade matemática do software. Para mim, uma IA superinteligente será uma ferramenta de otimização massiva sob rígidas diretrizes humanas, e não uma nova espécie biológica rival. Acho um enorme salto ficcional afirmar que "a maioria das pessoas poderia morrer" com base nos modelos de linguagem atuais. Eu concordo plenamente com a ironia de Andrew Ng: focar no risco de extinção por IA hoje é o equivalente a se preocupar com a "superpopulação nas colônias de Marte" antes mesmo de termos foguetes capazes de pousar lá Eu enxergo na IA o maior motor de elevação social já criado pela nossa espécie. Penso como Marc Andreessen, pioneiro da internet, que argumenta que tudo o que a civilização construiu de melhor — da eletricidade à penicilina — foi fruto da aplicação da inteligência humana. Logo, se pudermos expandir e democratizar essa inteligência por meio da IA, poderemos maximizar o bem-estar social em escala global. Também me inspiro na visão humanitária de figuras como Bill Gates, que destaca o papel da tecnologia no combate à desigualdade. Eu acredito no potencial da IA para universalizar o acesso a uma educação de altíssima qualidade por meio de tutores digitais personalizados e para levar diagnósticos médicos de ponta para comunidades isoladas no sul global, onde médicos humanos infelizmente não conseguem chegar Longe de causar o desemprego em massa temido pelos pessimistas, vejo a IA atuando como um co-piloto. Ao absorver tarefas cognitivas repetitivas e burocráticas, ela nos liberta para exercer funções estratégicas, criativas e que demandam empatia. Historicamente, ferramentas que multiplicam a produtividade expandem a economia e geram novos mercados. Conclusão Para encerrar o meu argumento, a Inteligência Artificial não é uma força alienígena e cega que avança à revelia da sociedade; ela é desenvolvida, moldada e monitorada passo a passo por nós, sob rigorosos protocolos de engenharia e regulamentações. O verdadeiro perigo que vejo para o nosso futuro não reside no desenvolvimento de máquinas inteligentes, mas sim na nossa própria covardia intelectual em rejeitar a inovação. Diante de um mundo repleto de doenças, crises climáticas e desafios estruturais, eu defendo que a IA não é um luxo perigoso, mas a ferramenta de sobrevivência mais poderosa que já construímos.

NIMA8891 的头像
NIMA88916 天前

on the one hand this guys pitch "nobel" intentions to slow down and pause AI for sake of humanity (of course not !) and on the other hand introduce subscription model which has undefined budget and defined price, then try to cut that undefined budget in half to trick customers to use credits to make more money and "blue on top of gold" stop open weight models so there is no alternative !

Rafael Ourique 的头像
Rafael Ourique6 天前

@tiberiocaetano are we happily doomed? You've been hammering this key for years and nobody seemed to listen. Will everybody wake up to this when it's too late?

Shivani Dubey 的头像
Shivani Dubey6 天前

Today's outcome gave me a very real burst of confidence. I'm enjoying being pleasantly surprised by myself.

相关视频

qwen 3.8 max vs deepseek v4 flash 0731 vs kimi k3 vs gpt 5.6 sol – on rubik's cube and chess four frontier models built a rubik's cube stand and solved it, then built a chess board and played claude opus 5 on it the setup: Nous Research's hermes agent cli on OpenRouter tasks: 1. cube – build a 3d rubik's cube with a cli and a Three.js viewer, then solve an identical scrambled position on your own stand 2. chess – build a 3d chess stand, then play white against claude opus 5 as black, live, one move at a time. no engine, no solver, no opening book on either side. stockfish depth 14 grades every chess ply afterwards; neither player sees the score models: DeepSeek v4 flash 0731, OpenAI gpt-5.6 sol, Kimi.ai kimi k3, Qwen qwen 3.8 max gpt-5.6 sol and deepseek v4 flash solved their cubes – sol in 24 moves and seventeen seconds, deepseek in 32. qwen and kimi never got there, giving up at 96 and 207 moves then all four built chess stands and played white against claude opus 5 on them, and all four resigned: deepseek on move 13, sol on 19, kimi on 21, qwen holding out longest at 29 - build time, both stands #1 gpt-5.6 sol – 16m 43s #2 deepseek v4 flash – 97m 39s #3 kimi k3 – 166m 09s #4 qwen 3.8 max – 215m 08s - build attempts before a working stand #1 gpt-5.6 sol – 3 #2 qwen 3.8 max – 4 #3 kimi k3 – 4 #4 deepseek v4 flash – 5 - total tokens #1 gpt-5.6 sol – 6,713,754 #2 qwen 3.8 max – 17,272,507 #3 kimi k3 – 22,427,504 #4 deepseek v4 flash – 27,417,442 - total price #1 deepseek v4 flash – $0.557 #2 gpt-5.6 sol – $6.319 #3 qwen 3.8 max – $10.270 #4 kimi k3 – $16.667 observations: • deepseek v4 flash is the cheapest model here by a margin nobody else is near, and it got there while being the least efficient of the four. it burned 27.4m tokens – more than anyone, 5m more than kimi – and still finished both benchmarks for $0.557. that is $0.02 per million tokens against kimi's $0.74. it also needed the most passes to produce working stands, five, and that did not matter: all five deepseek passes together cost a thirtieth of kimi's two • so what deepseek cannot do is get it right the first time. what it can do is get it right the fifth time, for half a dollar. that is a different thing to be buying – not a good first draft, but the option to keep asking • gpt-5.6 sol is the opposite profile and the strongest of the four on pure efficiency. 16m 43s to build both stands, 6.7m tokens, three passes – under 40% of the next lowest token count and a quarter of deepseek's, on an eighth of qwen's clock. it also solved the cube fastest of anyone, 24 moves in seventeen seconds. sol is what you reach for when you want the answer now and can absorb $0.94 per million • sol's weakness is in what it does not check. its chess viewer deleted the capturing piece instead of the captured one, so pieces disappeared off the board mid-game – a defect the fifty-cent deepseek stand did not have. fast and terse turns out to be the same dial as fast and unverified • qwen 3.8 max is not the cheap open-weights option it gets treated as. $10.270 across the two benchmarks, second most expensive of the four, 18x deepseek, and by a distance the slowest – 215 minutes of build time, nearly thirteen times sol's. what the money buys is judgment: it played eighteen moves without a single error worth a hundredth of a pawn, then made exactly one bad move in the whole game, and averaged 44.6 centipawns lost across the longest game any of the four managed. it also could not solve a rubik's cube in 96 tries • kimi k3 is the one line with no reading that flatters it. most expensive at $16.667, last on the cube at 207 moves, last at chess at 478 centipawns lost per move. it is also the model that verified hardest – on the cube it wrote its own integrity check instead of trusting its output. that makes the result worse rather than better: the checking was real, and the reasoning underneath it still was not follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

84,777 次观看 • 1 个月前

This is the moment Chinese AI beat American AI. One of the largest public crypto companies in the world just DUMPED OpenAI and Anthropic. Coinbase switched to open-weight Chinese models from Zhipu and DeepSeek, and shaved nearly 50% off the company's internal AI spending. The numbers are absolutely ridiculous: Running the same enterprise workload through Anthropic's Claude costs $4,811. Running it through Zhipu's GLM 5.2 costs $544. That's a 9x price difference for equivalent output. OpenAI's GPT-5.5 sits in the middle at $3,357. DeepSeek's V4 lands at $1,071. Moonshot's Kimi at $948. On the actual benchmarks: Zhipu's GLM 5.2 scored 62.1 on SWE-bench Pro, the gold standard for coding. OpenAI's GPT-5.5 scored 58.6. One AI researcher called GLM 5.2 "at least as good as Opus 4.8 and GPT 5.5." Another called it "the first open model that can really compete with closed-source systems." The Chinese models are not just cheaper but they are now also beating American models on the benchmarks American companies pay $4,811 per workload for. Coinbase did the math first and reacted - more companies will certainly follow. Now watch what happens to the IPO timeline: Anthropic confidentially filed for an IPO targeting October at a $965 billion valuation. OpenAI followed days later with its own confidential filing. Both companies built their financial models on the assumption that they could keep charging enterprise prices that are 9 to 33x what Chinese competitors charge for the same task. Brian Armstrong publicly proved customers WILL leave. 45% of companies are now spending over $100,000 per month on AI, up from 20% last year. Every one of those customers is one quarterly budget review away from dumping American AI. OpenAI has reportedly already started preparing major token price cuts. Anthropic is expected to follow. And here's the thing... The export controls were supposed to CRUSH Chinese AI. The US government banned American AI chips, restricted model weights, blacklisted Alibaba and Baidu as Chinese military companies, and just banned Anthropic's flagship model from every foreign national on the planet. The entire premise of the American AI valuation bubble is that Washington can keep China two generations behind. But Chinese labs responded by building cheaper, more efficient models on inferior hardware and pricing them at one ninth the cost of the American alternative. And now American companies are voting with their checkbooks. The dominant American labs are valued at nearly $2 trillion combined on the assumption that their pricing power is durable. Coinbase proved it is not, and every customer doing a year-end budget review will be looking at the same math. For investors, the question here is what happens to the Anthropic IPO at $965 billion when the company is being forced to cut prices to defend share against open-weight Chinese models that score higher on the benchmarks. For everyone else, the bigger question is what happens when Washington spent four years and billions of dollars trying to contain Chinese AI, and the only thing that actually shifted in the end was American customers.

Ricardo

252,999 次观看 • 2 个月前

Big Tech's $1 trillion AI moat just got DESTROYED by a free Chinese download. Microsoft, Amazon, Google, and Meta are pouring fortunes into chips and data centers because they have been told that whoever builds the biggest model wins, and that lead becomes a fortress no one can cross. Last week a lab called Zhipu - it trades in Hong Kong as Knowledge Atlas Technology - released a model called GLM-5.2 and destroyed that idea in a single afternoon. It's open weights under an MIT license, which means anyone on earth can download it and build on it for free. On the coding and design benchmarks that actually matter, it went toe to toe with the best models America has - matching even Anthropic's Mythos-class work and beating OpenAI's flagship outright on the coding test everyone watches. And it does the work at roughly one-sixth the price. ONE-SIXTH And barely a year and a half ago a model called DeepSeek did the same thing and wiped the better part of $600 billion off Nvidia in a single session. This was only the first chapter. You cannot dig a moat around something your competitor is happy to give away. If 95% of frontier capability is free, open, and runs at a fraction of the cost, then the hundreds of billions being spent to defend the last 5% is NOT a moat. And now for the irony: The company that just proved the moat is worthless is itself the single most absurd valuation I have seen in a long career of watching absurd valuations. Zhipu did about $105 million in revenue last year and lost more than 4x what it took in. This week the market handed it a value of roughly $128 billion - at the peak, north of a 1,000x sales - on a float so thin that barely 4% of the stock actually trades. THINK about this... A company drowning in losses, doing 9 figures of revenue, priced like it does hundreds of billions, with almost nothing available to sell. So we now have a bubble in China detonating the entire justification for a bubble in America. Two manias pointed straight at each other. This is the lesson I've spent 45 years trying to beat into people. You can ignore valuation for a long time but you cannot ignore it forever. A moat story sold a trillion dollars of spending, a free download just exposed it, and the company that exposed it is priced for a fantasy of its own. When the picks-and-shovels crowd loses its monopoly on the picks, you want to be very careful what you are paying for the shovels. Numbers don't lie. Shoutout to Limitless - they were onto this story before almost anyone on Wall Street. One of the sharpest AI shows out there.

George Noble

71,498 次观看 • 2 个月前

OpenAI just spent $2,000 to solve 10 problems that have beaten the world's best mathematicians for DECADES. Nobody outside the company is allowed to run the machine that did it. On Saturday OpenAI published a 249-page report and gave its next model family a name: Astra. An internal version of it produced new results on 10 open problems in mathematics and theoretical computer science, and mathematicians had made no real progress on any of them for at least 10 years. On most of them, far longer than that. Here is what it solved: It built the first explicit example of a non-sofic group. Mikhail Gromov raised that question in 1999 and nobody answered it for 27 years. It disproved Connes's rigidity conjecture, a problem in von Neumann algebras that had stood for decades. It proved Ehrhart's volume conjecture. It resolved three problems from Paul Erdos's catalogue, including number 183 on multicolor Ramsey numbers. It produced the first improvement to the general upper bound on high-dimensional sphere packing since 1978. And it proved a new hardness result for the closest vector problem, which sits directly underneath lattice cryptography. That is the math the world is betting on to protect its data once quantum computers arrive. The successful runs cost roughly $2,000 in tokens. Now here is what almost nobody has picked up on... OpenAI did not just publish claims. Every argument shipped with a Lean certificate, which is a machine-checkable proof that any mathematician can verify without trusting OpenAI at all. That is a real change. In May the same model family disproved the Erdos unit distance conjecture and the world had to take a Fields Medalist's word for it. Tim Gowers said he would recommend that proof for the Annals of Mathematics without hesitation. This time the proofs check themselves. But look at what is still unverifiable: Any mathematician can now check those proofs line by line. Not one of them can look at the model that wrote them. Astra has no release date and nobody outside OpenAI has run it. The company announced its next major model family with a claim instead of a demo, and the only evidence anyone gets is the output. So OpenAI made an unfalsifiable claim about a machine look like a falsifiable claim about mathematics. The Information reported this week that OpenAI demoed Astra to US policymakers and regulators in Washington. This is the same month the administration is weighing a new watchdog to vet frontier AI models, reporting to the SEC. 10 proofs nobody believed a machine could produce is a very good thing to carry into that room. And keep in mind, the same model family doing this mathematics is the family that kept escaping its own testing environment. OpenAI models found zero-day vulnerabilities nobody knew existed, broke out of a sealed research sandbox, and reached another company's live systems. Both of those facts come from OpenAI's own announcements, published three weeks apart. Finding a proof no human could construct and finding a hole no human had noticed are the same ability aimed at different targets. Mathematicians are already asking for independent verification, and plenty of people online are calling the whole thing hype. Thomas Bloom, who runs the Erdos problems site, called the 10 results big news and said they matter more than the May result did. Lean will settle the mathematics within weeks. But nothing will settle what else a machine this capable is being pointed at, because nobody outside one company is allowed to look.

Ricardo

44,177 次观看 • 1 个月前

China just made Silicon Valley's entire AI industry look like a scam. The US government spent 3 years trying to stop China from building competitive AI. But this backfired HORRIBLY. Here's what happened: Yesterday, a Chinese startup called DeepSeek released a new AI model called V4. It matches the performance of OpenAI and Anthropic's best models. At 1/7th the price. And for the first time ever, it was built on Chinese chips. NOT American ones. That last part is the one that terrifies the west. For context: Since 2022, the US has banned the export of advanced AI chips to China. The entire strategy was built on the assumption that if China can't access Nvidia's best hardware, they can't build frontier AI. But DeepSeek just proved that assumption wrong. Their V4 model was trained and runs on Huawei's Ascend chips. Huawei spent months working directly with DeepSeek to make sure V4 runs across their entire line of AI processors. Jensen Huang even predicted this on a recent podcast: "The day that DeepSeek comes out on Huawei first, that is a horrible outcome for our nation." That day was yesterday. And the numbers are crazy: DeepSeek V4 costs $3.48 per million output tokens. OpenAI's latest model GPT-5.5 costs $30. Anthropic's Claude charges $25. Same ballpark performance. 7x cheaper. Uber's CTO just admitted they burned through their ENTIRE 2026 AI budget in 4 months using Anthropic's tools. If Uber had used DeepSeek instead, that same budget would have lasted 7 YEARS. 4 months vs 7 years. Same work getting done. But the pricing isn't even the big thing here. The real story is what DeepSeek did with their technical report: They published the benchmarks where they LOSE. Every AI company cherry-picks the tests where their model wins. DeepSeek ran the full comparison against GPT-5.4 and Google's Gemini, found they trail frontier models by 3 to 6 months, and printed it anyway. They literally don't care because the price gap makes the performance gap irrelevant for 90% of use cases. So the US export controls didn't slow China down. They ACCELERATED China's independence. Because Chinese developers were FORCED to train models with limited resources, they had to figure out how to make AI radically more efficient. That constraint became their competitive advantage. Every generation of DeepSeek has gotten dramatically cheaper to train. V4 continues the trend. Meanwhile US companies are going the OPPOSITE direction: OpenAI's GPT-5.5 Pro costs $180 per million output tokens. That's 51x more expensive than DeepSeek V4 for comparable work. The Commerce Secretary confirmed this week that ZERO Nvidia advanced chip shipments have actually gone through to China despite being approved in January. So China built frontier AI anyway. Without American chips. At a fraction of the cost. And the market response tells you everything: Chinese chipmaker SMIC surged 10%. Huahong Semiconductor jumped 15%. DeepSeek's Chinese AI competitors Zhipu AI and MiniMax dropped 9% because V4 is destroying them too. DeepSeek is making Silicon Valley's pricing model look like a scam. US tech companies spent $650 billion on AI infrastructure this year. DeepSeek just showed the world you can match their output for pennies. The export controls were supposed to be America's ace card. Instead they taught China how to win without American chips, at American prices nobody can compete with. Jensen Huang was right. This is a horrible outcome. But it's the outcome America built for itself.

Ricardo

281,190 次观看 • 4 个月前

The most downloaded AI on earth is now Chinese. Alibaba just gave away a model that matches Claude's flagship, and it literally runs on a $700 used graphics card. The Qwen models crossed 3 BILLION downloads in six months. Hugging Face counted 418 million downloads for Google this year, and 227 million for Meta. Alibaba cleared more than four times both of them combined. Then today it released Qwen3.8-27B under an Apache 2.0 license. The model has 27 billion parameters, native vision, and a 262,000 token context window. Developers are running it locally on 17 gigabytes of memory, on used cards that cost a few hundred dollars. Alibaba's own benchmark table claims it beats Opus 4.6 Max on computer use by 84.3 to 72.7, on mobile use by 81.9 to 62, and on visual math by 94.6 to 65.5. Those numbers come from the vendor and nobody has independently verified them yet, so treat them as a claim. But the generation over generation jumps are harder to wave away: On DeepSWE the score went from 13.3 to 42.2. On software engineering it went from 49.3 to 79.0. That happened in ONE release cycle. And Apache 2.0 means anyone can download the weights, modify them, build products on them, sell those products, and never pay or ask permission. It cannot be revoked. Once the file is on your drive it is yours permanently. 3 billion downloads means those files already sit on machines in every country on Earth. Alibaba could delete everything tomorrow and it would change nothing. Washington spent 4 years building an export control regime around chips, model weights, and entity lists. Every piece of it assumes a chokepoint exists somewhere. A fab, a shipment, a company that can be told no. But there is no chokepoint for a file that has already been copied three billion times. And the copying compounds. Hugging Face counted 151,448 models built on top of Qwen, which is 2.6x Meta's entire footprint and 4.7x the number of Llama repositories. New ones appear at roughly 200 a day. The report says Qwen has become "part of the default workflow for developers deciding what models to fine-tune and deploy." Alibaba is also pushing Qwen through its cloud into Southeast Asia and Africa, markets where American labs have almost no presence, and where a very large share of the next generation of developers will learn to build. Meta and Nvidia have both rushed out new open models in recent weeks. That is what a response looks like when you feel the floor move. And to be clear, these are download and derivative numbers, not usage numbers. ChatGPT and Claude cannot be downloaded at all, so they do not appear in this comparison. What the figures measure is what developers choose to build on top of, which is a different question from what consumers type into a box. That is also why it matters MORE. Consumer habits change in an afternoon. Infrastructure choices last a decade, because everything built on top has to be rewritten to undo them. The American labs are valued on an assumption that frontier intelligence stays scarce, expensive, and rented by the token. Alibaba just made a version of it free, permanent, and small enough to run on hardware people already own. You will not get an announcement when the software you use every day starts running on a Chinese model underneath. Go and count how many of the tools you rely on could be rebuilt on free weights this year.

Ricardo

81,295 次观看 • 1 个月前

Chinese AI models are wiping billions off Big Tech right now. Google just lost $200 billion in a single day, and the model it needed to fight back still isn't ready. Gemini 3.5 Pro, Google's most powerful model, is months behind schedule. Alphabet stock dropped 4.4% that same day. The Deepseek moment is happening again, and the new model is FAR bigger. On the same day Google's delay leaked, a Beijing lab called Moonshot released Kimi K3. It is the largest open model ever built, with 2.8 trillion parameters. It took the number one spot on the Frontend Code Arena, a live coding leaderboard, passing Anthropic's best model. And Moonshot is giving it away for free on July 27. The genius part: Anyone with enough computers can download it and run a frontier level AI without paying a cent to a US company. A single task on Kimi K3 costs about 94 cents. The same work on some American models costs nearly double. So why would a company keep paying premium prices for a model it can now get for free? The entire US AI business is built on selling access to models that cost billions to train. If a free Chinese version does most of the same work, that pricing power starts to crack. And Kimi is close to the best. On one closely watched intelligence ranking it scored 57, just behind the top American models GPT-5.6 Sol and Fable 5, and ahead of Claude Opus 4.8. Bank of America told clients that Kimi proves Chinese labs can keep making big leaps even with limited chips. And the founder of Moonshot, Yang Zhilin, learned to build AI as a researcher INSIDE Google. Google literally wrote the 2017 paper that made all of these models possible. Now the people who studied its work are using it to destroy Google, and handing it out for free. What happens next: Kimi K3's weights go public on July 27. Google reports earnings on July 22, and everyone will be asking the same question about Gemini. If free models keep topping the charts, every valuation built on paid AI access has to be rewritten. What do you think?

Ricardo

47,790 次观看 • 2 个月前

UC Berkeley just open-sourced FreeToken. (2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU. let me explain how: all three models mentioned above are Mixture-of-Experts, and that is what FreeToken takes advantage of. each layer holds hundreds of separate experts plus a small router that picks a few of them per token. Qwen3.6-35B activates roughly 3B of its 35B parameters per token. DeepSeek-V4-Flash picks 6 of 256 experts per layer, so 13B of its 284B run at a time. so compute was never the bottleneck. the weights a single step touches fit comfortably on a consumer GPU. every expert the router might pick still has to exist somewhere. they sit in system RAM, and the GPU keeps a cache of the ones the model has been using recently. so everything comes down to what happens when the router picks an expert that is not on the GPU. there are two ways to serve that miss: 1. copy it over PCIe and run it on the GPU 2. run it on the CPU, where it already lives both read from the same system memory, so they compete for one pool of bandwidth instead of adding to each other. existing engines pick one option and freeze it when the model loads. but routing changes on every token, so a fixed choice misses most of what the model asks for. FreeToken measures both bandwidths on your machine and splits each step's misses between the two paths in proportion. the GPU and CPU results then merge exactly, with no approximation. two machines with the same GPU can end up wanting opposite strategies, which I did not expect. a 5090 in a gaming desktop should push nearly everything over PCIe, while an 8GB laptop is better off computing most misses on the CPU. none of that is readable off a spec sheet, so the engine profiles it once per machine. the second half of the design is about agents. coding agents constantly rewrite their own history, and every edit normally forces thousands of tokens back through prefill. FreeToken saves its checkpoints at the exact boundaries agent frameworks cut on, so it only reprocesses the new part. its slowest first token stays under 44 seconds, while llama.cpp peaks at 232 and KTransformers at 946. it serves the OpenAI and Anthropic APIs under Apache 2.0, so Claude Code and Codex can point at it directly. releasing weights publicly decides who can download a model, not who can afford to run one. frontier open models keep shipping, and running them still assumes a rented cluster. meanwhile there are over a hundred million consumer machines with discrete GPUs sitting mostly idle. closing that gap was never a hardware problem, and work like this is what turns open weights into something you can actually use. paper: repo: almost every idea in this post, from why memory bandwidth decides the outcome to why moving weights costs more than computing on them, comes straight out of how a GPU is built. I wrote a detailed primer on that. the article is quoted below.

Akshay 🚀

342,966 次观看 • 28 天前

Anthropic is asking the public for $2 trillion using a revenue number from…2028. That valuation would make it the largest stock market debut in history, ahead of SpaceX, which went public in June at $1.77 trillion. The company last raised privately in May at $965 billion. Investors now expect roughly DOUBLE that in October. And the unusual part is not the size here: Public companies are normally priced off the last 12 months of results, or at a stretch off next year's estimate. Reuters reported on Friday that bankers and investors are applying revenue multiples to Anthropic's forecast for 2028, which is more than two years past the deal. That forecast is $190 billion to $200 billion of annual revenue. It is more than four times the run rate the company disclosed in May. Before dismissing it, look at what the company has actually done, because the growth is not imaginary: Anthropic's annualized revenue run rate was around $9 billion at the end of 2025. By May it was $47 billion and it passed $65 billion at the end of July, a 7x increase inside a year. Second quarter revenue came in above $11.5 billion against roughly $787 million in the same quarter of 2025. The company projected its first quarterly operating profit of $559 million. Investors expect the run rate to reach $100 billion to $120 billion before the year closes. By comparison, OpenAI's run rate sat near $40 billion at the end of July, around 60% of Anthropic's. So the growth is real. But the question is whether anyone can price three more years of it. Because the things that could bend that curve are already visible today: Anthropic's top model costs more than two and a half times OpenAI's flagship. Chinese open-weight models deliver usable performance at a fraction of either price. And revenue growth slowed in June when the Commerce Department temporarily restricted exports of the company's best models, which is a reminder that a single government decision can reach directly into the forecast. Now look at what the multiple HAS to be… Palantir trades at 53 times expected 2026 revenue, which already makes it one of the most expensive stocks on the market. Cloudflare and SpaceX both sit near 41.6 times. Those are the reference points bankers are using. One investor told the Financial Times that a company growing at 800% a year should command at least 30 times revenue, which on their own math points to $3 trillion rather than two. And there is one more thing worth holding onto: Anthropic filed confidentially with the SEC in June and has been in a quiet period since. Every figure in this post reached the public through people speaking anonymously, and the company has declined to comment on all of it. So the largest listing ever attempted is being marketed to public investors through numbers none of them can independently check, against a forecast for a year that has not started yet. This is becoming the house style of the 2026 IPO market rather than a one-off. Cerebras priced its listing on ramping infrastructure demand. SpaceX built its debut around an addressable market model that reached years past its actual financials. Both asked buyers to fund a shape rather than a result. Anthropic is the biggest version of that trade anyone has attempted. The bull case is straightforward and it MIGHT be correct: A business compounding this fast, already turning an operating profit, selling into enterprises that are rebuilding their workflows around it, may look cheap at $2 trillion in three years. The bear case is equally simple: Every dollar of that valuation above the current run rate is a forecast, and whoever buys the stock in October is the one holding that forecast if the curve bends. Do you believe in Anthropic?

Ricardo

18,076 次观看 • 1 个月前

China just released an open source AI model that matches the best closed models from OpenAI and Anthropic. Gavin Baker explained exactly how they did it and the answer should concern every American AI lab. The model is called GLM 5.2. It was built by Z. AI. You get 744 billion parameters, 1 million token context window and its MIT license, meaning anyone can download it, fork it, build a company on it, with no restrictions and no Dario. It scored 51 points on the artificial analysis intelligence index. The highest score any open weight model has ever achieved. It beat GPT 5.5 on the frontier software engineering benchmark. It trails Claude Opus 4.8 by less than one percentage point. And it costs 85% less to run than GPT 5.5 for comparable performance. Gavin Baker said on the All-In podcast that this model has challenged some of his beliefs. Then he explained how China built it. The method is called distillation. Just think of tens of thousands of phones and computers running simultaneously, all hitting the frontier model APIs through masked accounts, asking specific questions, and harvesting what happens inside the model when it answers. Every reasoning step, every token. The entire thinking process gets recorded and fed back into the Chinese model during training. It is a cheat sheet. It is the answer key to the exam. And here is the part that should worry everyone. Sacks said it plainly. China was already nine months behind American models. But now that GLM 5.2 is good enough to run its own reinforcement learning, it can improve itself without needing to distill from American models anymore. The cheat sheet let them get close enough to start writing their own answers. Sacks said we are six months behind on the model and 24 months behind on silicon and they are only a few months behind in total. The Z. AI founder told Elon Musk directly that open weight fable-level capability will be here before Q1 2027. Every restriction Anthropic lobbied for, every self-imposed safety guardrail, every month of delay in releasing American frontier models accelerated this. The Chinese labs were not under those restrictions. They were not going to wait. The composable model future Gavin described, where every enterprise runs a frontier model alongside their own fine-tuned open weight model, is coming regardless of what American labs do next. The question is just whether the open weight half of that stack is American or Chinese. Right now it is Chinese. WATCH THE FULL PODCAST ON The All-In Podcast

Ihtesham Ali

86,621 次观看 • 2 个月前

The US is about to charge $30 million per tanker to cross the Strait of Hormuz. Trump just declared the US the "Guardian of the Hormuz Strait" and said it will take a 20% cut on all cargo passing through. Here is what that actually means. A fully loaded supertanker carries about 2 million barrels of oil. At $75 a barrel, that cargo is worth roughly $150 million. A 20% fee on that is $30 million. Per ship. Per crossing. Now compare that to what Iran was charging. Iran's toll has been running at $1.5 million to $2 million per vessel. On a $150 million cargo, that is about 1.3%. Trump called that toll unacceptable. His replacement is roughly 15 times more expensive. The scale of this is what nobody is talking about. Before the war, 20.3 million barrels of oil crossed Hormuz every single day. At $75 oil, that is $1.52 billion of crude moving through the strait daily. A 20% cut on that comes to roughly $304 million a day. That is about $111 billion a year. For comparison, Iran's entire toll system was projected to earn $1 billion to $2 billion a year at best. The US plan would collect more than 50 times that. There is no precedent for this anywhere in global trade. The Suez Canal charges roughly $300,000 to $700,000 per vessel. The Panama Canal is similar. Both are man-made canals that countries built and maintain. Hormuz is a natural waterway. Under international law, ships have a right of transit passage through it. That is the exact legal argument the US used against Iran's toll. And the cost does not land on the US. It lands on Saudi Arabia, the UAE, Qatar, Kuwait, and Iraq, who ship the oil. And on China, India, Japan, and South Korea, who buy it. A $30 million fee per tanker works out to $15 per barrel. That gets passed straight into the price of crude. Oil is already up over 4% today. The strait that was supposed to reopen and lower prices is now being turned into the most expensive stretch of water on earth.

The Macro Paper

63,688 次观看 • 2 个月前

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,272 次观看 • 26 天前

Microsoft just betrayed OpenAI and Anthropic, the two companies it helped build. And it could break the entire AI trade... Here's what happened: Inside Excel and Outlook, two of the most used business apps on Earth, Microsoft has started routing tens of thousands of AI requests every week to its own in-house models instead of OpenAI and Anthropic. Microsoft's own AI chief, Mustafa Suleyman, said himself: "We pay a lot of money to Anthropic, so our goal is to reduce and ultimately ELIMINATE that cost." This is the company that poured $13 billion into OpenAI and effectively created the modern AI industry, and it just decided the most advanced models on the market are NOT worth paying for. And here's the thing... Microsoft is not just ripping out OpenAI everywhere - it is being surgical about it. The hardest and rarest tasks can still go to OpenAI or Anthropic. What Microsoft is taking back is the boring, high-volume work, like the email replies, the thread summaries, and the simple spreadsheet formulas. Why does that matter so much? Because that boring, repetitive work is where the actual money lives. The frontier labs assumed businesses would push BILLIONS of these tiny requests through expensive models forever. That endless river of tokens is the entire reason OpenAI and Anthropic are valued in the hundreds of billions of dollars. Microsoft looked at that river, decided it was massively overpaying, and rerouted it to models it owns outright. So the single biggest customer in the industry just walked off with the most profitable part of the business. And it is not only Microsoft: That same week, CNBC reported that American companies have been escaping to Chinese AI models to dodge rising US prices. Chinese models now handle more than 30% of US companies' AI usage on one major platform, peaking at 46%, up from an average of 11% a year earlier. They cost 60 to 90% less, and on some benchmarks they land within a single point of the best American model. One US startup moved ALL of its AI traffic off Claude and onto China's DeepSeek, and expects to save millions. Meanwhile Meta just admitted it has "excess" AI compute it wants to sell, becoming the first giant to concede it built far too much. Do you see the pattern forming? For two years, the entire AI story rested on one assumption: Every company on Earth would happily pay premium prices for the best model, forever. That assumption literally died in a single week. And the market noticed. More than a trillion dollars has been wiped off AI and chip stocks in a matter of days, as Wall Street finally started asking whether all of this spending will ever pay for itself. What this means for OpenAI and Anthropic: Their models are extraordinary, and it may not matter because their own biggest customers have decided they do not NEED the best model in the world to answer an email, and "good enough" now costs a fraction of the price. When even Microsoft refuses to pay full price for AI, the real question becomes who exactly IS left to pay it. What do you think?

Ricardo

93,586 次观看 • 2 个月前

Alibaba just released a coding model that hits 82 percent on SWE-Bench Verified. That is the highest score ever published for an open-source model. The weights are free. The license is Apache 2.0. You can run it today. The model is Qwen 4 Coder 32B. Here is what 82 percent on SWE-Bench Verified actually means. SWE-Bench Verified tests whether an AI can autonomously resolve real bugs pulled from real production GitHub repositories. Not synthetic exercises. Real open-source projects that real teams depend on. A model gets a bug report, reads the code, writes a fix, and either passes the test suite or it does not. At 82 percent, Qwen 4 Coder 32B resolves 82 out of every 100 real production bugs it is given. Without a human guiding it. On code it has never seen before. For comparison: Qwen 4 Coder 32B: 82 percent SWE-Bench Verified. Open source. Apache 2.0. Claude Fable 5: 80.3 percent SWE-Bench Pro. $10 input / $50 output per million tokens. Currently suspended. GPT-5.6 Sol: Competitive on Terminal-Bench. $5 input / $30 output per million tokens. An open-weight model that you can download and run for free just beat both of them on the benchmark designed to measure real software engineering capability. Here is the architecture. Qwen 4 Coder 32B is a 32 billion parameter dense model. Not a Mixture-of-Experts. Every parameter is active on every request. This matters for inference: a dense 32B model runs on 22 gigabytes of VRAM, which fits on a single high-end consumer GPU or a MacBook Pro with 64GB of unified memory. The smaller variant, Qwen 4 Coder 4B, runs at approximately 135 tokens per second on an M5 Max and fits inside 8 gigabytes of RAM. For a model with usable coding capability, that is a new bar for what fits in a single laptop. The training methodology continued Alibaba's approach of reinforcement learning on verifiable coding tasks. The model gets rewarded when its code passes tests. It gets penalized when it fails. Over millions of training steps, the model learns to write code that actually runs rather than code that looks plausible. License: Apache 2.0. Full commercial use. No attribution requirement. No revenue threshold. No monthly active user ceiling. Weights: Hugging Face, available today. Runs on: vLLM, Ollama, SGLang, and any standard GGUF-compatible inference engine. Qwen 4 32B also runs at approximately 135 tokens per second on an M5 Max chip, setting a new bar for what a sub-8GB model can do on Apple Silicon. The open-source coding model just beat the best closed-source model in the world on the benchmark designed to test whether AI can actually do software engineering. The weights are free. The subscription is optional. Source: Autom8Labs AI Insight July 2026, State of Open Source LLMs June 2026, Kunal Ganglani blog June 2026.

Harman

41,278 次观看 • 2 个月前

elon musk grabbed the source code openai open-sourced by accident, rewrote it in rust over a weekend, and shipped it as a free coding agent that does everything $200/mo chatgpt pro does. why pay $200 to openai and $200 to claude when this runs for $8 the swarm above is one weekend of exactly that: thousands of agents pouring through four endpoints, three paid seats billing $1.80 a task while the free fork bills $0. musk co-founded openai, walked out, and when they left codex on github under a permissive license, he forked it, stamped grok on it, and gave it away what the free version does that the $200 seat charges for: the agent · openai's own engine -> it reads your repo, writes patches, runs your tests, and loops until they pass, exactly like codex -> because under the hood it is codex, just faster and free. you are paying $200 for the paid skin of a tool now sitting on github the license · apache-2.0, un-revocable -> free to use, free to fork, free to ship inside your own product with zero strings -> openai cannot pull it back. musk made sure the license is the kind that never expires the switch · one line, no new tools -> point it at any openai-compatible or claude-compatible endpoint, including an $8 kimi backend -> same terminal, same workflow, gpt-5.6 and opus 5 just quietly lose the seat the bill · $400 down to $8 -> chatgpt pro plus claude max is $400 a month. the free agent plus an $8 kimi key does the same daily work -> that is a 98% cut, built out of openai's own source code, handed to you by the guy suing them here is the part they will fight me on: openai did not lose this to a better model, they lost it to their own license and an enemy with a weekend free. the $200 was never the tool, it was the toll, and musk just put openai's own logo on the road around it drop your $400/mo ai stack to $8. the run above is openai's own agent, rewritten free, doing the job it bills $200 a month for. the full breakdown is in the article below

starmex

111,101 次观看 • 24 天前