Loading video...

Video Failed to Load

Go Home

Matt Maher tested frontier models in Cursor v. other harnesses. Cursor boosted model performance by 11% on average: Gemini: 52% → 57% GPT-5.4: 82% → 88% Opus: 77% → 93% His benchmark measures how well models implement a 100-feature PRD. Cursor consistently outperformed.

912,650 views • 4 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Elon Musk just made one if the biggest moves in taking over the programming industry “SpaceX just bought Cursor for $60 billion. Do you realize how big this is? SpaceX went public — the biggest IPO in history. $75 billion raised, almost a $2 trillion valuation and the first thing to do with that money? Buy the most popular AI coding tool on the planet. Here's why that changes everything. Elon now owns 3 layers: the compute, Colossus data centers, the models, Grok through xAI, and now the tool that developers actually use every day. It's the full stack. And here's what makes Cursor different from Claude Code or Codex. Cursor is model agnostic. You can run Claude in it, GPT, Gemini, whatever model you want. It's not locked to any one company, and now it has SpaceX's resources behind it. Cursor said they were bottlenecked by compute. Well, that bottleneck has just been removed. $4 billion in annual revenue, over half the Fortune 500 already uses it, and now it's backed by a $2 trillion company. OpenAI has Codex, Anthropic has Claude Code, and now Elon has Cursor.” Let me break this down in simple terms Elon Musk now controls more of the full AI picture: - Massive computers, power (data centers like Colossus) - Smart AI models (Grok from xAI) - The actual tool millions of developers use every day (Cursor) For every day users this means Faster and smarter apps and websites in the future. More developers using powerful AI tools means new apps, games, websites, and features get built quicker and cheaper. This means better video games, smoother streaming, smarter phone apps and better programs For Developers they can describe what they want in plain English (“make a feature that does X”) and the AI handles more of the heavy lifting

Wall Street Apes

213,267 views • 1 month ago

Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring Federico Cassano (Composer Lead at Cursor) and Dmytro Dzhulgakov (Co-Founder at Fireworks AI). The Cursor team trained Composer 2 on Fireworks by starting with a strong base model (Kimi 2.5) and performing large-scale mid-training on code tokens and web data to learn common patterns and libraries, followed by a large-scale Reinforcement Learning run to learn how to navigate the Cursor harness, call tools, and write correct code. Today's episode dives into the systems and infrastructure challenges of making that large RL run happening, and there were many (!!), from numerical mismatch to global distribution to synchronizing rollouts across asynchronous pipelines to keeping track of expert activation across runs and more. Extremely nerdy in-the-weeds challenges that Federico and Dima were delighted to nerd out on together :) Beyond RL infra, we also discussed Online vs Simulated rollouts, self-summarization for long-horizon agents, environment design ("the most powerful RL environment is the product itself"), and other technical nuggets. PS: We filmed this episode before the SpaceX news, while the Cursor team was still compute-constrained. While Cursor now has *all* the flops, the takeaways and hurdles crossed ring true for any serious application-level company that is racing to post-train their own models. I believe that more serious application companies will go the way of Cursor and post-train their own models. 00:00 Introduction 00:53 Why Cursor Trained Composer 2 04:55 Specialization vs Bitter Lesson 06:16 Composer 2 Training Recipe 16:32 Scaling RL Infrastructure Globally 23:32 Floating Point Drift 25:11 MoE Sensitivity Explained 26:25 Router Replay Fix 27:19 Real Time RL Loop 31:49 Long Horizon Agents 34:29 Why RL Everywhere 37:34 LLM as Judge Rewards 39:14 RL in Hard Domains 40:13 Build Your Own Environments 44:34 Closing Thoughts

Sonya Huang 🐥

79,302 views • 2 months ago

I got to try Grok 4.5 in early access in Cursor for the past few days and I absolutely enjoyed it. It feels like Opus 4.8 at 2x the speed at a much cheaper price point. I tasked it to brainstorm > plan > implement a big feature for my game (this act 1 boss fight) and it did not disappoint. - It is much smarter than Composer 2.5, during planning mode, it is able to think through my request more robustly, ensuring that edge cases are covered and makes sure to ask the right questions to confirm with me first. - It is much better at brainstorming ideas/suggestions, similar to Opus 4.8, though I think Fable still edges out a little when it comes to brainstorming ideas and suggestions - It is FAST. probably the fastest of all frontier models (Opus 4.8, GPT 5.5 etc), which makes it a joy to build with, because I can stay in the flow - It has much improved visual/animation capabilities than Composer 2.5, it can code up animations (i wanted an explosion animation with particle effects) with much, much better visuals, animation movement and timing. This is a big leap and I was so happy to see this improvement. - The best part for me is that I can just use the same model from planning down to execution without switching to a lower cost model because the price point is cheaper than other frontier models. I'll be testing this model with more challenging tasks in the next few days but I think this is going to be my main driver for vibe coding for a while. Also, its nice to see Grok back in the race. 🙌

Danny Limanseta

1,410,316 views • 20 days ago

🚨A 25 YEAR OLD BUILT THE FASTEST GROWING SOFTWARE COMPANY IN HISTORY.. WITH ZERO MARKETING SPEND.. AND SPACEX JUST OFFERED $60 BILLION TO BUY IT.. His name is Michael Truell.. He started coding at 11.. Interned at Google at 18.. Dropped out of MIT to start a company that built AI tools for mechanical engineering.. That company failed.. So he pivoted.. And built Cursor.. An AI-powered code editor that writes software for you.. Here's how fast it grew.. $100 million in annual revenue in 12 months.. Fastest in SaaS history.. Broke every record ever set by Slack, Zoom, and Wiz.. $500 million by month 21.. $1 billion by November 2025.. $2 billion by February 2026.. Projected to hit $6 billion by end of year.. Zero marketing spend.. Not a single dollar.. Pure word of mouth from developers who couldn't stop talking about it.. Over 1 billion lines of code accepted per day.. Used by 70% of Fortune 1000 companies.. Every single one of Nvidia's 40,000 engineers uses it.. Coinbase hit 100% adoption among their developers.. And he did this with a team of four MIT co-founders.. One of them was a three-time International Math Olympiad competitor from Pakistan.. Another was a college squash captain with zero startup experience who built the entire product strategy.. They spent zero on sales.. Zero on ads.. Zero on growth hacking.. The product sold itself.. But here's where the story takes a turn nobody expected.. Even at $50 billion valuation.. Even generating billions in revenue.. They hit a wall.. Not a market wall.. A physics wall.. They couldn't get enough GPUs to train their next AI model.. The physical chips didn't exist in sufficient quantities for them to buy.. Money couldn't solve the problem.. Enter Elon Musk.. On April 21.. SpaceX announced a deal to potentially acquire Cursor for $60 billion.. The largest acquisition option in tech history.. The structure is insane.. SpaceX gives Cursor immediate access to Colossus.. xAI's supercomputer equivalent to one million Nvidia H100 GPUs.. For nine months of joint development.. At the end.. SpaceX can buy the company for $60 billion.. If they don't buy it.. They owe Cursor a $10 billion breakup fee.. The largest breakup fee in corporate history.. Think about what that means for Cursor.. Either they get acquired for $60 billion.. Or they walk away with $10 billion in cash and nine months of free training on the most powerful supercomputer on earth.. There is no losing scenario.. And here's why Musk wants it.. SpaceX is preparing for an IPO at $1.75 trillion.. The biggest IPO ever.. But aerospace alone can't justify that number.. By merging xAI into SpaceX.. And now acquiring Cursor.. Musk transforms SpaceX from a rocket company into an AI empire that owns the compute, the models, and the developer tools.. Cursor is the missing piece.. The application layer that puts xAI's models into the daily workflow of every Fortune 500 engineering team.. Oh and one more thing.. In 2022.. FTX's trading firm Alameda Research made a seed investment in Cursor.. During the FTX bankruptcy.. Liquidators sold that stake for $200,000.. That stake is now worth approximately $3 billion.. Sam Bankman-Fried called it the worst liquidation decision in venture capital history.. From a prison cell.. A failed mechanical engineering startup.. Pivoted by four kids from MIT.. Zero marketing.. Zero sales team.. Built the fastest growing software company in history.. And now SpaceX is writing a $60 billion check for it.. This is the most insane founder story in Silicon Valley history.. And most people haven't even heard of Michael Truell.

Evan Luthra

988,417 views • 3 months ago

The most overlooked part of the SpaceX IPO thesis is the model and most people are completely missing it (Save this) Everyone has been focused on the Anthropic compute deal and the Colossus revenue because those are numbers you can put in a spreadsheet. Six months ago, xAI was competing reasonably well on model performance but was not clearly on the frontier. Then SpaceX exercised its option to acquire Cursor for $60 billion, the largest startup acquisition in history just days after completing the largest IPO in history at $75 billion. Cursor is a team of 700 to 800 people, was on track to exit 2026 at up to $10 billion in revenue, had millions of professional developers using it daily, and had already built a team with the genuine potential to compete at the frontier, the one thing holding them back was compute. SpaceX just gave them the largest GPU cluster in the world to work with. Grok 4.3, a 1.5 trillion parameter model, is currently training with Cursor's proprietary coding data being injected directly into pre-training, not just fine tuning which is a fundamentally more powerful integration than anything the market is currently modeling. The prior version, Grok 4, was already on the Pareto frontier as of 10 to 12 days ago, the most intelligent 500 billion parameter model in the world, sitting alongside Google Gemini, Anthropic, and OpenAI as one of only four systems at the true frontier. Composer 2.5, the previous Cursor model was Pareto dominant in coding tasks just before the acquisition closed, meaning SpaceX inherited a model that was already best-in-class in the highest-value AI use case in the market. The AWS parallel is the one everyone keeps missing. Bezos built data center capacity for Black Friday, sat on idle infrastructure the rest of the year, and monetized it into what was at the time the most profitable technology business in history and investors hated it in 2009 and 2010 because he was burning free cash flow on capacity that had no obvious revenue yet. SpaceX is in exactly that position, it built Colossus for xAI's own training needs, is monetizing excess capacity to Anthropic at $1.25 billion per month across 220,000 Nvidia GPUs, and has reportedly secured up to 20% of Nvidia's early Vera Rubin allocation, giving it the most powerful and scarcest GPU infrastructure in the world during the critical window when those chips are hardest to get. The $60 billion Cursor acquisition closed at a moment when SpaceX had essentially unlimited compute, a team already at the frontier, and a product with deep enterprise distribution, three things no other model lab had simultaneously when it was at this stage. The market is pricing the compute business conservatively and ignoring the model call option entirely, and coding is the fastest path to AGI, once you are on the Pareto frontier with that compute, revenue scales fast. Anthropic went from negligible revenue to $30 billion annualized in under 18 months and that is the existence proof. Bullish on SpaceXAI and Elon Musk

Milk Road AI

69,446 views • 1 month ago

China just released an open source AI model that matches the best closed models from OpenAI and Anthropic. Gavin Baker explained exactly how they did it and the answer should concern every American AI lab. The model is called GLM 5.2. It was built by Z. AI. You get 744 billion parameters, 1 million token context window and its MIT license, meaning anyone can download it, fork it, build a company on it, with no restrictions and no Dario. It scored 51 points on the artificial analysis intelligence index. The highest score any open weight model has ever achieved. It beat GPT 5.5 on the frontier software engineering benchmark. It trails Claude Opus 4.8 by less than one percentage point. And it costs 85% less to run than GPT 5.5 for comparable performance. Gavin Baker said on the All-In podcast that this model has challenged some of his beliefs. Then he explained how China built it. The method is called distillation. Just think of tens of thousands of phones and computers running simultaneously, all hitting the frontier model APIs through masked accounts, asking specific questions, and harvesting what happens inside the model when it answers. Every reasoning step, every token. The entire thinking process gets recorded and fed back into the Chinese model during training. It is a cheat sheet. It is the answer key to the exam. And here is the part that should worry everyone. Sacks said it plainly. China was already nine months behind American models. But now that GLM 5.2 is good enough to run its own reinforcement learning, it can improve itself without needing to distill from American models anymore. The cheat sheet let them get close enough to start writing their own answers. Sacks said we are six months behind on the model and 24 months behind on silicon and they are only a few months behind in total. The Z. AI founder told Elon Musk directly that open weight fable-level capability will be here before Q1 2027. Every restriction Anthropic lobbied for, every self-imposed safety guardrail, every month of delay in releasing American frontier models accelerated this. The Chinese labs were not under those restrictions. They were not going to wait. The composable model future Gavin described, where every enterprise runs a frontier model alongside their own fine-tuned open weight model, is coming regardless of what American labs do next. The question is just whether the open weight half of that stack is American or Chinese. Right now it is Chinese. WATCH THE FULL PODCAST ON The All-In Podcast

Ihtesham Ali

86,163 views • 1 month ago

Chamath is making one of the most important business arguments of 2026. Half of large US companies right now cannot generate returns that exceed their cost of capital, which has normalized back to its long run average of 8 to 11%. Another one in seven companies globally is stuck generating persistent returns between 1 and 5% and most businesses don't have room for error and in this environment walks every frontier AI lab saying the same thing, give us your data, your workflows, your processes and our model will make everything better. And companies by the millions said yes. What they didn't fully account for is what happens on the other side of that door. Every time an employee runs a query through a frontier model API, the prompt goes through external servers, workflows, customer data, pricing logic, internal processes, all of it transmitted through a third party. As Alex Karp said companies are spending on tokens while handing over the exact proprietary advantages that make their business worth owning. Microsoft blocked internal use of Anthropic's Claude Fable 5 but over its 30-day data retention policy and the largest software company in the world decided a frontier model's data handling was too risky for its own employees. A US government action revoked access to another frontier model for foreign nationals overnight. Now here's where the cost math becomes impossible to ignore. Deutsche Bank calculated a roughly 65x cost gap between frontier models like Claude Fable 5 at ~$3.25 per task and open-source alternatives at ~$0.05. For 90% of everyday enterprise tasks, performance is comparable. Open-weight models now match closed frontier systems on core agent tasks at roughly one-tenth the cost, a high-volume deployment that costs $250/day on Claude runs at $12/day on an open-source equivalent. Chamath Palihapitiya tested this directly by running a standard enterprise code migration task through an orchestration layer wrapping an open-source model came in 16.4x cheaper than using a frontier model directly.

Milk Road AI

280,929 views • 24 days ago

Aman Sanger on scaling Cursor to $100M in revenue in 12 months The four co-founders started working on Cursor in January 2023. “It was about 2-4 months of experimenting, trying different things and finding something that fit,” Aman recalls. “Then we shipped it and it had an okay launch.” The initial version of Cursor got some initial buzz because they shipped it with GPT-4 at a time where very few products were using Open AI’s latest model. But then usage tanked. The team began to doubt their approach after the failed launch. Aman explains: “The entirety of that summer was just incredibly slow growth, and that was somewhat demoralizing. The big question in our minds was, ‘Are we being too ambitious?’ We were trying to build this general purpose thing for all engineers, but with this really small team, maybe we should focus more narrowly on some particular use case like tests or bug detection.” But the fact that they were users of the product gave the team the confidence to keep going down their initial path: “The really magical thing about this product was that we were users of it. So we could iterate incredibly quickly, and we tried all these different things that summer. Then we found this core set of features that worked incredibly well.” The two key features were Command K for instructed edit ability and code-based indexing that let you ask questions about your whole code base. “After we integrated those two features and launched them, growth kind of just took off.” Aman continues: “A lot of the work of Cursor has been just experimenting with what is possible. For everything you see in the product, there’s like 10 failed experiments that didn’t work… All of our work for the first 6 months to a year was trying to find new ways to harness these models and make these models better for programming.” Another factor that contributed to Cursor’s success was their willingness to ship half-finished features: “We released these half-finished things, which a lot of our competitors refused to do. The first version of Copilot++ and Cursor Tab sucked. But once you release it to the world and see how people react to it, you can improve on it a ton… We biased toward releasing as soon as something shows signs of usefulness to the team.” Video source: Peak XV Partners (2025)

Startup Archive

56,166 views • 1 year ago

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,319 views • 2 months ago