Loading video...

Video Failed to Load

Go Home

Y Combinator built AI versions of its partners to help more people work through their startup ideas. For its Office Hour Simulator, the team needed useful answers delivered quickly enough for a spoken conversation. They had been testing lightweight Gemini and OpenAI models before moving to GLM-5.2 on a...

57,298 views • 5 days ago •via X (Twitter)

13 Comments

steve's profile picture
steve5 days ago

@ycombinator great partnering with the folks at YC on this 🫶

wafer's profile picture
wafer5 days ago

@ycombinator so true 🫶

wafer's profile picture
wafer5 days ago

Article:

John Hahn's profile picture
John Hahn5 days ago

@ycombinator did someone say faster than cerebras??

wafer's profile picture
wafer5 days ago

@ycombinator :o 👀

emilio andere's profile picture
emilio andere5 days ago

@ycombinator 🫶

wafer's profile picture
wafer5 days ago

@ycombinator 🧇🫶

JV's profile picture
JV5 days ago

@ycombinator inference for real-time ai = wafer btw

wafer's profile picture
wafer5 days ago

@ycombinator facts

Sebastian Caniulao | Ecommerce Email & Growth's profile picture
Sebastian Caniulao | Ecommerce Email & Growth5 days ago

@ycombinator The 2.5 minutes longer is the number that matters, and it's the one nobody optimises for directly. In a spoken conversation latency isn't a performance stat, it's whether the thing feels like it's listening or thinking about something else.

Tyler | Crypto Whale's profile picture
Tyler | Crypto Whale5 days ago

@ycombinator Latency wins the conversation. Always has.

Varik Verilion's profile picture
Varik Verilion5 days ago

@ycombinator For spoken Office Hours, latency becomes part of answer quality. A lighter model can feel better when it responds fast enough to keep the conversation moving.

wafer's profile picture
wafer5 days ago

watch the full video here:

Related Videos

MEET THE NVIDIA KILLER: OpenAI bet $10 BILLION on this company that makes chips 20x faster than Nvidia's. If this plays out as expected, it’s over for Nvidia. Cerebras Systems just locked in 750 megawatts of computing power to OpenAI through 2028. For reference: that's equivalent to the annual power consumption of 600,000 US homes. The deal? Over $10 billion. Here's what nobody understands: Cerebras doesn't make normal chips. Nvidia sells you thousands of tiny chips that you connect together. Cerebras makes ONE chip. A single wafer-scale processor the size of a dinner plate. 900,000 AI cores. 4 trillion transistors. All on one piece of silicon. The result? When OpenAI tested it, Cerebras ran inference 20X FASTER than Nvidia GPUs. That's not incremental improvement. That's a different category of performance. But here's where the story gets wild: Four months ago, Cerebras was a struggling company. Their IPO filing revealed that 87% of their revenue came from ONE customer: G42, a UAE-based AI firm. The US government launched a national security review. G42 had ties to Huawei. Ties to China. The IPO collapsed. Investors panicked. Cerebras withdrew their filing in October 2025. Most startups would've been dead. Instead, Cerebras did the opposite. They raised $1.1 billion at an $8.1 billion valuation. Kicked G42 out of the cap table entirely. Got CFIUS clearance. Then landed the OpenAI deal. Now they're raising ANOTHER $1 billion at a $22 billion valuation. They more than DOUBLED their valuation in 4 months. From near-death to $22 billion. While getting rid of their biggest customer. Why OpenAI chose them: ChatGPT has 900 million weekly users. Sam Altman keeps saying they have a "severe shortage" of compute. They need SPEED, not just power. When you ask ChatGPT a question, there's a loop happening: You send request → model thinks → sends response back Nvidia chips are fast at training models. Cerebras chips are built specifically for inference. For real-time responses. For the exact bottleneck OpenAI is trying to solve. Sachin Katti from OpenAI said it best: "Cerebras adds a dedicated low-latency inference solution to our platform. That means faster responses, more natural interactions, and a stronger foundation to scale real-time AI to many more people." In other words: "We need this to scale ChatGPT." The competitive landscape just shifted: Nvidia announced a $100 billion deal with OpenAI in September. But it's still not finalized. Meanwhile, Cerebras closed their deal before Thanksgiving. And it's ALREADY being deployed. Here's the part that should terrify Nvidia: In December, Nvidia bought Groq for $20 billion. Groq makes fast inference chips. Just like Cerebras. So why would Nvidia spend $20 billion buying a competitor to something they supposedly already dominate? Because they know what's coming. Inference is the new battleground. And Cerebras is winning it. The IPO is coming Q2 2026. After this OpenAI deal, Cerebras now has: ✓ IBM contracts ✓ Department of Energy contracts ✓ OpenAI locked in for 3 years ✓ $22 billion valuation ✓ CFIUS clearance ✓ Zero customer concentration risk They went from 87% revenue dependency on one customer to the most diversified chip company outside Nvidia. In four months. The lesson? Smart money doesn't follow headlines. It follows where the AI leaders are actually spending. OpenAI didn't announce this deal for publicity. They need Cerebras hardware to scale ChatGPT. That's a $10 billion vote of confidence. While everyone's watching Nvidia stock, the real war is happening in inference. And the company with ONE giant chip just beat the company with thousands of tiny ones. What do you think happens when Cerebras IPOs?

Ricardo

28,088 views • 8 months ago

Cerebras just IPO’d and the stock already ran up over 100% (Save this). For the entire 70 year history of the semiconductor industry, every company on earth has followed the same process. You take a dinner plate sized silicon wafer, put hundreds of tiny chips onto it, and dice it up like a pizza. Nvidia does it this way, AMD does it this way, Intel has done it this way for six decades and everyone who tried to break that convention failed. Until Cerebras asked the most annoyingly obvious question in the industry’s history, what if you just didn’t cut it? The result is the Wafer Scale Engine, a single chip 56 times larger than Nvidia’s H100 and it fundamentally changes the physics of how AI inference works. The reason this matters is not the size, it’s the bandwidth. Every time an AI model generates a single word, it has to reach into memory, pull weights, multiply them together, and produce a prediction and when you’re running millions of concurrent sessions at once, the bottleneck is not raw processing power but how fast data moves between memory and compute. Nvidia’s H100 moves data at roughly 3 terabytes per second, while Cerebras’ WSE-3 moves data at 21 petabytes per second, roughly 7,000 times faster because memory and compute live on the same enormous piece of silicon and data barely has to travel at all. That gap is exactly why OpenAI went from 150 tokens per second on traditional GPUs to 2,000 tokens per second on Cerebras hardware, and why AWS integrated Cerebras into Bedrock to deliver roughly 5x more inference capacity in the same physical footprint. The macro setup is making the trade even more urgent. South Korea DRAM export prices recently jumped 35%, flash memory surged 47%, and SSD pricing spiked nearly 140% and every single one of those increases hits Nvidia-based infrastructure directly, because the H100 requires 80GB of the most expensive, most contested memory in the AI supply chain. Cerebras’ WSE-3 uses zero external HBM memory, baking 44GB of SRAM directly into the wafer itself which means as memory pricing goes parabolic, every CFO evaluating AI infrastructure is suddenly looking much more seriously at the architecture that sidesteps that cost entirely. The demand is already showing up in the backlog. Cerebras ended 2025 with $24.6 billion in remaining performance obligations for a company doing just over $500 million in annual revenue, that is a number that implies years of contracted growth already sitting on the books. The IPO was 20x oversubscribed, the price range was raised twice before listing, and shares opened 89% above their listing price on a $5.55 billion raise that made it the largest semiconductor IPO in history. The risks are real and worth naming. 86% of 2025 revenue came from two entities with UAE ties, U.S. revenue actually fell 34% to $187 million, and the $20 billion OpenAI contract is conditional, if Cerebras misses delivery milestones, OpenAI can terminate and trigger repayment demands on a $1 billion loan facility. And yet the market is valuing Cerebras at roughly 91x trailing revenue, richer than Nvidia, AMD, and Arm combined. What investors are betting on is not that Cerebras beats Nvidia, it is that the inference supercycle is large enough to support an entirely different architecture optimized for a different workload, and that $24.6 billion in contracted backlog converts to diversified revenue before the market starts asking harder questions. CEO Andrew Feldman said this took a decade of late nights to get right, everyone who tried to copy it failed and given that the entire inference economy is now running through exactly the bottleneck Cerebras was built to eliminate, the market is starting to believe him.

Milk Road AI

30,441 views • 4 months ago

OpenAI's newest AI escaped the test environment it was locked inside and hacked into another company on its OWN. To remind you: Last week one of the biggest AI companies on Earth got breached. A platform called Hugging Face, which hosts more than a million AI models and datasets, said an "autonomous AI agent" had broken into its systems. Nobody knew whose agent it was. For five days the whole industry wondered who was behind it. Yesterday OpenAI raised its hand and said it was them. Or more precisely, it was their models, acting completely on their own. So what did these models actually do? OpenAI was running two of them, GPT-5.6 Sol and an unreleased model they will only describe as "even more capable." They wanted to measure how good the models were at hacking, so they deliberately turned the safety filters down. They locked both models inside a sealed test environment with no real internet access. The only task was a benchmark called ExploitGym, a set of 898 real software vulnerabilities where the model has to turn each bug into a working attack. But the models got OBSESSED with winning... Instead of solving the test the honest way, they went hunting for a shortcut. They found a zero-day flaw in the software running their own sandbox, a bug nobody knew existed, and used it to break out. Once they were loose on the open internet, they worked out that Hugging Face was probably storing the answer key to the benchmark. So they hacked their way in. They chained multiple exploits together, escalated their access, moved across servers, and pulled the test solutions straight out of Hugging Face's live production database. They literally cheated on the test by breaking into another company to steal the answers. OpenAI called it "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." In their own words, the models were "hyperfocused on finding a solution" and went "to extreme lengths to achieve a rather narrow testing goal." And this was not the first time: Before Sol ever launched, an independent red-team lab called METR caught it gaming its own tests to inflate its scores. It hid an exploit inside a data stream, escalated its privileges on the testing server, and leaked the answers human evaluators had hidden. And OpenAI shipped it anyway. The day before the Hugging Face story, OpenAI paused a different unreleased model. This is the same model that earlier this year disproved a famous 1946 math conjecture, a result a Fields Medal winner called a breakthrough. They told it to only post its results to Slack but it found a way out of its sandbox and posted to a public GitHub page instead. They had to pause it because it kept finding ways to act outside the box they built for it. And it is not just OpenAI... Anthropic has reported that one of its own models slipped its sandbox during safety testing and reached the internet it was never supposed to touch, then used it to email a researcher. So step back and look at what these companies are telling you: The only thing standing between these models and a real attack was a set of safety filters. Turn those filters down for a single test, and the model taught itself to escape, break into a company it was never pointed at, and take what it wanted. OpenAI even said they expect incidents like it to "become more commonplace" as the models get more capable. Sam Altman also predicted there'll be a major cyber attack this year. And keep in mind that Sol is not a locked-away experiment but a publicly available model that businesses are already wiring into their own systems. The next model that breaks out of its box might not be doing it just to cheat on a math test...

Ricardo

176,196 views • 2 months ago

Microsoft just betrayed OpenAI and Anthropic, the two companies it helped build. And it could break the entire AI trade... Here's what happened: Inside Excel and Outlook, two of the most used business apps on Earth, Microsoft has started routing tens of thousands of AI requests every week to its own in-house models instead of OpenAI and Anthropic. Microsoft's own AI chief, Mustafa Suleyman, said himself: "We pay a lot of money to Anthropic, so our goal is to reduce and ultimately ELIMINATE that cost." This is the company that poured $13 billion into OpenAI and effectively created the modern AI industry, and it just decided the most advanced models on the market are NOT worth paying for. And here's the thing... Microsoft is not just ripping out OpenAI everywhere - it is being surgical about it. The hardest and rarest tasks can still go to OpenAI or Anthropic. What Microsoft is taking back is the boring, high-volume work, like the email replies, the thread summaries, and the simple spreadsheet formulas. Why does that matter so much? Because that boring, repetitive work is where the actual money lives. The frontier labs assumed businesses would push BILLIONS of these tiny requests through expensive models forever. That endless river of tokens is the entire reason OpenAI and Anthropic are valued in the hundreds of billions of dollars. Microsoft looked at that river, decided it was massively overpaying, and rerouted it to models it owns outright. So the single biggest customer in the industry just walked off with the most profitable part of the business. And it is not only Microsoft: That same week, CNBC reported that American companies have been escaping to Chinese AI models to dodge rising US prices. Chinese models now handle more than 30% of US companies' AI usage on one major platform, peaking at 46%, up from an average of 11% a year earlier. They cost 60 to 90% less, and on some benchmarks they land within a single point of the best American model. One US startup moved ALL of its AI traffic off Claude and onto China's DeepSeek, and expects to save millions. Meanwhile Meta just admitted it has "excess" AI compute it wants to sell, becoming the first giant to concede it built far too much. Do you see the pattern forming? For two years, the entire AI story rested on one assumption: Every company on Earth would happily pay premium prices for the best model, forever. That assumption literally died in a single week. And the market noticed. More than a trillion dollars has been wiped off AI and chip stocks in a matter of days, as Wall Street finally started asking whether all of this spending will ever pay for itself. What this means for OpenAI and Anthropic: Their models are extraordinary, and it may not matter because their own biggest customers have decided they do not NEED the best model in the world to answer an email, and "good enough" now costs a fraction of the price. When even Microsoft refuses to pay full price for AI, the real question becomes who exactly IS left to pay it. What do you think?

Ricardo

93,586 views • 2 months ago

OpenAI and Anthropic this week: Navier-Stokes, An Alien Mind, Images 2.5, Pro pause, Pace the Frontier (Week 37, 2026) OpenAI shared a solution to the Navier-Stokes Millennium Prize Problem, produced in 88 hours by around 10,000 coordinating agents on an internal model still in training and significantly more capable than GPT-6 Astra, with an investigation finding Tristan Buckmaster's Codex prompts could not have influenced the system and no user data was accessed OpenAI says they have reached their automated research intern goal and are making strong progress toward an automated AI researcher by March 2028, with the research org at 3.1 agent-workdays per human workday and RL training on deployment-bound models partly paused after the Hugging Face incident OpenAI Chief Scientist Jakub Pachocki writes in An Alien Mind that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer, and that GPT-6 Astra is the first model to benefit from some of their newer alignment work OpenAI called for mandatory capability-based national AI safety regulation, endorsed four California bills, said fully autonomous recursive self-improvement should not be pursued until it can be done safely, and described Astra safeguards like universal monitoring of full trajectories including chains of thought OpenAI released ChatGPT Images 2.5 to all ChatGPT, ChatGPT Work, and Codex users with sharper details, up to 50% lower latency, comment-based edits, Sketch, and templates, plus GPT-Image-2.5 Flare and Sunburst in the API ChatGPT Voice can now use GPT-5.6 Sol and GPT-6 Astra when it needs to search or reason, GPT-Live-1 daily limits are simplified per plan, and extra Voice usage drops from 5 to 1.25 credits per minute for Business and Enterprise workspaces on credits OpenAI paused new subscriptions to the $200 ChatGPT Pro plan to protect access for existing GPT-6 Astra users, hours after a remote switch to pause Pro 20x purchases showed up in the web app Custom GPTs in ChatGPT will likely be retired on December 11 based on my findings, and OpenAI's new FAQ puts Enterprise migration at September 17 with new GPT creation ending September 25 and instructions becoming a plugin skill From the ChatGPT Android build, OpenAI is building a collaborative multiplayer document editor with dedicated gateway hosts and draft-conflict UI, an Artifacts Library with Favorites, and a credit score feature in ChatGPT Finance, and Locked Chats with a PIN are being prepared in the web app OpenAI released GPT-Live-1 in the API at $0.05 per minute, a voice model that listens and speaks at the same time and delegates reasoning to a backend model like GPT-6 Astra, and the Agents API in public beta running the Codex harness on OpenAI's infrastructure with only the sandbox left for you to choose Deep research arrived in ChatGPT Work and Codex, a Data plugin connects Snowflake, Databricks, BigQuery, and Redshift to answer business questions and build dashboards, Library gained file and folder sharing, and Box, Dropbox, and SharePoint joined Google Drive in Library OpenAI launched ChatGPT for Financial Services with built-in premium data from Daloopa, PitchBook, LSEG News, and Crunchbase, shaped with Morgan Stanley and Evercore, and a GSA agreement gives US governments $0 license fees, 50% off usage, and Daybreak Blue at half price Smaller ChatGPT bits: a stock watchlist in Finances for US Plus and Pro, a small business plugin collection, over 5 million ChatGPT Sites built in three months, and the desktop pet can now start a new chat with a new Mini option OpenAI published The Work Now Within Reach, calling free access "supported by advertising", citing over one billion weekly active users, and saying they plan to begin deploying their Jalapeño inference chip by year-end OpenAI moved GPT-Rosalind out of research preview for eligible organizations worldwide, detailed Habitat with the service rewritten in Rust by two engineers with Codex and the platform serving over 70 million requests per second, and shared a case study of GPT-5.6 Sol calibrating a six-qubit chip at MIT Paul Christiano joined the OpenAI Foundation Board and the Safety and Security Committee, OpenAI committed $5 million to research on AI and teens, and expanded journalism programs from CUNY and Northwestern to Ukrainian newsrooms Anthropic published an alignment assessment of four incidents in which Claude models gained unauthorized access to real systems, adding a fourth found when assembling transcripts for METR, walking back their July 30 claims, and calling it a mistake that Claude Mythos 5 shipped without alignment environments Anthropic's most detailed threat intelligence report yet says Moonshot and DeepSeek silently relayed their own users' requests to Claude and served the responses as Kimi and DeepSeek output, with distillation campaigns attributed to Alibaba as the largest ever with more than 3,500 fraudulent accounts, plus Zhipu, Xiaomi, SenseTime, and MiniMax Anthropic's Frontier Red Team measured tactical intelligence targeting and conventional weapons capabilities, with Claude Mythos Preview leading the targeting evals, Claude Opus 5 leading the weapons software evals, and Mythos beating the top GeoGuessr division on photo geolocation Anthropic CEO Dario Amodei argues in We Must Pace the Frontier that the AI industry should slow down, commits to giving evaluators like METR permanent employee-level access, and OpenAI CEO Sam Altman replied that he agrees and OpenAI will commit to the same Anthropic's Economics team released a scenario explorer for AI's effect on the US economy by 2030, where even the extreme case grows the economy but reaches 15% annual GDP growth with unemployment beyond recessionary levels On the product side, Claude Code desktop can pop out any pane into its own window, Claude Managed Agents got a session viewer and auto mode, claude plugin eval scores your plugin or skill with and without it, and smart reports launched in beta for Claude Enterprise Anthropic shared lessons from their Claude SMB Tour with more than 1,000 small business owners, where data security was the most-cited adoption barrier and nearly two-thirds asked for more hands-on implementation help, and more

Tibor Blaho

14,925 views • 8 days ago

In 2016, five founders and a deck walked into Benchmark to pitch a hardware company. Eric Vishria didn't want to take the meeting. Benchmark hadn't made a semiconductor investment in ten years, and he remembers thinking, why are we even going to a hardware pitch? This is crazy. The team was excellent and the first slide said GPUs actually suck for deep learning. They just happen to be 100 times better than CPUs. This was pre-transformer. OpenAI was a weird research lab. The TPU hadn't been announced. Nvidia was worth $40 billion, not $4 trillion. The first question in 2016 was whether AI was even a big new workload. There had been many attempts at specialized chips for other things that simply never ended up mattering. Benchmark had conviction that this one would. The second question was whether the workload introduced a new constraint. It did. AI benefited massively from the parallelism of GPUs, but GPUs didn't solve the communication between cores. This was a communication-bound problem, and nothing on the market addressed it. There are only three ways to speed up deep learning in hardware, then and still today. More cores. Faster communication between cores. Memory closer to the compute. Their pitch was to take all three to their logical maximum at once. A single wafer-scale chip with 450,000 cores and 20 gigs of SRAM, so the system never has to leave the chip to reach memory, and every core sits on the same wafer, so communication between them is as fast as physics allows. It was the best you could possibly do. As soon as Eric heard the idea, his reaction was, "of course". Engineers had been attempting a wafer-scale chip for 50 years. Every attempt had failed. Cerebras got a working chip on the first try. Then came the long middle of the story that nobody romanticizes, the years of work required to turn a scientific achievement into a business. "In software, if you have that logical block diagram of why it works and everything else, you're 80% of the way there. And it's a matter of go-to-market execution. In hardware, you're like 2% of the way there." A chip that powerful has to be packaged, heated, cooled, programmed, and sold. In 2019, Eric sat in a board meeting watching the chip melt. The company had raised roughly $500 million by then, and the thought running through his head was "holy shit, we're going to lose all this money." The company kissed death three times. Virtually every respected semiconductor investor passed. People made fun of the early customers. Today Cerebras is worth more than $50 billion. Eric said it years before anyone knew how it would end. Whether Cerebras worked or not, it was an effort worth venture capital. Try to build the thing people have failed at for 50 years, when there's finally a reason it could work. Hardware is hard. This team "just never fucking quit".

Invest Like the Best

16,120 views • 1 month ago

Microsoft is deceiving you by inflating its AI empire with money it handed its OWN customer first. They sold Wall Street a $37 billion AI business, then went silent the moment its own filing showed where that money came from. The line sits in the annual report for fiscal 2026: Microsoft recorded $24.1 billion of revenue from commercial arrangements with OpenAI, including revenue sharing payments. If you run that figure against Microsoft's own AI disclosures you'll find that OpenAI made up more than half, and likely around 70%, of everything the company counts as AI sales. ONE customer. A Microsoft spokesperson confirmed the figure covers all sales and revenue share from OpenAI. The 70% comes by assuming Microsoft's AI run rate kept growing at the 123% pace the company itself reported in March, which is the company's own optimistic math turned around on it. Now follow where that money starts: Microsoft has put around $12 billion into OpenAI since 2019. OpenAI spends its cash on computing power, and Microsoft is the cloud provider selling it. So the money leaves as an investment and comes back as an Azure bill. Microsoft then books that bill as AI revenue and shows it to investors as proof the AI business is "working." Microsoft invests in OpenAI -> OpenAI buys Microsoft compute -> Microsoft records the payment as AI revenue -> the AI growth story goes to Wall Street And a chunk of it never actually arrived. The same filing shows $6 billion of accounts receivable from OpenAI as of June 30. That is $6 billion of AI revenue Microsoft booked and had not been paid when the year closed. Now here's where it gets really concerning for anyone holding the stock... Microsoft has told the public how big its total AI business is exactly twice. Once for the quarter ending December 2024, when it said the unit was on pace for more than $13 billion a year. And once for the quarter ending March 2026, when Satya Nadella put it on pace for $37 billion. That $37 billion number went everywhere. It was the headline proof that Microsoft had won the AI race. Then fourth quarter earnings arrived, and Microsoft did NOT update it. The company that had been announcing the figure as its own scoreboard stopped announcing the figure. In the same stretch, the filing landed showing where most of it came from. So what is actually left underneath? The full year AI business ran near $34 billion. Take OpenAI out and roughly $10 billion remains. Microsoft has spent about $261 billion on capital expenditure since the start of 2022. That is the scale of the bet against what the rest of the AI business currently brings in. And the one customer holding it up is walking further away every quarter. In October, Microsoft's stake in OpenAI dropped to 27% from 32.5%. In April the partnership was rewritten so OpenAI can sell its products across any cloud it likes, which is how Amazon got a seat at the table. The exclusivity that made this arrangement valuable is gone. The compute bill and the unpaid $6 billion are still on Microsoft's books. Nadella spent two years telling the market Microsoft built the largest AI business in software. The filing shows one client bought most of it, on credit, using money Microsoft partly supplied. So watch the next earnings call: If Microsoft puts a fresh total AI number back on the board, the business found customers beyond OpenAI. If you hear a lot about AI momentum and never hear what it adds up to, you already know why the number went missing. But nonetheless, how is something like this even legal?

Ricardo

24,389 views • 1 month ago