Loading video...
Video Failed to Load
Y Combinator built AI versions of its partners to help more people work through their startup ideas. For its Office Hour Simulator, the team needed useful answers delivered quickly enough for a spoken conversation. They had been testing lightweight Gemini and OpenAI models before moving to GLM-5.2 on a... show more
57,298 views • 5 days ago •via X (Twitter)
13 Comments

@ycombinator great partnering with the folks at YC on this 🫶

@ycombinator so true 🫶

Article:

@ycombinator did someone say faster than cerebras??

@ycombinator :o 👀

@ycombinator 🫶

@ycombinator 🧇🫶

@ycombinator inference for real-time ai = wafer btw

@ycombinator facts

@ycombinator The 2.5 minutes longer is the number that matters, and it's the one nobody optimises for directly. In a spoken conversation latency isn't a performance stat, it's whether the thing feels like it's listening or thinking about something else.

@ycombinator Latency wins the conversation. Always has.

@ycombinator For spoken Office Hours, latency becomes part of answer quality. A lighter model can feel better when it responds fast enough to keep the conversation moving.

watch the full video here:
