Loading video...

Video Failed to Load

Go Home

What does it take to run AI mental wellbeing support at scale, safely? For Sword Health: a jump from 30B to 200B+ on Token Factory, dedicated NVIDIA Blackwell endpoints, and production tail latency cut from 20+ seconds to under 12.

355,285 views • 15 days ago •via X (Twitter)

3 Comments

punished car leak's profile picture
punished car leak14 hours ago

у вас бутилка із очка торчить, товаріщ яндекс

Dangcing Chan's profile picture
Dangcing Chan14 hours ago

under 12 seconds is the real flex here

Dangcing Chan's profile picture
Dangcing Chan1 day ago

Latency is just one piece. Real safety at scale needs robust guardrails against hallucinations, not just faster GPUs. Otherwise you are just scaling up bad advice efficiently.

Related Videos