Loading video...
Video Failed to Load
What does it take to run AI mental wellbeing support at scale, safely? For Sword Health: a jump from 30B to 200B+ on Token Factory, dedicated NVIDIA Blackwell endpoints, and production tail latency cut from 20+ seconds to under 12.
355,285 views • 15 days ago •via X (Twitter)
3 Comments

punished car leak14 hours ago
у вас бутилка із очка торчить, товаріщ яндекс

Dangcing Chan14 hours ago
under 12 seconds is the real flex here

Dangcing Chan1 day ago
Latency is just one piece. Real safety at scale needs robust guardrails against hallucinations, not just faster GPUs. Otherwise you are just scaling up bad advice efficiently.




