Loading video...
Video Failed to Load
Try now Hyperfast LLM running on custom built GPUs Answers in miliseconds, not seconds How? 🤯
602,378 views • 2 years ago •via X (Twitter)
10 Comments

Via @mattshumer_

Hyper-optimizing a particular program for a TPU can yield 100×+ speedup. Speaking from experience

"An LPU Inference Engine, with LPU standing for Language Processing Unit™, is a new type of end-to-end processing unit system that provides the fastest inference for computationally intensive applications with a sequential component to them, such as AI language applications (LLMs)."

Answer: Custom hardware and great software (I work at Groq)

Seriously impressed. At this speed I think it's possible to build a human-like conversation experience where the AI can even interrupt you while you're speaking.

Not affiliated btw, it's made by @JonathanRoss321

waiting for someone to get confused between grok and groq

@levelsio No confusion, we trademarked our name in 2016.

They developed their own hardware that are built using LPUs instead of GPUs!

This is insane. What are some novel use-cases now possible because of this speed?
