Loading video...

Video Failed to Load

Go Home

Try now Hyperfast LLM running on custom built GPUs Answers in miliseconds, not seconds How? 🤯

602,378 views • 2 years ago •via X (Twitter)

10 Comments

@levelsio's profile picture
@levelsio2 years ago

Via @mattshumer_

Beff – e/acc's profile picture
Beff – e/acc2 years ago

Hyper-optimizing a particular program for a TPU can yield 100×+ speedup. Speaking from experience

Suhail's profile picture
Suhail2 years ago

"An LPU Inference Engine, with LPU standing for Language Processing Unit™, is a new type of end-to-end processing unit system that provides the fastest inference for computationally intensive applications with a sequential component to them, such as AI language applications (LLMs)."

ben sima's profile picture
ben sima2 years ago

Answer: Custom hardware and great software (I work at Groq)

Tony Dinh 🎯's profile picture
Tony Dinh 🎯2 years ago

Seriously impressed. At this speed I think it's possible to build a human-like conversation experience where the AI can even interrupt you while you're speaking.

@levelsio's profile picture
@levelsio2 years ago

Not affiliated btw, it's made by @JonathanRoss321

Justin / Get Impeached 2!'s profile picture
Justin / Get Impeached 2!2 years ago

waiting for someone to get confused between grok and groq

Groq Inc's profile picture
Groq Inc2 years ago

@levelsio No confusion, we trademarked our name in 2016.

Jay Scambler's profile picture
Jay Scambler2 years ago

They developed their own hardware that are built using LPUs instead of GPUs!

Krishiv's profile picture
Krishiv2 years ago

This is insane. What are some novel use-cases now possible because of this speed?

Related Videos