4/ to achieve maximum memory efficiency, we quantize model... show more

Alexandr Wang
67,151 views • 12 days ago
We took a 30B model and split it in... show more

NVIDIA AI
761,258 views • 1 month ago
We’re releasing model weights for our 8B- parameter Dynamic... show more

AI at Meta
195,607 views • 1 year ago
🚀 Self-speculation brings 6.75x real speedup for LLM generation... show more

Pavlo Molchanov
66,604 views • 2 months ago
Day 11/90 of Inference Engineering How does vLLM work... show more

max fu
70,543 views • 1 month ago
The first natively trained 1-bit model: BitNet 2B. Trained... show more

Md Ismail Šojal 🕷️
43,786 views • 5 months ago
We are excited to introduce Mercury, the first commercial-grade... show more

Inception
1,916,613 views • 1 year ago
LLaDA (the first Large Language Diffusion Model) is *just*... show more

apolinario (poli)
82,599 views • 1 year ago
You don’t have to use one model (or one... show more

Workshop AI
962,487 views • 4 months ago
Our model continues to be aggressive an wet with... show more

Tim Buckley
33,959 views • 2 months ago
Day 12/90 of Inference Engineering What is chunked prefill... show more

max fu
29,417 views • 1 month ago
A peanut-sized Chinese model just dethroned Gemini at reading... show more

AlphaSignal
92,071 views • 4 months ago
A peanut-sized Chinese model just dethroned Gemini at reading... show more

Jafar Najafov
13,630 views • 4 months ago
OPEN ZOE VTUBER MODEL We prepared a Zoe Vtuber... show more

Monster Con is OUT NOW!
53,553 views • 1 year ago
I wanted to test a walten files model, by... show more

𝗘𝗟𝗙𝗭𝗚𝗔𝗠𝗘𝗥
165,214 views • 8 days ago
1/ Gemini 2.5 is here, and it’s our most... show more

Sundar Pichai
864,579 views • 1 year ago
the fact that i can take an image of... show more

Jan
131,884 views • 9 months ago
All the big language models under one roof for... show more

Shubham Saboo
299,259 views • 3 years ago
Many of you asked for code & weights for... show more

Physical Intelligence
441,382 views • 1 year ago