正在加载视频...

视频加载失败

We've run thousands of LLM inference serving benchmarks at Modal. We're releasing the results so you don't have to. We're releasing the code so that you can. Introducing: The LLM Engineer's Almanac. Just in time for the AI Engineer 🔜 Paris 🇫🇷 World's Fair.

23,037 次观看 • 1 年前 •via X (Twitter)

11 条评论

Charles 🎉 Frye 的头像
Charles 🎉 Frye1 年前

tl;dr: go to to explore our benchmarking results for popular open weights models run on the @vllm_project, @lmsysorg SGLang, and @NVIDIA TensorRT-LLM frameworks in their out-of-the-box configuration. With code so you can deploy!

Charles 🎉 Frye 的头像
Charles 🎉 Frye1 年前

We distilled our learnings from this work and from discussions with teams serving LLMs into an executive summary designed to answer the questions technical leaders are asking: - when do I build vs buy LLM inference? - how do I estimate cost? - which framework should I use?

Charles 🎉 Frye 的头像
Charles 🎉 Frye1 年前

On that last question: our results indicate that, absent any tuning, vLLM and SGLang achieve comparable performance. Other factors, most of them non-technical, will end up dictating your decision. We like and use them both!

Charles 🎉 Frye 的头像
Charles 🎉 Frye1 年前

TensorRT-LLM is a whole 'nother beast. There's not really such a thing as out-of-the-box configuration with TRT-LLM -- the defaults are not intended to be run in production. So we get worse performance there, even though tuning can lead to superior perf.

Charles 🎉 Frye 的头像
Charles 🎉 Frye1 年前

(more details and nuance in the executive summary)

Charles 🎉 Frye 的头像
Charles 🎉 Frye1 年前

All benchmarks are wrong, but some are useful. They are more useful if you can read about the methodology and understand its assumptions, its purposes, and its limitations. So we wrote that up too!

Charles 🎉 Frye 的头像
Charles 🎉 Frye1 年前

Of course, words only get you so far. And benchmarking on Modal can be highly productive, due to the high parallelism you can achieve. So we're also releasing the framework we wrote, stopwatch. Kudos to @jackcookjack for some incredible work on this.

Charles 🎉 Frye 的头像
Charles 🎉 Frye1 年前

Stopwatch is built on top of great work by the team from Neural Magic (acq. @RedHat_AI). Their framework is called guidellm, and if you want to run LLM engine benchmarks outside of Modal, you should check it out!

Charles 🎉 Frye 的头像
Charles 🎉 Frye1 年前

These are just the first few entries in the LLM Engineer's Almanac, our follow-up to the popular GPU Glossary. We hope to build multiple useful tools and documents for engineers building LLM applications and share them here. Let us know what you'd find most useful!

Charles 🎉 Frye 的头像
Charles 🎉 Frye1 年前

We're incredibly grateful to everyone who gave us feedback as we worked on this project. On Twitter (follow them!): @mgoin_, @0xishand, @moinnadeem, @cdpierse, @NikhilKMurthy, @sgdescent Not on Twitter: Yineng Zhang of SGLang

Rainmaker 的头像
Rainmaker2 年前

Can Machine Learning beat the market? Check out this post on my free Substack where I share code and commentary for an XGBoost model and a Random Forest model that both deliver powerful performances.

相关视频