
BentoML
@bentomlai • 2,577 subscribers
🍱 Inference Platform built for speed and control. Acquired by @Modular Join our community 👉 https://t.co/ZrgHxWh8AH
Videos

🚀 OpenLLM 0.6 is out providing fast inference speed! Check out our video comparing #OpenLLM and #Ollama handling concurrent requests on the Llama 3 8B model. Ollama is ideal for local LLM deployment, but it is not designed for high concurrency scenarios essential for deployments in the cloud, powering applications with a large number of users. OpenLLM enables you to run any open-source model as OpenAI-compatible endpoints with just a single command. It supports a broad array of models like Llama 3, Mistral, and Qwen 2, including those fine-tuned with proprietary datasets and quantized versions. Ready to speed up your AI deployments? See the OpenLLM repo #BentoML #OpenSource
BentoML18,798 次观看 • 2 年前
没有更多内容可加载