Загрузка видео...

Не удалось загрузить видео

На главную

Open source/weight models are often used in regulated industries like Health Care or Financial Services, where they handle personally identifiable data, and can't send it to proprietary LLM providers. We recently chatted to Vaibhav (VB) Srivastav about the partnership VS Visual Studio Code and Hugging Face inference providers have...

18,215 просмотров • 10 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

learned a lot from this conversation with Simon Mo and Matt Bornstein. biggest takeaways for me: -there are a lot of reasons why we should like open-weight models. a lot of these arguments stop at handwavy things like "what if the labs stop releasing frontier models to the public" or "it's lower cost." but simon's position as lead maintainer of vLLM and CEO of Inferact give him authority to talk about some of the other, more interesting and concrete reasons to pay attention to open-weight models, namely that they allow end-users to calibrate latency / other performance metrics with way more customizability than what any of the frontier closed-source labs offer (and without the fear that your job might be met with a refusal at some random point where you're deep in a 2 hour job) -re: the above point...for this reason, a lot of US companies (inferact included!) choose to use open-weight models over their closed-source alternatives. this also isn't limited to internal workloads / research - on a recent a16z podcast the team at Decagon spoke about how something like 90% of their customer service ai agents run on open-weight models that they've fine-tuned. -we should really appreciate how many companies/teams came out researchers fascinated by the wave of very small open-weight models that were being distilled from e.g. gpt-3.5 and earlier models in 2022/2023 (prior to the release of chatGPT!). these small models motivated the development of pagedattention, which then led to vlmm/inferact (at other layers of the stack with similar origin stories, you can look at teams like openrouter or ollama). in other words, we have open-weight models to thank for a bunch of the orchestration infra we now rely on. i think yet another, indirect, way we can point to open-source/weight infra pushing the frontier forward. anyway, a lot more in this convo, it was a lot of fun!

Elena

12,922 просмотров • 1 месяц назад

.Josh Wolfe: Anybody Using DeepSeek App Is 'Absolute Fool' "Anybody using the DeepSeek app is an absolute fool. If you're using DeepSeek on companies like Together Compute, one of Lux's companies, which can get rid of the CCP censorship, then it's probably okay. But remember, the open-source movement is something we deeply believe in. Most great technologists, entrepreneurs, and venture capitalists are on the side of open source. The closed-source models that have consumed tens of billions of dollars are the ones that are really going to be at risk. When you look at Hugging Face, a major repository, or Together Compute, Runway ML, and a lot of Lux's companies, they have been pioneers in open source. Now, why am I not worried about open source, even with the DeepSeek model? As long as you don't have the CCP censorship on it, the models with their open weights allow people to run on their proprietary data. This means companies like pharma or defense companies that have their own siloed, proprietary data—think about Bloomberg with their proprietary longitudinal data, or Meta with their data—are the ones who will have the edge. Even as open source takes hold, these companies will still dominate. I’m not worried about open source being the problem. I’m more concerned about people overfunding closed models with no proprietary source. A lot of capital is going to be burned there, and we’re already seeing that with people worried about OpenAI in some aspects."

Josh Caplan

40,039 просмотров • 1 год назад

NOBODY wants to send their data to Google or OpenAI. Yet here we are, shipping proprietary code, customer information, and sensitive business logic to closed-source APIs we don't control. While everyone's chasing the latest closed-source releases, open-source models are quietly becoming the practical choice for many production systems. Here's what everyone is missing: Open-source models are catching up fast, and they bring something the big labs can't: privacy, speed, and control. I built a playground to test this myself. Used CometML's Opik to evaluate models on real code generation tasks - testing correctness, readability, and best practices against actual GitHub repos. Here's what surprised me: OSS models like MiniMax-M2, Kimi k2 performed on par with the likes of Gemini 3 and Claude Sonnet 4.5 on most tasks. But practically MiniMax-M2 turns out to be a winner as it's twice as fast and 12x cheaper when you compare it to models like Sonnet 4.5. Well, this isn't just about saving money. When your model is smaller and faster, you can deploy it in places closed-source APIs can't reach: ↳ Real-time applications that need sub-second responses ↳ Edge devices where latency kills user experience ↳ On-premise systems where data never leaves your infrastructure MiniMax-M2 runs with only 10B activated parameters. That efficiency means lower latency, higher throughput, and the ability to handle interactive agents without breaking the bank. The intelligence-to-cost ratio here changes what's possible. You're not choosing between quality and affordability anymore. You're not sacrificing privacy for performance. The gap is closing, and in many cases, it's already closed. If you're building anything that needs to be fast, private, or deployed at scale, it's worth taking a look at what's now available. MiniMax-M2 is 100% open-source, free for developers right now. I have shared the link to their GitHub repo in the next tweet. You will also find the code for the playground and evaluations I've done.

Akshay 🚀

50,323 просмотров • 10 месяцев назад