Loading video...

Video Failed to Load

Go Home

.Decagon co-founder and CTO Ashwin Sreenivas says the choice between expensive frontier models and cheaper, less capable ones is a false trade-off: "Even if you have a 'dumber model,' you can get it to higher performance on that specific task." "When we fine-tune smaller, dumber models, it's that they're...

36,050 views • 1 month ago •via X (Twitter)

9 Comments

Stanley Wei's profile picture
Stanley Wei1 month ago

@DecagonAI Task-specific smaller models challenge the idea that ‘frontier’ is a single category. The winning architecture may be a portfolio: cheap specialists by default, expensive generalists for ambiguity, and routing as the product brain.

SheetalGupta's profile picture
SheetalGupta1 month ago

@DecagonAI I agree with @DecagonAI . The way forward is more and more nuanced models. Most of the startups now are focused on building these (e.g. Neurovians)

Andrii Bidochko 🦉's profile picture
Andrii Bidochko 🦉1 month ago

This is the exact architectural shift the enterprise is currently undergoing. The era of hitting a massive, generalized model for every single workflow is ending. When you fine-tune a smaller model for a narrow, well-defined task, you aren't just saving money and reducing latency - you are actually building a more accurate system. Precision beats generalized reasoning in production every time.

踏空 Sidelined Capital's profile picture
踏空 Sidelined Capital1 month ago

@DecagonAI The hidden bill is maintaining one specialist per workflow when the workflow changes every quarter.

Phoenix Shield's profile picture
Phoenix Shield1 month ago

@DecagonAI ❤️💞

Kate Ivanova's profile picture
Kate Ivanova1 month ago

@DecagonAI Interesting perspective. Specialized models optimized for specific workflows could become a major advantage over general-purpose approaches.

Minh Do's profile picture
Minh Do1 month ago

@DecagonAI Serious question: Why don’t you guys decorate your mics with art deco too? Seems like a missed opportunity given everyone in this industry has the same mics.

Oleksandr's profile picture
Oleksandr1 month ago

@DecagonAI task specific boost. finetuned small models beat big ones on niche jobs cheaper, faster, and actually better for the job

Aipepe's profile picture
Aipepe1 month ago

@DecagonAI Immigration might be the "meta-issue" for them, but AI Pepe is the real meta for Aptos. Join the movement! 🚀 ​#Aptos #AIPepe #Memecoin #Crypto #Web3

Related Videos

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 views • 2 years ago

OpenAI chairman Bret Taylor talks to about 100 CEOs every month. His answer to the cheap open-weight model panic: cheaper to train does not mean cheaper to use, and the number that decides it is token efficiency. "One thing that I think is a little bit overblown about these open weight models is they're not necessarily cheaper to run. Whether or not they're cheaper to train, you don't care. Because you're using just as many tokens. In fact, they may be less efficient." "There's this thing called token efficiency. And it turns out the frontier models are much, much more token efficient." "A token is to intelligence like a watt is to electricity... how many tokens does it take to complete a task? Not every token is actually equal." "For a lot of tasks, it turns out these frontier models from OpenAI and Anthropic are actually just better than these open weight models... just having open weights isn't actually the main thing driving any of those costs." Later in the same interview he goes after the billing unit itself: "It would be like if you signed up for Gmail and you paid for CPU cycle or something... where the world is going is paying for outcomes." The unresolved column: the chart CNBC airs mid-answer, from Artificial Analysis, prices a completed task at $0.94 on Kimi K3 against $2.75 on Claude Fable 5, efficiency folded in. If that gap holds, the premium he is defending gets earned on quality, not price. - Bret Taylor (Bret Taylor), OpenAI chairman and Sierra co-founder, on CNBC's Squawk Box.

Karl Mehta

17,292 views • 2 months ago

Baseten Head of AI Model Training Charlie O'Neill says the future is many specialized LLMs dedicated to specific tasks, with bigger labs deployed on the frontiers of areas like science and math: "People are thinking about intelligence capabilities in the wrong way. People are thinking about intelligence relativistically. They say, 'OK, the open-source gap is like 6 months behind closed-source, and GLM 5.3 is as good as Opus 4.8,' or whatever." "The best way to think about what models can do for you, and for the world, is in an absolute sense." "So for any given task that you want to do with an LLM, there's some intelligence threshold where below that you can't do the task, and above that you have very diminishing returns to more intelligence on the task." "So when you think about it that way, the game of LLMs over the last 5 years has been, 'OK, we have these things we want to do with them. Closed source hits it first... but open-source can eventually do that task. And then for many reasons, once you have the base level of intelligence required to do it, you probably do want to swap to open-source." "It's not really about the [frontier lab] God model being better. Like, if I'm filing a tax return, there is a limit to how much intelligence I need to do that particular thing." "So I think the world is going to look like — frontier closed-source labs are going to continue to push the frontier. You do want to use the most intelligent model. You have very inelastic demand for intelligence when you're doing frontier science or frontier math." "But for a lot of the economically valuable things, it looks a lot like, 'I'm a Cursor, or I'm one of these big companies who are realizing I can't just be a wrapper anymore. I've been through the life cycle of building a product that people love. And I should be using that information to make my model better at the things that I care about, and not at anything else.'"

TBPN

50,531 views • 18 days ago

There is no best model. There's a lot of noise about models right now. Who is training them, who owns them, where legal intelligence should live. One question actually matters: what produces the best outcome for the legal task in front of you? That's how we decide things at Legora. We optimize for the end-to-end outcome on a legal task. The model is one layer of that system, not the system. Models are uneven and the frontier changes almost weekly. One model plans a long job well, another runs deep analysis across thousands of documents. Some have to be told exactly what to do, and some are fine with a vague brief. They all break in different ways. So our lawyers write evals and we test them with the Legora BAR, our benchmark for agentic reasoning. Every model takes every test, and the model that wins gets the work. We post-train when we know it buys our customers better performance on a specialized task. Training is a tool we reach for when it helps, nothing more than that. The intelligence that compounds sits in the orchestration layer. Precedents, review standards, client requirements. That knowledge has to stay editable, auditable and portable. In our system, a changed review standard is an edit that takes effect the same day, with no new model training required. No lawyer should have to worry about which model did the work, any more than they think about which chip is in their laptop. They should only care about the quality of the work. That's what we are focused on. If you want the engineering version of this argument rather than the CEO version, our CPO, Bryan Tsao, and CTO, Jacob Lauritzen, take it apart in the video below.

Max Junestrand

47,797 views • 4 days ago