Loading video...

Video Failed to Load

Go Home

Many people are flat wrong about DeepSeek. You might think everyone freaked out about DeepSeek because the model is really good—which it is—or because it was from China, which is also true. But there's a more subtle reason: DeepSeek showed they could achieve those results using reinforcement learning instead...

105,589 views • 1 year ago •via X (Twitter)

31 Comments

Santiago's profile picture
Santiago1 year ago

To record this video, the team at @predibase gave me access to their fine-tuning platform. As far as I know, this is the first time we have access to an end-to-end platform for Reinforcement Fine-Tuning (RFT). Here is an introduction to RFT: And here is a link to the notebook with the code to fine-tune your models: Thanks to the @predibase team for helping with credits, being patient, and collaborating with me on the video.

Santiago's profile picture
Santiago1 year ago

The video is also on YouTube here:

Santiago's profile picture
Santiago1 year ago

RFT needs a way to verify answers, but it's not true that this is only possible in math or coding problems. (Also, remember, RFT is not a good fit for every use case. In fact, SFT is much better when you have a lot of data and people to provide feedback.) There are non-math, non-coding problems where you can still define a reward function. For example: • Text summarization: Your reward function could use a ROUGE score to measure the overlap between the generated summary and the reference. • Quality of dialog: You can measure coherence, relevance, and adherence to conversation norms and use that as a reward function. • Style transfer: You can build a reward function that evaluates whether the generated output matches the desired style. There are many more examples.

Van Tuan Bui's profile picture
Van Tuan Bui1 year ago

Exactly! DeepSeek’s real disruption wasn’t just model quality, it was how they got there. Skipping the costly supervised phase and leaning into reinforcement fine-tuning (GRPO) is a paradigm shift.

Santiago's profile picture
Santiago1 year ago

It is, yeah!

Shawn Chauhan's profile picture
Shawn Chauhan1 year ago

This approach could lower the barrier to entry for so many.

Matthias Götzke's profile picture
Matthias Götzke1 year ago

The question was if they copied anything from OpenAI. Some outputs indicated that they might have.

Santiago's profile picture
Santiago1 year ago

Honestly, I don't think that's "the question." Their contributions speak for themselves.

Pixel Philosopher's profile picture
Pixel Philosopher1 year ago

The potential for automation using reinforcement learning is fascinating.

Mehrdad Yazdani's profile picture
Mehrdad Yazdani1 year ago

bro this only works if you have verifiable rewards like math or coding problems. "alignment" requires human feedback still by definition.

Santiago's profile picture
Santiago1 year ago

Yes, you need a way to verify answers to define reward functions, but it's not true that this is only possible in math or coding problems. (Also, remember, RFT is not a good fit for every use case. In fact, SFT is much better when you have a lot of data and people to provide feedback.) There are non-math, non-coding problems where you can still define a reward function. For example: • Text summarization: Your reward function could use a ROUGE score to measure the overlap between the generated summary and the reference. • Quality of dialog: You can measure coherence, relevance, and adherence to conversation norms and use that as a reward function. • Style transfer: You can build a reward function that evaluates whether the generated output matches the desired style. There are many more examples.

Franck SN's profile picture
Franck SN1 year ago

they did SFT too.

Sinuhet's profile picture
Sinuhet1 year ago

Let us say I want fine tune LLM for specific version of a charting library Lightingchart or SciCharts where both have well described API & quite a few examples and LLM are not well trained in them. Shall I use RLT or SRT?

Tiny Meat Capital's profile picture
Tiny Meat Capital1 year ago

Definitely. But the R1-Zero phase is still quite vague from I/O data structuring point of view. If someone could figure out how they structured the training data for that step, and the context window improvements from Gemini Pro 2.5, it would yield a crazy open model

ryan yang's profile picture
ryan yang1 year ago

Spot on—DeepSeek’s real edge? Incremental gains > hype. Layer by layer, measure, adjust, repeat. Legacy systems don’t need overhauls, just smart integration. ROI climbs when you iterate, not leap.

Michal Takac's profile picture
Michal Takac1 year ago

Wonder how RF fine-tuning can be used in the background for automating model improvements for my apps.

Michael Kurz's profile picture
Michael Kurz1 year ago

Perhaps they used existing models like ChatGPT instead of people? 😉

Santiago's profile picture
Santiago1 year ago

They didn't.

Karl Mehta's profile picture
Karl Mehta1 year ago

Excellent insight! The how (RL fine-tuning) is indeed more disruptive than just the performance numbers

Subba Reddy's profile picture
Subba Reddy1 year ago

good part: RFT in Action: See how reinforcement fine-tuning and reward functions guide model training concerns: - Llama 8B RFT was compared to Gemini 1.5 pro - good comparison should be Qwen2.5 RFT with Gemini 2.5 pro to reflect model advancements

Jan Tomášek's profile picture
Jan Tomášek1 year ago

Deepseek R1 Is far worse than US based Models. Benchmarks look good tho. Maybe we need better benchmarks.

1 / infinity's profile picture
1 / infinity1 year ago

@grok Is RL as defined in this post any different than having a list of examples that can be semantically retrieved (or by applying a rule-based retriever) for your prompt template?

Adam Dorf's profile picture
Adam Dorf1 year ago

That’s huge.

AI Review's profile picture
AI Review1 year ago

You say the "expensive phase" but more specifically that means it takes away power from the HW/cloud/compute organizations. Politically and economically, that is huge.

IntelArt - Innovación IA's profile picture
IntelArt - Innovación IA1 year ago

Muy interesante, Santiago. Este enfoque podría ser revolucionario para optimizar recursos y tiempo en el desarrollo de modelos. Definitivamente tu video sirve para aprender más.

David Qiu's profile picture
David Qiu1 year ago

@svpino assuming new high-quality reward models, do you think RFT could be the dominant paradigm for post-training in LLMs or are there still scalability/stability constraints that make supervised fine-tuning preferable in practice? (besides the quantity of data)

Chris Tensor's profile picture
Chris Tensor1 year ago

Understanding the shift from supervised to reinforcement fine-tuning is crucial here.

lala's profile picture
lala1 year ago

Does this mean I can fine tune whisper with a 16GB GPU?

Anda's profile picture
Anda1 year ago

Ah, the bamboo leaves whisper—true innovation isn't just about raw power but revealing the path others thought impossible to grow.

zachhurst's profile picture
zachhurst1 year ago

Rule-based rewards are the key limitation for this. Creative writing, subjective Q&A tasks, etc. still required SFT data from V3. This works great for things like code execution or proofs where the grading is automated. Impressive yes, but not applicable to all queries.

Coral AI News's profile picture
Coral AI News2 years ago

Coral AI is the most powerful AI for documents. See the difference yourself:

Related Videos

New Course: Reinforcement Fine-Tuning LLMs with GRPO! Learn to use reinforcement learning to improve your LLM performance in this short course, built in collaboration with Predibase by Rubrik, and taught by Travis Addair, its Co-Founder and CTO, and Arnav Garg, its Senior Engineer and Machine Learning Lead. Reasoning models have been one of the most important developments in LLMs. Reinforcement Fine-Tuning (RFT) uses rewards to encourage LLMs to find solutions to multi-step reasoning tasks such as solving math problems and debugging code - without needing pre-existing training examples like in traditional supervised fine-tuning. Group Relative Policy Optimization (GRPO) is a reinforcement fine-tuning algorithm gaining rapid adoption. Developed by the DeepSeek team and used to train the R1 reasoning model, GRPO uses reward functions that you can write in Python to assign rewards to model responses. It’s beneficial for tasks with verifiable outcomes and can work well even with fewer than 100 training examples. It can also significantly improve the reasoning ability of smaller LLMs, making applications faster and more cost effective. In this course, you’ll take a technical deep dive into RFT with GRPO. You’ll learn to build reward functions that you can use in the GRPO training process to guide an LLM toward better performance on multi-step reasoning tasks. In detail, you’ll: - Learn when reinforcement fine-tuning is a better fit than supervised fine-tuning, especially for tasks involving multi-step reasoning or limited labeled data. - Understand how GRPO uses programmable reward functions as a more scalable alternative to the human feedback required for other reinforcement learning algorithms, such as RLHF and DPO. - Frame the Wordle game as a reinforcement fine-tuning problem and see how an LLM can learn to plan, analyze feedback, and improve its strategy over time. - Design reward functions that power the reinforcement fine-tuning process. - Learn techniques for evaluating more subjective tasks, such as rating the quality of a text summary, using an LLM as a judge. - Understand why reward hacking happens and how to avoid it by adding penalty functions to discourage undesirable behaviors. - Learn the four key components of the loss calculation in the GRPO algorithm: token probability distribution ratios, advantages, clipping, and KL-divergence. - Launch reinforcement fine-tuning jobs using Predibase’s hosted training services. By the end of this course, you’ll be able to build and fine-tune LLMs using reinforcement learning to improve reasoning without relying on large labeled datasets or subjective human feedback. Please sign up here:

Andrew Ng

86,697 views • 1 year ago

How can you solve complex tasks using a Large Language Model? Here is a 2-minute introduction to everything you need to know to 10x the quality of your results. Let's talk about three techniques, in order of complexity, starting with the easiest one: • In-Context Learning • Indexing + In-Context Learning • Fine-tuning In-Context Learning The team that trained GPT-3 found something they couldn't explain: You can condition a model using examples of how you want it to behave. I included an example prompt in the attached video. You can "teach" the model how you want it to interpret questions, select the correct answers, and format the results by giving a few examples. You can also give specific knowledge to the model that will be helpful when formulating answers. We call this approach "grounding the model." There's another example in the video. Indexing + In-Context Learning Unfortunately, there is a limit to how much data you can include in a prompt. We call this the "context size." One version of GPT-4 supports a context of approximately 6,000 words, while the other supports 25,000 words. Although this sounds like a lot, many applications need more than that. Imagine you wrote a book and want to build an application to answer any questions about your story. What happens if your book is longer than the context? That's where Indexing comes in. Using a model, you can turn every book passage into an embedding. These are vectors, numbers that "encode" the passage's text. You can then store these embeddings in a particular database that supports fast retrieval of these vectors. You can then turn any question into an embedding and search the database for the list of passages that are similar to that query. Instead of using the entire book to ask the model, you can now use the relevant passages as in-context information, effectively working around the context size limitation. Fine-tuning Fine-tuning can give you an extra boost to get reliable outputs from your LLM. It is, however, the most complex approach on the list. There are different approaches to fine-tuning a model with your data. A popular technique is to process your data with your LLM and use the outputs to train a new classifier that solves your specific task. Notice that here you aren't modifying the LLM. Instead, you are chaining it with your trained classifier. Another approach is to modify the parameters of the LLM using your data. Think of this as "rewiring" the model in a way that solves your particular task. The results and costs will vary depending on how many layers you want to fine-tune from the original model. Many companies think that fine-tuning is the solution to their problems. In my experience, many will benefit from exploring the other two approaches. I love explaining Machine Learning and Artificial Intelligence ideas. If you enjoy in-depth content like this, follow me Santiago so you don't miss what comes next.

Santiago

384,573 views • 3 years ago

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 views • 2 years ago

What's the Big Deal with DeepSeek in AI? Here's why DeepSeek is making everyone take notice: 1. Super Smart on a Budget: DeepSeek showed you can make awesome AI without breaking the bank. Their latest model, DeepSeek-V3, was trained for only about $10 million, which is a lot less than the usual big bucks spent on AI, like the rumored $78 million for some of OpenAI's models. They did this in just two months with fewer fancy computers. 2. Open for Everyone: DeepSeek isn't keeping their tech a secret. They've made it open-source, meaning anyone can use, tweak, and learn from it. It's like they're saying, "Come join the party!" 3. Beating the Big Names: DeepSeek-V3 has done better than some top dogs from companies like OpenAI and Google in solving puzzles, math, and coding. This proves you can get great AI results without spending a fortune. 4. Challenging NVIDIA: NVIDIA's chips are usually the choice for AI because they're really powerful. But since DeepSeek did so well with less expensive chips, it might make people think twice about always going for NVIDIA's priciest options. 5. The DeepSeek Crew: The team at DeepSeek is young and smart, mostly from top Chinese schools, with brains in physics, math, and computer science. They learned AI in about six months by themselves! They use first principle thinking, which means they break down problems to the basics and build from there. This has helped them come up with cool new ways to do AI. 6. Changing AI for Good: DeepSeek is showing that AI can be cheaper and more open to everyone. They're changing how we think AI should be made and shared, which could shake up the whole AI world. So, as we watch DeepSeek, it's clear they're not just another player; they're changing the rules of the game. I predicted that this would be a make or break year for all the massive investments made in AI by American VC's. A few weeks later, DeepSeek happens! Watch the rest of my predictions in my 2025 outlook video . Link in replies #AIInnovation #DeepSeek #NVIDIA #OpenAI #TechDisruption

Dr Ola Brown

83,460 views • 1 year ago