Loading video...
Video Failed to Load
Many people are flat wrong about DeepSeek. You might think everyone freaked out about DeepSeek because the model is really good—which it is—or because it was from China, which is also true. But there's a more subtle reason: DeepSeek showed they could achieve those results using reinforcement learning instead... show more
105,589 views • 1 year ago •via X (Twitter)
31 Comments

To record this video, the team at @predibase gave me access to their fine-tuning platform. As far as I know, this is the first time we have access to an end-to-end platform for Reinforcement Fine-Tuning (RFT). Here is an introduction to RFT: And here is a link to the notebook with the code to fine-tune your models: Thanks to the @predibase team for helping with credits, being patient, and collaborating with me on the video.

The video is also on YouTube here:

RFT needs a way to verify answers, but it's not true that this is only possible in math or coding problems. (Also, remember, RFT is not a good fit for every use case. In fact, SFT is much better when you have a lot of data and people to provide feedback.) There are non-math, non-coding problems where you can still define a reward function. For example: • Text summarization: Your reward function could use a ROUGE score to measure the overlap between the generated summary and the reference. • Quality of dialog: You can measure coherence, relevance, and adherence to conversation norms and use that as a reward function. • Style transfer: You can build a reward function that evaluates whether the generated output matches the desired style. There are many more examples.

Exactly! DeepSeek’s real disruption wasn’t just model quality, it was how they got there. Skipping the costly supervised phase and leaning into reinforcement fine-tuning (GRPO) is a paradigm shift.

It is, yeah!

This approach could lower the barrier to entry for so many.

The question was if they copied anything from OpenAI. Some outputs indicated that they might have.

Honestly, I don't think that's "the question." Their contributions speak for themselves.

The potential for automation using reinforcement learning is fascinating.

bro this only works if you have verifiable rewards like math or coding problems. "alignment" requires human feedback still by definition.

Yes, you need a way to verify answers to define reward functions, but it's not true that this is only possible in math or coding problems. (Also, remember, RFT is not a good fit for every use case. In fact, SFT is much better when you have a lot of data and people to provide feedback.) There are non-math, non-coding problems where you can still define a reward function. For example: • Text summarization: Your reward function could use a ROUGE score to measure the overlap between the generated summary and the reference. • Quality of dialog: You can measure coherence, relevance, and adherence to conversation norms and use that as a reward function. • Style transfer: You can build a reward function that evaluates whether the generated output matches the desired style. There are many more examples.

they did SFT too.

Let us say I want fine tune LLM for specific version of a charting library Lightingchart or SciCharts where both have well described API & quite a few examples and LLM are not well trained in them. Shall I use RLT or SRT?

Definitely. But the R1-Zero phase is still quite vague from I/O data structuring point of view. If someone could figure out how they structured the training data for that step, and the context window improvements from Gemini Pro 2.5, it would yield a crazy open model

Spot on—DeepSeek’s real edge? Incremental gains > hype. Layer by layer, measure, adjust, repeat. Legacy systems don’t need overhauls, just smart integration. ROI climbs when you iterate, not leap.

Wonder how RF fine-tuning can be used in the background for automating model improvements for my apps.

Perhaps they used existing models like ChatGPT instead of people? 😉

They didn't.

Excellent insight! The how (RL fine-tuning) is indeed more disruptive than just the performance numbers

good part: RFT in Action: See how reinforcement fine-tuning and reward functions guide model training concerns: - Llama 8B RFT was compared to Gemini 1.5 pro - good comparison should be Qwen2.5 RFT with Gemini 2.5 pro to reflect model advancements

Deepseek R1 Is far worse than US based Models. Benchmarks look good tho. Maybe we need better benchmarks.

@grok Is RL as defined in this post any different than having a list of examples that can be semantically retrieved (or by applying a rule-based retriever) for your prompt template?

That’s huge.

You say the "expensive phase" but more specifically that means it takes away power from the HW/cloud/compute organizations. Politically and economically, that is huge.

Muy interesante, Santiago. Este enfoque podría ser revolucionario para optimizar recursos y tiempo en el desarrollo de modelos. Definitivamente tu video sirve para aprender más.

@svpino assuming new high-quality reward models, do you think RFT could be the dominant paradigm for post-training in LLMs or are there still scalability/stability constraints that make supervised fine-tuning preferable in practice? (besides the quantity of data)

Understanding the shift from supervised to reinforcement fine-tuning is crucial here.

Does this mean I can fine tune whisper with a 16GB GPU?

Ah, the bamboo leaves whisper—true innovation isn't just about raw power but revealing the path others thought impossible to grow.

Rule-based rewards are the key limitation for this. Creative writing, subjective Q&A tasks, etc. still required SFT data from V3. This works great for things like code execution or proofs where the grading is automated. Impressive yes, but not applicable to all queries.

Coral AI is the most powerful AI for documents. See the difference yourself:
