Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Understanding OpenAI o1: Noam Brown on integrating reasoning into the model. Takeaways: - Avoid MCTS and current paradigm of using processes outside of the model during inference - Think about how to directly integrate reasoning into the model architecture

313,371 görüntüleme • 1 yıl önce •via X (Twitter)

10 Yorum

Casper Hansen profil fotoğrafı
Casper Hansen1 yıl önce

This is why entropix or Quiet-STaR represents as an interesting way to directly build in better reasoning.

Casper Hansen profil fotoğrafı
Casper Hansen1 yıl önce

My take is that online planning is integrated at the model level. On top of that, they have a unique RL technique that truly enables the model to reason. I think @natolambert has some good posts on the RL in o1, also has a high-level blog:

Teknium (e/λ) profil fotoğrafı
Teknium (e/λ)1 yıl önce

He didnt say to avoid it he said to go broader then just mcts

Casper Hansen profil fotoğrafı
Casper Hansen1 yıl önce

Avoid MCTS was kind of inferred from the context. In general, I think what Noam is trying to convey is to think outside of the easy to-go-to methods like MCTS. Instead, think about it differently - how can we directly integrate search into the model?

You Jiacheng profil fotoğrafı
You Jiacheng1 yıl önce

Let me repeat:

Think_Different_ profil fotoğrafı
Think_Different_1 yıl önce

Nowhere did he imply to avoid MCTS…

Minh Pham profil fotoğrafı
Minh Pham1 yıl önce

Did he mention to avoid MCTS?

Casper Hansen profil fotoğrafı
Casper Hansen1 yıl önce

Yes, in the 1 minute clip above, he mentions it at the end

Tianhao Wu profil fotoğrafı
Tianhao Wu1 yıl önce

I don’t think he mean avoid MCTS

Casper Hansen profil fotoğrafı
Casper Hansen1 yıl önce

I provided some clarification here on what I meant. Emphasis on model-level search.

Benzer Videolar

OpenAI just announced API access to o1 (advanced reasoning model) yesterday. I'm delighted to announce today a new short course, Reasoning with o1, built with OpenAI, and taught by Colin Jarvis, Head of AI Solutions at OpenAI, to show you how to use this effectively! Unlike previous language models which generate output directly, o1 “thinks before it responds,” and generates many reasoning tokens before returning a more thoughtful and accurate response. It is great at complex reasoning -- including planning for agentic workflows, coding, and domain-specific reasoning in STEM fields like law. But how you should use it is quite different from other LLMs. I think o1 will be a game changer for many AI applications; and in this course, you'll learn how to use it effectively. In detail, you’ll: - Learn to recognize what tasks o1 is suited for, and when to use a smaller model, or combine o1 with a smaller model - Understand the new principles of prompting reasoning models: Be simple and direct; no explicit chain-of-thought required; use structure; show rather than tell - Implement multi-step orchestration in which o1 plans, and hands tasks over to gpt-4o-mini to execute specific steps; this illustrates a design pattern to optimize intelligence (accuracy) and cost - Use o1 for a coding task to build a new application, edit existing code, and test performance by running a coding competition between o1-mini and GPT 4o - Use o1 for image understanding and learn how it performs better with a "hierarchy of reasoning," in which it incurs the latency and cost upfront, preprocessing the image and indexing it with rich details so it can be used for Q&A later - Learn a technique called meta-prompting, in which you use o1 to improve your prompts. Using a customer support evaluation set, you'll iteratively use o1 to modify a prompt to improve performance You'll also learn about how OpenAI used reinforcement learning to produce a model that uses "test-time compute" to improve performance. I think you'll find this course enjoyable and valuable. Please sign up for it here:

Andrew Ng

357,661 görüntüleme • 1 yıl önce