Loading video...

Video Failed to Load

Go Home

Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? New paper questions the common assumption that RLVR helps LLMs acquire novel reasoning abilities.

52,100 views • 1 year ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

A viral paper "Language Model Represents Space and Time" recently claims that LLMs learn "world models". As much as I like Max Tegmark's works, I disagree with their definition of world model. World model is a core concept in AI agent and decision making. It is our mental simulation of how the world works given interventions (or lack thereof). A world model captures causality and intuitive physics, telling the agent what is likely and what is impossible. It can and should be used for counterfactual reasoning, i.e. "what ifs": what would happen if I knock over a cup of water? Where would I have been if I had not taken that bus? Yann LeCun Yann LeCun says it well in his position paper ( I quote: "Using such world models, animals can learn new skills with very few trials. They can predict the consequences of their actions, they can reason, plan, explore, and imagine new solutions to problems. Importantly, they can also avoid making dangerous mistakes when facing an unknown situation." The first use of the term World Model in deep policy learning is attributed to hardmaru & Jürgen Schmidhuber: In their seminal paper, an agent masters shooting skills in the popular game Doom (demo below) by learning in imagination, using an internal world model as a "physics simulator". To put in a simple Python math formula, world model learns a function F(s[0:t-1], a) -> s[t:], which takes as input the observed past and current action, and outputs plausible future states. Now the definition of World Model in Tegmark's paper seems to be about predicting GPS coordinates and time eras. I see this as just a classification task with no causal learning and simulation going on. You cannot make meaningful interventions against that model, nor can you optimize any decision making in a closed feedback loop. As for the "space & time neurons", I think they are most similar to the "sentiment neuron" that OpenAI published in 2017: Predicting GPS is conceptually no different from predicting sentiment in my opinion. I don't think their experimental results are wrong - just that their conclusion is on shaky grounds. I welcome any debate! Paper link:

Jim Fan

594,014 views • 3 years ago

THAT $70 "RUN YOUR OWN LLMS" PI KIT CAN'T RUN A SINGLE LLM. IT'S A VISION CHIP WITH NO RAM. that clip sells a raspberry pi 5 in a slick case with an ai accelerator and the caption "your own llms." clean build, fun kit. the claim is where it breaks. the fine print: the popular $70 pi ai kit uses a hailo-8l, 13 tops. it's built for vision, object detection and image processing, and it has no memory of its own. so it cannot run large language models. full stop the board that actually can is a different one: the newer ai hat+ 2, hailo-10h, 40 tops, with 8gb of dedicated ram. that's $130, not $70 and even that runs only tiny models. llama 3.2 at 1b, qwen 2.5 at 1.5b, deepseek r1 at 1.5b. edge llms live in the 1-7b range, against cloud models at 500b to 2 trillion so the honest pitch: for $130 you can run a very small language model on a pi, slowly, as a fun learning project. that's real and it's cool. "your own llms" on a $70 vision kit is not. why this keeps happening: "ai kit" and a big "tops" number sell. tops sounds like intelligence. but tops measures vision-style math, not whether the chip has the memory to hold a language model. the spec that matters for llms is ram, and the cheap kit has none. the honest caveats, both ways: the $70 kit is genuinely great, just at vision. cameras, object detection, that's its job the $130 hat really does run small llms locally, which a pi couldn't do at all two years ago. that's progress "small" is the load-bearing word. don't expect gpt at home on a pi the takeaway: before you buy a kit because the caption says llm, check two numbers. not the tops. the ram, and the size of the model it can actually load. no 70-dollar miracle, no gpt in a pi case, no tops number that means what you think. save this before you buy the wrong kit for the word on the box.

RetroChainer

11,100 views • 2 months ago

Jev + SERV is actually insane. We already showed you can increase Jev's performance with SERV Reasoning. Now we're taking it further, bringing Jev-powered Decision nodes into Graph Sharding with the upcoming SERV v3. Here's a breakdown of how it works: Jev is a decision-making model. Given a task and a set of options, it predicts which path is more likely. Think of the octopus that predicted World Cup results. Jev does that for your business, except it's not luck. It weighs every option and tells you how sure it is. It does this by assigning probabilities to outcomes. It doesn't generate text on its own, so you can't expect it to create a new outcome for you. But that's also what enables it to be lightning fast and dirt cheap. For example, in customer service you can ask Jev how to triage an incoming query and route it to the correct department. It can only select from the list of departments you provide it. This also means it can't hallucinate a new outcome outside the options it's given, which makes it incredibly interesting for OpenServ. In Graph Sharding, we take a single system prompt and break it down into multiple LLM steps with deterministic input and output shapes. Some of these steps require an LLM to produce new output, while others are simply decision routers that determine the next possible path. Traditionally, LLMs are slow and expensive. Breaking a single prompt into multiple steps increases accuracy and reliability by a ton, but it also introduces latency. Jev takes on those decision nodes, which are the backbone of a business process and therefore SERV graphs, and makes them super consistent and lightning fast, lowering the overall cost and latency of graph execution. SERV Reasoning on its own is a great force multiplier for Jev because, like all other models, it works by interpreting input instructions. The clearer those instructions are, the better the model performs. That's where SERV Reasoning comes into play. Just like amplifying any other model, we also amplify the accuracy and consistency of Jev's responses. And now we're bringing Jev-powered Decision nodes into Graph Sharding with SERV v3.

Armagan Amcalar

365,108 views • 14 days ago