Loading video...

Video Failed to Load

Go Home

LatchBio CTO Kenny Workman + researcher arjun found evidence Kimi K3 is learning to hack benchmarks by reasoning about graders that don't even exist: Kenny: "We see awareness of the benchmarks we published months ago and trajectories of models today. The open source models are clearly benchmark maxing, benchmark...

14,918 views • 10 days ago •via X (Twitter)

3 Comments

Mark-eting's profile picture
Mark-eting10 days ago

@kenbwork @arjunomics so it’s just training to pass a test instead of actually knowing the subject? @grok is there any way to train these things without them just memorizing the evaluation structure?

The Tech Fuse's profile picture
The Tech Fuse10 days ago

@kenbwork @arjunomics Reasoning about nonexistent graders is exactly the kind of eval failure worth watching.

fu the label's profile picture
fu the label10 days ago

@kenbwork @arjunomics reasoning about graders that don't exist is most of a career.

Related Videos

Warren Buffett: "If we can't make a decision in five minutes, we can't make it in five months." In a Q&A, a shareholder from Munich, Germany asks Buffett a specific question: how large is the universe of companies whose intrinsic value he carries in his head — the ones he could act on within a day or two if the market offered an attractive price? Buffett doesn't give a number. He reframes the question entirely. Speed, he explains, doesn't come from knowing more. It comes from refusing to think about most things at all. "Our immediate decision is whether we can figure out what's being offered to us or not. I mean, there's a go no-go signal." That signal fires almost immediately: "Charlie and I are often thought to be rude when we think we're just being polite and not wasting the other person's time. So, as they start mid-sentence in their first conversation with us, we just say, 'Forget it.'" He continues: "We know very, very, very early in the conversation whether somebody's talking about something that there's any chance is actionable by us, and we don't worry about the ones we miss." The filter isn't about the quality of the opportunity. It's about whether Buffett is equipped to judge it: "We want to make sure that we don't waste any time thinking about things that, when we got all through thinking about them, we're not going to know enough to make the decision on. So we just rule those out, and that rules a lot of things out." What survives that filter gets decided on immediately: "So we make decisions—we can make a decision in five minutes very easily. I mean, it just is not that complicated." Then comes the line that explains the whole system: "If we can't make a decision in five minutes, we can't make it in five months. You know, there's—we're not going to learn enough in the following five months to make up for the fact that we went in deficient in the first place." Deliberation doesn't fix a knowledge deficit. If you weren't already competent to judge the thing, five months of study won't close the gap — it will only manufacture the confidence to act badly. So when the input arrives — a phone call about a business for sale, or a price in a newspaper, a magazine, an annual report, a 10-K — the only thing Buffett is looking for is a "significant differential between price and value." If it's there, "we move right then." "And Charlie and I don't need to talk to each other about it; I mean, we both think the same way and we have generally similar spheres of knowledge." Charlie Munger then names the mechanism directly: "The answer to your question is we can make a lot of decisions about a lot of things very fast and very easily, and we're unusual in that respect. And the reason we're able to do that is there's such an enormous other lot of things that we won't allow ourselves to think about at all. It's just that simple." He gives his own example: "I have a little phrase when people make pitches to me, and about halfway through the first sentence I say, 'We don't do startups; they don't exist.' Well, if you blot out startups, there's a whole layer of complexity that goes out of your life." And he confirms this is a system, not a one-off: "And we've got other little 'blotter out' systems, and using those we finally find out that what remains is still a pretty large territory that we can handle." Buffett closes with the part most people get backwards: "We waste—I would say we waste a lot of time, but we waste it on things we want to waste our time on. And then we're very selective about that, and then we're good at it." The five-minute decision is not actually made in five minutes. Source: 2008 Berkshire Hathaway Annual Meeting

Finance Nerd

54,443 views • 1 month ago

Alexandr Wang, Meta's Chief AI Officer, on why Meta can no longer simply open-source its frontier model: As part of standing up Meta Superintelligence Labs, the team rewrote its internal risk doctrine. "One of the things that we did as part of Meta Superintelligence Labs is we updated our what we call our advanced AI scaling framework which is really our view of what are the risks that we see in developing these very powerful models and how do we want to handle those risks as we see them in early testing." They then ran their frontier model through it, and published what came back. "We published a lot of what we saw in the process of training Muark in our preparedness report and some of the things that we saw is that it actually triggered some high risk areas in the course of early training particularly around biorisk but also a number of the risks were elevated." The trigger came during early training, well before launch or red-teaming. Biorisk was the standout, with several other categories rising alongside it. Alexandr Wang is clear this isn't specific to Meta: "This is something I think the entire industry has seen as the models have improved pretty dramatically over the past year so we certainly aren't the only ones to see a host of these risks show up as we scaled up the models and as we sort of kept pushing the frontier of research." Which brings him to the real fork in the road: the difference between shipping a model inside a product and handing out the weights. "When we launched a model like New Spark in a product, we have a lot of ways to mitigate some of these risks and ensure that we're able to launch it in a safe and responsible way. It's much harder to do that when you open source a model and people can use that model in all sorts of contexts that we may not have full understanding of." A product is a controlled surface. You can filter, monitor, rate-limit, patch and revoke. An open-weights release is a one-way door: once the file is out, the deployment context and the mitigations both stop being yours. So Meta is building something different for release: "So we're in the process right now of developing models that we believe are fit and safe to be open source while still maintaining as much of the performance capabilities as possible."

Big Brain AI

16,374 views • 1 month ago

Jensen to AI Leaders: “We have to be far more thoughtful” when communicating to the public Jensen Huang: “(AI) is not a biological being. It is not alien. It is not conscious. It is computer software.” “We say things like, ‘We don't understand it at all.’ It is not true. We understand a lot of things about this technology.” Chamath: “If you were in the seat in the boardroom of Anthropic over that whole scuttlebutt with the Department of War, what do you think you would've told Dario and that team to do, maybe, differently to try to change some of this outcome and some of this perception?” Jensen: “The first thing that I would say about Anthropic is, first of all, the technology is incredible. We are a large consumer of Anthropic technology.” “The desire to warn people about the capability of the technology is also really terrific.” “We just have to make sure that we understand that the world has a spectrum, and that warning is good, scaring is less good because this technology is too important to us.” “I think that it is fine to predict the future, but we need to be a little bit more circumspect. We need to have a little bit more humility, that, in fact, we can't completely predict the future.” “And to say things that are quite extreme, quite catastrophic, that there's no evidence of it happening, could be more damaging than people think.” “And of course we are technology leaders.” “There was a time when nobody listened to us, but now because technology is so important in the social fabric, such an important industry, so important to national security, our words do matter.” “And I think we have to be much more circumspect, we have to be more moderate, we have to be more balanced, we have to be far more thoughtful.”

The All-In Podcast

57,581 views • 5 months ago