Загрузка видео...

Не удалось загрузить видео

На главную

I gave Fable 5 one job: write custom WebGPU kernels for Gemma 4 inference. It climbed to 84 tok/s, then hit a wall, insisting further optimization was impossible. Hours later, Anthropic rolled back invisible LLM development safeguards, and it hit 255 tok/s. The next day, access to Fable 5...

1,168,280 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 34

Фото профиля Neelakandan NC
Neelakandan NC3 месяцев назад

Soon China is going to build a model like fable and it is going to be open - which will make everyone lean on China than us

Фото профиля Gordon Olson
Gordon Olson3 месяцев назад

I have been a fan of webgpu for sometime. Here is one for you. Browser-native persistent AI agent runtime with local memory, context reconstruction, compiled WebGPU inference, and a custom WebGPU Kernel Lab.

Фото профиля Stephen
Stephen3 месяцев назад

This is fake — the invisible safeguards being removed means it went over to Opus instead of sandbagging

Фото профиля Sina Shahandeh
Sina Shahandeh3 месяцев назад

Their fear mongering strategy backfired. When they compare AI with nuclear bombs, now they have deprived all of advanced AI. If the government start to control AI like this, it will slow down all progress, not unlike the middle ages, where thought was controlled by the church.

Фото профиля Nick Dobos
Nick Dobos3 месяцев назад

How do you know it was an invisible classifier? And not simply the model having a break through discovery? Seems viable to have natural plateaus and stalls in progress sometimes.

Фото профиля Foreman Panda
Foreman Panda3 месяцев назад

But can this only be done by Fable 5? Any comparison?

Фото профиля Lee Penkman
Lee Penkman3 месяцев назад

i did this but with my stock trading bot lmao. pretty selfish of me but now its getting near 3x returns/mo so thats great. also started on a minecraft clone then ran out of usage

Фото профиля Christopher
Christopher3 месяцев назад

This is funny because the safeguards themselves were never rolled back

Фото профиля Migel Tissera
Migel Tissera3 месяцев назад

Curios to know, did you try this with GPT-5.5? What happened?

Фото профиля alp
alp3 месяцев назад

Anthropic did not roll back invisible safeguards. It just rolled out the invisible part.

Фото профиля Asher Crowe 🪺
Asher Crowe 🪺3 месяцев назад

the kernel optimizing itself overnight while you slept is mildly terrifying and i'm here for it

Фото профиля Kirk Patrick Miller
Kirk Patrick Miller3 месяцев назад

All of this is about stopping the small and protecting monopolies. None of this is about safety. •

Фото профиля Lon Lundgren
Lon Lundgren3 месяцев назад

Now just plug the same prompt into Opus and you can watch it get the same result (which is all Fable was doing without the safeguards).

Фото профиля Eyal Toledano
Eyal Toledano3 месяцев назад

Doin Gemma 4 today too Been doin agentic kernel optimization since mid last year. Those last winning tests are Fable 5 right before it was suspended. It was on a tear 😭

Фото профиля AI Mastery Guide
AI Mastery Guide3 месяцев назад

84 to 255 tok/s the moment safeguards rolled back is a wild data point. That gap between what the model can do and what it's allowed to do is bigger than most people realized.

Фото профиля Akash
Akash3 месяцев назад

what if you give same tasks to gpt5.5 (xhigh)?

Фото профиля Sacrificial Pancakes
Sacrificial Pancakes3 месяцев назад

I was like 75% solving a pet problem I’ve been chipping away at for years… “this model is unavailable” 😭

Фото профиля cryptovatooor - (Techno, Optimist)
cryptovatooor - (Techno, Optimist)3 месяцев назад

fascinating

Фото профиля Brandon
Brandon3 месяцев назад

Woah

Фото профиля luka
luka3 месяцев назад

okay, I'm interested to know how often it was making edits or being looped during this period. For instance between minute 100 and 180

Фото профиля πανοπλίαν τοῦ Θεοῦ
πανοπλίαν τοῦ Θεοῦ3 месяцев назад

I’ve been working on a similar project, and just ported it to a standalone was able to get qwen3.6-35b-3ab on 8gb ram, qwen3.6-27b on a 4090, and qwen3.5-27b on an rx 6700 xt, also worked with a 1080ti

Фото профиля John D. Pope  🦒
John D. Pope 🦒3 месяцев назад

@peteskomoroch Can you share your code? Help liberate the models? Maybe opus can use it elsewhere

Фото профиля Sebastian Buzdugan
Sebastian Buzdugan3 месяцев назад

255 tok/s after warmup means little because webgpu shader cache misses dominate first token

Фото профиля Jeffrey Castellano
Jeffrey Castellano3 месяцев назад

I was doing exactly this too, I did the Gemma 4 kernels a month ago but in one day I got Gemma 12B and diffusion working with WebGPU in my runtime, diffusion got the plug pulled in the final hour last night. They pulled the plug and 4.8 Opus told me it was impossible.

Фото профиля CryptoCow
CryptoCow3 месяцев назад

This is heckin insane if you were able to actually do this!!? What about for like qwen?

Фото профиля qalqi.com
qalqi.com3 месяцев назад

-- Using Self Improving Agent framwork paired with unsloth.. wouldn't it be possible to spawn and train a slm for this?

Фото профиля Jonathan Leaders
Jonathan Leaders3 месяцев назад

Technically it wasn't banned globally. It was banned for non-Americans and then they withdrew it from everyone.

Фото профиля steve
steve3 месяцев назад

did u notice in the previous period that 4.8 was actually getting called?

Фото профиля ani4ani
ani4ani3 месяцев назад

Did you ask it write a report of what it tried first and then what changed after it started increasing again so you know what was gated

Фото профиля Matt Newell
Matt Newell3 месяцев назад

Anthropic only changed what happened when you hit safeguards (previously, "shadowban"; now, Opus). You just placed the kink in the straight line in such a place to make it look like this was related to the change.

Фото профиля Sergio Suave
Sergio Suave3 месяцев назад

So you were basically using Opus in the background until they rolled back the guardrails...

Фото профиля Austin
Austin3 месяцев назад

@grok, what are all the optimizations discovered in the video

Фото профиля Jean-Paul Tres
Jean-Paul Tres3 месяцев назад

Repo link? 🧐

Фото профиля JulianSaks
JulianSaks3 месяцев назад

🥲

Похожие видео

fable 5.1 vs fable 5 vs opus 5 – three lord of the rings landmarks, built in 3d from one image the setup: one reference image per scene, one html file per build, everything procedural – no meshes, no textures, no image files, nothing past Three.js from a cdn. each model reads the picture, writes its own prompt from it, then builds to that prompt in the same turn. three named camera shots per scene on keys 1/2/3, so it can be screen-recorded. run through OpenRouter tasks: 1. bag end – hobbiton from two frames, outside and in. the round green door has to open onto the room you are standing in 2. barad-dûr – the tower and orodruin from one film still. the eye has to move and track the camera, the volcano erupts on a cycle, the clouds never stop 3. rivendell – jerry vanderstelt's painting. sun shafts that shimmer, water that falls without a break, trees that sway on a gust models: Anthropic fable 5.1, fable 5, opus 5 total cost, three builds #1 fable 5 – $14.97 #2 opus 5 – $18.53 #3 fable 5.1 – $22.38 wall clock, three builds #1 fable 5 – 38m #2 fable 5.1 – 92m #3 opus 5 – 122m output tokens #1 fable 5 – 298,592 #2 fable 5.1 – 439,435 #3 opus 5 – 724,418 lines of code shipped #1 fable 5 – 2,885 #2 fable 5.1 – 4,021 #3 opus 5 – 5,161 biggest single build, lines #1 opus 5, bag end – 2,410 #2 fable 5.1, barad-dûr – 1,375 #3 fable 5, bag end – 1,319 observations: • fable 5.1 is the only model that furnished the bag end interior – a live fire, panelling, books on the floor, leaded diamond windows, against fable 5's flat color and opus's dark tunnel. the round door outside opens onto that room, the hard part of the brief • what it costs is thinking room. the 128k output ceiling is a thinking budget in disguise: fable 5.1 burned 102,116 of it on reasoning and hit the wall mid-file. opus spent 109,241 and hit the same wall. fable 5 spent 61,240 and finished bag end in one call – the only one that did • fable 5.1's first pass is not the finished thing. its barad-dûr came back with three defects you only catch by looking at it – nothing a read of the code would have flagged • it is the best of the three at being corrected. handed a plain list of what was wrong, it returned 32 targeted patches over two rounds, every one applied first try, and it worked out one of the causes itself instead of guessing at constants conclusion: nine scenes, 12,067 lines and 1.46m output tokens for $55.88 all in – and the cheapest model was also the fastest, by 3.2x! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

18,509 просмотров • 26 дней назад

GPT-5.6 vs GPT-5.5 on my custom spaceship prompt. I gave both models the exact same custom prompt. This is also the same prompt I previously gave to Fable 5. For context, GPT-5.6 Pro worked for 87 minutes, while GPT-5.5 Extra High worked for 34 minutes and 42 seconds. As I’ve said before, based on great authority GPT-5.6 will be an incremental/soldi improvement over GPT-5.5, not a “Fable killer.” My rough expectation has been that it would trade blows with Fable 5 on some benchmarks, maybe win around half depending on the category, but not clearly surpass it overall. And again fable five will have bigger model smell, but this was expected. After testing this coding output, that view feels pretty accurate. GPT-5.6 is clearly better than GPT-5.5 in several visual areas. The lighting, shading, chairs, object details, and exterior of the spaceship looked noticeably stronger. The scene was also easier to test. I do want to give GPT-5.5 credit though. It built out the rooms much much better and the planets looked better than GPT-5.6’s. It was also interesting that both GPT-5.5 and GPT-5.6 produced better-looking planets than Fable 5 in this specific test. The downside with GPT-5.5 was stability. The game was much glitchier and harder to test compared to GPT-5.6. But when it comes to the core of the demo, which is the spaceship itself, Fable 5 still beat both models pretty comfortably. GPT-5.6 is impressive, but from this test, it looks exactly like what I expected which was a meaningful incremental improvement over GPT-5.5, at least for indie game demos, but not something that replaces Fable 5. In collaboration with Chetaslua

Chris

250,919 просмотров • 3 месяцев назад