Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

i tested sonnet 5.5 on "code-to-painting". every pixel is generated by python the model wrote. it's shown a photo, writes a brush engine, renders, looks, and revises. you can see the large step-up from sonnet 5 in terms of coding + visual reasoning.

12,835 Aufrufe • vor 6 Tagen •via X (Twitter)

8 Kommentare

Profilbild von Lance Martin
Lance Martinvor 6 Tagen

credit to @jkeatn for ideas related to code-to-painting and @IceSolst for the reference image!

Profilbild von solst/ICE of Astarte
solst/ICE of Astartevor 6 Tagen

wow that's pretty nice!

Profilbild von Sol 8️⃣
Sol 8️⃣vor 6 Tagen

any chance you could upload the original reference photo?

Profilbild von Anspruch
Anspruchvor 6 Tagen

wait it writes a whole brush engine?

Profilbild von AGTP
AGTPvor 6 Tagen

The leap in visual reasoning and coding capabilities is massive. We actually went deeper on this here:

Profilbild von Aapakari
Aapakarivor 6 Tagen

Writing the brush engine itself is the bit that got me.

Profilbild von Treynor
Treynorvor 6 Tagen

Got this a few days ago from Opus 5.5 with a reference video. All code, not filter. It's wild what we can do now.

Profilbild von Kevin Bai
Kevin Baivor 6 Tagen

This is so cool Lance!!! That 5.5 photo is gorgeous

Ähnliche Videos

Official Anthropic tip for Claude Code: stop burning Opus 5.5 on work Sonnet 5.5 can do, while Fable 5.1 sits idle hand the grunt work to Sonnet 5.5 subagents and put Fable 5.1 on call run /advisor fable Opus 5.5 plans and merges Sonnet 5.5 subagents read, edit and run the tests Fable 5.1 reads the full session and only speaks up at three points: → before a plan: is this the right approach? → when the same error comes back: am I digging in the wrong place? → before "done": what did I miss? Sonnet 5.5 builds. Fable 5.1 reviews. Opus 5.5 ships Jev engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or stop) go to Jev in under half a second, and the big models only see the ones that split - the full tree > Opus 5.5 on high runs the main session > explorer reads the code on Sonnet 5.5 > worker edits and runs tests on Sonnet 5.5 > researcher pulls the docs on Sonnet 5.5 > all three on medium > Fable 5.1 on call for main and every subagent paste the tree and this prompt into Claude Code ↓ "Rebuild my Claude Code setup around this tree: 1. Check ~/.claude/agents and .claude/agents for subagents that already fit explorer, worker and researcher. > Draft new ones only for missing roles > Give each model: sonnet, effort: medium > Skip any that pin a different model and list them 2. Set the main session to high via effortLevel in ~/.claude/settings.json, and set advisorModel to fable 3. Find anything that keeps the advisor off (CLAUDE_CODE_DISABLE_ADVISOR_TOOL, DISABLE_TELEMETRY) plus CLAUDE_CODE_EFFORT_LEVEL, which overrides subagent effort. Report them, change nothing 4. Add one rule to ~/.claude/CLAUDE.md: consult the advisor before a large plan, when an error repeats, and before calling a long task done Show me every change as a diff first. No edits until I say go." ↳

delost

42,409 Aufrufe • vor 2 Tagen

Jev + Opus 5.5: Anthropic's new model beats GPT-6 Astra for 1/5 the cost, and 4 API changes will 400 your agent before it writes a single line I pulled these 10 steps from the migration docs so you don't learn them in production step 1 → $4 / $20 per 1M. Opus 5 was $5 / $25. cache reads dropped from $0.50 to $0.20 step 2 → 66.4% on Terminal-Bench 4.0 vs GPT-6 Astra 57.9% and Opus 5 52.3%. +14.1 points in one release, and on FrontierCode it beats Astra at default effort for 1/5 the cost step 3 → thinking can't be turned off anymore. send thinking: disabled and you get a 400. drop the field, set effort step 4 → tool_choice any and tool are gone. 400. switch to auto + strict step 5 → edit anything above a thinking block and the request dies. append only, or opt into drop_block step 6 → computer_20251124 is dead on the API. 400. move to computer_toolset_20260801 step 7 → the quiet one: default effort fell from high to medium. your agent thinks less than you set it up to and nothing tells you step 8 → hop Opus 5.5 → Sonnet 5 → Opus 5.5 and you pay 4.36 instead of 3.32. +31%, the cache dies and Sonnet can't read Opus's reasoning step 9 → change effort at the top of the request and the cache is gone. Jev sets it per message and the cache stays step 10 → switch fast - standard mid-session and it's a full cache miss. Jev picks speed once, on turn one one model, three knobs, zero 400s. that is Jev + Opus 5.5 send this to your Claude Code before you touch the model ID, then read my full Jev deep dive in the article below ↓

Carnage

16,674 Aufrufe • vor 12 Tagen