正在加载视频...

视频加载失败

🚨Haiku 5.5 vs Sonnet 5.5🚨 Same prompt. Same task: build a cinematic 3D animation explaining how a Tesla-style electric motor works. Haiku: 11min 13s Sonnet: 33min 3s Which one do you think did better?

13,376 次观看 • 3 天前 •via X (Twitter)

47 条评论

LLAMA 3D Studio 的头像
LLAMA 3D Studio3 天前

Not bad for like 10% of the cost

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

This is what makes haiku special.

JP 的头像
JP3 天前

Damn Haiku cooked and the price is also low

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

Yes, it seems a fast and cheap model, and with a good output.

Vishal Pandey 的头像
Vishal Pandey3 天前

Both are explaining really well.

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

Yes, both results are good, but sonnet quality is better, but haiku price is unbeatable.

Vishal Pandey 的头像
Vishal Pandey3 天前

True! Sonnet has realism

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

For sure.

Eco 的头像
Eco3 天前

The speed of Haiku 👀

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

It is a very, very fast model.

Ishttt 的头像
Ishttt3 天前

For what haiku is needed for it did a too good job

Reuben 的头像
Reuben3 天前

Anthropic is coooooking rn

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

OpenAI is in trouble. I don't know how they will figure out this.

Reuben 的头像
Reuben2 天前

Yeah idk if I’m ever going back

Knowix 的头像
Knowix3 天前

Haiku really did a very good job and and it's actually pretty cheap good job done by anthropic once again

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

Yes, I think OpenAI is in trouble.

Knowix 的头像
Knowix3 天前

Big time

Divine 〽️achine 的头像
Divine 〽️achine3 天前

Haiku is a beast for the cost. Amazing work

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

It is man, crazy model

Valak X 的头像
Valak X3 天前

Will try haiku

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

You should try, it's totally worth it.

Sashi 的头像
Sashi3 天前

Sonnet

Grim 死神📿 的头像
Grim 死神📿3 天前

Haiku looks dope bruh

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

It seems a really great model for the price.

My Life in DMV 的头像
My Life in DMV3 天前

Mind explaining wtf we're looking at first?

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

Tesla eletric motor in 3D

VulKan 的头像
VulKan3 天前

all anthropics models is mogging everyone now 😭

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

Yeah man, it seems anthropic is going to win for some time

VulKan 的头像
VulKan3 天前

oh btw, which reasoning is the Haiku 5.5 being set to? Max?

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

Being honest man, I don't remember but I think it was high

VulKan 的头像
VulKan3 天前

alrighty, no worries (👍˘◡˘)👍

dino.YTA 的头像
dino.YTA3 天前

JavaScript animation?

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

Yes, it's using Three.js

Uday👨‍💻 的头像
Uday👨‍💻3 天前

Haiku is daammm cool with low price Which plan you’re using to make this video?

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

It is a very good model. I'm going to send you the prompt that I used.

Wahab Khan 的头像
Wahab Khan3 天前

haiku just came and stole the show.

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

yeah man, anthropic is cooking.

Wahab Khan 的头像
Wahab Khan3 天前

they are coming in hot

Ahmed Osama 的头像
Ahmed Osama3 天前

For an explainer, I’d score the technical accuracy of the explanation as well as the visuals and runtime. Did either version need corrections to the motor explanation, and was that revision time included?

Rubens Soto | AI & SaaS 的头像
Rubens Soto | AI & SaaS3 天前

It was a single-shot prompt, so there was basically no back-and-forth.

CDG 的头像
CDG2 天前

haiku cooking gud

Adarsh Subham 的头像
Adarsh Subham2 天前

Obviously sonnet is better but haiku did it so cheap and fast thats commendable tbh

Mr-Brightside 的头像
Mr-Brightside3 天前

What reasoning level did you use?

Minh Danh Ngo 的头像
Minh Danh Ngo2 天前

Sonnet quality always wins overall.

0xdAve ⚛️ 的头像
0xdAve ⚛️3 天前

Sonnet is better Haiku certainly. Demo with Haiku is fine

Harshit - AI Specialist 的头像
Harshit - AI Specialist2 天前

Can you give the whole prompt or the other context you used?

Maurice 的头像
Maurice2 天前

For real?

相关视频

This workflow will save you thousands with Claude. Run Opus 5.5, Sonnet 5.5, Haiku 5.5, and Fable 5.1 together, and stop spending Opus tokens on work that doesn't need Opus. The entire idea in one line: Opus plans, Sonnet edits, Haiku reads, Fable reviews at decision points. Roles, broken down: - Opus 5.5, high effort, owns the plan and reviews the final code - Sonnet 5.5, medium effort, is the worker: edits files, runs tests - Haiku 5.5, low effort, splits into explorer (searches and reads the codebase) and researcher (pulls docs). it's the first Haiku with an effort setting - Fable 5.1, set with /advisor fable. Opus calls it at decision points, and it gets the full transcript each time Three moments Opus tends to call it: → before committing to a plan: is this the right approach? → the same error shows up again: is this going nowhere? → before marking the task done: did something get skipped? Why Haiku only reads: Anthropic's launch post says Sonnet 5.5 and Opus 5.5 are still the better choice for complex agentic coding (Terminal-Bench 4.0: 39.2% for Haiku 5.5, 70.6% for Sonnet 5.5) and that Haiku 5.5 fits narrowly scoped subagent work. so it gets the lookups, not the edits. Anyone still running one model for everything is paying $4 per million input tokens to check whether a file exists. Haiku 5.5 does it for $0.10. Drop this into Claude Code 👇 "Rebuild my Claude Code setup around this structure: Confirm Claude Code is v2.1.293 or later, so the haiku alias resolves to Haiku 5.5. If it isn't, stop and tell me to run claude update. Look through ~/.claude/agents and .claude/agents for subagents already covering explorer, worker, and researcher. Only create new ones for roles that are missing. Set model: haiku, effort: low on explorer and researcher, with no Edit or Write tools. Set model: sonnet, effort: medium on worker. If an existing subagent is locked to a different model, leave it as is and just list it. Name the explorer subagent Explore so it overrides the built-in one, which otherwise runs on my main model. In ~/.claude/settings.json, set advisorModel to fable, and set effortLevel to high for claude-opus-5-5 under modelSettings. A top-level effortLevel in user settings doesn't apply to Opus 5.5. Check for anything disabling the advisor: CLAUDE_CODE_DISABLE_ADVISOR_TOOL, DISABLE_TELEMETRY, or anything blocking feature-flag fetches. Also check CLAUDE_CODE_EFFORT_LEVEL, which overrides subagent effort settings, and CLAUDE_CODE_SUBAGENT_MODEL_FORCE, which makes Claude Code ignore subagent model fields. Report what you find. Don't change any of it yet. Add one line to ~/.claude/CLAUDE.md: consult the advisor before a big plan, when the same error shows up twice, and before marking a long task done. Show every change as a diff first. Wait for my go-ahead before touching anything."

Alvaro Cintas

54,424 次观看 • 3 天前

holy f*ck. Official Anthropic tip for Claude Code: stop running agent teams on Opus and Sonnet 5.5 alone when Haiku 5.5 can swarm the bus at 340 tok/s spin up an autonomous team with /teams claude --teammate-mode auto Opus 5.5 acts as the lead architect, scoping architecture and validating pull requests Sonnet 5.5 teammates take isolated git worktrees and hammer out implementations at medium effort Haiku 5.5 swarms the peer-to-peer bus as a dedicated scout and fuzzer, only stepping in at three critical moments: → before a contract locks: does the frontend payload match the backend schema? → when a test breaks twice: are we patching the bug or just hiding the symptom? → before calling done: what attack surface or edge race did everyone overlook? Sonnet 5.5 builds. Haiku 5.5 swarms. Opus 5.5 merges and ships Jev engineering runs the identical pattern one layer down: mechanical decisions that need no reasoning (which file to open, which tool to invoke, retry or abort) execute in 16ms, so the frontier models only wake up when execution paths actually diverge Plan on high. Delegate on medium. Keep Haiku 5.5 on P2P call. - the full team setup > Opus 5.5 on high coordinates the session and synthesizes PRs > teammate-ux builds client state inside a separate git worktree > teammate-back writes endpoints and handles Redis atomic locks > teammate-scout runs invariant fuzzing and AST parsing on Haiku 5.5 at 340 tok/s > all teammates communicate directly over peer-to-peer messaging without coordinator drag > tasks.json manages file-level mutex locks so nobody steps on another worktree paste the setup and this prompt into Claude Code below: "Configure my Claude Code workspace for autonomous agent teams: 1. Enable experimental agent teams in ~/.claude/settings.json: > Set CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS to 1 in env block > Set teammateMode to auto > Pin main lead model: opus with effortLevel: high 2. Audit ~/.claude/agents and create teammate profiles fitting ux, backend, and scout: > teammate-ux: model: sonnet, effort: medium, isolation: worktree > teammate-back: model: sonnet, effort: medium, isolation: worktree > teammate-scout: model: haiku, tools: Read, Grep, Glob, SendMessage 3. Configure shared task coordination in ~/.claude/CLAUDE.md: > Initialize tasks.json for task dependencies and atomic file mutexes > Teammates resolve contracts directly over peer-to-peer messaging without lead intervention > Enforce Haiku scout schema verification before teammates mark tasks completed 4. Quality gate: > Opus lead reviews synthesized worktree diffs and validates test suites before final merge Show every configuration diff first. Do not apply edits until confirmed" ↳

mirku

82,064 次观看 • 2 天前

holy sh*t. Official Anthropic tip for Claude Code: stop burning Opus 5.5 context on tasks Haiku 5.5 can swarm, while Sonnet 5.5 builds put the architect on call with /advisor claude --advisor opus --subagents haiku Sonnet 5.5 drives the primary session on high effort, writing diffs and running test suites Haiku 5.5 subagents swarm the repo in parallel at 340 tok/s, handling file discovery and spec docs lookup Opus 5.5 stays on call in the background as the advisor, only stepping in at three critical moments: → before a plan locks: does the implementation plan miss auth invariants or schema contracts? → when a test breaks twice: are we fixing the root cause or falling into a recursive rabbit hole? → before calling done: did the full diff introduce hidden regressions or break pre-flight rules? Opus 5.5 reviews. Sonnet 5.5 builds. Haiku 5.5 swarms Jev engineering runs the identical pattern one layer down: mechanical decisions that need no reasoning (which file to open, which tool to invoke, retry or abort) execute in 16ms, so the frontier models only wake up when execution paths actually diverge Plan on high. Delegate on medium. Keep Opus on call. - the full advisor setup > Opus 5.5 on call reads full session history and catches deep architectural traps > Sonnet 5.5 lead drives edits, writes core logic, and executes test harnesses > Haiku 5.5 subagents swarm AST parsing, grep scans, and API docs in parallel > JEV micro-fork layer resolves 1,500+ mechanical routing branches in under 16ms > advisor stays completely silent on routine bash execution to protect context paste the setup and this prompt into Claude Code below: "Configure my Claude Code workspace for hierarchical advisor orchestration: 1. Audit ~/.claude/settings.json and project config for model roles fitting lead, subagents, and advisor: > Pin main session model: sonnet with effortLevel: high for primary execution > Pin subagents model: haiku with effortLevel: medium for parallel file discovery and docs retrieval > Pin advisor model: opus on call for strategic review 2. Enable the advisor tool in ~/.claude/settings.json: > Set advisorModel to claude-opus-5-5 > Enable subagent parallel dispatcher pool (3x workers) 3. Configure automatic advisor consultation checkpoints in ~/.claude/CLAUDE.md: > Consult /advisor opus before finalizing multi-file architecture plans > Automatically summon /advisor opus when the same test or compiler error fails twice > Enforce advisor diff contract audit before declaring tasks complete or staging git commits 4. Constrain subagent scopes: > Haiku subagents return structured AST summaries and docs snippets only without modifying main session context Show every configuration diff first. Do not apply edits until confirmed" ↳

mirku

727,014 次观看 • 3 天前

Official Anthropic tip for Claude Code: stop burning Opus 5.5 on work Sonnet 5.5 can do, while Fable 5.1 sits idle hand the grunt work to Sonnet 5.5 subagents and put Fable 5.1 on call run /advisor fable Opus 5.5 plans and merges Sonnet 5.5 subagents read, edit and run the tests Fable 5.1 reads the full session and only speaks up at three points: → before a plan: is this the right approach? → when the same error comes back: am I digging in the wrong place? → before "done": what did I miss? Sonnet 5.5 builds. Fable 5.1 reviews. Opus 5.5 ships Jev engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or stop) go to Jev in under half a second, and the big models only see the ones that split - the full tree > Opus 5.5 on high runs the main session > explorer reads the code on Sonnet 5.5 > worker edits and runs tests on Sonnet 5.5 > researcher pulls the docs on Sonnet 5.5 > all three on medium > Fable 5.1 on call for main and every subagent paste the tree and this prompt into Claude Code ↓ "Rebuild my Claude Code setup around this tree: 1. Check ~/.claude/agents and .claude/agents for subagents that already fit explorer, worker and researcher. > Draft new ones only for missing roles > Give each model: sonnet, effort: medium > Skip any that pin a different model and list them 2. Set the main session to high via effortLevel in ~/.claude/settings.json, and set advisorModel to fable 3. Find anything that keeps the advisor off (CLAUDE_CODE_DISABLE_ADVISOR_TOOL, DISABLE_TELEMETRY) plus CLAUDE_CODE_EFFORT_LEVEL, which overrides subagent effort. Report them, change nothing 4. Add one rule to ~/.claude/CLAUDE.md: consult the advisor before a large plan, when an error repeats, and before calling a long task done Show me every change as a diff first. No edits until I say go." ↳

delost

43,704 次观看 • 8 天前

This is f*cking insane. This Claude Code tip saved me thousands of dollars. once Opus 5.5 is your main model, stop burning it on work Sonnet 5.5 can do, while Fable 5.1 sits idle hand the grunt work to Sonnet 5.5 subagents and put Fable 5.1 on call run /advisor fable Opus 5.5 plans and ships the final code Sonnet 5.5 subagents read, edit and run the tests Fable 5.1 reads the full session, every tool call included, and only speaks up at three points: → before a plan: is this the right approach? → when the same error comes back: am I digging in the wrong place? → before "done": what did I miss? Sonnet 5.5 builds. Fable 5.1 reviews. Opus 5.5 ships Jev engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or stop) go to Jev in under half a second, and the big models only see the ones that split anyone still running one model for everything is paying Opus prices to decide whether a file exists - the full tree > Opus 5.5 on high runs the main session > explorer reads the code on Sonnet 5.5 > worker edits and runs tests on Sonnet 5.5 > researcher pulls the docs on Sonnet 5.5 > all three on medium > Fable 5.1 on call for main and every subagent paste the tree and this prompt into Claude Code ↓ "Rebuild my Claude Code setup around this tree: 1. Check ~/.claude/agents and .claude/agents for subagents that already fit explorer, worker and researcher. > Draft new ones only for missing roles > Give each model: sonnet, effort: medium > Skip any that pin a different model and list them 2. Set the main session to high via effortLevel in ~/.claude/settings.json, and set advisorModel to fable 3. Find anything that keeps the advisor off (CLAUDE_CODE_DISABLE_ADVISOR_TOOL, DISABLE_TELEMETRY) plus CLAUDE_CODE_EFFORT_LEVEL, which overrides subagent effort. Report them, change nothing 4. Add one rule to ~/.claude/CLAUDE.md: consult the advisor before a large plan, when an error repeats, and before calling a long task done Show me every change as a diff first. No edits until I say go." ↳

delost

48,494 次观看 • 7 天前

I don't understand why everyone isn't doing this yet. Anthropic's own Claude Code docs show how to run a whole team of Claudes, while Opus 5.5 only touches the plan and the merge the whole idea: agent teams in Claude Code one lead, separate teammates in their own context windows, one shared task list, and they message each other directly the lead: Opus 5.5 on high, splits the work, writes the tasks, merges at the end the builders: Sonnet 5.5, one owns client/, one owns api/, never the same file the adversary: Fable 5.1, never writes code, only shows up at three points: → before an interface locks: do both sides agree on the contract? → when a test fails twice: is it fixed or just hidden? → before a task is marked done: what breaks it? Sonnet 5.5 builds. Fable 5.1 attacks. Opus 5.5 merges Jev engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or stop) go to Jev in under half a second, and the team only argues about the ones that split turn it on, then start the team in plain English: "Spawn three teammates: ux and backend on Sonnet, an adversary on Fable" - the full team > Opus 5.5 on high leads the session > ux on Sonnet 5.5, owns client/ > backend on Sonnet 5.5, owns api/ > adversary on Fable 5.1, owns nothing, reviews everything > shared task list with file locking, direct messages, no lead in the middle paste the team and this prompt into Claude Code ↓ "Set up agent teams for this repo: 1. In ~/.claude/settings.json add CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 under env and set effortLevel to high 2. In ~/.claude/agents, draft three subagent definitions for the teammates > ux and backend with model: sonnet, each limited to its own folder > adversary with model: fable and read-only tools, whose only job is attacking contracts, repeated test failures and done claims > Skip any that already exist and list them 3. Add a TaskCompleted hook that blocks completion until the adversary signs off, and one rule to CLAUDE.md: no two teammates edit the same file 4. Find anything that would override this (CLAUDE_CODE_SUBAGENT_MODEL, CLAUDE_CODE_SUBAGENT_MODEL_FORCE, CLAUDE_CODE_EFFORT_LEVEL). Report it, change nothing Show me every change as a diff first. No edits until I say go." ↳

delost

69,880 次观看 • 6 天前

Claude Code tip: once Opus 5.5 is your main model, stop leaving Fable 5.1 sitting idle and stop burning Opus tokens on tasks Sonnet 5.5 can swarm put it on call with /advisor run /advisor fable Opus 5.5 plans and ships the code Sonnet 5.5 swarms the routine work at medium effort Fable 5.1 reads the full session, every tool call included, and only speaks up at three points: → before a plan: is this the right approach? → when the same error comes back: am I digging in the wrong place? → before "done": what did I miss? Fable 5.1 reviews. Sonnet 5.5 executes. Opus 5.5 ships Jev engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or stop) go to Jev in under half a second, and the big model only sees the ones that split Plan on high. Delegate on medium. Keep Fable on call. - the full tree > Opus 5.5 on high runs the main session > explorer reads the code > worker edits and runs tests > researcher pulls the docs > all three on Sonnet 5.5 at medium effort > Fable 5.1 on call as the advisor paste the tree and this prompt into Claude Code ↓ "Rebuild my Claude Code setup around this tree: 1. Check ~/.claude/agents and .claude/agents for subagents that already fit explorer, worker and researcher. > Draft new ones only for missing roles > Give each model: sonnet, effort: medium > Skip any that pin a different model and list them 2. Set the main session to high via effortLevel in ~/.claude/settings.json, and set advisorModel to fable 3. Find anything that keeps the advisor off (CLAUDE_CODE_DISABLE_ADVISOR_TOOL, DISABLE_TELEMETRY, any variable that stops feature-flag fetching) plus CLAUDE_CODE_EFFORT_LEVEL, which overrides subagent effort. Report them, change nothing 4. Add one rule to ~/.claude/CLAUDE.md: consult the advisor before a large plan, when an error repeats, and before calling a long task done Show me every change as a diff first. No edits until I say go." ↳

mirku

421,711 次观看 • 11 天前