正在加载视频...

视频加载失败

We’re excited to add support for Opus 4.6. With adaptive thinking, Opus 4.6 will decide the right amount of reasoning for your query, reducing latency and improving quality. Internally, we've already resolved bugs that Opus 4.5 couldn't tackle. Try it out on the latest!

12,404 次观看 • 5 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

BREAKING: Anthropic just dropped Opus 4.8—and it is a MONSTER We've been testing for about a week Every 📧 and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check: - Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus 4.8 scores a 63—a hair higher than GPT-5.5's score of 62, and a full 30 points higher than Opus 4.7. It tackled a ground-up rewrite of a production codebase, and actually built something that works. HOWEVER: Coding performance varied a lot at different reasoning levels. We recommend using it on xhigh for best results. - Incredibly good writer. Opus 4.8 scored a 79.6 on our writing benchmark—measuring models on real-world writing tasks we do all of the time like essay writing, promo email writing, and more. It beats GPT-5.5 by 6 points. It produces well-written prose with fewer "AI-isms". It's also very good at writing in your voice given the right context. HOWEVER: Writing performance also varied with reasoning levels. Medium reasoning had higher incidence of AI-isms—we found best results with high. - Beast at knowledge work. Opus 4.8 is very good at general knowledge work tasks like report creation, research and more. It produced the best PowerPoint one-shot we've ever seen on our deck generation benchmark. - Emotionally intelligent, willing to question the frame. I've also found it to be quite good at talking through psychological or interpersonal issues. It has a high EQ, and it's also good at not glazing and helping to expand your perspective. Its thought process feels extremely rich and dynamic. THE BAD: These days a model is only as good as its harness, and Codex is still a far superior harness to the Claude Desktop app. This has kept me using Codex + GPT-5.5 as my daily driver, but I am flipping back and forth a lot more between Codex and Claude. Anthropic is back baby! Read the rest on Every 📧:

Dan Shipper 📧

354,033 次观看 • 2 个月前

let me explain what Anthropic just did they built an AI model so good at finding security vulnerabilities that they have refused to release it meet Claude Mythos → it’s Anthropic’s newest frontier model and it’s not available to the public. not because it’s not ready. because it’s too dangerous → Mythos found tens of thousands of zero day vulnerabilities across every major operating system and web browser… many of them 1 to 2 decades old. for context… Opus 4.6 found about 500. Mythos found tens of thousands → it found vulnerabilities in the Linux kernel. a 27 year old vulnerability in OpenBSD. a 16 year old vulnerability in FFmpeg → it doesn’t just find bugs. it writes the exploits too. that’s the part that scared them → so instead of releasing it… Anthropic has created Project Glasswing. a cybersecurity initiative where they hand picked 40+ companies to use Mythos for defense only → the partner list reads like a who’s who of tech… Amazon, Apple, Microsoft, Google, Nvidia, Broadcom, Cisco, CrowdStrike, Palo Alto Networks, JPMorgan, the Linux Foundation → Anthropic is giving up to $100 million in usage credits to these partners and $4 million to open source security organizations → they’re briefing CISA and the Commerce Department on how to handle this → the benchmarks are truly insane… Mythos hit 77.8% on SWE-bench Pro where Opus 4.6 scored 53.4%. hit 93.9% on SWE-bench Verified where Opus 4.6 scored 80.8% → Anthropic’s head of frontier red team said this is “the first time a model is this good that we decided to approach release in a very different way” this is the first time an AI company has held back a model because it was too capable not too expensive. not too slow. too dangerous and instead of locking it in a vault they weaponized it for defense and gave it to the companies that run the internet that’s either the most responsible thing an AI company has ever done… or the scariest only time will tell

klöss

21,270 次观看 • 3 个月前