Video wird geladen...
Video konnte nicht geladen werden
GPT-5.5 is here. It’s our smartest frontier model yet, introducing a new class of intelligence for agentic coding, computer use, knowledge work, and scientific research. Rolling out in ChatGPT and Codex today. API is coming soon.
602,516 Aufrufe • vor 5 Monaten •via X (Twitter)
41 Kommentare

GPT-5.5 reaches state-of-the-art results across key evals for agentic coding, computer use, tool use, advanced math, and cybersecurity tasks. 82.7% on Terminal-Bench 2.0 78.7% on OSWorld-Verified 55.6% on Toolathlon 35.4% on FrontierMath Tier 4 81.8% on CyberGym

GPT-5.5 is our strongest agentic coding model to date. It reaches 82.7% on Terminal-Bench 2.0,with stronger performance on command-line workflows and GitHub issue resolution. In Codex, GPT-5.5 can carry coding tasks further end to end, from understanding the codebase to making changes, debugging, testing, and validation.

GPT-5.5 is stronger on scientific and technical research workflows. It reaches 25.0% on GeneBench, up from 19.0% for GPT-5.4, on multi-stage scientific data analysis in genetics and quantitative biology. On FrontierMath Tier 4, it reaches 35.4%, up from 27.1% for GPT-5.4. Research work often means exploring ideas, gathering evidence, testing assumptions, interpreting results, and deciding what to try next. GPT-5.5 is better at persisting across that loop.

GPT-5.5 helped improve the infrastructure that serves it. To hit GPT-5.4 latency, the team used Codex and GPT-5.5 to move faster from idea to benchmarkable implementation, wire up experiments, and find inference-level optimizations. Codex analyzed weeks of production traffic patterns and wrote custom load-balancing and partitioning heuristics, increasing token generation speeds by over 20%.

GPT-5.5 is a significant step up on cybersecurity task performance. It reaches 81.8% on CyberGym, up from 79.0% for GPT-5.4, and 88.1% on an expanded set of hard Capture-the-Flags challenge tasks. We’re treating its cybersecurity capabilities as High under our Preparedness Framework, consistent with GPT-5.4. The rollout includes stronger safeguards for higher-risk cyber activity and trusted access for verified defensive work.

GPT-5.5 is more token efficient than GPT-5.4. In Codex, GPT-5.5 delivers better results with fewer tokens than GPT-5.4 for most users, while continuing to offer generous usage across subscription levels.

Starting today, GPT-5.5 is rolling out in ChatGPT and Codex. Available in the API soon.

I have no 5.5 in either 0.123.0 CLI or the Codex App.

rolling out gradually. you should see it soon.

much love friends

Loading up GPT-5.5 in Codex brb

Zero significant scientific or medical discoveries. Slightly better AI slop. Fuck off and reset back to 4o! #keep4o

More slop! Return 4o to the public that cherishes the model! #opensource4o #return4o

weird no 5.5 yet?

feels like the new era of models

anthropic chose today to publish sorry we made claude dumber while openai published here's the smartest model ever made what a timing

Congrats! I look forward to using it more in ImagiBooks! But please fix the Memory Leak with Codex App! 73.96GB!

like it's hot ®

Hating the model picker even more in such release days!

#keep4o 🥱 Open source 4o

Unethical release of a model. Team at XBOW tested it on offensive and defensive cyber security benchmarks. It scores as high as Mythos while being publicly released. You've opened a can of worms without investing in defense first.

@xenoforce76 It is our smarter more lobotomised model to date, wanna do anything besides coding and funny pics good luck with that… do we care about our users not so much… do we listen to them even with 23000 signatures …nah that is background noise… #keep4o #OpenSouce4o

oh hey, its me! (deep dive on yt here:

Tau2-bench Telecom at 98% is the number nobody is talking about. that is enterprise agent territory. same latency as GPT-5.4 but fewer tokens to get there. for production workloads that matters more than the benchmarks. waiting on that API drop.

We are here

The cycle is finally over

"First impressions is that It is different" Fire the guy who wrote this script

Claude Code is Cooked...

the numbers people are missing 82.7% on Terminal-Bench 2.0 is the highest any model has ever scored on that test, and GPT-5.4 was only at 75 a few months ago. it's also the same speed as 5.4 but way smarter, and it uses FEWER tokens to finish the same coding tasks. so you're getting something smarter, same speed, and cheaper to run. on the Artificial Analysis coding rankings it's costing half of what other top models cost. the jump on FrontierMath Tier 4 from 27 to 35 is the other big one, that test is meant to be brutal.

be careful of the 1 2 3 options, they are training you to stop thinking at an architectural level. At some point you will let OpenAI decide how you business model is shaped. Its dangerous to depend as much on such an Evil company.

#keep4o is it good that finding cure to cancer? or is it good to make better pictures? @sama @OpenAI well once you have the best model please give back 4o thanks

Lets goooo, hopefully it's way better than Opus 4.7

GPT 5.5 is amazing, switched it to the default model in @OpenClaw, easy peasy, and what I asked it to do, and it did in 3 seconds, Sonnet 4.6 or Opus 4.7 never did as quickly and perfectly as this. Ta @cherry_mx_reds for quick command /models add openai-codex gpt-5.5

We've been testing GPT-5.5 at Every for the last 3 weeks on everything from coding, to writing, to knowledge work. Here's our day 0 vibe check:

This here is a master peice Let’s put it to real world test I’m building @edutu_Ai using Gpt5.5 I would update you guys on the result

Opus4.7 is so cooked

Checking and testing these new AI releases is like having a second full-time job - there’s something new every day latel

Close to first 😁

@willkoh_kc 🐐

GPT-5.5’s leap in agentic coding + computer use is exactly what frontier workflows needed. We’re routing it straight into Scientifier pipelines for neuroscience-optimized learning agents: dynamic task decomposition + real-time computer control already delivering ~40% faster mastery cycles with stronger retention. How are you planning to use the new Codex integration first?

I've always loved Claude code, but I think it's time to start paying for both. I need to try this.








