Loading video...
Video Failed to Load
We are super excited to launch the in-app browser inside Codex with comment mode! View any web pages & iterate with your agent quickly with just point and click. Codex will automatically capture a screenshot, the DOM element, and feed it as precise context to your next chat. No... show more
405,205 views • 4 months ago •via X (Twitter)
58 Comments

thanks for sharing james! so to confirm, the integrated browser is NOT live yet right?

Sorry it's hard to find right now! We'll be giving Codex the ability to drive the browser autonomously soon, so it will be much easier to discover :)

Any advice on getting Codex to trigger it? It still wants to use the Playwright tool even when the browser tab is open.

We haven't plugged in the browser-use agent yet, so the only agent workflow that Codex knows for the browser is iterating based on your comments! We'll ship an update next week that lets Codex fully drive the browser.

Thanks! Just having a built in browser is great! It really makes it the everything app.

❓

It's on Windows!

@LittleBrainz Definitely isn't. The app itself is but the new capabilities arent

@LittleBrainz looking into this now

@LittleBrainz should be available, can you try updating?

@LittleBrainz If not, can you give me your specs and the version number of your Codex?

You guys should just make gpt5.4 good at frontend. This would be bomb!

we've got something cooking for you

🥹

@ajambrosino Amazing! Does it work like a regular browser when it comes to cookies etc? A lot of apps that I'm building need auth so they have to be logged in - this is often a hassle with playwright sessions, it never figures it out

@ajambrosino Hey! You should be able to auth provided there is email + password flow. The browser today is a bit limited because it doesn't support a few things like multiple tabs, web permissions, etc. We are incrementally bringing the full capabilities of a browser into this product!

Hello, so on the Codex mac app, the file tree browser needs a splitter so you can drag it so its wider on the horizontal axis. If you have deep file trees or long file names this view is much too constricted and it hurts usability.

thanks for the feedback, passing it to the team!

The other thing is for the excel file viewer, we need dark mode, its way too bright with all that white

@jxnlco I can’t make codex spin up the browser, and for computer use I see a message relayed to the fact that I don’t have a plugin. I’m confused

@jxnlco sorry about that! We are hooking up the agent with the browser next week, so you'll have to click on a localhost or local file link to open up the browser. For computer use, click settings, computer use and install the plugin :)

@jxnlco Here

@jxnlco Trying to see if this is because of EU. Are you based in Europe?

@jxnlco Yes, sorry, I read later that is not available in EU? Don’t understand why, Anthropic do released in EU too

i dont see computer use or any updates and im not in the EU @JamesZmSun

This feature is fantastic! I can't express it enough much this will help my workflow. THANK YOU!

@DamiDina do agents really need visuals though?

This is one of those small-feeling features that quietly changes how you build. Turning the browser into structured context (DOM + screenshot + intent) removes so much friction between “seeing” and “fixing.” Feels less like prompting and more like pair programming with the web itself.

in-app browser is the right primitive. the failure mode i kept hitting was agents losing the exact page state after a refresh and then repeating bad clicks. are you persisting page context per task or mostly relying on screenshot history?

we don't pass the page context to the agent yet. We have some safety work that we are doing this week to prevent against prompt injection, etc. Right now, the agent can only see screenshots if you comment, and the page title!

waiting for linux support 👀

My Macbook Pro with 16GB ram was already struggling w/ my AI use, I don't think I can try this out reliably, but I am very impressed with what you've done here and will eagerly try it when upgrading to a new Mac w/ Ram maxed out this time.

Does this browser trigger an anti-crawler mechanism? For example, log in to reddit and leave a comment

Can't use in app browser in windows codex app

can you try updating? if it's not working, can you give me your version number & spec?

Reinstalling it did work for me, thanks

How to use it?

If you open a localhost or local file link in Codex, it will open the browser! Otherwise, you can also use Cmd + Shift + B to open it up directly

Point and click productivity just leveled up!

thats pretty cool, gives me a reason to get back on codex

@sama Fucking yes, I knew you were doing something on UI

Does it still suck at front end design? Nobody wants ugly web pages

This looks genuinely useful. Real question though — how does this compare to what Claude Opus 4.7 is doing with its vision upgrade? They just shipped 3x image resolution and 98.5% visual acuity for computer use. Different approach but same problem: get the AI to actually see what you're seeing. What's Codex doing here that Opus can't?

This is great! Can you all do something like this for markdowns/plans? It’d be great if one could review those almost like word/google docs and leave comments / annotations for the agent to address and then see/manage updates in a single spec/plan file.

wow

Good job, respect

Awesome. Less screenshots will be taken

@soumitrashukla9 wow this changes everything - finally i can point it to and make it update its own website lol :p

Interesting! Does the DOM capture use something like a headless browser to get the full, post-JS-rendered state, or is it more of a static snapshot? Also, how does it handle dynamic SPAs where the DOM mutates frequently?

Tried it out and.. not having full agent integration makes this rough. "No more switching between browsers or dragging screenshots" I.. already dont' have to do this with playwright, and not only that, the agent can actually see my screen. Hopefully we can get real benefits soon

@sama Can I use it to browse a project website, capture screenshots and put together an initial user manual structure?

@sama Yep! You can use it on remote websites as well :)

@sama Great. Will give it a test!

Here’s my variant of the Codex demo, as I didn’t quite grasp what was happening in the initial screencast. In my demo, I ask about the Virtuoso home page using its URL, then use the sidebar (very important) to open the built-in HTML browser. Once that’s in place, I’m able to perform operations such as deriving page descriptions in various structured data formats (JSON-LD, RDF-Turtle etc) for upload to my knowledge base. This knowledge base manifests as a Semantic Web, courtesy of Linked Data principles (i.e., naming entities and relationships using hyperlinks, and using those hyperlinks to construct RDF statements in subject–predicate–object, or entity–attribute–value form). Ultimately, this is just another step in the progression of knowledge base generation and management—supporting improved recall and reuse of data, information, and knowledge. From a modern business model perspective, in the age of AI, I also have the ability to monetize these knowledge bases using fine-grained, attribute-based access control (ABAC), where access is determined by evaluating attributes across the target knowledge base and the AI agents seeking access on behalf of human operations.

@sama BROOOOOOOOOOOOOOOOOOO

This looks really cool James, looking forward to playing around with it!

@ajambrosino Let’s goooo

Wow 🤯 this is really great


