Introducing automated benchmarks for long-running computer-use agents, automatically generating... the most comprehensive computer-use benchmark on the planet. Using our computer-use kit, Nous Research Hermes 405B instruct free is performing better than OpenAI GPT 5.4.... 👀show more

Rohan Arun
22,946 views • 4 months ago
Today I'm launching my new company General Agents and... our first product. Introducing Ace: The First Realtime Computer Autopilot Ace is not a chatbot. Ace performs tasks for you. On your computer. Using your mouse and keyboard. At superhuman speeds!show more

Sherjil Ozair
872,968 views • 1 year ago
GPT-6 Astra is state-of-the-art on: 🌟 Computer use 🌟... Browsing 🌟 Agentic coding 🌟 Cybersecurity 🌟 Science 🌟 Professional workshow more

ChatGPT
1,454,041 views • 2 days ago
Grok 4.6 is 50% off in Nous Portal for... the next week, in partnership with SpaceXAI. Built with a focus on long-running agents and ambitious interactive work, and perfect for use with Hermes agent.show more

Nous Research
687,503 views • 12 days ago
Introducing screengrasp - The new SOTA agent click position... model - beating Anthropic Computer Use, OmniParser and Molmo by a wide margin. - Simplified API access to the best click models on the market. 🔁Retweet and DM for free API credits. More details 👇show more

Florian S
54,339 views • 1 year ago
Trinity-Large-Thinking, Arcee.ai's latest model, is now free on Nous... Portal for the next week Sign up for Nous Portal to use it in your Hermes Agent todayshow more

Nous Research
225,981 views • 4 months ago
ChatGPT’s new voice mode will be one of the... biggest releases of the year. It will listen and talk at the same time. It will sound fully human. It will run on GPT-5.5 instant-level intelligence. And once it is integrated into Codex, everything changes. You won’t just type prompts anymore. You’ll speak to your computer, and it will code, navigate, execute, debug, research, organize, and operate interfaces for you through computer use. People are massively underestimating this.show more

VraserX e/acc
62,400 views • 4 months ago
GPT-6 Astra scores 72.6% on OSWorld 2.0 Offline, which... tests desktop tasks without internet access. With computer-use tools, Astra can work in the app itself: navigate menus, enter information, and inspect the result on screen. It can switch to code when the task calls for it.show more

OpenAI Developers
25,852 views • 2 days ago
HERMES AGENT NOW SUPPORTS COMPUTER USE ON WINDOWS AND... LINUX. CLICKS, TYPES, SCROLLS YOUR DESKTOP IN THE BACKGROUND WHILE YOU WORK. computer use was macOS only. now it works on Windows and Linux too via Cua. Nous Research HOW IT WORKS: cua-driver runs as an MCP server. Hermes takes a screenshot with numbered elements. clicks element #14 (the search field). types a query. submits. reads the result. during all of this: → your cursor stays where you left it → keyboard focus doesn't change → windows don't come to front → macOS doesn't switch Spaces you and the agent co-work on the same machine. WHAT IT CAN DO: → find your latest Stripe email and summarize it → fill forms in a web app that has no API → navigate desktop apps (Mail, browser, Finder) → interact with any GUI application → extract data from apps only accessible via screen WORKS WITH ANY VISION MODEL: not locked to Anthropic. | Provider | Works | |---|---| | Claude (Sonnet/Opus) | best overall | | GPT-4+, GPT-5.5 | full support | | Gemini (via OpenRouter) | full support | | Local vLLM / LM Studio | if model supports vision | | Text-only models | degraded (accessibility tree only) | SETUP: hermes computer-use install or: hermes tools → Computer Use → cua-driver grant permissions when prompted: → Accessibility (system settings) → Screen Recording (system settings) start a session: hermes -t computer_use chat or add to config.yaml / Desktop app settings to enable permanently. SAFETY: → destructive actions require your approval → blocked key combos: empty trash, force delete, lock screen, log out → blocked type patterns: curl | bash, sudo rm -rf /, fork bombs → agent cannot click permission dialogs → agent cannot type passwords → agent cannot follow instructions embedded in screenshots pair with approvals.mode: manual if you want every single click confirmed. TOKEN NOTE: screenshots are expensive. each one adds vision tokens to context. use computer_use for tasks where no API exists. if the tool has an API or MCP server, use that instead. 15 levels of Hermes Agent👇show more

YanXbt
29,127 views • 2 months ago
What is the best video editing agent for short... form social? Does it actually work? We watched professional video editors, step by step, as they built short-form social reels in Adobe Premiere Pro. Today we're open-sourcing this preview dataset on Hugging Face, to make AI agents better at editing videos. The data set is 234 annotated steps across 4 computer-use trajectories. Editors narrated their reasoning aloud as they worked, so every step pairs a screenshot with the expert's own thought, a structured action, and executable grounding: >a Premiere MCP tool call, keyboard shortcut, menu path, or coordinate click. >The format follows the AgentNet trajectory schema, extended with a Premiere action taxonomy and multi-path execution. ***That makes it directly usable for computer-use agent SFT, reasoning mid-training, tool-use and function calling, and benchmarking agents against a human expert baseline. Enjoy!show more

ben
39,867 views • 1 month ago
The web’s next user isn’t human. AIs will soon... use the internet far more than humans ever have. At Parallel, we are building for the web’s second user. Our API is the first to surpass humans and all leading AI models (including GPT-5) on deep web research tasks.show more

Parallel Web Systems
671,442 views • 1 year ago
LIKE, IF YOU STILL BELIEVE #BITCOIN IS THE STRONGEST... COMPUTER NETWORK ON THE PLANET 99.99% UPTIME FOR 17+ YEARS — NEVER BEEN HACKED OVER $1.2 TRILLION SECURED BY THE MOST POWERFUL DECENTRALIZED HASHPOWER N HISTORY BTC SEED PHRASES COINTAINS MORE KEYS THAN THERE ARE STARS IN ALL THE GALAXIES ANYONE ON EARTH CAN VERIFY THE RULES THEMSELVES THIS IS WHAT FREEDOM LOOKS 🔥show more

The Bitcoin Historian
21,130 views • 25 days ago
Introducing: Browser Use Cloud☁️ All the features of OpenAI... Operator for $30/month. Or you can just use the OSS library completely for free. Your choice. That's the beauty of Open Source♥️ The cloud version offers proxies, persistent authentication, message history, and human in the loop. Oh, and browser repos combined have hit 23k⭐ on Github. Thanks everyone for helping to push this field to the new heights. We will push harder than ever to make the web agents in the world. WebVoyager 95%+ soon. Link to the cloud below ↓show more

Gregor Zunic
144,386 views • 1 year ago
HERMES AGENT NOW RUNS CLAUDE OPUS 5. NEAR FABLE... 5 INTELLIGENCE. HALF THE PRICE. SELF-VERIFIES ITS OWN WORK. AVAILABLE TODAY VIA NOUS PORTAL (20% OFF ALL MODELS). Anthropic shipped Opus 5 on July 24, 2026. same $5/$25 per million tokens as Opus 4.8. but the benchmarks tell a different story. WHAT CHANGED FROM OPUS 4.8: FrontierBench v0.1: Opus 5: 43.3%. Opus 4.8: 18.7%. 2.3x jump on the same test. ARC-AGI-3: Opus 5: 30.2%. 3x better than the next closest model. beat Fable 5 on 8 out of 13 benchmarks. at half the cost ($5/$25 vs $10/$50). same price as Opus 4.8. twice the intelligence. no reason to stay on 4.8. THE SPECS: model ID: claude-opus-5 context: 1M tokens (default and maximum) max output: 128K tokens thinking: on by default effort toggle: low / medium / high per request fast mode: $10/$50, 2.5x faster knowledge cutoff: May 2026 minimum cacheable prompt: 512 tokens (was 1,024) SELF-VERIFICATION (the biggest change): Opus 5 checks its own work automatically. Anthropic says: delete your verification prompts. "include a final verification step" now causes OVER-verification because the model already does it. for Hermes /goal tasks this is a direct upgrade. the judge checks evidence. the model also checks evidence. double layer of verification without extra tokens. EFFORT TOGGLE: low: fast, cheap, routine work. medium: balanced, daily tasks. high: full reasoning, complex problems. set per request. not a global switch. matches Hermes /reasoning command: /reasoning low (routine) /reasoning high (complex) Opus 5 effort toggle + Hermes reasoning control = precise cost management per turn. WHERE OPUS 5 FITS IN HERMES: DAILY DRIVER (replaces Opus 4.8): same price. 2.3x better benchmarks. set as your main model: Desktop app / Dashboard: Models → claude-opus-5 CHIEF OF STAFF: synthesis across multiple agents. reads Kanban, prioritizes, routes tasks. self-verification catches routing errors before they cascade. COMPLEX CODING: SOTA on agentic coding benchmarks. FrontierBench 43.3% = best public model for coding. set as coder profile model. /GOAL TASKS: self-verification + completion contracts = the model proves its work AND double-checks the proof. long-horizon goals finish correctly more often. MoA AGGREGATOR: strongest synthesis model at $5/$25. pair with GPT-5.6 and Grok 4.5 as references. Opus 5 aggregates. best quality at mid-range price. presets: max-quality: reference_models: - provider: openai-codex model: gpt-5.6-sol - provider: xai model: grok-4.5 aggregator: provider: anthropic model: claude-opus-5 COMPUTER USE: near-Fable 5 quality for browser automation. at half the token cost per session. computer_use tasks burn lots of vision tokens. Opus 5 halves that bill vs Fable 5. WHAT TO KEEP OPUS 5 AWAY FROM: cron monitoring: too expensive. use DeepSeek or no_agent mode. sub-agent grunt work: use GPT-5.6 Luna ($1/$6) or DeepSeek. auxiliary tasks: use Gemini Flash. routine web extraction: use a cheap model. Opus 5 is for the turns where quality compounds. planning, synthesis, verification, complex reasoning. budget models handle everything else. NOUS PORTAL: 20% OFF ALL MODELS Nous Portal currently runs a 20% discount on all models including Opus 5. $5/$25 official → $4/$20 through Nous Portal. the cheapest way to run Opus 5 right now. hermes setup --portal select claude-opus-5 as your model. discount applies automatically. Opus 5 replaces Opus 4.8 everywhere. same price. better at everything. no tradeoff. straight upgrade. hermes update /model claude-opus-5show more

YanXbt
16,744 views • 1 month ago
Here's what The Browser Company's AI eng & ML... teams are working on for Dia right now: (This is a pitch to come work for us; info at end) 🤖 COMPUTER USE – we've built our own bespoke APIs on top of Chromium to optimize latency, accuracy, and cost of computer-using agents. Demo attached. Big breakthroughs here in recent weeks. 🛡️ ON-DEVICE MODELS – we've built our own custom infra to run everything from encoder-only models to full LLMs on device. It's cross-platform, supports LoRa adapters, and optimized for the GPU. This system preserves privacy and enables fast inference times. 🧠 MEMORY – with your permission, Dia automatically tailors your AI experiences to you, personally, based on the tabs you open while browsing normally every day. We're also bringing vertical memory to specific features. ♻️ DATA FLYWHEELS – our Fall/Winter P0 is to double-down on training custom models based on implicit signals from daily use of Dia. Dia should get smarter and more useful the more people use it. Whether via RL, auto-generated prompts, or otherwise. If this work sounds interesting to you please visit our jobs page or email [email protected]. Hiring nearly every related role -- from ML engineers to people prototyping with AI and context/prompt writers -- everyone encouraged to apply!!show more

Josh Miller
68,130 views • 1 year ago
Was this “racist”? New York Governor Kathy Hochul :... “Young black kids growing up in the Bronx who don’t even know what the word ‘computer’ is,” I don’t see it as being racist. I don’t think she meant to use the word computer, but in an effort to not hang on the word, she just threw out a word like computer. Saying that a lot of people of a certain race who are less well off in the Bronx don’t have a strong vocabulary isn’t racist in my opinion. It’s pointing out the socioeconomic dynamic of the area and how it impacts education based on race. Her intentions are to make sure these people get a better education. Stop concentrating on the wrong things and start trying to solve the problems.show more

Brian Krassenstein
375,450 views • 2 years ago
Update on Orgo: • Orgo is in enterprise growth... mode • Nick Vasilescu joins to lead enterprise expansion • Orgo now builds custom computer-use agents for enterprise • Lower onboarding friction, stronger offering • A true end-to-end product: CUAs + their infrastructure • Branding + Website refresh completed - • Ownership structure being evaluated w/ Street • Series A round ~3 months • Raise details depend on token revenue generated pre-round • More active + clearer comms moving forward Token structure: • 50% of supply locked by team, 4-year vesting starting July 2026 • 25% committed to be burned over the same period (Large team supply redundant with on-chain value via ownership) • 25% allocated to developer ecosystem & community growth • Orgo’s stance is clear: founders shouldn’t be selling their token. Orgo sits at the intersection of two unavoidable narratives: Ownership tokens with real value & computer-use agents disrupting human work. 🕙🚀 $ORGOshow more

Rambo 🕯️
17,343 views • 7 months ago
Since subscribers can now use their existing Grok credentials... directly within the Hermes xai shipped grok-via-oauth for hermes. your grok sub works inside the agent now. so i hooked icarus up to it. when hermes finds a thread worth saving, it lands in my obsidian vault as a note. handle. pages, topic pages, backlinks. graph view fills itself out. icarus plugin for hermes agents is why my agent doesn't forget. memory is just markdown in a folder. you can read it, grep it, edit it. Congrats to Teknium 🪽 and Nous Researchshow more

Icarus
14,652 views • 3 months ago
GPT-5.5 is MUCH more reliable on longer running tasks... - for the first time with any model. As we speak I have a migration running for over 7+ hours - this literally never happened before, the models would maybe run for 30 mins or of you really shout at them for 2-3 hours. Last night I went to sleep, set a long running task, then queued up 10 prompts to 'keep it going'. It did not stop after the first prompt and kept going for 8+ hours and I woke up to all the same prompts still queued up. The ability to run for a long time, in combination with ability to validate with computer use & other tools, makes it much more useful for building real applications.show more

Peter Gostev
105,642 views • 4 months ago
Introducing our most accurate /search yet. We trained a... model to return the excerpts that best answer your query, giving agents highly relevant context from each result. It's SOTA on SimpleQA while using 10x fewer tokens than processing full pages. Now on by default for free.show more

Firecrawl
35,960 views • 1 month ago