this is literally f**king insane i just 2x my... usage limits on $200/mo codex i figured out how to use deepseek v4.1-flash on $10/mo opencode go for routine subagent work... while gpt-6 astra directs the project and sol handles implementation. here's how to set it up in 3 mins: → connect opencode go to codex through model-router → install quota flow, including its skill and agent profiles → open a fresh astra task and paste: quota flow tells flash to handle discovery and checks, sol to implement, and luna to review when needed. the same implementation agent keeps its context through the build → test → fix loop. every extra agent should earn its call.show more

Avid
61,962 Aufrufe • vor 15 Tagen
this is the best trick to maximum usage limits... on chatgpt codex codex's best kept secret is that your main agent doesn't have to do everything... custom agents are just files in ~/.codex/agents, and one file gives you a second worker on deepseek v4 flash > create ~/.codex/agents/deepseek-worker.toml > set model = "opencode-go/deepseek-v4-flash" with model_reasoning_effort = "max" > keep it bounded: one task packet, no scope creep, report back ```toml name = "deepseek_worker" description = "bounded implementation, testing, and cleanup on deepseek v4 flash" model = "opencode-go/deepseek-v4-flash" model_reasoning_effort = "max" ``` then @ deepseek_worker in the composer... your root agent plans while the worker ships the implementation planning on the main model, execution on the flash lane... that's the whole trick (we run this exact file, last i checked it keeps the heavy turns off the main thread)show more

Avid
49,981 Aufrufe • vor 1 Monat
codex usage got f**king nerfed. so i'm switching from... gpt-5.6 sol (max) to mimo-v2.6-pro on opencode... it matches the performance for 1/10th price and almost 2x speed [here is how to set it up in codex in 1 min] 1. model-router → model-picker → toggle the... 2. Cmd+Q Codex 3. ready to goshow more

Avid
34,982 Aufrufe • vor 6 Tagen
this is f**king insane someone figured out how to... use fable 5.1 with free gpt 5.6 luna subagents and never hit usage limits [here is how to set it up in 3 mins] 1. install 'fable-orchestrator' repo 2. type /fable [task] 3. done save this and give it to your agentshow more

Avid
26,229 Aufrufe • vor 25 Tagen
holy sh*t. this is f**king insane. muse spark 1.3... beats fable 5 and gpt 5.6 at 1/20th price i cancelled my $200/mo Claude subscription for this. i replaced fable 5 with muse spark 1.3 for only $10/mo on opencode go [it takes 3 minutes to set up:] 1/ install the model-router repo 2/ toggle opencode free to green 3/ that’s it.show more

Avid
69,306 Aufrufe • vor 24 Tagen
holy sh*t this is f**king dangerous i just figured... out how to run Opencode in Codex It has Deepseek V4 Flash which replaces Opus 5 at 1/4th price You can get 10,000 request in only $10/month along with Kimi K3 and Qwen 3.5 Max [here is how you set it up] 1. install the 'codex-router' 2. Put in the Opencode Go API key 3. Done that's it Save this no matter what. This will be the best thing you do this weekshow more

Avid
117,270 Aufrufe • vor 1 Monat
holy sh*t this is f**king dangerous. a Chinese developer... figured out how to run Fable 5.1 for FREE through GPT-5.6 Luna subagents on basically every task. that means you can keep using Fable without constantly worrying about hitting agent usage limits. the setup takes like 3 minutes: > install the fable-orchestrator repo > run /fable orchestrator [task] that’s it. your agent handles the rest. give this to your agent immediately. this might be the most useful setup you’ll find this week.show more

MIKE
47,222 Aufrufe • vor 25 Tagen
this is f**king insane i cancelled my $20/mo claude... subscription for this. a dev found a way around OpenAI's usage limits and got unlimited GPT-5.6 Luna on Codex for FREE ($0)😳 He built a three step local workflow on a single machine. This involved no servers, shared accounts, or a team. Instead he used a GitHub repo, one prompt file, and a local coding agent. Traditional users pay full price or hit usage limits. His cost: basically $0. Here's the set up (3min): → Repo layer: Installed 'sol-advisor' from GitHub. → Prompt layer: Copied the 'savings-prompt' out of the repo. → Agent layer: Pointed codex at gpt 5.6 luna with the prompt loaded. → Activation: Loaded it into codex. Usage limits gone. usage after activation: normal codex access. unlimited. one trick. one window. use this before it gets patched.show more

painn
98,056 Aufrufe • vor 1 Monat
You can now orchestrate Fable 5, Sol, and any... model inside Codex with one plugin. It's called Codex-Orchestration. Assign Fable 5 as the advisor, Sol as the executor, or any model to any role. Then define the order they work in. Codex handles the routing. I ran Fable 5 High as planner with GPT-5.6 Sol Extra High as executor on a set of issues Opus and GPT-5.5 always struggled with. Done in 30 minutes. 40% fewer limit hits. 2x faster implementation. Install it by pasting this into Codex: "Install Codex Orchestration: codex plugin marketplace add Cjbuilds/Codex-Orchestration codex plugin add codex-orchestration@codex-orchestration Verify the installation, then tell me to start a new task." Then assign your models: @ codex-orchestration advisor: Claude Fable 5 High, Executor: GPT-5.6 Sol High Open source. Tweak the routing however you want.show more

Alvaro Cintas
92,149 Aufrufe • vor 2 Monaten
I went a little overboard with Codex last week... and burned through my entire weekly allowance in two days. Luckily, my quota reset today. Otherwise, I’m not sure what I would’ve done. It got me thinking: instead of asking one large model to handle everything from start to finish, why not let a stronger model plan the project and review the work, while a model built for execution handles the day-to-day implementation? So I tried it. The result was better than I expected. I used GPT-5.6 Sol in Codex as the decision-maker, then ran Ling-3.0-flash from Ant Ling inside OpenCode as the execution engine. Together, they built a small 3D farming game. Before writing any code, I had Codex create four documents: SPEC.md defined the product scope and the lines we couldn’t cross. ARCHITECTURE.md laid out the isometric coordinate system, state machine, and module boundaries. TASKS.md broke the project into small jobs Ling could tackle one at a time. ACCEPTANCE.md explained how each step would be tested and what “done” actually meant. Then I gave Ling a very straightforward role: You are the execution model for this project. Read all four documents before you begin. Work only on the task assigned for this round. When you’re done, run typecheck, test, and build. If anything fails, read the error, fix it, and run the checks again. Do not move on to the next task early. Ling handled dependency installation, project structure, strict TypeScript configuration, test setup, and a production build in 6 minutes and 3 seconds. It ran into issues with the Vite test config, a TS6310 error, and a missing jsdom dependency along the way. Instead of stopping at the first error, it kept reading the logs and fixing the problems until all three checks passed. The speed was honestly hard to believe. If you exclude the time spent waiting on tools, it was producing more than 100 tokens per second. That made the whole development loop feel noticeably faster. After this experiment, I’m planning to keep using the same workflow. If the task is small, there’s no reason to call an expensive planning model for every single step. If the task is large, handing the entire project to a Flash model in one prompt isn’t a great idea either. The setup that makes more sense to me is: Use a more capable model such as Codex to explore the project, make architectural decisions, and break the work down. Put the constraints into specs, schemas, types, and tests instead of leaving them buried in chat history. Give Ling-3.0-flash a steady stream of clear, verifiable implementation tasks. Report bugs with structured context and actual error logs, rather than saying, “It still doesn’t work.” Bring Codex back in for architecture reviews, visual checks, and changes that affect multiple parts of the project. The point of this setup isn’t to give AI a big “build the whole project” button. It’s to turn software development into a pipeline with a much more sensible cost structure: Codex figures out the plan, sets the boundaries, and catches problems. Ling-3.0-flash moves quickly, calls tools reliably, and works through well-defined tasks at scale. For agent workflows that involve lots of repetitive edits, production tasks, and tool calls, this may be a more practical answer than simply using the biggest model for everything.show more

雪踏乌云
23,107 Aufrufe • vor 2 Monaten
ChatGPT Web is now inside Codex 😲 this open-source... project has already crossed 2.7k stars instead of using a separate workflow, it lets you use ChatGPT Web models directly from Codex's model picker what you get: - GPT-5.6 Pro for eligible accounts - free Luna access - ChatGPT Web quota - Codex tools + context - images, streaming and reasoning - open-source + MIT licensed getting started: 1. go to 2. install the launcher 3. sign in with your ChatGPT account 4. run the browser checks 5. install the models 6. restart Codex and select ChatGPT Web the interesting part? you can keep using Codex normally while routing the selected model through ChatGPT Web no separate API key for the ChatGPT model 2.7k+ stars and still actively updated worth checking if you already use Codex and want to experiment with ChatGPT Web modelsshow more

K2S
100,249 Aufrufe • vor 28 Tagen
OpenCode Go is now wired into Codex!! The pricing... is insane. $10 gets you 10,000 DeepSeek requests every 5 hours (no weekly limits). I converted that into DeepSeek API dollars because I thought I was reading it wrong, and the same 5 hours of usage would run somewhere between $10 and $30 depending on how big your context gets. So one afternoon of use already covers the whole sub. It comes with Kimi K3 as well. Both sit in the picker next to my other models now. Next time we hit a limit in the middle of a loop, we can grab it and keep going.show more

Ziwen
434,241 Aufrufe • vor 1 Monat
Codex can now run Deepseek-v4- flash! There's a catch... though. Deepseek's official setup switches your entire codex over to them, so your GPT models stop showing up at all. This is exactly what Codex Router is for. It adds models to the list instead of replacing them, so sol, grok, kimi and deepseek all sit in the same picker and i just grab whichever one suits the job. Deepseek v4-flash is $0.28 per million output tokens. opus 4.8 is $25. same picker, 89x apart. Links in the comment. setup's in the video 👇show more

Ziwen
145,018 Aufrufe • vor 1 Monat
I still think Hermes agent is the most slept-on... AI tool of 2026. For literally $6/mo, you can launch multiple subagents that work for you 24/7. Most people don't know you can do this, but it's a complete game-changer. Instead of one Hermes assistant doing everything sequentially, you run specialized agents in parallel, each with its own job, its own context, and its own memory. Practical example: → Research agent: scans your watchlist and competitors overnight, delivers a morning brief → Content agent: drafts and schedules your posts based on what's trending in your niche → Ops agent: manages your inbox, flags anything urgent, drafts replies for your review All three can run simultaneously and improve over time. How to start: 1. Install Hermes Terminal command: curl -fsSL | bash (can also download desktop) 2. Prompting Simply tell Hermes directly: "I want to run separate subagents for [task 1], [task 2], and [task 3]. Set them up to run independently and report back to me." For the cheapest setup, you can use a $4/month VPS with Hostinger, plug in DeepSeek V4 Flash as your default model. There isn't another AI tool with this much value in 2026. Hermes is still so underrated.show more

Miles Deutscher
82,643 Aufrufe • vor 2 Monaten
Making OpenCode as lean as Pi agent? Just trimmed... 25k out of OpenCode's system prompt (from 30k to 4-5k tokens) How? Just disable skills and get rid of massive skill definition bloat. Who needs skills anyway? Just kidding, this is the not the way. It makes the agent lame and defeats the point of using one. But it sets a precedent: Find a way to use skills without their definitions pre-loaded into the system prompt every single turn. Another interesting stuff: Upon testing this temporary "no skill setup" with two of hottest OpenCode Zen free models, Mimo V2.5 vs DeepSeek V4 Flash: One thinks more and talks less One thinks less and talks more Check the video to see which is which If you made it here, I'm finding a way to leanest OpenCode setup that I can get I simply don't believe that OpenCode can't be as lean as Pi Upon tinkering, I made a plugin that temporarily extracts the system prompt while I test, and noticed the hundreds of definitions in it from my .agents/skills directory which is shared across all my coding agents (Cursor, Antigravity, Claude, etc.) Of course disabling skills is not the answer, but it just proved that there is a way to strip the system prompt of these massive skill defs Aside from the system prompt hierarchy that injects confusion imo if you have a conflicting and redundant AGENTS.md which I discovered upon digging into OpenCode's source code Apparently it has prompt.ts/system.ts/instruction.ts/llm.ts and loads base .txt prompts based on model family (claude/gpt-o/gpt-5/codex/gemini/others) that all work together to make OpenCode aware of who it was and how it should use tools and become a "coding agent" Gotta find the most minimal mix that fits right into my workflow Make OpenCode as lean as Pi? We'll see. All inshow more

raymel 👋
37,939 Aufrufe • vor 4 Monaten
I optimized my Dream Loop skill to give you... Astra's 3D magic for 80% cheaper! This demo was one-shot, in 25 minutes, for 7% quota on ChatGPT Plus. Yes, the $20/mo one. 3 days of banging my head on the wall and burning Astra tokens. Sharing it for free. Y'all better build some cool stuff with this Link: Details, demo prompt, and cost breakdown are in replies. To be clear: It's not perfect. The demo is a small interactive scene, not a full game. For full games, your best bet is to build it with smaller models, then use this skill to improve the graphics. This is the best I could do, but it's not going to fix the fact that Astra is expensive and Plus quotas are tinyshow more

Anshu
348,943 Aufrufe • vor 18 Tagen
Something just launched that makes VS Code feel ancient.... And im all in for it: Cline Desktop app, an open-source app for open-weight models (and you know how much i love open source) brings its coding agent into a standalone Cline Desktop app for Mac + Windows. Just open a project and tell it what you want to change. No need to open VS Code first. In Cline's demo, it adds priority filters to a task board, works through the code and runs the build. Glad to see the model choice stays open, too. You can bring your own API keys, use supported open-weight models and even switch models halfway through a project!show more

Chubby♨️
34,235 Aufrufe • vor 13 Tagen
This is my MacBook Air M2 (16GB) with Hermes... Agent powered by MiniCPM5-2B. 96k context, and a very respectable output of 20-24 tok/s. This is totally usable, and MiniCPM5 can do all the tool calls to manage the basics with Hermes in a personal agent context. I wouldn't use it to build the next Salesforce, but for the home enthusiast this is a powerful and fast local model that can really punch above its weight, and runs on almost anything. Look at it go.. Amazing!show more

GooGZ AI
23,255 Aufrufe • vor 13 Tagen
Vorflux has opened its new cloud platform, allowing users... to run tasks from a plan to a merged PR on a dedicated machine. It can plan, build, test, and review on its own. > Diff reviews are done by a different model family than the one that wrote it by default. > First boot is saved as a snapshot, with all subsequent sessions waking up from that same starting point. > A browser agent walks through the real user flow, and the recording is included in the pull request. Video proof in the PR 👀show more

🚨 AI News | TestingCatalog
12,051 Aufrufe • vor 1 Monat
▣ Introducing Endless: infinite inference (kinda). An experimental harness... to milk every ounce out of your Codex subscription. Since Codex can let an in-progress turn keep going even after your usage hits 100%, why not put that to the test? Endless starts one Codex turn and gives the agent a wait_for_user_input tool. Once it finishes a task, it calls that tool and waits. Your next message becomes the tool result, keeping the entire session inside the same turn. It runs through Codex’s own app server using your existing ChatGPT login. Native tools, automatic compaction, context tracking, and quota tracking still work as usual. ⚠️ NOTE: I CAN’T CONFIRM THAT YOU WON’T GET BANNED OR PUNISHED FOR USING THIS TOOL. USE IT AT YOUR OWN RISK.show more

maria
254,287 Aufrufe • vor 1 Monat