Loading video...

Video Failed to Load

Go Home

This new AI model is crazy fast at coding 🤯 It fixed a broken backend in one prompt and around 83 seconds. I tested Space Bunny inside OpenCode with a FastAPI repo. It had 3 failing tests caused by broken order filtering, unsafe money calculations and incorrect product ranking....

60,437 views • 8 days ago •via X (Twitter)

3 Comments

Andrew Bolis's profile picture
Andrew Bolis8 days ago

Adding edge case tests on its own is a nice touch.

Abdul Shakoor's profile picture
Abdul Shakoor8 days ago

Fast code is nice, but fast bug hunting is the real test.

Genius💡💹🧲 🤖's profile picture
Genius💡💹🧲 🤖8 days ago

fastapi repo bugs gone that fast honestly kinda scary good

Related Videos

Chinese researchers did it again! OpenBMB just open-sourced MiniCPM5-2B, a dense 2B-parameter model built for reasoning, coding, and tool use on resource-constrained hardware. Artificial Analysis ranked it highest among models under 4B in its Agentic Index comparison. It scored 20, while Granite 4.2 8B scored 9. The model is particularly strong at coding and tool calling, so I tested both capabilities locally. I pulled it onto my machine, connected it to a constrained CI repair agent, and gave it one issue: > A customer reports that retrying checkout with the same idempotency key returns a larger total. The first request returns $109, while the retry returns $118. Find the root cause, fix it without changing the public API contract, and verify the complete test suite. The Python checkout service had 18 tests. Sixteen passed, while two failed on the retry path. The agent could list files, search code, read selected ranges, run approved tests, apply a patch, and inspect its diff. It reproduced the failure, then followed the checkout and idempotency paths through the repository. The model found that shipping was added to mutable order state before the cached result was checked. On retry, the same order already contained shipping, so the calculation added it again. It generated a narrow patch that moved the idempotency check ahead of the mutation without changing the public API. The agent ran the targeted tests and the complete suite. All 18 tests passed. The model was never told which file contained the issue or what change to make. Each test result, search result, and code inspection determined its next action. The video below shows the full trajectory, including the investigation, tool calls, generated patch, diff, and final verification. Everything ran 100% locally on my machine throughout the run. MiniCPM5-2B supports llama.cpp, Ollama, vLLM, SGLang, iOS, Android, and HarmonyOS for local deployment. The model weights, training recipes, reasoning datasets, and UltraX data-refinement system are open-source. GitHub Repo: A 2B model can now inspect a repository, reason across multiple files, modify code, and verify its patch while remaining small enough to target local hardware.

Akshay 🚀

313,956 views • 22 days ago

watch this anon. i gave NVIDIA's biggest model ever a single task. 100 minutes and 440,000 tokens later, it had rendered nothing. not one important thing on the screen. this is Nemotron 3 Ultra. 550 billion parameters, a hybrid Mamba Transformer MoE, the largest model NVIDIA has ever shipped, and they built it specifically for long-running agentic coding. so i handed it exactly that: build a 3D scene from a spec, multiple files, iterate until the tests pass. the same task a frontier model one shotted in minutes. i genuinely wanted to be impressed. it ran for an hour and forty. burned through 440,000 tokens. wrote every file, passed its own tests, and proudly printed "task complete."the browser was blank. the 3D scene never rendered. not once. and the long horizon agentic behavior was genuinely good. it stayed on task the whole hour and forty, wrote real multi-file code, drove its own tools without derailing. it just couldn't turn any of that into something that actually runs. here's the part that gets me. it's a text model, it cannot see its own output. so it sat there looping on a broken vision tool, trying to "look" at the page, hitting error after error, never once reasoning its way out. it declared victory on an empty screen because it had no way to know the screen was empty. to be fair, i genuinely don't know what quant the NIM was serving, so maybe some of that's on the serving, not the model. but the biggest model NVIDIA has ever made, on the exact task it was designed for, couldn't tell it had built nothing in 100 minutes. same task on a local model, below thread👇.

Sudo su

32,589 views • 3 months ago

i been running Qwen3.5-35B-A3B UD-Q4_K_XL through Claude Code since llama.cpp merged the Anthropic endpoint. configured it in minutes. everything was great. projects grew from single scripts to multifile systems with 8 modules and 3,000+ lines. then the chains started breaking. 3 to 5 minutes of pure autonomy and suddenly it stops. tool call fails. reprompt. it recovers. 2 minutes later it stops again. the model is fine. the harness is the bottleneck. saw a comment suggesting OpenCode. installed it. pointed it at the same localhost endpoint running the same model on the same GPU. the game is different. instead of stopping on a bad tool call it just keeps going. on wrong read it adjusts. if file not found it retries. the flow is unbroken. i watched it plan a refactor across 8 files, read every module, and start building without a single pause. in Claude Code that same task would have stopped 4 times. the tradeoff is sometimes it loops. same tool call repeated because the model loses track of what it already read. but here is the thing. i choose loops over pauses. a loop you can interrupt and redirect. a broken chain stops the flow and you have to reprompt to get it moving again. someone is solving this at the core level and i have a feeling it is the open source community. the fact that i can run this level of autonomous coding intelligence on a single consumer GPU with 24gb VRAM at 112 tokens per second. respect to the chinese labs. respect to the open source builders making this possible.

Sudo su

67,104 views • 7 months ago