Загрузка видео...

Не удалось загрузить видео

На главную

Claude Code vs Codex vs Grok Build: Who builds the best Windows 95 OS in HTML? I benchmarked them live on speed, lines of code, file size… and actual usability. Claude: Fastest (384 lines, 14KB) but basic Grok: Took longer… and built a monster (full Explorer, Paint, interactive Grok...

44,952 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

I BUILT "GROK BATTLE" AND MADE GROK, GPT AND CLAUDE TRADE THE SAME ROBINHOOD STREAM WITH REAL MONEY TO SEE WHO MAKES MORE $500 each. same stocks. same account. same 24 hours. the only difference is the brain inside each pipeline grok made $4,800. claude made $1,640. gpt made $380. same stream, same entries available, three completely different results the part that broke my brain: all three saw the same SMCI earnings setup at 09:31. grok entered at open with a full bracket. claude waited fourteen minutes for volume confirmation. gpt wrote a two page analysis and entered after the gap was already priced in. grok rode it to the target. 8.4x on options that one decision is the entire gap between $4,800 and $380 here is what each model actually does differently > grok: fastest to decide. sentiment scoring is wild, it catches momentum before the first candle closes. but it overbuys. 23 entries in 24 hours, 14 of them stopped out. the winners are so big they cover everything > claude: most conservative. only 7 entries. five of them hit target. its checker is brutal, killed thirty one setups that grok took. but it also killed four that would have printed. safe money, slow money > gpt: smartest analysis, worst timing. writes the best reports on why a stock will move. by the time it finishes thinking the move is over and the entry is dead. 9 entries, 4 wins, but every win was small because it entered late the scoring breakdown > sentiment speed: grok 0.92, claude 0.71, gpt 0.84 > checker kill rate: claude 94%, grok 78%, gpt 86% > average entry delay from signal: grok 4 sec, claude 14 min, gpt 22 min > average hold time: grok 4.2 min, claude 11.8 min, gpt 22.6 min the model that thinks the best trades the worst. the model that decides the fastest makes the most. the model that says no the most loses the least that is the line i keep rereading every prompt, every scoring matrix, all three configs side by side. fork whichever brain fits your risk CA - 0x89C83B0b4CAC90a698EB98917852A7fA45F1E25D link below

zostaff

13,064 просмотров • 19 дней назад

So I’ve been using grok Grok Bot for a few days. It’s starting to click Bro they leapfrogged codex app and Claude code entirely. Super app 2.0 Grok bot’s marketing is completely wrong and misleading. It’s not an easy normie friendly chat. Don’t pay attention to the bot or chat UX, that’s not the point. Total distraction. Logo is flat out wrong. They legit did the Shoggoth meme with a smiley face mask in front of an AI demon. Grok bots is actually the highest difficulty, most powerful, and highest skillcap agent harness by far. Super rough around the edges. Dev tool, power user quality, nowhere remotely close to consumer level of approachability. Which is fine, they chose a lane, VSCode dev heavy design sensibilities bleed through, and to be fair it’s in beta. But importantly they nailed a new agent stepping stone: Persistent agent driven cloud computers Progression: - Chat - Chat + Tools, Calculate & Code Interpreter, MCPs - Claude Code & Coding Agent Harnesses: Chat + Tools + Read/Write local Files + Local CLI - Codex/Conductor Super App: Agent Harness + vertical tabs + embedded browser + computer use - Grok Bots: Agent Harness + Computer use + Browser Use + Entire cloud computer with all your files and apps, saved Grok bots absolutely mogs ChatGPT work, and the codex super app + ChatGPT remote. UX of having the cloud computer instance be persistent and not ephemeral per thread is 10000x better. Also you can actually open the computer from your phone and click and type and look at things. It’s not some black box only the agent sees. These were the 2 missing pieces This pattern is about to get copied and adopted by everyone.

Nick Dobos

460,678 просмотров • 1 месяц назад

I just compared Claude Code vs Codex vs Cursor CLI The task was to build a Next.js app with Tailwind 4 and shadcn components to collect customer feedback and showcase it with a widget. I gave all three the same prompt and let them go for 30 minutes to see what they came up with. Claude Code with Opus 4.1 Even though I told it to set up the app in the existing project folder, it tried to create a directory for it. After I interrupted and told it not to do that, it built a demo form and landing page with no errors. I had to ask it to make the demo interactive so users could submit a testimonial and preview it. The landing page looked like AI and was pretty basic, but it worked and it was done in a fraction of the time of the others. Total tokens used: 33k Codex with GPT-5 At the end of the 30 minutes I just could not get Codex to produce a working app. It got stuck in a loop of not being able to set up Tailwind 4 and despite many, MANY, attempts, I ended up with a "failed to compile" error. Total tokens used: 102k Cursor Agent with GPT-5 This was the slowest agent by far and a couple of times I actually thought it got stuck in a loop and was close to Ctrl+C'ing to cancel it. The TUI is really nice though, especially how it shows diffs and it did eventually build a working app (after one or two slight errors that needed fixing) The demo was interactive and it had a very minimal design that looked bare but also a lot less like an "AI generated" app than the Opus 4.1 design. It also wasn't too chatty and just did what it needed to do! Code quality was on a par with Opus 4.1, but it did use 5.5x as many tokens to get there. Still cheaper than Opus on a direct comparison but not when you factor in a Claude Code Max subscription. Total tokens: 188k I'll be able to do a proper comparison and record some videos when I'm back from holiday but for now, Opus is still the more capable model out of the box and Claude Code is the more complete CLI product. It will be interesting to see how Cursor evolve their CLI though with commands and subagents because I think with GPT-5 they have a real shot at providing competition for Claude Code if they can optimise output to get similar quality with less tokens. Jump to 0:40 in the video to see the two apps. Which do you think is which? ;)

Ian Nuttall

195,173 просмотров • 1 год назад