GPT-6.1 Sol Max vs Opus 5.5 just created an... animation about OpenAI dots and how these agents actually work here is a simple promt that was created in less than one minute which model looks bettershow more

kepo
18,001 görüntüleme • 1 gün önce
i just created an entire video with gpt-5.6 sol... for site every scene, camera angle, animation - is a line of code in react zero traditional editing. gpt-5.6 and @remotion are legendaryshow more

Suhail Kakar
27,207 görüntüleme • 2 ay önce
Salute to the Qwen team 🫡 We tested Qwen... 3.7-Max, Gemini 3.5 Flash, GPT-5.5, and Claude Opus 4.7. The biggest shock came from Qwen. In less than a month (3.6 Max dropped April 20), Qwen went from the worst multimodal output on our sakura tree test, barely keeping up with Gemini, GPT, and Claude , to matching Gemini 3.5 Flash frame for frame on this soccer test, and outperforming GPT-5.5 and Claude Opus 4.7. It rendered a perfectly proportioned soccer player and the most lifelike ball in the entire test. Remarkable spatial reasoning. Also: Gemini 3.5 Flash is now faster than GPT-5.5, which used to be the fastest in our past tests.show more

GMI Cloud
75,785 görüntüleme • 4 ay önce
Gemini 4 argon vs GPT 6.1 sol both models... were tested with same prompt but results came out really different > sol was tested by me in devin cloud at default reasoning > gemini was tested by lentils earlier this month on pre release checkpoint argon isn’t publicly available yet, so can’t say how good offical checkpoint is, but this one looks pretty good already which one did better?show more

J A Z I I
14,442 görüntüleme • 6 gün önce
Qwen 3.7-max beats Opus 4.7 and GPT-5.5 We tested... three frontier models on a real agentic task: write a Tetris bot that plays the game and trains itself. Each model could read its own code, run benchmarks, and rewrite itself across 10 iterations. Then we compared the final bots head to head. Qwen 3.7-Max: training cost $1.32, bot improvement +56% Claude Opus 4.7: training cost $12.15, bot improvement +28% GPT-5.5: training cost $2.85, bot improvement +7% Qwen won on every dimension - biggest jump, 9× cheaper than Claude, 2× cheaper than GPT. Long agentic loops is where Qwen Max actually delivers.show more

atomic.chat
870,326 görüntüleme • 4 ay önce
I tested Kimi K3 vs Claude Opus 4.8 Same... prompt, an armory bay with lighting, props, and detail. Top is Kimi K3, bottom is Opus 4.8. It's not even close. Kimi K3 built a full scene with textures, proper lighting, ammo crates, weapon racks, working detail everywhere. Opus 4.8 gave me a near empty room with a couple of floating tables. No doubt it beats Opus 4.8. Kimi K3 is Fable 5 level, and it's clearly better than GPT-5.6 Sol at 3D and games. An open weight model just matched the best closed models on the market. Let that sink in.show more

Bhavy☄️
479,920 görüntüleme • 2 ay önce
a moonshot engineer leaked the benchmark anthropic, openai and... xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article belowshow more

starmex
33,133 görüntüleme • 1 ay önce
The xAI API is incredible. I just created an... AI assistant that can fetch news content from URLs and write a post about it on my own writing style. Super simple to set up and you can try free. Here’s how:show more

Alvaro Cintas
15,385,921 görüntüleme • 1 yıl önce
Opus 4.6 vs GPT 5.4 (High) (1/9) prompt: Build... a single-file HTML/CSS/JS (no libs) demo that uses SVG to simulate a plant growing: stem extends, leaves sprout + unfurl with springy/windy “physics”, then seamlessly loops forever. For the initial impressions I'm really impressed by GPT, for speed they both felt about the same, but gpt was still half cheaper than opus. I also much prefer the design and animation that gpt produced, physics on the leaves are super cool and it also loops pretty nicely whilst opus just fades out the plant. Still got a bunch of tests to run but this is really exciting.show more

Dev Ed
660,154 görüntüleme • 7 ay önce
Our @Grammarly AI agents are here! Today, we’re launching... eight new AI agents designed for students and professionals. We created many of these agents with students in mind because they’re the first generation entering a job market where employers expect both subject expertise AND AI fluency. These agents help with everything from finding credible sources to predicting reader reactions. One agent we’ve gotten great feedback on is AI Grader (I wish I had this in school), which you can see in the video below. It looks at your assignment rubric and gives you suggestions like your professor would, and a grade prediction before you submit your work. And these agents are available in docs, our new AI-native writing surface! I’m deeply proud of this launch—docs is powered by Coda (Superhuman Docs) technology and is a great integration moment between Grammarly and Coda. This is just the beginning of Grammarly’s journey to offering agents that work everywhere people work and collaborate. I’ve been loving using these agents, and I’m excited for our customers to get access. Try them for yourself here and let me know what you think:show more

Shishir
13,564 görüntüleme • 1 yıl önce
This is wild, I just gave Kimi K3, Grok... 4.5, GPT 4.6 Sol, and Claude Opus 5 a starting cue, then asked them to finish the drawing themselves I also told them to be CREATIVE in their own way These are the results, and ngl I genuinely can't decide which one hits the bestshow more

Ann Nguyen
419,610 görüntüleme • 2 ay önce
Sakura vs Jinx, created with the AI Seedance 2.0.... Artificial intelligence is evolving at an impressive speed. It is already capable of generating a fight between a 2D character and a 3D character in a way that feels natural and believable. What is even more surprising is that this can be created in just a few minutes with a single prompt. Something that previously required a large team, significant resources, and a long production time can now be generated with just one well-crafted prompt.show more

nachos2d
240,509 görüntüleme • 7 ay önce
Grok 4.5 in Grok Build created an FPS game... in under an hour. The prompt was simple. I told it to write a game design document and pull free assets from the web. Then I had it create a TODO.md with implementation phases and run a loop to build out each phase. SpaceXAI and Cursor cooked here. This model is insane.show more

tetsuo
1,817,673 görüntüleme • 3 ay önce
KIMI K3 VS OPUS 4.8 SIMULATING A 3D PARTICLE... SYSTEM same prompt, same target: thousands of particles reacting to gravity and to each other in real time → K3: more organic distribution, smoother motion, particles actually cluster and drift like something physical → Opus: more rigid pattern, less natural, movement reads more like a grid than a simulation K3: $0.65 in API. Opus 4.8: ~$1.30 this is simulated physics, not just aesthetics, and it's the kind of detail that separates a demo from something you'd actually shipshow more

Jade
32,766 görüntüleme • 2 ay önce
BREAKING: Anthropic just dropped Opus 4.8—and it is a... MONSTER We've been testing for about a week Every 📧 and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check: - Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus 4.8 scores a 63—a hair higher than GPT-5.5's score of 62, and a full 30 points higher than Opus 4.7. It tackled a ground-up rewrite of a production codebase, and actually built something that works. HOWEVER: Coding performance varied a lot at different reasoning levels. We recommend using it on xhigh for best results. - Incredibly good writer. Opus 4.8 scored a 79.6 on our writing benchmark—measuring models on real-world writing tasks we do all of the time like essay writing, promo email writing, and more. It beats GPT-5.5 by 6 points. It produces well-written prose with fewer "AI-isms". It's also very good at writing in your voice given the right context. HOWEVER: Writing performance also varied with reasoning levels. Medium reasoning had higher incidence of AI-isms—we found best results with high. - Beast at knowledge work. Opus 4.8 is very good at general knowledge work tasks like report creation, research and more. It produced the best PowerPoint one-shot we've ever seen on our deck generation benchmark. - Emotionally intelligent, willing to question the frame. I've also found it to be quite good at talking through psychological or interpersonal issues. It has a high EQ, and it's also good at not glazing and helping to expand your perspective. Its thought process feels extremely rich and dynamic. THE BAD: These days a model is only as good as its harness, and Codex is still a far superior harness to the Claude Desktop app. This has kept me using Codex + GPT-5.5 as my daily driver, but I am flipping back and forth a lot more between Codex and Claude. Anthropic is back baby! Read the rest on Every 📧:show more

Dan Shipper
354,876 görüntüleme • 4 ay önce
Systems created from racism cannot be reformed. That is... why we keep seeing the same thing happen over and over again… with police and with ICE. An unarmed Black man, in Milwaukee, was shot in the head three times after being tased and restrained by five officers. Meanwhile, just 80 miles away, a white man was actively on top of a police officer, assaulting him… and was taken into custody without being shot. We have seen this same disparity play out over and over again. And we will continue to see this happen because… American policing is deeply rooted in systems created to control and hunt down Black populations... and ICE was created from a system rooted in racial exclusion, and the criminalization of immigrants. And we see the consequences of that history every single day. The problem is not just a few bad officers/agents, or a lack of training. The problem is not just that the wrong person was hired… These systems were built on racism. They were built to use state violence against certain people. You cannot reform a system whose foundation is rotten. You cannot train racism out of a system created by racism. You cannot reform your way out of a structure that is operating how it was designed to operate. At some point, we have to stop asking how to make these systems work better… And we have to start asking why we keep defending systems that were never designed to protect people in the first place.show more

Jesus Freakin Congress
48,945 görüntüleme • 2 ay önce
A lesson for every Polymarket bot developer: I built... a strategy that looked perfect on paper. Backtested it. Looked like a winner. Almost went live. Then i actually measured real costs. Strategy was dead before the first trade. And this is what bot building on Polymarket actually looks like. Here is what happened (and what you MUST know): Backtested mean reversion on crypto dips. SOL came back at +44% return and 70.5% win rate. Beautiful clean curve. Looked ready to ship. Then i measured real round-trip costs on SOL flash dips. Backtest assumed 0.45% in fees and slippage. Reality was 1.44%. Strategy stops working at 0.70%. Starting again. But the lesson was worth more than any profit the strategy could have made. Here is what i actually learned: Taking a dip with a market order means you eat the spread the dip just created. The volatility making your signal is the same volatility destroying your fill. You see the opportunity. You enter. You already lost. But resting a limit order below market and letting the dip come to you? You collect the maker rebate instead. Same thesis. Completely opposite execution. One bleeds money, one prints it. That one realization changed how i think about bot strategy entirely. 180 strategies tested to get there. 179 dead. That is not failure. That is how you find the 5 that actually work. Building a bot on Polymarket is not about finding a magic strategy. It is about eliminating every wrong answer until only the right one is left.show more

Oracle Boar
14,379 görüntüleme • 5 ay önce
This is f*cking insane. This tip saved me thousands... of dollars. run Opus 5.5, Sonnet 5.5, and Fable 5.1 together, and stop burning Opus on work it was never needed for. the whole idea in one line: the strong model plans, the mid-tier model executes, Fable stays quiet until it's actually needed. roles, broken down: Opus 5.5, high effort, owns the plan and ships the final code Sonnet 5.5, medium effort, splits into explorer (reads the codebase), worker (edits files, runs tests), researcher (pulls docs) Fable 5.1, called through /advisor fable, reads everything happening in the session but stays silent unless something's actually wrong three moments where Fable speaks: → a plan goes out: is this actually the right call? → the same failure shows up again: is the search going nowhere? → the task gets marked finished: did something get skipped? Jev engineering does the same thing one level down. the forks that don't need real thought, which file, which tool, keep going or stop, go straight to Jev and come back in under half a second. the big models only ever see the forks that genuinely need a decision. anyone still running one model for everything is paying Opus prices to decide whether a file exists. drop this into Claude Code: "Rebuild my Claude Code setup around this structure: Look through ~/.claude/agents and .claude/agents for subagents already covering explorer, worker, and researcher. Only create new ones for roles that are missing. Set model: sonnet, effort: medium on each. If an existing subagent is locked to a different model, leave it as is and just list it. In ~/.claude/settings.json, set effortLevel to high and advisorModel to fable. Check for anything disabling the advisor, CLAUDE_CODE_DISABLE_ADVISOR_TOOL, DISABLE_TELEMETRY, or anything blocking feature-flag fetches, plus CLAUDE_CODE_EFFORT_LEVEL, which can override subagent effort settings. Report what you find. Don't change any of it yet. Add one line to ~/.claude/CLAUDE.md: check in with the advisor before a big plan, when the same error shows up twice, and before marking a long task done. Show every change as a diff first. Wait for my go-ahead before touching anything."show more

rvaniaaa
249,004 görüntüleme • 5 gün önce
GPT Image 2 + Seedance 2 = cartoon chase... scene. I wanted to see how far I could push storyboard-to-video consistency, so I created a full 15-second cat vs mouse sequence inside a completely destroyed house. The process was actually simple: - generated a 3x3 cinematic storyboard sheet in GPT Image 2 - designed every frame like a real animation pre-production sketch - added camera directions, motion cues, timing notes, and destruction progression - used Seedance 2 to animate the entire sequence into one continuous cinematic shot flow What surprised me most was how well Seedance 2.0 understood the visual continuity between frames. The flying debris, motion blur, exaggerated jumps, and camera movement felt surprisingly cohesive. . . .show more

Beginnersblog
38,595 görüntüleme • 5 ay önce
Opus 4.6 vs GPT-5.4 (4/9) prompt: Build a production-quality... 3D flight-tracking web app using React + Vite + Three.js (react-three-fiber + drei) that visualizes live OpenSky aircraft data on a rotatable 3D Earth, with real-time plane motion, smooth interpolation, altitude-accurate positioning, and polished lighting/post-processing. Both models did really well on this one and honestly I’m impressed with both. GPT-5.4 had the nicer post-processing out of the box. I really liked the subtle light shimmer on the airplanes when rotating the planet, and the camera work when clicking a plane felt better overall. Opus 4.6 though had a few details I liked more. It automatically went and found a much nicer Earth texture on GitHub, while with GPT-5.4 I had to reprompt it to go look for a better one. I also preferred Opus’s plane model overall, it just looked more polished, whereas GPT-5.4’s plane asset looked a bit funny. One thing I noticed with GPT-5.4 is that when you click into the plane, the camera sometimes clips through the planet, which breaks the effect a bit. Opus handled that part more cleanly. Overall this felt like a strong result from both, just with different strengths. GPT-5.4 felt better on presentation and post-processing, while Opus had better asset choices and a more premium-looking Earth/plane combo.show more

Dev Ed
279,565 görüntüleme • 7 ay önce