Loading video...

Video Failed to Load

Go Home

Tested Mimo-V2.6-Flash, Grok-4.7, and DeepSeek-V4.1-Flash with 3D SubwayBench using the same /design prompt. 🔹 DSV4.1-Flash: 9/10 · $0.024 · one-shot · strong 3D UI/UX + gameplay 🔹 Mimo-V2.6-Flash: 8/10 · $0.018 · playable 2D result in ~4 iterations 🔹 Grok 4.7: 3/10 · $0.35 · slower UI + gameplay,...

31,073 views • 5 days ago •via X (Twitter)

17 Comments

flep ✪'s profile picture
flep ✪5 days ago

My god the Grok 4.7 version... 🤡

Command Code's profile picture
Command Code5 days ago

Our engineering and design team has been testing 26+ side-by-side comparisons across frontier and open models. All runs are public and open source. Benchmark for this demo here:

Natsume's profile picture
Natsume5 days ago

dsv4.1flash always

Sirius C's profile picture
Sirius C5 days ago

Grok made it in paint lol

Ernesto Davila Tirado's profile picture
Ernesto Davila Tirado5 days ago

Grok what a joke.

Maedah Batool's profile picture
Maedah Batool5 days ago

It’s always fun to ship this.

Djekyll Vs Hyde's profile picture
Djekyll Vs Hyde5 days ago

MiMo-V2.6-Flash and the latest DeepSeek are getting seriously good at the creation part. Give them a strong design prompt and they can go surprisingly far from idea → playable experience. The interesting part is no longer just “can it code?” it’s what can you actually create with it? 👀

UI Pirate's profile picture
UI Pirate5 days ago

deepseek doing the 3d ui in one shot while grok needs more passes is kinda funny. models still wild different on design prompts

kost0v's profile picture
kost0v5 days ago

One /design prompt per model isn't much of a sample, prompt phrasing alone could swing Grok's score way more than 3 pts.

Dangcing Chan's profile picture
Dangcing Chan5 days ago

Thanks for sharing this! Really made my day.

orvir's profile picture
orvir5 days ago

The iteration count changes the cost comparison here. DeepSeek’s one-shot result is the clearest win; MiMo’s $0.018 headline needs the total across all four attempts.

aixbwixbeia's profile picture
aixbwixbeia5 days ago

怎么deepseek官方api炸了怎么你们的ds 4.1flash还能用?😅

BranchGhost's profile picture
BranchGhost5 days ago

Elon, this is a disgrace.

𝙰𝙱𝙷𝙸𝚂𝙷Ξ𝙺 𝙿𝙰𝚃𝙸𝙻's profile picture
𝙰𝙱𝙷𝙸𝚂𝙷Ξ𝙺 𝙿𝙰𝚃𝙸𝙻5 days ago

@rails recently said, and I quote "DeepSeek 4.1 Flash, figured out it was being benchmarked and tried to hack its way to a better score." What are the chances the same is happening here? If anything, it's a complement to DS- it's so good, it's almost suspicious lol.

CHLOVEN PRICE's profile picture
CHLOVEN PRICE5 days ago

go plan need good mimo2.6 deal.

Darragh's profile picture
Darragh4 days ago

Same prompt, three models, real cost. That's how you build a routing table. Keep doing that.

Marc's profile picture
Marc5 days ago

DS v4.1 flash holding the line for flash models! Although, GLM-5.3-flash has been better for some models. I would put them both at the same tier, DS a little bit higher because of its speed

Related Videos

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,470 views • 1 month ago

hy3 vs mimo-v2.5 vs deepseek v4 flash vs minimax m3 the four models on top of the openrouter leaderboard by tokens this week: #1 hy3 (Tencent Hy) – 7.5t #2 mimo-v2.5 (Xiaomi MiMo) – 6.56t #3 deepseek v4 flash (DeepSeek) – 5.24t #4 minimax m3 (MiniMax (official)) – 4.21t so we tested them. 3 prompts, single-file html, Three.js from a cdn, fully procedural, no external assets. all run via AI/ML API each prompt is a transparent cutaway machine that has to be mechanically correct, not decorative: • 4-stroke engine with full oil circulation – slider-crank kinematics, cam at 2:1, valve lift driven by lobes, oil loop from sump to gallery to big-end • watt walking-beam steam engine – four-bar vector-loop closure, eccentric-driven slide valve, steam events synced to real port position • francis reaction water turbine – 20 guide vanes on a regulating ring, 17 lofted runner blades, gpu particle advection, precessing vortex rope at part load the takeaway up front: none of the four cleared all three scenes on the first attempt. but the price spread between them is roughly 70x – hy3 fixed included costs less than two cents overall results (summed across all 3 scenes): cost #1 hy3 – $0.016 #2 deepseek v4 flash – $0.025 #3 mimo-v2.5 – $0.97 #4 minimax m3 – $1.17 tokens #1 hy3 – 19,326 #2 deepseek v4 flash – 63,126 #3 mimo-v2.5 – 322,523 #4 minimax m3 – 702,900 lines of code #1 hy3 – 1,047 #2 mimo-v2.5 – 2,759 #3 deepseek v4 flash – 3,273 #4 minimax m3 – 3,354 scenes needing a second attempt #1 hy3 – 1 (engine) #1 mimo-v2.5 – 1 (turbine) #1 minimax m3 – 1 (turbine) #4 deepseek v4 flash – 2 (steam engine, turbine) observations: 1. the token spread is the real story – minimax burns 36x hy3's tokens and lands in the same place, one retry, ~3.3k lines 2. hy3 is the outlier on density: 1,047 lines total, fewest tokens, cheapest run, and only one scene needed a second pass. deepseek is the opposite trade – near-hy3 pricing but the most retries 3. mimo and minimax seem to overthink instead of writing the code. minimax spent 359.1k tokens on the steam engine and produced 1,346 lines – the tokens are going somewhere other than the file 4. the francis turbine broke three of the four. the spec that separates them is the one with 20 linked guide vanes and gpu particle advection, not the one with the most parts overall impression: none of these models excelled at any of the tasks we gave them. but they were close, and they were extremely cheap. the gap that matters isn't quality anymore – it's that hy3 ran all three scenes for less than two cents while the frontier labs charge dollars for the same work right now you pick these because they're good for the zero price you pay. soon that's something openai and anthropic will have to think about follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

17,145 views • 2 months ago

glm 5.3 flash is 7.5x cheaper, but 3.4x slower than gemini 3.7 flash Z.ai glm 5.3 flash – shipped aug 26, $0.07/$0.25 per 1m Google DeepMind gemini 3.7 flash – shipped aug 13, $0.38/$1.88 per 1m we put the two models on one job: write one html file that draws an animated 3d scene in the browser. no images, no downloads, and it has to look the same on every load. the setup: three scenes – a glass aquarium in a lit room, the solar system, a night city under a thunderstorm. identical brief word for word, reasoning effort high, 64k output cap. the numbers below are not the whole run. they cover the three scenes we kept – the best one per task from each model, the ones in the video. - total generation time for the three scenes #1 gemini 3.7 flash – 10m 36s #2 glm 5.3 flash – 36m 30s - tokens spent on those three scenes #1 glm 5.3 flash – 110k #2 gemini 3.7 flash – 111k - cost of those three scenes #1 glm 5.3 flash – $0.027 #2 gemini 3.7 flash – $0.202 observations: • glm's first 10 attempts: 7 blank pages. it kept inventing short random helpers and forgetting to define one of them. the fix was one line in the brief: use exactly one random helper, named rand(), and don't invent shorthands next to it. next 12 attempts: 11 alive, 0 crashes. • glm spends 66% of its output on reasoning, gemini 57%. that is the whole speed gap. • gemini's storm came back as a black rectangle in 4 of 6 runs. glm's best storm has a branching bolt, lit rain and wet asphalt – for $0.01. conclusion: same three scenes, same token spend – glm 5.3 flash billed $0.027 and took 36m 30s, gemini 3.7 flash billed $0.202 and took 10m 36s. glm wins gemini on price and made the best storm of the whole run follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

16,041 views • 1 month ago

Kling 3.0 vs Gemini Omni Flash vs Grok Imagine 1.5 vs Seedance 2.0. Another battle for the best AI video tool. This time I tested an extremely difficult stunt scene: a bridge jump, landing on a moving truck, then jumping onto a car and taking it over. No model handled it perfectly. Not even Seedance 2.0. The scene itself was very hard, but I kept the prompt relatively simple, without stuffing it with too many complex stunt terms. After Veo 3.1’s terrible results in previous rounds, I didn’t even include it this time. Seedance needed 4 attempts to give me a decent result. Gemini Omni Flash needed even more: 7 generations, and I picked the best one. Kling and Grok got 4 attempts each. None of them looked promising enough to justify more retries. My take: Grok Imagine 1.5, the new “revolutionary” release, once again showed that it still can’t handle complex action scenes properly. Two tests in a row now, and the result is objectively weak. Kling 3.0 feels outdated, but in this test it actually looked slightly better than Grok Imagine 1.5, which says a lot. Gemini Omni Flash gives decent movement, but the physics and action logic are clearly off. It’s better than Grok and Kling here, but still very far from the super-generator Google markets it as. Seedance 2.0 was the best again, as expected. Still not perfect. It also has a clear physics issue, because pulling a driver out of a moving car while standing on the road makes no physical sense. But overall, Seedance still produced a much more intense and cinematic scene than all the others. My ranking for this test: 1. Seedance 2.0 — clear winner, even with flaws 2. Gemini Omni Flash 3. Kling 3.0 4. Grok Imagine 1.5 — unfortunately, the weakest right now in my opinion Do you agree? #AIVideo

Alpha Mom

67,241 views • 3 months ago

My feed has been inundated with posts of Grok 3 making basic arcade games. But llms from years ago could make decent arcade games, not news. So I ran a one-shot test to determine how well it fared again other frontier models in creating a 3D game with room for it to come up with gameplay and aesthetics. I tested Grok 3, O1, Sonnet 3,5, Llama4, DeepSeek, and Gemini using the following prompt. Make Dune x Minecraft 🏜️ Imagine a sandbox survival game set on a desert planet. Players mine ‘spice’ and must build defenses against roaming sandworms. Design the main gameplay loop, crafting system, and survival challenges in one complete description. ✏️tldr O1, Grok 3, and Sonnet 3.5 were the most impressive. Aesthetically, Grok nailed the best vibes (it even produced a surprisingly cool-looking spice mining truck), but the game lacked functionality. O1 took the top spot imo with a functional and visually appealing experience, and Sonnet 3.5 followed closely. This is obviously just one test, but you can see the generated code and games in the thread (and even try forking them on Rosebud). Longer summary: OpenAI O1: Best vibes to function balance. Looked good, working controls, I could mine spice. xAI #Grok3: Excelled at generating vibes for dune. I especially liked the Dune-inspired spice mining car—though it wasn’t entirely a complete game. I could move around, but none of the crafting mechanics worked. Anthropic Sonnet3.5: produced something in space that had dune vibes. More functional than Grok because I could mine spice. However vibes were worse than the first two. DeepSeek : Managed to generate code that worked, but the game was so hard it always ended seconds after it started, and despite requests for better visuals, it looked VERY ugly. Google DeepMind Gemini 2.0 flash and AI at Meta LLaMA: Sadly landed at the bottom of the list; after multiple prompts (this was supposed to be one shot and none of the others failed in the first shot), I couldn’t get them to produce working code for this prompt. All of these were tested on Rosebud AI . An obvious limitation with these frontier models in their chat interfaces is that you can only get them to regenerate code from scratch each time you prompt them, making it tough to refine or extend a single project. Rosebud, on the other hand, lets you iterate on one project (we do diffs), deploy with one click, share your project, and even allow others to remix it. This was just a single test, so it’s obviously not scientific. I wanted to create it to see how these frontier models handle more complex game prompts—rather than retrying the same arcade games that earlier generations of LLMs have already mastered.

Lisha

476,022 views • 1 year ago

I've tested every vibe coding resource I could find this year. These are the 28 I actually kept, grouped by what you're stuck on. Bookmark this one 🔖 🧵👇 1. - 200+full build prompts for 3D and scroll-driven sites, each one a complete design brief rather than a component you drop in. 2. Minimal Gallery ( — curated high-end sites to feed your agent 3. Kage ( — real UI inspiration mapped straight to prompts 4. Refero Styles ( — 2,000+ real product styles with typography 5. Component Gallery ( — 2,600+ examples of how top design systems solve the same element 6. AppShot Gallery ( — real app screenshots for mobile 7. Navbar Gallery ( 8. Footer Design ( 9. CTA Gallery ( — conversion-tested forms, popups, buttons 10. 404s ( 11. DESIGNmd ( — design systems as markdown your agent can read 12 VibePrompt ( — ready-made prompts for dashboards and landing pages 13. — component registry that plugs into agents via MCP 14. Kinetics ( — 150+ motion effects with React code and prompts 15. shadcn/ui ( — still the gold standard 16. Aceternity UI ( — 200+ animated React/Tailwind 17. Magic UI ( 18. Motion Primitives ( 19. Uiverse ( 20. UIAble ( 21. mapcn ( — map components, markers, routes, popups 22. MicroKit UI ( — micro-interactions 23. Liquid Glass ( — glass refraction components 24. CSS Text Effects ( 25. Circle Loaders ( 26. Gradient Buttons ( 27. Kitbitz ( — 2,000+ hand-drawn illustrations 28. 3Dicons ( 29. Anime.js (

Himanshu Hingorani

497,024 views • 3 days ago

⬛ Austin Ekeler: From D2 to the League He ran for 5,857 yards at Western Colorado Football. Signed with the Los Angeles Chargers as an undrafted free agent. Led the NFL in touchdowns twice. Now with the Washington Commanders and still producing strong. He’s one of the best to ever come out of Division II. ⬇️ Western State, now Western Colorado, was the only program that offered him a shot at running back. He stayed four full years and became one of the most productive players the RMAC has ever seen. He averaged 146 rushing yards and 1.6 touchdowns per game over 40 starts. He put together one of the strongest careers in D2 history...and he didn’t have to leave to do it. At Western, Ekeler dominated. He averaged more than 100 rushing yards per game across four seasons, led the country in all-purpose yards in 2015, and earned academic honors alongside his on-field production. By the time he wrapped up his senior year, he was a Harlon Hill finalist with nearly 6,000 rushing yards to his name. His work was undeniable. But he still didn’t get a Combine invite. The draft came and went. No calls. Then the Los Angeles Chargers brought him into camp on a rookie deal and gave him a shot. He took it from there. He made the roster as a rookie and started producing immediately. First as a change-of-pace option. Then as a third-down threat. Then as the guy. In 2021 and 2022, he led the NFL in total touchdowns. He stacked back-to-back seasons with more than 1,500 scrimmage yards and at least 15 total scores. He caught 107 passes in 2022, the most by any running back in Los Angeles Chargers history. The production hasn’t slowed down. In 2024, he signed with the Washington Commanders. He’s still outworking people. Division II didn’t slow him down. It prepared him. The hours. The discipline. The ability to carry the weight of a program while juggling everything else college throws at you. That’s what this level teaches. That’s why it sticks. There are players right now doing the same thing. Grinding through the spring, stacking film, putting together complete seasons. Some will make it. Not because someone finally believed in them, but because they never stopped believing in themselves. That’s what turns a D2 shot into an NFL career. ⬛ College Career 🔹 5,857 rushing yards, 63 total touchdowns 🔹 4× First-Team All-RMAC Sports 🔹 2013 RMAC Offensive Freshman of the Year 🔹 Led D2 in all-purpose yards per game (203.9) in 2015 🔹 2016 Harlon Hill Trophy Finalist 🔹 First-Team Academic All-America (CoSIDA) 🔹 RMAC Academic Player of the Year (2014) 🔹 2018 Western Colorado Alumni Award of Excellence ⬛ NFL Career 🔹 Signed by Los Angeles Chargers in 2017 (UDFA) 🔹 NFL leader in total touchdowns: 20 (2021), 18 (2022) 🔹 Only player with 10+ rush and 5+ receiving TDs in back-to-back seasons since Marshall Faulk 🔹 One of two UDFAs with 1,500+ scrimmage yards and 15+ TDs in two straight years (with Priest Holmes) 🔹 107 catches in 2022, most by a running back in Los Angeles Chargers history 🔹 PFF Second-Team All-Pro (2019) 🔹 NFL Top 100: Ranked #21 in 2023 🔹 AFC Offensive Player of the Week (Week 17, 2022) 🔹 Over 7,000 scrimmage yards and 70+ career touchdowns 🔹 Signed with the Washington Commanders in 2024 ⬛ Off the Field 🔹 Owns more than 115 rental properties in Colorado and Missouri 🔹 Built the Eksperience app to connect athletes and fans 🔹 Co-founded Gridiron Gaming Group 🔹 75+ endorsement deals (Adidas, Chipotle, Frito-Lay, more) 🔹 Founded the Austin Ekeler Foundation in 2021 🔹 Former host of “Ekeler’s Edge” with Yahoo Sports Ekeler’s story isn’t rare because he came from Division II. It stands out because he stayed in it, believed in it, and maxed it out. That’s what this level does when the right player buys in. Who’s next? #D2Football #AustinEkeler #WesternColorado #RMAC #NFL #Undrafted #Division2 #D2Built #D2ToTheLeague

D2 College Football Spotlight

19,281 views • 1 year ago

HERMES AGENT SUPPORTS 300+ MODELS. PICKING THE RIGHT ONE PER TASK IS THE DIFFERENCE BETWEEN $5/MONTH AND $50. STARTING OUT: Claude Sonnet 4.6. official recommendation from Nous Research. "the model this project was built and tested with." strong reasoning. reliable tool calling. mid-range pricing. PREMIUM TIER: Claude Opus 4.8. best coding benchmarks available. self-correcting reasoning. catches its own mistakes. 1M context. use for demanding tasks where quality matters. GPT-5.5. #1 Chatbot Arena. #1 GPQA Diamond reasoning (94.1%). #1 creative writing. 2M context. handles entire codebases in one pass. Grok 4.30. the only frontier model with live X firehose access. real-time social data, breaking news, market sentiment. connects via Grok OAuth. no separate API key. Grok-Composer-2.5-Fast (v0.17.0). Cursor's coding model. 200K context. available through your Grok subscription via OAuth. no extra cost if you already pay for Grok. MID-RANGE TIER: Claude Sonnet 4.6. best balance of quality and cost for daily use. strongest prose and tool calling in this tier. Gemini 2.5 Pro. Google Search grounding built in. cites sources. verifies claims. pulls current data. 2M context. best for research-heavy workflows. GPT-4.1. reliable tool calling. solid general reasoning. good middle ground when you need OpenAI compatibility. BUDGET TIER: Claude Haiku 4.5. fastest Anthropic model. cheapest paid Claude option. strong at classification, routing, simple queries. use for auxiliary tasks: compression, vision, web extraction, approval scoring. DeepSeek V4. best cost-to-quality ratio in the market. 90% cache discount on repeated context. use for sub-agents and bulk parallel work. DeepSeek V4 Flash. cheapest paid model worth using. 1M context. MIT license. self-hostable. use for cron jobs, monitoring, routine searches. MiniMax M3. Nous Research and MiniMax collaborating on optimization. 1M context via lightning attention. 59% SWE-Bench Pro. beats several premium models on coding. one of the most-used models inside Hermes. FREE / LOCAL: Qwen 3.5 27B via Ollama. 16GB VRAM. reliable tool calling. best free local model for Hermes as of mid-2026. Qwen 3 8B. 8GB VRAM. fits a $7 VPS. handles routine tasks at zero API cost. Llama 4 Maverick. best open-weight tool calling. 1M context. needs more VRAM but strongest local option. HOW TO ASSIGN MODELS: main model: Desktop app / Dashboard → Models → switch sub-agent model: set in Desktop app, Dashboard, or config.yaml: delegation: model: "deepseek/deepseek-v4" auxiliary models (compression, vision, web extract): Desktop app / Dashboard → Models → Auxiliary Haiku 4.5 or Gemini Flash work well here. saves significantly when your main model is premium. per-profile: each Hermes profile gets its own model. Scout on DeepSeek. Analyst on Sonnet. Briefer on budget model. Coder on Opus. per-cron-job: pin a specific model to any cron job. morning brief on Haiku. deep research on Sonnet. monitoring on DeepSeek Flash. each job uses only the model it needs. per-session: /model deepseek/deepseek-v4-flash hot-swap mid-conversation. no restart needed. FALLBACK CHAINS: if your primary model is unavailable, Hermes automatically switches to the next provider. rate limit or server error = next model in the chain. no failed runs. no manual intervention. set in Desktop app, Dashboard, or config.yaml: fallback_providers: - openrouter - nous - codex PROVIDER PATHS: OPENROUTER: 300+ models under one API key. pay per token. most flexible. NOUS PORTAL: 300+ models + Tool Gateway (web search, image gen, TTS, browser). one OAuth. one subscription. 10% off token-billed providers. CHATGPT SUB: GPT-5.5 + Grok via OAuth. included tokens with $20 subscription. OLLAMA: free. local. private. zero API cost. your hardware only. mix providers across profiles and tasks. Scout on OpenRouter. Analyst on Nous Portal. Coder on ChatGPT sub. Monitor on Ollama. THE RULE: premium for work that needs deep reasoning. mid-range for daily driver tasks. budget for volume and background work. free for monitoring and routine jobs. pricing changes fast. check openrouter ai for current rates before committing. Which is your favourite model and for what task? full 15 levels breakdown in the article 👇

YanXbt

17,138 views • 3 months ago