Loading video...

Video Failed to Load

Go Home

deepseek v4.1 flash vs secret model tested both with the same prompt at the highest reasoning available > v4.1 flash took 2 hours to finish both tests and cost $2.60 > secret model finished both in 30 minutes, can’t reveal cost yet v4.1 flash spent a lot of time...

28,586 views • 3 days ago •via X (Twitter)

47 Comments

J A Z I I's profile picture
J A Z I I3 days ago

voxel pagoda test for both models

K2S's profile picture
K2S3 days ago

Bro how u use deepseek v4.1 flash in Ai hub mix last time I used 😒 hit the limit without even building something Are u using paid or free lmk?

J A Z I I's profile picture
J A Z I I3 days ago

i have credits, load like 5 to 10$ and you should be good for testing almost all chinese models

Alina Fomina's profile picture
Alina Fomina3 days ago

2 hours is wild. but still deepseek v4.1 flash looks solid

J A Z I I's profile picture
J A Z I I3 days ago

Agreed but missed alot stuff

Dorian's profile picture
Dorian3 days ago

And The secret model is (IMO) next Gemini flash model. Only Gemini have true flash Decode speed. Before i thought that's Grok 4.7, but quality is not good for Frontier model, and elon said that is slower than 4.6, but better overall

J A Z I I's profile picture
J A Z I I3 days ago

hehe

Dorian's profile picture
Dorian3 days ago

Deepseek 4.1 flash Has better Water detail, and i think overall deepseek is better

Hussain Hashim | Building SundayBack's profile picture
Hussain Hashim | Building SundayBack3 days ago

@notjazii Always the trade-off between speed and depth. Hit this wall when testing models for my own stuff. A balance is crucial!

Knowix's profile picture
Knowix3 days ago

having the same feeling as when Astra launched, might as well test this out myself

Yohaku's profile picture
Yohaku3 days ago

deepseek goated bro, love it

Shubh's profile picture
Shubh3 days ago

Cooking that well in 30 mins is really impressive legend 👀

Chii's profile picture
Chii3 days ago

high feels less overthinking

J A Z I I's profile picture
J A Z I I3 days ago

wait fr? I tested on max in Hermes so lemme try with minimax test

Chii's profile picture
Chii3 days ago

yeah try opencode

J A Z I I's profile picture
J A Z I I3 days ago

not working in opencode no idea why

Chii's profile picture
Chii3 days ago

need to add custom model

xabz's profile picture
xabz3 days ago

so when can we know the secret model's cost

J A Z I I's profile picture
J A Z I I3 days ago

maybe in few days

Shwepsik2121's profile picture
Shwepsik21213 days ago

Secret model is flash or pro?

J A Z I I's profile picture
J A Z I I3 days ago

it's fast fast I can't say flash or pro or air or lite or GPT 6-7

WolfSWAP | SWAP & WIN's profile picture
WolfSWAP | SWAP & WIN3 days ago

Secret model is way better hoping that it costed less

balega_dev's profile picture
balega_dev3 days ago

What is the secret model?

silent's profile picture
silent3 days ago

it's obvious grok 4.7 better

J A Z I I's profile picture
J A Z I I3 days ago

😂😭

Deepak Reddy's profile picture
Deepak Reddy3 days ago

ig we're cooked with Grok 4.7

Veee's profile picture
Veee3 days ago

what's the secret model tho? Qwen?

Varik Verilion's profile picture
Varik Verilion3 days ago

highest reasoning setting does that, it keeps re-checking an answer it already had.

AGI Pulse's profile picture
AGI Pulse3 days ago

ngl DS one looks more detailed

ahs's profile picture
ahs3 days ago

left one is definitely better

J A Z I I's profile picture
J A Z I I3 days ago

yeah it did better but some details are missing 100%

KXLAD3's profile picture
KXLAD33 days ago

Is deepseek good at coding? Also what's the cheapest way to get access? 😭😭

J A Z I I's profile picture
J A Z I I3 days ago

nothing for now soon it'll be opencode but wait for now

Farhan's profile picture
Farhan3 days ago

this is why speed matters so much for agent workflows. 4x faster changes the whole experience

Schubert Santos 🇺🇸|🇧🇷's profile picture
Schubert Santos 🇺🇸|🇧🇷3 days ago

This secret model seems really decent at 3D. Could you prompt some other stuff? Could be simple. There’s a thing I like to test in every model - I ask for a three js representation of a temperature and umidity reader using a esp32. So far Astra is winning, but still quite bugged

J A Z I I's profile picture
J A Z I I3 days ago

I ran 3d ship tests so I'll post it later tonight

kitik kote's profile picture
kitik kote3 days ago

try high thinking , 4.1v better high

Haleemah's profile picture
Haleemah3 days ago

secret hmm

EDDY VU's profile picture
EDDY VU3 days ago

Did the extra 90 minutes actually translate to a cleaner pagoda, or was it just burning tokens stuck in thinking loops?

BreezeOg🐉's profile picture
BreezeOg🐉3 days ago

Impressive speed upgrade, that’s a game changer in ai models of today

Deepu's profile picture
Deepu3 days ago

dsv4 looks better imo but 2 hrs is a lot btw jazi, how about not labelling the output next time and revealing in the comment instead, less bias and more fun 🙃

Atma Blabber's profile picture
Atma Blabber3 days ago

deepseek clearly better output. but time.....

DEV's profile picture
DEV3 days ago

Overthinking can tank efficiency. Speed without depth often leads to shallow results.

Dunamix™️®️'s profile picture
Dunamix™️®️3 days ago

which do you prefer?

J A Z I I's profile picture
J A Z I I3 days ago

I like both, as one did faster and got good details and other did way better

David's profile picture
David3 days ago

降价了

Super530's profile picture
Super5302 days ago

Its behavior appear to be different for language used for prompting. I don't see over think issue when poking (戳戳) in Chinese.

Related Videos

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,250 views • 14 days ago

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,202 views • 18 days ago