Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

deepseek v4.1 flash vs secret model tested both with the same prompt at the highest reasoning available > v4.1 flash took 2 hours to finish both tests and cost $2.60 > secret model finished both in 30 minutes, can’t reveal cost yet v4.1 flash spent a lot of time...

28,586 Aufrufe • vor 3 Tagen •via X (Twitter)

47 Kommentare

Profilbild von J A Z I I
J A Z I Ivor 3 Tagen

voxel pagoda test for both models

Profilbild von K2S
K2Svor 3 Tagen

Bro how u use deepseek v4.1 flash in Ai hub mix last time I used 😒 hit the limit without even building something Are u using paid or free lmk?

Profilbild von J A Z I I
J A Z I Ivor 3 Tagen

i have credits, load like 5 to 10$ and you should be good for testing almost all chinese models

Profilbild von Alina Fomina
Alina Fominavor 3 Tagen

2 hours is wild. but still deepseek v4.1 flash looks solid

Profilbild von J A Z I I
J A Z I Ivor 3 Tagen

Agreed but missed alot stuff

Profilbild von Dorian
Dorianvor 3 Tagen

And The secret model is (IMO) next Gemini flash model. Only Gemini have true flash Decode speed. Before i thought that's Grok 4.7, but quality is not good for Frontier model, and elon said that is slower than 4.6, but better overall

Profilbild von J A Z I I
J A Z I Ivor 3 Tagen

hehe

Profilbild von Dorian
Dorianvor 3 Tagen

Deepseek 4.1 flash Has better Water detail, and i think overall deepseek is better

Profilbild von Hussain Hashim | Building SundayBack
Hussain Hashim | Building SundayBackvor 3 Tagen

@notjazii Always the trade-off between speed and depth. Hit this wall when testing models for my own stuff. A balance is crucial!

Profilbild von Knowix
Knowixvor 3 Tagen

having the same feeling as when Astra launched, might as well test this out myself

Profilbild von Yohaku
Yohakuvor 3 Tagen

deepseek goated bro, love it

Profilbild von Shubh
Shubhvor 3 Tagen

Cooking that well in 30 mins is really impressive legend 👀

Profilbild von Chii
Chiivor 3 Tagen

high feels less overthinking

Profilbild von J A Z I I
J A Z I Ivor 3 Tagen

wait fr? I tested on max in Hermes so lemme try with minimax test

Profilbild von Chii
Chiivor 3 Tagen

yeah try opencode

Profilbild von J A Z I I
J A Z I Ivor 3 Tagen

not working in opencode no idea why

Profilbild von Chii
Chiivor 3 Tagen

need to add custom model

Profilbild von xabz
xabzvor 3 Tagen

so when can we know the secret model's cost

Profilbild von J A Z I I
J A Z I Ivor 3 Tagen

maybe in few days

Profilbild von Shwepsik2121
Shwepsik2121vor 3 Tagen

Secret model is flash or pro?

Profilbild von J A Z I I
J A Z I Ivor 3 Tagen

it's fast fast I can't say flash or pro or air or lite or GPT 6-7

Profilbild von WolfSWAP | SWAP & WIN
WolfSWAP | SWAP & WINvor 3 Tagen

Secret model is way better hoping that it costed less

Profilbild von balega_dev
balega_devvor 2 Tagen

What is the secret model?

Profilbild von silent
silentvor 3 Tagen

it's obvious grok 4.7 better

Profilbild von J A Z I I
J A Z I Ivor 3 Tagen

😂😭

Profilbild von Deepak Reddy
Deepak Reddyvor 3 Tagen

ig we're cooked with Grok 4.7

Profilbild von Veee
Veeevor 3 Tagen

what's the secret model tho? Qwen?

Profilbild von Varik Verilion
Varik Verilionvor 3 Tagen

highest reasoning setting does that, it keeps re-checking an answer it already had.

Profilbild von AGI Pulse
AGI Pulsevor 3 Tagen

ngl DS one looks more detailed

Profilbild von ahs
ahsvor 3 Tagen

left one is definitely better

Profilbild von J A Z I I
J A Z I Ivor 3 Tagen

yeah it did better but some details are missing 100%

Profilbild von KXLAD3
KXLAD3vor 2 Tagen

Is deepseek good at coding? Also what's the cheapest way to get access? 😭😭

Profilbild von J A Z I I
J A Z I Ivor 2 Tagen

nothing for now soon it'll be opencode but wait for now

Profilbild von Farhan
Farhanvor 3 Tagen

this is why speed matters so much for agent workflows. 4x faster changes the whole experience

Profilbild von Schubert Santos 🇺🇸|🇧🇷
Schubert Santos 🇺🇸|🇧🇷vor 3 Tagen

This secret model seems really decent at 3D. Could you prompt some other stuff? Could be simple. There’s a thing I like to test in every model - I ask for a three js representation of a temperature and umidity reader using a esp32. So far Astra is winning, but still quite bugged

Profilbild von J A Z I I
J A Z I Ivor 3 Tagen

I ran 3d ship tests so I'll post it later tonight

Profilbild von kitik kote
kitik kotevor 3 Tagen

try high thinking , 4.1v better high

Profilbild von Haleemah
Haleemahvor 3 Tagen

secret hmm

Profilbild von EDDY VU
EDDY VUvor 3 Tagen

Did the extra 90 minutes actually translate to a cleaner pagoda, or was it just burning tokens stuck in thinking loops?

Profilbild von BreezeOg🐉
BreezeOg🐉vor 3 Tagen

Impressive speed upgrade, that’s a game changer in ai models of today

Profilbild von Deepu
Deepuvor 3 Tagen

dsv4 looks better imo but 2 hrs is a lot btw jazi, how about not labelling the output next time and revealing in the comment instead, less bias and more fun 🙃

Profilbild von Atma Blabber
Atma Blabbervor 3 Tagen

deepseek clearly better output. but time.....

Profilbild von DEV
DEVvor 3 Tagen

Overthinking can tank efficiency. Speed without depth often leads to shallow results.

Profilbild von Dunamix™️®️
Dunamix™️®️vor 3 Tagen

which do you prefer?

Profilbild von J A Z I I
J A Z I Ivor 3 Tagen

I like both, as one did faster and got good details and other did way better

Profilbild von David
Davidvor 3 Tagen

降价了

Profilbild von Super530
Super530vor 2 Tagen

Its behavior appear to be different for language used for prompting. I don't see over think issue when poking (戳戳) in Chinese.

Ähnliche Videos

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,250 Aufrufe • vor 13 Tagen

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,202 Aufrufe • vor 17 Tagen