正在加载视频...

视频加载失败

deepseek v4.1 flash vs secret model tested both with the same prompt at the highest reasoning available > v4.1 flash took 2 hours to finish both tests and cost $2.60 > secret model finished both in 30 minutes, can’t reveal cost yet v4.1 flash spent a lot of time...

28,586 次观看 • 3 天前 •via X (Twitter)

47 条评论

J A Z I I 的头像
J A Z I I3 天前

voxel pagoda test for both models

K2S 的头像
K2S3 天前

Bro how u use deepseek v4.1 flash in Ai hub mix last time I used 😒 hit the limit without even building something Are u using paid or free lmk?

J A Z I I 的头像
J A Z I I3 天前

i have credits, load like 5 to 10$ and you should be good for testing almost all chinese models

Alina Fomina 的头像
Alina Fomina3 天前

2 hours is wild. but still deepseek v4.1 flash looks solid

J A Z I I 的头像
J A Z I I3 天前

Agreed but missed alot stuff

Dorian 的头像
Dorian3 天前

And The secret model is (IMO) next Gemini flash model. Only Gemini have true flash Decode speed. Before i thought that's Grok 4.7, but quality is not good for Frontier model, and elon said that is slower than 4.6, but better overall

J A Z I I 的头像
J A Z I I3 天前

hehe

Dorian 的头像
Dorian3 天前

Deepseek 4.1 flash Has better Water detail, and i think overall deepseek is better

Hussain Hashim | Building SundayBack 的头像
Hussain Hashim | Building SundayBack3 天前

@notjazii Always the trade-off between speed and depth. Hit this wall when testing models for my own stuff. A balance is crucial!

Knowix 的头像
Knowix3 天前

having the same feeling as when Astra launched, might as well test this out myself

Yohaku 的头像
Yohaku3 天前

deepseek goated bro, love it

Shubh 的头像
Shubh3 天前

Cooking that well in 30 mins is really impressive legend 👀

Chii 的头像
Chii3 天前

high feels less overthinking

J A Z I I 的头像
J A Z I I3 天前

wait fr? I tested on max in Hermes so lemme try with minimax test

Chii 的头像
Chii3 天前

yeah try opencode

J A Z I I 的头像
J A Z I I3 天前

not working in opencode no idea why

Chii 的头像
Chii3 天前

need to add custom model

xabz 的头像
xabz3 天前

so when can we know the secret model's cost

J A Z I I 的头像
J A Z I I3 天前

maybe in few days

Shwepsik2121 的头像
Shwepsik21213 天前

Secret model is flash or pro?

J A Z I I 的头像
J A Z I I3 天前

it's fast fast I can't say flash or pro or air or lite or GPT 6-7

WolfSWAP | SWAP & WIN 的头像
WolfSWAP | SWAP & WIN3 天前

Secret model is way better hoping that it costed less

balega_dev 的头像
balega_dev3 天前

What is the secret model?

silent 的头像
silent3 天前

it's obvious grok 4.7 better

J A Z I I 的头像
J A Z I I3 天前

😂😭

Deepak Reddy 的头像
Deepak Reddy3 天前

ig we're cooked with Grok 4.7

Veee 的头像
Veee3 天前

what's the secret model tho? Qwen?

Varik Verilion 的头像
Varik Verilion3 天前

highest reasoning setting does that, it keeps re-checking an answer it already had.

AGI Pulse 的头像
AGI Pulse3 天前

ngl DS one looks more detailed

ahs 的头像
ahs3 天前

left one is definitely better

J A Z I I 的头像
J A Z I I3 天前

yeah it did better but some details are missing 100%

KXLAD3 的头像
KXLAD33 天前

Is deepseek good at coding? Also what's the cheapest way to get access? 😭😭

J A Z I I 的头像
J A Z I I3 天前

nothing for now soon it'll be opencode but wait for now

Farhan 的头像
Farhan3 天前

this is why speed matters so much for agent workflows. 4x faster changes the whole experience

Schubert Santos 🇺🇸|🇧🇷 的头像
Schubert Santos 🇺🇸|🇧🇷3 天前

This secret model seems really decent at 3D. Could you prompt some other stuff? Could be simple. There’s a thing I like to test in every model - I ask for a three js representation of a temperature and umidity reader using a esp32. So far Astra is winning, but still quite bugged

J A Z I I 的头像
J A Z I I3 天前

I ran 3d ship tests so I'll post it later tonight

kitik kote 的头像
kitik kote3 天前

try high thinking , 4.1v better high

Haleemah 的头像
Haleemah3 天前

secret hmm

EDDY VU 的头像
EDDY VU3 天前

Did the extra 90 minutes actually translate to a cleaner pagoda, or was it just burning tokens stuck in thinking loops?

BreezeOg🐉 的头像
BreezeOg🐉3 天前

Impressive speed upgrade, that’s a game changer in ai models of today

Deepu 的头像
Deepu3 天前

dsv4 looks better imo but 2 hrs is a lot btw jazi, how about not labelling the output next time and revealing in the comment instead, less bias and more fun 🙃

Atma Blabber 的头像
Atma Blabber3 天前

deepseek clearly better output. but time.....

DEV 的头像
DEV3 天前

Overthinking can tank efficiency. Speed without depth often leads to shallow results.

Dunamix™️®️ 的头像
Dunamix™️®️3 天前

which do you prefer?

J A Z I I 的头像
J A Z I I3 天前

I like both, as one did faster and got good details and other did way better

David 的头像
David3 天前

降价了

Super530 的头像
Super5302 天前

Its behavior appear to be different for language used for prompting. I don't see over think issue when poking (戳戳) in Chinese.

相关视频

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,250 次观看 • 14 天前

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,202 次观看 • 18 天前