Загрузка видео...

Не удалось загрузить видео

На главную

deepseek v4.1 flash vs secret model tested both with the same prompt at the highest reasoning available > v4.1 flash took 2 hours to finish both tests and cost $2.60 > secret model finished both in 30 minutes, can’t reveal cost yet v4.1 flash spent a lot of time...

28,586 просмотров • 3 дней назад •via X (Twitter)

Комментарии: 47

Фото профиля J A Z I I
J A Z I I3 дней назад

voxel pagoda test for both models

Фото профиля K2S
K2S3 дней назад

Bro how u use deepseek v4.1 flash in Ai hub mix last time I used 😒 hit the limit without even building something Are u using paid or free lmk?

Фото профиля J A Z I I
J A Z I I3 дней назад

i have credits, load like 5 to 10$ and you should be good for testing almost all chinese models

Фото профиля Alina Fomina
Alina Fomina3 дней назад

2 hours is wild. but still deepseek v4.1 flash looks solid

Фото профиля J A Z I I
J A Z I I3 дней назад

Agreed but missed alot stuff

Фото профиля Dorian
Dorian3 дней назад

And The secret model is (IMO) next Gemini flash model. Only Gemini have true flash Decode speed. Before i thought that's Grok 4.7, but quality is not good for Frontier model, and elon said that is slower than 4.6, but better overall

Фото профиля J A Z I I
J A Z I I3 дней назад

hehe

Фото профиля Dorian
Dorian3 дней назад

Deepseek 4.1 flash Has better Water detail, and i think overall deepseek is better

Фото профиля Hussain Hashim | Building SundayBack
Hussain Hashim | Building SundayBack3 дней назад

@notjazii Always the trade-off between speed and depth. Hit this wall when testing models for my own stuff. A balance is crucial!

Фото профиля Knowix
Knowix3 дней назад

having the same feeling as when Astra launched, might as well test this out myself

Фото профиля Yohaku
Yohaku3 дней назад

deepseek goated bro, love it

Фото профиля Shubh
Shubh3 дней назад

Cooking that well in 30 mins is really impressive legend 👀

Фото профиля Chii
Chii3 дней назад

high feels less overthinking

Фото профиля J A Z I I
J A Z I I3 дней назад

wait fr? I tested on max in Hermes so lemme try with minimax test

Фото профиля Chii
Chii3 дней назад

yeah try opencode

Фото профиля J A Z I I
J A Z I I3 дней назад

not working in opencode no idea why

Фото профиля Chii
Chii3 дней назад

need to add custom model

Фото профиля xabz
xabz3 дней назад

so when can we know the secret model's cost

Фото профиля J A Z I I
J A Z I I3 дней назад

maybe in few days

Фото профиля Shwepsik2121
Shwepsik21213 дней назад

Secret model is flash or pro?

Фото профиля J A Z I I
J A Z I I3 дней назад

it's fast fast I can't say flash or pro or air or lite or GPT 6-7

Фото профиля WolfSWAP | SWAP & WIN
WolfSWAP | SWAP & WIN3 дней назад

Secret model is way better hoping that it costed less

Фото профиля balega_dev
balega_dev3 дней назад

What is the secret model?

Фото профиля silent
silent3 дней назад

it's obvious grok 4.7 better

Фото профиля J A Z I I
J A Z I I3 дней назад

😂😭

Фото профиля Deepak Reddy
Deepak Reddy3 дней назад

ig we're cooked with Grok 4.7

Фото профиля Veee
Veee3 дней назад

what's the secret model tho? Qwen?

Фото профиля Varik Verilion
Varik Verilion3 дней назад

highest reasoning setting does that, it keeps re-checking an answer it already had.

Фото профиля AGI Pulse
AGI Pulse3 дней назад

ngl DS one looks more detailed

Фото профиля ahs
ahs3 дней назад

left one is definitely better

Фото профиля J A Z I I
J A Z I I3 дней назад

yeah it did better but some details are missing 100%

Фото профиля KXLAD3
KXLAD33 дней назад

Is deepseek good at coding? Also what's the cheapest way to get access? 😭😭

Фото профиля J A Z I I
J A Z I I3 дней назад

nothing for now soon it'll be opencode but wait for now

Фото профиля Farhan
Farhan3 дней назад

this is why speed matters so much for agent workflows. 4x faster changes the whole experience

Фото профиля Schubert Santos 🇺🇸|🇧🇷
Schubert Santos 🇺🇸|🇧🇷3 дней назад

This secret model seems really decent at 3D. Could you prompt some other stuff? Could be simple. There’s a thing I like to test in every model - I ask for a three js representation of a temperature and umidity reader using a esp32. So far Astra is winning, but still quite bugged

Фото профиля J A Z I I
J A Z I I3 дней назад

I ran 3d ship tests so I'll post it later tonight

Фото профиля kitik kote
kitik kote3 дней назад

try high thinking , 4.1v better high

Фото профиля Haleemah
Haleemah3 дней назад

secret hmm

Фото профиля EDDY VU
EDDY VU3 дней назад

Did the extra 90 minutes actually translate to a cleaner pagoda, or was it just burning tokens stuck in thinking loops?

Фото профиля BreezeOg🐉
BreezeOg🐉3 дней назад

Impressive speed upgrade, that’s a game changer in ai models of today

Фото профиля Deepu
Deepu3 дней назад

dsv4 looks better imo but 2 hrs is a lot btw jazi, how about not labelling the output next time and revealing in the comment instead, less bias and more fun 🙃

Фото профиля Atma Blabber
Atma Blabber3 дней назад

deepseek clearly better output. but time.....

Фото профиля DEV
DEV3 дней назад

Overthinking can tank efficiency. Speed without depth often leads to shallow results.

Фото профиля Dunamix™️®️
Dunamix™️®️3 дней назад

which do you prefer?

Фото профиля J A Z I I
J A Z I I3 дней назад

I like both, as one did faster and got good details and other did way better

Фото профиля David
David3 дней назад

降价了

Фото профиля Super530
Super5302 дней назад

Its behavior appear to be different for language used for prompting. I don't see over think issue when poking (戳戳) in Chinese.

Похожие видео

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,250 просмотров • 14 дней назад

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,202 просмотров • 18 дней назад