Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

deepseek v4.1 flash vs secret model tested both with the same prompt at the highest reasoning available > v4.1 flash took 2 hours to finish both tests and cost $2.60 > secret model finished both in 30 minutes, can’t reveal cost yet v4.1 flash spent a lot of time...

28,586 görüntüleme • 3 gün önce •via X (Twitter)

47 Yorum

J A Z I I profil fotoğrafı
J A Z I I3 gün önce

voxel pagoda test for both models

K2S profil fotoğrafı
K2S3 gün önce

Bro how u use deepseek v4.1 flash in Ai hub mix last time I used 😒 hit the limit without even building something Are u using paid or free lmk?

J A Z I I profil fotoğrafı
J A Z I I3 gün önce

i have credits, load like 5 to 10$ and you should be good for testing almost all chinese models

Alina Fomina profil fotoğrafı
Alina Fomina3 gün önce

2 hours is wild. but still deepseek v4.1 flash looks solid

J A Z I I profil fotoğrafı
J A Z I I3 gün önce

Agreed but missed alot stuff

Dorian profil fotoğrafı
Dorian3 gün önce

And The secret model is (IMO) next Gemini flash model. Only Gemini have true flash Decode speed. Before i thought that's Grok 4.7, but quality is not good for Frontier model, and elon said that is slower than 4.6, but better overall

J A Z I I profil fotoğrafı
J A Z I I3 gün önce

hehe

Dorian profil fotoğrafı
Dorian3 gün önce

Deepseek 4.1 flash Has better Water detail, and i think overall deepseek is better

Hussain Hashim | Building SundayBack profil fotoğrafı
Hussain Hashim | Building SundayBack3 gün önce

@notjazii Always the trade-off between speed and depth. Hit this wall when testing models for my own stuff. A balance is crucial!

Knowix profil fotoğrafı
Knowix3 gün önce

having the same feeling as when Astra launched, might as well test this out myself

Yohaku profil fotoğrafı
Yohaku3 gün önce

deepseek goated bro, love it

Shubh profil fotoğrafı
Shubh3 gün önce

Cooking that well in 30 mins is really impressive legend 👀

Chii profil fotoğrafı
Chii3 gün önce

high feels less overthinking

J A Z I I profil fotoğrafı
J A Z I I3 gün önce

wait fr? I tested on max in Hermes so lemme try with minimax test

Chii profil fotoğrafı
Chii3 gün önce

yeah try opencode

J A Z I I profil fotoğrafı
J A Z I I3 gün önce

not working in opencode no idea why

Chii profil fotoğrafı
Chii3 gün önce

need to add custom model

xabz profil fotoğrafı
xabz3 gün önce

so when can we know the secret model's cost

J A Z I I profil fotoğrafı
J A Z I I3 gün önce

maybe in few days

Shwepsik2121 profil fotoğrafı
Shwepsik21213 gün önce

Secret model is flash or pro?

J A Z I I profil fotoğrafı
J A Z I I3 gün önce

it's fast fast I can't say flash or pro or air or lite or GPT 6-7

WolfSWAP | SWAP & WIN profil fotoğrafı
WolfSWAP | SWAP & WIN3 gün önce

Secret model is way better hoping that it costed less

balega_dev profil fotoğrafı
balega_dev3 gün önce

What is the secret model?

silent profil fotoğrafı
silent3 gün önce

it's obvious grok 4.7 better

J A Z I I profil fotoğrafı
J A Z I I3 gün önce

😂😭

Deepak Reddy profil fotoğrafı
Deepak Reddy3 gün önce

ig we're cooked with Grok 4.7

Veee profil fotoğrafı
Veee3 gün önce

what's the secret model tho? Qwen?

Varik Verilion profil fotoğrafı
Varik Verilion3 gün önce

highest reasoning setting does that, it keeps re-checking an answer it already had.

AGI Pulse profil fotoğrafı
AGI Pulse3 gün önce

ngl DS one looks more detailed

ahs profil fotoğrafı
ahs3 gün önce

left one is definitely better

J A Z I I profil fotoğrafı
J A Z I I3 gün önce

yeah it did better but some details are missing 100%

KXLAD3 profil fotoğrafı
KXLAD33 gün önce

Is deepseek good at coding? Also what's the cheapest way to get access? 😭😭

J A Z I I profil fotoğrafı
J A Z I I3 gün önce

nothing for now soon it'll be opencode but wait for now

Farhan profil fotoğrafı
Farhan3 gün önce

this is why speed matters so much for agent workflows. 4x faster changes the whole experience

Schubert Santos 🇺🇸|🇧🇷 profil fotoğrafı
Schubert Santos 🇺🇸|🇧🇷3 gün önce

This secret model seems really decent at 3D. Could you prompt some other stuff? Could be simple. There’s a thing I like to test in every model - I ask for a three js representation of a temperature and umidity reader using a esp32. So far Astra is winning, but still quite bugged

J A Z I I profil fotoğrafı
J A Z I I3 gün önce

I ran 3d ship tests so I'll post it later tonight

kitik kote profil fotoğrafı
kitik kote3 gün önce

try high thinking , 4.1v better high

Haleemah profil fotoğrafı
Haleemah3 gün önce

secret hmm

EDDY VU profil fotoğrafı
EDDY VU3 gün önce

Did the extra 90 minutes actually translate to a cleaner pagoda, or was it just burning tokens stuck in thinking loops?

BreezeOg🐉 profil fotoğrafı
BreezeOg🐉3 gün önce

Impressive speed upgrade, that’s a game changer in ai models of today

Deepu profil fotoğrafı
Deepu3 gün önce

dsv4 looks better imo but 2 hrs is a lot btw jazi, how about not labelling the output next time and revealing in the comment instead, less bias and more fun 🙃

Atma Blabber profil fotoğrafı
Atma Blabber3 gün önce

deepseek clearly better output. but time.....

DEV profil fotoğrafı
DEV3 gün önce

Overthinking can tank efficiency. Speed without depth often leads to shallow results.

Dunamix™️®️ profil fotoğrafı
Dunamix™️®️3 gün önce

which do you prefer?

J A Z I I profil fotoğrafı
J A Z I I3 gün önce

I like both, as one did faster and got good details and other did way better

David profil fotoğrafı
David3 gün önce

降价了

Super530 profil fotoğrafı
Super5302 gün önce

Its behavior appear to be different for language used for prompting. I don't see over think issue when poking (戳戳) in Chinese.

Benzer Videolar

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,250 görüntüleme • 14 gün önce

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,202 görüntüleme • 18 gün önce