Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Five frontier models. One simple prompt. No retries. No blender. Just a simple frontend web design. Fable 5.1 vs. GPT-6 Astra vs. DeepSeek v4.1 Flash vs. GLM 5.3 Flash vs. Qwen3.8 Flash Next Which model built the best looking luxury watch landing page? IMO Fable 5.1 and GLM 5.3...

16,501 Aufrufe • vor 21 Tagen •via X (Twitter)

28 Kommentare

Profilbild von thewhitemage.
thewhitemage.vor 21 Tagen

I really like GLM 5.3 Flash's, for its size, it's pretty damn close to frontier.

Profilbild von Mia
Miavor 21 Tagen

I think it's only second to fable. Trying complicated stuff now to test their limits.

Profilbild von Promptyx
Promptyxvor 21 Tagen

They all mostly look same??

Profilbild von Mia
Miavor 21 Tagen

The watch is different on each, and some are more immersive

Profilbild von irony.somi
irony.somivor 21 Tagen

glm > deepseek > astra > fable > qwen

Profilbild von Aj
Ajvor 21 Tagen

Fable

Profilbild von Dan C.
Dan C.vor 21 Tagen

I think I prefer Deepseek. Incredible for the size of model.

Profilbild von Tom Maiaroto
Tom Maiarotovor 21 Tagen

Holy crap I don't like Astra or Fable 5.1. The others are somehow better. Didn't expect that! The "Reserve" button text is too small in the Astra version, plus some weird rule (longer than an em dash) above "Meridian 38" title. The tracking for the text "calibre" in the Fable 5.1 is too weird and the selected color way out on the right which has no label that it's the current selection. Yes, Fable 5.1 and Astra suck here.

Profilbild von Carlos Ajoy
Carlos Ajoyvor 21 Tagen

I think Astra looks More attractive

Profilbild von Geazi Anc
Geazi Ancvor 21 Tagen

how about Tencent Hy4? I'm not seeing anyone talking about this model, and it looks be great as well.

Profilbild von Philip McBride
Philip McBridevor 21 Tagen

Impressed with them all, but Fable 5.1 wins.

Profilbild von sgiath
sgiathvor 21 Tagen

They are too similar to tell. That tells me that it is too simple task so it does not properly differentiate the models. Try something more complex or with more wiggle room so the differences are more apparent.

Profilbild von David Grant
David Grantvor 21 Tagen

Astra > GLM 5.3 > others Qwen last

Profilbild von Mia
Miavor 21 Tagen

Yes

Profilbild von Tim Messerschmidt
Tim Messerschmidtvor 21 Tagen

The slight moving up and down of the watches was a little unsettling for me 😁 Love the back-to-back comparison, though!

Profilbild von Eco
Ecovor 21 Tagen

I think Astra

Profilbild von Guillem
Guillemvor 20 Tagen

I tried to look at the watch only, ignoring who had made it, and just remember my first impression. For me: Astra is the best, DeepSeek and Fable are tied at the "good enough" level, the others are too cartoony for me and personally I wouldn't use them.

Profilbild von Olivia Parker
Olivia Parkervor 21 Tagen

One prompt is a draft contest. Live listing is the score.

Profilbild von Sergio
Sergiovor 21 Tagen

How is it that all five models ended up using the same price for the watch?

Profilbild von Divine 〽️achine
Divine 〽️achinevor 21 Tagen

Deepseek looks great, i assume it was much cheaper than sol/fable for comparable output

Profilbild von ilyas
ilyasvor 21 Tagen

Hello, how do you write these kinds of commands, or do you use example prompts? Could you please provide information?

Profilbild von Knowix
Knowixvor 21 Tagen

I feel GLM 5.3 Flash did a better job here

Profilbild von Arman
Armanvor 21 Tagen

Bijan-coded benchmark

Profilbild von curdmudgeon
curdmudgeonvor 21 Tagen

they all seem fine

Profilbild von Tango Uniform
Tango Uniformvor 21 Tagen

all went for 4850? “Build a single-file HTML landing page for Halden — Meridian 38, a fictional minimalist luxury watch brand from Oslo. Dark, quiet, premium. The watch itself must be drawn with CSS/SVG and show the real local time. Include specs, price and a reserve button.”

Profilbild von Daniel B. - AI & Tech
Daniel B. - AI & Techvor 21 Tagen

Fable gets it for me, but DeepSeek is a close second, very close!

Profilbild von Tyler Folkman
Tyler Folkmanvor 21 Tagen

Agreed Qwen is the worst but for how well it runs on 2 sparks, locally it's a very powerful model. I liked GLM the most.

Profilbild von Carmelo schepis
Carmelo schepisvor 20 Tagen

fable’s spacing is elite but the typography feels a bit mid. still a massive win for the frontier tier.

Ähnliche Videos

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,360 Aufrufe • vor 1 Monat

glm 5.3 flash is 7.5x cheaper, but 3.4x slower than gemini 3.7 flash Z.ai glm 5.3 flash – shipped aug 26, $0.07/$0.25 per 1m Google DeepMind gemini 3.7 flash – shipped aug 13, $0.38/$1.88 per 1m we put the two models on one job: write one html file that draws an animated 3d scene in the browser. no images, no downloads, and it has to look the same on every load. the setup: three scenes – a glass aquarium in a lit room, the solar system, a night city under a thunderstorm. identical brief word for word, reasoning effort high, 64k output cap. the numbers below are not the whole run. they cover the three scenes we kept – the best one per task from each model, the ones in the video. - total generation time for the three scenes #1 gemini 3.7 flash – 10m 36s #2 glm 5.3 flash – 36m 30s - tokens spent on those three scenes #1 glm 5.3 flash – 110k #2 gemini 3.7 flash – 111k - cost of those three scenes #1 glm 5.3 flash – $0.027 #2 gemini 3.7 flash – $0.202 observations: • glm's first 10 attempts: 7 blank pages. it kept inventing short random helpers and forgetting to define one of them. the fix was one line in the brief: use exactly one random helper, named rand(), and don't invent shorthands next to it. next 12 attempts: 11 alive, 0 crashes. • glm spends 66% of its output on reasoning, gemini 57%. that is the whole speed gap. • gemini's storm came back as a black rectangle in 4 of 6 runs. glm's best storm has a branching bolt, lit rain and wet asphalt – for $0.01. conclusion: same three scenes, same token spend – glm 5.3 flash billed $0.027 and took 36m 30s, gemini 3.7 flash billed $0.202 and took 10m 36s. glm wins gemini on price and made the best storm of the whole run follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

16,041 Aufrufe • vor 1 Monat