Загрузка видео...

Не удалось загрузить видео

На главную

We tested GLM 5.3 Flash and Qwen 3.8 Flash in Pi. Same task for both: use ego lite to build an interactive 3D solar system. Qwen cost 3.6x more and ran 1.7x longer, but the output tells a different story. Comment "prompt" below and we'll drop it.

24,535 просмотров • 25 дней назад •via X (Twitter)

Комментарии: 17

Фото профиля ego
ego25 дней назад

Qwen 3.8 Flash: 1h 13m, $0.18. This one surprised us. Smooth orbital animation, thoughtful lighting, a few small details we didn't ask for but were glad to see. Clearly spent the extra time refining before calling it done. Results:

Фото профиля ego
ego25 дней назад

GLM 5.3 Flash: 43min, $0.05. Solid and quick. Planets in place, orbits working, nothing broken. It just didn't have the extra layer of detail that made Qwen's build feel finished. Results:

Фото профиля Aymane
Aymane25 дней назад

You know what I'm gonna say ego don't you :)

Фото профиля ego
ego25 дней назад

lol bro did u join our discord server 😭 you’d be the first one to get windows version

Фото профиля Aymane
Aymane25 дней назад

Joining right now

Фото профиля fate
fate25 дней назад

prompt

Фото профиля Yovani
Yovani24 дней назад

cost and runtime tell only half the story. if the more expensive run produces a meaningfully better workflow, the useful metric is value per completed task rather than tokens alone.

Фото профиля xcPenny
xcPenny25 дней назад

love it 🥰

Фото профиля Space Time
Space Time25 дней назад

Tpmorp

Фото профиля ego
ego25 дней назад

Build the most visually stunning and coherent single HTML interactive experience: an accelerated 3D simulation of the solar system. First, use ego lite to find and study one or two exceptional interactive space visualizations for inspiration. Create a cinematic star field with the Sun, eight planets, realistic orbital motion, visible orbit paths, adjustable time speed, pause, reset, camera rotation, zoom, and clickable planet information. Add a polished modern HUD with clear controls and elegant visual design. Use Three.js. Keep everything in one HTML file and avoid unnecessary features. After the core experience works, use ego lite for one brief validation pass and fix obvious issues.

Фото профиля 无我
无我25 дней назад

You must let me join our broken system, and I will give you the feedback you need like crazy.Please private message me, my friend.

Фото профиля Mehmet Odabaş
Mehmet Odabaş25 дней назад

Prompt

Фото профиля Balaji K
Balaji K25 дней назад

Prompt

Фото профиля ego
ego25 дней назад

Build the most visually stunning and coherent single HTML interactive experience: an accelerated 3D simulation of the solar system. First, use ego lite to find and study one or two exceptional interactive space visualizations for inspiration. Create a cinematic star field with the Sun, eight planets, realistic orbital motion, visible orbit paths, adjustable time speed, pause, reset, camera rotation, zoom, and clickable planet information. Add a polished modern HUD with clear controls and elegant visual design. Use Three.js. Keep everything in one HTML file and avoid unnecessary features. After the core experience works, use ego lite for one brief validation pass and fix obvious issues.

Фото профиля Ekkarat Prasongsap
Ekkarat Prasongsap25 дней назад

prompt

Фото профиля ego
ego25 дней назад

Build the most visually stunning and coherent single HTML interactive experience: an accelerated 3D simulation of the solar system. First, use ego lite to find and study one or two exceptional interactive space visualizations for inspiration. Create a cinematic star field with the Sun, eight planets, realistic orbital motion, visible orbit paths, adjustable time speed, pause, reset, camera rotation, zoom, and clickable planet information. Add a polished modern HUD with clear controls and elegant visual design. Use Three.js. Keep everything in one HTML file and avoid unnecessary features. After the core experience works, use ego lite for one brief validation pass and fix obvious issues.

Фото профиля why
why25 дней назад

Interesting benchmark. Curious what made Qwen costlier — is it token usage or model behavior?

Похожие видео

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,360 просмотров • 25 дней назад

glm 5.3 flash is 7.5x cheaper, but 3.4x slower than gemini 3.7 flash Z.ai glm 5.3 flash – shipped aug 26, $0.07/$0.25 per 1m Google DeepMind gemini 3.7 flash – shipped aug 13, $0.38/$1.88 per 1m we put the two models on one job: write one html file that draws an animated 3d scene in the browser. no images, no downloads, and it has to look the same on every load. the setup: three scenes – a glass aquarium in a lit room, the solar system, a night city under a thunderstorm. identical brief word for word, reasoning effort high, 64k output cap. the numbers below are not the whole run. they cover the three scenes we kept – the best one per task from each model, the ones in the video. - total generation time for the three scenes #1 gemini 3.7 flash – 10m 36s #2 glm 5.3 flash – 36m 30s - tokens spent on those three scenes #1 glm 5.3 flash – 110k #2 gemini 3.7 flash – 111k - cost of those three scenes #1 glm 5.3 flash – $0.027 #2 gemini 3.7 flash – $0.202 observations: • glm's first 10 attempts: 7 blank pages. it kept inventing short random helpers and forgetting to define one of them. the fix was one line in the brief: use exactly one random helper, named rand(), and don't invent shorthands next to it. next 12 attempts: 11 alive, 0 crashes. • glm spends 66% of its output on reasoning, gemini 57%. that is the whole speed gap. • gemini's storm came back as a black rectangle in 4 of 6 runs. glm's best storm has a branching bolt, lit rain and wet asphalt – for $0.01. conclusion: same three scenes, same token spend – glm 5.3 flash billed $0.027 and took 36m 30s, gemini 3.7 flash billed $0.202 and took 10m 36s. glm wins gemini on price and made the best storm of the whole run follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

15,997 просмотров • 27 дней назад

Very quick comparison between Ornith 1.5 35B MoE and Qwen 3.8 27B Dense. 📣 Clearly it's not a fair one, but let's in any case see how it went. 397B download in progress! In the Videos below: - Brick SA -> Qwen 3.8 - Lego Streets -> Ornith 1.5 Context: - Pi agent used in both cases, prompt below - M3 Ultra 512GB - Ornith-1.5-35B-A3B-oQ8e MTP hosted on oMLX - mlx-community/Qwen3.8-27B-8bit DFlash 2 hosted on mlx-dspark - Ornith has been incredibly fast with speed from 75 t/s to 45 t/s (160K+ context) - Qwen3.8 suffered context much more reaching 5 t/s above 160K context, but it can be due to engine tested still work in progress - Prompt: using threejs, and cdn, create a lego like game that's inspired by gta san andreas, with beautiful aesthetics and graphics and ability to steal cars. it should be 3d and have a nice and large map and areas. the graphics should be decent and nice. it should feature iconic things from gta san andreas Notes: - Qwen 3.8 result is 0-shot, while Ornith 1.5 required 6 iterations - Ornith 1.5 is not at the same level of autonomy as Qwen 3.8 27B honestly. To try getting the same results I'm constantly nudging, steering and 🤬 at it. - Qwen 3.8 took 6 hours to complete, but it was more a problem of timeout of Chrome headless used for testing, real prefill/decode TBD. I'll test again now with oMLX - Pi agent has a nearly perfect Cache Hit ratio that for local models is MEGA important!

Ivan Fioravanti ᯅ

12,150 просмотров • 1 месяц назад