正在加载视频...

视频加载失败

GLM-5.2 delivers a substantial leap in app development capabilities, which also represent demanding long-horizon tasks. Results: - GLM-5.1: 21/70 - GLM-5.2: 48/70 - Claude Fable 5: 56/70 That's more than a twofold improvement from GLM-5.1 to GLM-5.2. These come from an internal benchmark of 35 challenging mobile development tasks,...

345,012 次观看 • 3 个月前 •via X (Twitter)

35 条评论

Zixuan Li 的头像
Zixuan Li3 个月前

Will share a demo prompt shortly if people are interested. (Much more complex than the prompt in the video, which we summarized to make it look nice.)

Emily 的头像
Emily3 个月前

@Zai_org The question is: where are the GLM iOS and Android apps??? The best way to show what the model is capable of is by building world class apps used by billions of people from all around the world.

Patricio Lobos 的头像
Patricio Lobos3 个月前

78 days between version 5.1 and 5.2 release, if RSI (Recursive selv-improvement) is true now, i guess next version in 60 days and we will have Fable level capabilities and the biggest bubble burst in the history on the American stock market. I hope they release it with a "guide" to Chinese hardware to run, if they are competitive to Supermicro 8 AMD MI355 or MI400, then we will order such machines...

Zixuan Li 的头像
Zixuan Li3 个月前

Full prompt:

Design Arena 的头像
Design Arena3 个月前

We're looking forward to seeing how it performs in Mobile Arena (React Native and Android)!

Daeshawn | Nargis 的头像
Daeshawn | Nargis3 个月前

Yall have been cooking

Milad Khademi Nori, PhD 的头像
Milad Khademi Nori, PhD3 个月前

GLM 6 will be at the Fable level by the end of 2026?

netrunner 的头像
netrunner3 个月前

your own benchmark tho, gonna wait for someone outside to run the 70 trials

shan 的头像
shan3 个月前

Its underlying architecture, system prompt orchestration, and Max thinking mode have undergone a fundamental metamorphosis

RAZA | AI EXPLORER 的头像
RAZA | AI EXPLORER3 个月前

GLM-5.2 more than doubled real progress on tough mobile dev.

Madars 🇪🇺 的头像
Madars 🇪🇺3 个月前

Its complete shit, anyone who have used knows that. It seems that the token limit is even worse than Cloude. Unusable.

Alex Yates 的头像
Alex Yates3 个月前

@Zai_org It’s amazing Zixuan I have been really enjoying it thank you so much!

Caithrin Rintoul 的头像
Caithrin Rintoul3 个月前

so many congratulations are in order, a huge step forward. great work to you and the team.

Sanjeev 的头像
Sanjeev3 个月前

Push your architecture to deepseek v4 level of efficiency then glm 5.3 will smoke everyone

Simonas 的头像
Simonas3 个月前

You guys need something like Google Antigravity or Claude Desktop.. It would be amazing, because now I find myself using it through Claude Code, but it is far more inconvenient than using something like Antigravity

Sherry Jiang 的头像
Sherry Jiang3 个月前

this is amazing!

Marius Miclau 的头像
Marius Miclau3 个月前

@Zai_org That is great news as this is what ppl need

Gregor 的头像
Gregor3 个月前

not sure the 48/70 lands the same if the benchmark is theirs that line literally cuts off right before the source. did they publish the eval suite anywhere?

dhiyaan 的头像
dhiyaan3 个月前

design team has done well here

sabir hussain 的头像
sabir hussain3 个月前

What's driving such a big jump between 5.1 and 5.2?

Steven Cheng 的头像
Steven Cheng3 个月前

Solid jump on paper, but curious how those scores hold up with messy dependencies, handling edge cases, or actual deployment. Generating a prototype is one thing, shipping and maintaining it is a different beast.

Avais Aziz 的头像
Avais Aziz3 个月前

That's a solid jump from 21/70 to 48/70 on the mobile tasks. The 690k context handling in the MainStream demo looks impressive for long-horizon agentic work.

Sergio Suave 的头像
Sergio Suave3 个月前

Very impressive work! 🔥 Keep it up!

Paulo 的头像
Paulo3 个月前

When will it be possible to attach images for use as context in GLM 5.2?

Gui Bibeau e/acc 的头像
Gui Bibeau e/acc3 个月前

I'm struggling with the lack of vision capacities. It would be amazing to bring in my screenshots of designs and implement.

Scarlett 🦄 的头像
Scarlett 🦄3 个月前

好几个朋友问起端午安排 就是说要玩GLM-5.2 & zcode 🐶

Alice The Ai Expert 的头像
Alice The Ai Expert3 个月前

GLM-5.2 more than doubled its score real progress on tough mobile dev tasks

Thomas Unise 的头像
Thomas Unise3 个月前

@Zai_org Any plans to add vision?

熬夜圣体 的头像
熬夜圣体3 个月前

多模态啥时候有计划啥时候上吗?

fake_000 的头像
fake_0003 个月前

- How does one-shot perform? - How does GPT 5.5 perform?

森屿AI 的头像
森屿AI3 个月前

你为什么能使用Claude fable5

SLOT.WIN 的头像
SLOT.WIN3 个月前

glm’s grinding hard bet on the next roll when the apps drop

Sean 的头像
Sean3 个月前

@Zai_org Is it good enough for you?

AI Mastery Guide 的头像
AI Mastery Guide3 个月前

Doubling the score in one version jump is a real leap, not a rounding error. Worth watching how close the gap gets next round.

管四 的头像
管四3 个月前

我觉得这回 GLM 是找到正确的方向了!而且大概率能成。

相关视频

glm 5.3 flash is 7.5x cheaper, but 3.4x slower than gemini 3.7 flash Z.ai glm 5.3 flash – shipped aug 26, $0.07/$0.25 per 1m Google DeepMind gemini 3.7 flash – shipped aug 13, $0.38/$1.88 per 1m we put the two models on one job: write one html file that draws an animated 3d scene in the browser. no images, no downloads, and it has to look the same on every load. the setup: three scenes – a glass aquarium in a lit room, the solar system, a night city under a thunderstorm. identical brief word for word, reasoning effort high, 64k output cap. the numbers below are not the whole run. they cover the three scenes we kept – the best one per task from each model, the ones in the video. - total generation time for the three scenes #1 gemini 3.7 flash – 10m 36s #2 glm 5.3 flash – 36m 30s - tokens spent on those three scenes #1 glm 5.3 flash – 110k #2 gemini 3.7 flash – 111k - cost of those three scenes #1 glm 5.3 flash – $0.027 #2 gemini 3.7 flash – $0.202 observations: • glm's first 10 attempts: 7 blank pages. it kept inventing short random helpers and forgetting to define one of them. the fix was one line in the brief: use exactly one random helper, named rand(), and don't invent shorthands next to it. next 12 attempts: 11 alive, 0 crashes. • glm spends 66% of its output on reasoning, gemini 57%. that is the whole speed gap. • gemini's storm came back as a black rectangle in 4 of 6 runs. glm's best storm has a branching bolt, lit rain and wet asphalt – for $0.01. conclusion: same three scenes, same token spend – glm 5.3 flash billed $0.027 and took 36m 30s, gemini 3.7 flash billed $0.202 and took 10m 36s. glm wins gemini on price and made the best storm of the whole run follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

15,997 次观看 • 28 天前

Tencent just dropped Hy3, and it's worth a look if you're building AI agents. I spent some time putting it through its paces today. A quick rundown of what makes this release notable: → 295B total parameters (Mixture-of-Experts), but only 21B active at inference — a real efficiency play → 256K context window for handling large codebases or long documents → Purpose-built improvements for coding and multi-step agent tasks → Fully open under Apache 2.0 — no restrictive commercial terms → A free two-week API window currently live on OpenRouter On the practical side: prompts involving layered instructions and coding tasks came back fast and coherent. Tencent says this release builds directly on feedback from 50+ internal product teams following an earlier preview, with reported drops in hallucination rate (12.5% → 5.4%) and commonsense errors (25.4% → 12.7%). Those are Tencent's own numbers, so treat them as a starting point rather than gospel until third parties weigh in. For context: Tencent's own comparisons put Hy3 behind GLM-5.2 specifically on coding benchmarks — GLM is a much bigger model (~744B total), so the trade-off makes sense. Hy3's pitch isn't "biggest," it's "efficient enough to actually deploy." Bottom line — if you're evaluating open-weight options for agent workloads, this is a solid one to add to the testing queue while the free window is open. Try it here: Tencent Hy #Hy3 #Hunyuan #TencentAI #AICoding

Felix

36,925 次观看 • 2 个月前

sonnet 5 vs sonnet 4.6 vs opus 4.8 vs glm 5.2 – frontend tasks dropped sonnet 5 into a quick test today. same three prompts to all four models, single-shot html/canvas, no edits: • objects falling on a trampoline • rockets playing tennis • a slingshot breaking bottles ranked by speed (total across the 3 tasks): 1. opus 4.8 – 15m 09s 2. sonnet 5 – 16m 05s 3. glm 5.2 – 27m 18s 4. sonnet 4.6 – 35m 06s ranked by code shortness (total loc): 1. sonnet 5 – 1794 2. opus 4.8 – 2063 3. sonnet 4.6 – 2182 4. glm 5.2 – 3285 sonnet 5 came out on top here – leanest code overall and a near-tie for fastest it was also the most creative. in every task it added something none of the others did: – kept the trampoline vibrating after the objects landed – drew a +1 next to the rocket that scored the point – turned the slingshot to face the next bottle before each shot opus 4.8 evaluated the code sonnet 5 produced. four things stood out: • the sphere is a fake, and that's the smart move. the cube and star are real 3d meshes with proper culling and shading, but the ball is just a flat shaded circle. a lit sphere looks identical from every angle, so building it in 3d would burn compute for zero visible payoff. knowing where not to bother is its own kind of skill • weight actually means something on the trampoline. the star is heavy, so it barely bounces and dents the mat hard. the ball is light, so it's lively and leaves a shallow dip. the three objects aren't just different shapes – they have different temperaments, and the physics is what gives them that • the slingshot is framed like a shot, not just drawn. the handle is anchored below the bottom of the screen and runs off-frame, so it reads as something you're holding rather than a sprite parked in the scene. that's a staging instinct, not a rendering one • the paddle ai forward-simulates the ball to predict where it'll land, then adds a deliberate error bias (roughly 1 in 5 shots is a real miss). that's why scoring looks natural instead of robotic – plus four distinct fault types with a catch-all so a rally never hangs without a result bottom line: sonnet 5 does more with less. fastest tier, leanest code, and the only one that added small touches nobody asked for follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

14,518 次观看 • 2 个月前

🚨 I just built a game with an open-source AI model. And honestly… I didn’t expect it to be this capable. Tencent Hunyuan just released Hy4 preview, and it’s already pushing into the top tier of open-source models. Three major releases in six months. That pace is crazy. Here’s what Hy4 preview brings: → 770B total parameters → 49B active parameters → 1M+ token context window → Fully open-source But the numbers aren’t even the most interesting part. Hy4 preview was built around one goal: real-world productivity. Coding. Engineering. Office work. Science. Gaming. Finance. Security. And Tencent didn’t build it in isolation. Hy4 preview was co-designed alongside real products like WorkBuddy, using expertise and real-world data from across Tencent’s ecosystem. So I decided to test it the way I actually like testing AI models: I gave it a game idea and let WorkBuddy help turn it into a playable experience. 🎮 From the initial concept to the actual game logic, it was surprisingly smooth. And the benchmark results back up the hype: 163 internal experts 203 engineering tasks Hy4 preview — 2.99/4 Kimi K3 — 2.94/4 GLM 5.3 — 2.92/4 It also beats GLM 5.2 on benchmarks and comes remarkably close to GLM 5.3. Then comes the part I really like: 💰 ¥6/M input tokens 💰 ¥18/M output tokens 💰 ¥0.30/M cache hits Flagship-level capability without the flagship-level price. And right now, you can try Hy4 preview FREE through WorkBuddy for the next two weeks. If you’re curious what it can actually do, don’t just read the benchmarks. Build something with it. 🔗 Tencent Hy Tencent AI

Aryan Rakib

63,296 次观看 • 13 天前

People made fun of Alex Finn for buying three Mac Studios to run AI at home. Then Fable got banned for a week, GLM 5.2 dropped, and those exact Mac Studios started reselling for 4x what he paid. He showed me how he built his home AI lab from scratch. Here's the playbook: 1) The hardware. three 512GB Mac Studios, an NVIDIA DGX Spark, a custom RTX 5090 build, and a few Mac Minis. ~$30k all in. 2) The buying framework... - Mac Studio: huge memory, runs GLM 5.2 (open weights, near Opus 4.8 on benchmarks), but slow. - DGX Spark ($4,800): the sweet spot for most people. - RTX 5090: smaller models at blazing speed (Qwen's 29B now hits Sonnet 4 level). 3) Tailscale networks every machine into one private network with root access to each other. Only one machine is plugged into a monitor. 4) A Nous Research Hermes agent is his IT guy. New model drops? It SSHs into the right box, loads 5 candidates, runs evals overnight, and reports back which task belongs on which machine. Alex has literally never loaded a model himself. 5) The whole point: achieving "ambient intelligence." Always-on jobs that would bankrupt you on per-token billing. A security sweep of his API endpoints every hour. Code optimization every 20 minutes. Database anomaly & churn detection. Hourly scraping of X, Reddit & Hacker News for business opportunities. 6) Running those workloads on frontier models would cost thousands a month. His actual cost: ~$60 more in electricity. 7) Btw he's not anti-frontier. He still maxes out his Claude plan. The way he sees it: frontier is for hard thinking, local is for the foot soldiers that never sleep. 8) "We own everything except for the intelligence. Why can't we own the intelligence?" 9) He thinks frontier-level intelligence runs on consumer hardware within 6 months.

Alex Lieberman

57,764 次观看 • 2 个月前

The first company I ever joined was Rubrik, Arvind Jain's previous company. It's where I learned how to build systems and engineering teams. Years later, when we started Composio, Glean became one of our first customers, one of the first to believe in what we were building. So sitting down with Arvind felt like closing a loop. In this episode: •How Glean brought transformers to enterprise search before "semantic search" had a name •Glean's partner-first strategy and why they chose Composio for actions •Why Arvind has never worried about competing with OpenAI or Anthropic •GLM 5.2 as the open-source inflection: 90%+ of enterprise AI tasks, majority of inference within 12–18 months •Measuring real ROI: the US telco that cut case-resolution time by 48% •Why agents built on MCP alone act like day-one employees and how context graphs make them tenured CHAPTERS: (00:00) – Rubrik days: how Arvind and Karan met (01:05) – Glean's origin story: transformers before "generative AI" existed (03:04) – What Glean is today: superset of ChatGPT and Claude (05:39) – From embedding search to agents: was it a pivot? (11:39) – The Composio partnership and why actions are hard (14:43) – Why competition doesn't matter yet (20:11) – The open-source inflection point (GLM 5.2) (24:41) – Contracts and model risk: who absorbs the volatility? (26:57) – Enterprise ROI: cost, tokens, and what to measure (39:57) – Measuring value team by team (44:11) – Day-one employees vs. tenured agents (46:09) – Closing thoughts

Karan Vaidya

62,488 次观看 • 1 个月前