Video wird geladen...
Video konnte nicht geladen werden
GLM-5.2 delivers a substantial leap in app development capabilities, which also represent demanding long-horizon tasks. Results: - GLM-5.1: 21/70 - GLM-5.2: 48/70 - Claude Fable 5: 56/70 That's more than a twofold improvement from GLM-5.1 to GLM-5.2. These come from an internal benchmark of 35 challenging mobile development tasks,... show more
345,012 Aufrufe • vor 3 Monaten •via X (Twitter)
35 Kommentare

Will share a demo prompt shortly if people are interested. (Much more complex than the prompt in the video, which we summarized to make it look nice.)

@Zai_org The question is: where are the GLM iOS and Android apps??? The best way to show what the model is capable of is by building world class apps used by billions of people from all around the world.

78 days between version 5.1 and 5.2 release, if RSI (Recursive selv-improvement) is true now, i guess next version in 60 days and we will have Fable level capabilities and the biggest bubble burst in the history on the American stock market. I hope they release it with a "guide" to Chinese hardware to run, if they are competitive to Supermicro 8 AMD MI355 or MI400, then we will order such machines...

Full prompt:

We're looking forward to seeing how it performs in Mobile Arena (React Native and Android)!

Yall have been cooking

GLM 6 will be at the Fable level by the end of 2026?

your own benchmark tho, gonna wait for someone outside to run the 70 trials

Its underlying architecture, system prompt orchestration, and Max thinking mode have undergone a fundamental metamorphosis

GLM-5.2 more than doubled real progress on tough mobile dev.

Its complete shit, anyone who have used knows that. It seems that the token limit is even worse than Cloude. Unusable.

@Zai_org It’s amazing Zixuan I have been really enjoying it thank you so much!

so many congratulations are in order, a huge step forward. great work to you and the team.

Push your architecture to deepseek v4 level of efficiency then glm 5.3 will smoke everyone

You guys need something like Google Antigravity or Claude Desktop.. It would be amazing, because now I find myself using it through Claude Code, but it is far more inconvenient than using something like Antigravity

this is amazing!

@Zai_org That is great news as this is what ppl need

not sure the 48/70 lands the same if the benchmark is theirs that line literally cuts off right before the source. did they publish the eval suite anywhere?

design team has done well here

What's driving such a big jump between 5.1 and 5.2?

Solid jump on paper, but curious how those scores hold up with messy dependencies, handling edge cases, or actual deployment. Generating a prototype is one thing, shipping and maintaining it is a different beast.

That's a solid jump from 21/70 to 48/70 on the mobile tasks. The 690k context handling in the MainStream demo looks impressive for long-horizon agentic work.

Very impressive work! 🔥 Keep it up!

When will it be possible to attach images for use as context in GLM 5.2?

I'm struggling with the lack of vision capacities. It would be amazing to bring in my screenshots of designs and implement.

好几个朋友问起端午安排 就是说要玩GLM-5.2 & zcode 🐶

GLM-5.2 more than doubled its score real progress on tough mobile dev tasks

@Zai_org Any plans to add vision?

多模态啥时候有计划啥时候上吗?

- How does one-shot perform? - How does GPT 5.5 perform?

你为什么能使用Claude fable5

glm’s grinding hard bet on the next roll when the apps drop

@Zai_org Is it good enough for you?

Doubling the score in one version jump is a real leap, not a rounding error. Worth watching how close the gap gets next round.

我觉得这回 GLM 是找到正确的方向了!而且大概率能成。
