Загрузка видео...

Не удалось загрузить видео

На главную

GLM-5.2 delivers a substantial leap in app development capabilities, which also represent demanding long-horizon tasks. Results: - GLM-5.1: 21/70 - GLM-5.2: 48/70 - Claude Fable 5: 56/70 That's more than a twofold improvement from GLM-5.1 to GLM-5.2. These come from an internal benchmark of 35 challenging mobile development tasks,...

345,012 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 35

Фото профиля Zixuan Li
Zixuan Li3 месяцев назад

Will share a demo prompt shortly if people are interested. (Much more complex than the prompt in the video, which we summarized to make it look nice.)

Фото профиля Emily
Emily3 месяцев назад

@Zai_org The question is: where are the GLM iOS and Android apps??? The best way to show what the model is capable of is by building world class apps used by billions of people from all around the world.

Фото профиля Patricio Lobos
Patricio Lobos3 месяцев назад

78 days between version 5.1 and 5.2 release, if RSI (Recursive selv-improvement) is true now, i guess next version in 60 days and we will have Fable level capabilities and the biggest bubble burst in the history on the American stock market. I hope they release it with a "guide" to Chinese hardware to run, if they are competitive to Supermicro 8 AMD MI355 or MI400, then we will order such machines...

Фото профиля Zixuan Li
Zixuan Li3 месяцев назад

Full prompt:

Фото профиля Design Arena
Design Arena3 месяцев назад

We're looking forward to seeing how it performs in Mobile Arena (React Native and Android)!

Фото профиля Daeshawn | Nargis
Daeshawn | Nargis3 месяцев назад

Yall have been cooking

Фото профиля Milad Khademi Nori, PhD
Milad Khademi Nori, PhD3 месяцев назад

GLM 6 will be at the Fable level by the end of 2026?

Фото профиля netrunner
netrunner3 месяцев назад

your own benchmark tho, gonna wait for someone outside to run the 70 trials

Фото профиля shan
shan3 месяцев назад

Its underlying architecture, system prompt orchestration, and Max thinking mode have undergone a fundamental metamorphosis

Фото профиля RAZA | AI EXPLORER
RAZA | AI EXPLORER3 месяцев назад

GLM-5.2 more than doubled real progress on tough mobile dev.

Фото профиля Madars 🇪🇺
Madars 🇪🇺3 месяцев назад

Its complete shit, anyone who have used knows that. It seems that the token limit is even worse than Cloude. Unusable.

Фото профиля Alex Yates
Alex Yates3 месяцев назад

@Zai_org It’s amazing Zixuan I have been really enjoying it thank you so much!

Фото профиля Caithrin Rintoul
Caithrin Rintoul3 месяцев назад

so many congratulations are in order, a huge step forward. great work to you and the team.

Фото профиля Sanjeev
Sanjeev3 месяцев назад

Push your architecture to deepseek v4 level of efficiency then glm 5.3 will smoke everyone

Фото профиля Simonas
Simonas3 месяцев назад

You guys need something like Google Antigravity or Claude Desktop.. It would be amazing, because now I find myself using it through Claude Code, but it is far more inconvenient than using something like Antigravity

Фото профиля Sherry Jiang
Sherry Jiang3 месяцев назад

this is amazing!

Фото профиля Marius Miclau
Marius Miclau3 месяцев назад

@Zai_org That is great news as this is what ppl need

Фото профиля Gregor
Gregor3 месяцев назад

not sure the 48/70 lands the same if the benchmark is theirs that line literally cuts off right before the source. did they publish the eval suite anywhere?

Фото профиля dhiyaan
dhiyaan3 месяцев назад

design team has done well here

Фото профиля sabir hussain
sabir hussain3 месяцев назад

What's driving such a big jump between 5.1 and 5.2?

Фото профиля Steven Cheng
Steven Cheng3 месяцев назад

Solid jump on paper, but curious how those scores hold up with messy dependencies, handling edge cases, or actual deployment. Generating a prototype is one thing, shipping and maintaining it is a different beast.

Фото профиля Avais Aziz
Avais Aziz3 месяцев назад

That's a solid jump from 21/70 to 48/70 on the mobile tasks. The 690k context handling in the MainStream demo looks impressive for long-horizon agentic work.

Фото профиля Sergio Suave
Sergio Suave3 месяцев назад

Very impressive work! 🔥 Keep it up!

Фото профиля Paulo
Paulo3 месяцев назад

When will it be possible to attach images for use as context in GLM 5.2?

Фото профиля Gui Bibeau e/acc
Gui Bibeau e/acc3 месяцев назад

I'm struggling with the lack of vision capacities. It would be amazing to bring in my screenshots of designs and implement.

Фото профиля Scarlett 🦄
Scarlett 🦄3 месяцев назад

好几个朋友问起端午安排 就是说要玩GLM-5.2 & zcode 🐶

Фото профиля Alice The Ai Expert
Alice The Ai Expert3 месяцев назад

GLM-5.2 more than doubled its score real progress on tough mobile dev tasks

Фото профиля Thomas Unise
Thomas Unise3 месяцев назад

@Zai_org Any plans to add vision?

Фото профиля 熬夜圣体
熬夜圣体3 месяцев назад

多模态啥时候有计划啥时候上吗?

Фото профиля fake_000
fake_0003 месяцев назад

- How does one-shot perform? - How does GPT 5.5 perform?

Фото профиля 森屿AI
森屿AI3 месяцев назад

你为什么能使用Claude fable5

Фото профиля SLOT.WIN
SLOT.WIN3 месяцев назад

glm’s grinding hard bet on the next roll when the apps drop

Фото профиля Sean
Sean3 месяцев назад

@Zai_org Is it good enough for you?

Фото профиля AI Mastery Guide
AI Mastery Guide3 месяцев назад

Doubling the score in one version jump is a real leap, not a rounding error. Worth watching how close the gap gets next round.

Фото профиля 管四
管四3 месяцев назад

我觉得这回 GLM 是找到正确的方向了!而且大概率能成。

Похожие видео

glm 5.3 flash is 7.5x cheaper, but 3.4x slower than gemini 3.7 flash Z.ai glm 5.3 flash – shipped aug 26, $0.07/$0.25 per 1m Google DeepMind gemini 3.7 flash – shipped aug 13, $0.38/$1.88 per 1m we put the two models on one job: write one html file that draws an animated 3d scene in the browser. no images, no downloads, and it has to look the same on every load. the setup: three scenes – a glass aquarium in a lit room, the solar system, a night city under a thunderstorm. identical brief word for word, reasoning effort high, 64k output cap. the numbers below are not the whole run. they cover the three scenes we kept – the best one per task from each model, the ones in the video. - total generation time for the three scenes #1 gemini 3.7 flash – 10m 36s #2 glm 5.3 flash – 36m 30s - tokens spent on those three scenes #1 glm 5.3 flash – 110k #2 gemini 3.7 flash – 111k - cost of those three scenes #1 glm 5.3 flash – $0.027 #2 gemini 3.7 flash – $0.202 observations: • glm's first 10 attempts: 7 blank pages. it kept inventing short random helpers and forgetting to define one of them. the fix was one line in the brief: use exactly one random helper, named rand(), and don't invent shorthands next to it. next 12 attempts: 11 alive, 0 crashes. • glm spends 66% of its output on reasoning, gemini 57%. that is the whole speed gap. • gemini's storm came back as a black rectangle in 4 of 6 runs. glm's best storm has a branching bolt, lit rain and wet asphalt – for $0.01. conclusion: same three scenes, same token spend – glm 5.3 flash billed $0.027 and took 36m 30s, gemini 3.7 flash billed $0.202 and took 10m 36s. glm wins gemini on price and made the best storm of the whole run follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

15,997 просмотров • 28 дней назад

Tencent just dropped Hy3, and it's worth a look if you're building AI agents. I spent some time putting it through its paces today. A quick rundown of what makes this release notable: → 295B total parameters (Mixture-of-Experts), but only 21B active at inference — a real efficiency play → 256K context window for handling large codebases or long documents → Purpose-built improvements for coding and multi-step agent tasks → Fully open under Apache 2.0 — no restrictive commercial terms → A free two-week API window currently live on OpenRouter On the practical side: prompts involving layered instructions and coding tasks came back fast and coherent. Tencent says this release builds directly on feedback from 50+ internal product teams following an earlier preview, with reported drops in hallucination rate (12.5% → 5.4%) and commonsense errors (25.4% → 12.7%). Those are Tencent's own numbers, so treat them as a starting point rather than gospel until third parties weigh in. For context: Tencent's own comparisons put Hy3 behind GLM-5.2 specifically on coding benchmarks — GLM is a much bigger model (~744B total), so the trade-off makes sense. Hy3's pitch isn't "biggest," it's "efficient enough to actually deploy." Bottom line — if you're evaluating open-weight options for agent workloads, this is a solid one to add to the testing queue while the free window is open. Try it here: Tencent Hy #Hy3 #Hunyuan #TencentAI #AICoding

Felix

36,925 просмотров • 2 месяцев назад

sonnet 5 vs sonnet 4.6 vs opus 4.8 vs glm 5.2 – frontend tasks dropped sonnet 5 into a quick test today. same three prompts to all four models, single-shot html/canvas, no edits: • objects falling on a trampoline • rockets playing tennis • a slingshot breaking bottles ranked by speed (total across the 3 tasks): 1. opus 4.8 – 15m 09s 2. sonnet 5 – 16m 05s 3. glm 5.2 – 27m 18s 4. sonnet 4.6 – 35m 06s ranked by code shortness (total loc): 1. sonnet 5 – 1794 2. opus 4.8 – 2063 3. sonnet 4.6 – 2182 4. glm 5.2 – 3285 sonnet 5 came out on top here – leanest code overall and a near-tie for fastest it was also the most creative. in every task it added something none of the others did: – kept the trampoline vibrating after the objects landed – drew a +1 next to the rocket that scored the point – turned the slingshot to face the next bottle before each shot opus 4.8 evaluated the code sonnet 5 produced. four things stood out: • the sphere is a fake, and that's the smart move. the cube and star are real 3d meshes with proper culling and shading, but the ball is just a flat shaded circle. a lit sphere looks identical from every angle, so building it in 3d would burn compute for zero visible payoff. knowing where not to bother is its own kind of skill • weight actually means something on the trampoline. the star is heavy, so it barely bounces and dents the mat hard. the ball is light, so it's lively and leaves a shallow dip. the three objects aren't just different shapes – they have different temperaments, and the physics is what gives them that • the slingshot is framed like a shot, not just drawn. the handle is anchored below the bottom of the screen and runs off-frame, so it reads as something you're holding rather than a sprite parked in the scene. that's a staging instinct, not a rendering one • the paddle ai forward-simulates the ball to predict where it'll land, then adds a deliberate error bias (roughly 1 in 5 shots is a real miss). that's why scoring looks natural instead of robotic – plus four distinct fault types with a catch-all so a rally never hangs without a result bottom line: sonnet 5 does more with less. fastest tier, leanest code, and the only one that added small touches nobody asked for follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

14,518 просмотров • 2 месяцев назад

🚨 I just built a game with an open-source AI model. And honestly… I didn’t expect it to be this capable. Tencent Hunyuan just released Hy4 preview, and it’s already pushing into the top tier of open-source models. Three major releases in six months. That pace is crazy. Here’s what Hy4 preview brings: → 770B total parameters → 49B active parameters → 1M+ token context window → Fully open-source But the numbers aren’t even the most interesting part. Hy4 preview was built around one goal: real-world productivity. Coding. Engineering. Office work. Science. Gaming. Finance. Security. And Tencent didn’t build it in isolation. Hy4 preview was co-designed alongside real products like WorkBuddy, using expertise and real-world data from across Tencent’s ecosystem. So I decided to test it the way I actually like testing AI models: I gave it a game idea and let WorkBuddy help turn it into a playable experience. 🎮 From the initial concept to the actual game logic, it was surprisingly smooth. And the benchmark results back up the hype: 163 internal experts 203 engineering tasks Hy4 preview — 2.99/4 Kimi K3 — 2.94/4 GLM 5.3 — 2.92/4 It also beats GLM 5.2 on benchmarks and comes remarkably close to GLM 5.3. Then comes the part I really like: 💰 ¥6/M input tokens 💰 ¥18/M output tokens 💰 ¥0.30/M cache hits Flagship-level capability without the flagship-level price. And right now, you can try Hy4 preview FREE through WorkBuddy for the next two weeks. If you’re curious what it can actually do, don’t just read the benchmarks. Build something with it. 🔗 Tencent Hy Tencent AI

Aryan Rakib

63,296 просмотров • 13 дней назад

People made fun of Alex Finn for buying three Mac Studios to run AI at home. Then Fable got banned for a week, GLM 5.2 dropped, and those exact Mac Studios started reselling for 4x what he paid. He showed me how he built his home AI lab from scratch. Here's the playbook: 1) The hardware. three 512GB Mac Studios, an NVIDIA DGX Spark, a custom RTX 5090 build, and a few Mac Minis. ~$30k all in. 2) The buying framework... - Mac Studio: huge memory, runs GLM 5.2 (open weights, near Opus 4.8 on benchmarks), but slow. - DGX Spark ($4,800): the sweet spot for most people. - RTX 5090: smaller models at blazing speed (Qwen's 29B now hits Sonnet 4 level). 3) Tailscale networks every machine into one private network with root access to each other. Only one machine is plugged into a monitor. 4) A Nous Research Hermes agent is his IT guy. New model drops? It SSHs into the right box, loads 5 candidates, runs evals overnight, and reports back which task belongs on which machine. Alex has literally never loaded a model himself. 5) The whole point: achieving "ambient intelligence." Always-on jobs that would bankrupt you on per-token billing. A security sweep of his API endpoints every hour. Code optimization every 20 minutes. Database anomaly & churn detection. Hourly scraping of X, Reddit & Hacker News for business opportunities. 6) Running those workloads on frontier models would cost thousands a month. His actual cost: ~$60 more in electricity. 7) Btw he's not anti-frontier. He still maxes out his Claude plan. The way he sees it: frontier is for hard thinking, local is for the foot soldiers that never sleep. 8) "We own everything except for the intelligence. Why can't we own the intelligence?" 9) He thinks frontier-level intelligence runs on consumer hardware within 6 months.

Alex Lieberman

57,764 просмотров • 2 месяцев назад

The first company I ever joined was Rubrik, Arvind Jain's previous company. It's where I learned how to build systems and engineering teams. Years later, when we started Composio, Glean became one of our first customers, one of the first to believe in what we were building. So sitting down with Arvind felt like closing a loop. In this episode: •How Glean brought transformers to enterprise search before "semantic search" had a name •Glean's partner-first strategy and why they chose Composio for actions •Why Arvind has never worried about competing with OpenAI or Anthropic •GLM 5.2 as the open-source inflection: 90%+ of enterprise AI tasks, majority of inference within 12–18 months •Measuring real ROI: the US telco that cut case-resolution time by 48% •Why agents built on MCP alone act like day-one employees and how context graphs make them tenured CHAPTERS: (00:00) – Rubrik days: how Arvind and Karan met (01:05) – Glean's origin story: transformers before "generative AI" existed (03:04) – What Glean is today: superset of ChatGPT and Claude (05:39) – From embedding search to agents: was it a pivot? (11:39) – The Composio partnership and why actions are hard (14:43) – Why competition doesn't matter yet (20:11) – The open-source inflection point (GLM 5.2) (24:41) – Contracts and model risk: who absorbs the volatility? (26:57) – Enterprise ROI: cost, tokens, and what to measure (39:57) – Measuring value team by team (44:11) – Day-one employees vs. tenured agents (46:09) – Closing thoughts

Karan Vaidya

62,488 просмотров • 1 месяц назад