Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Z.ai GLM 5.2 is now live on Eigent from day 0! threw a real long-horizon task: research 30 companies across 6 sectors of the AI infrastructure stack, structure it into JSON, then build an interactive HTML report. same prompt, 5.1 vs 5.2. where 5.2 pulls ahead: → plans deeper...

23,827 görüntüleme • 3 ay önce •via X (Twitter)

2 Yorum

RAZA | AI EXPLORER profil fotoğrafı
RAZA | AI EXPLORER27 gün önce

@Zai_org GLM-5.2 handles long-horizon tasks with deeper research, stronger planning, and 1M context.

Anika Patel profil fotoğrafı
Anika Patel2 ay önce

@Zai_org deeper planning and verification

Benzer Videolar

🚨 I just built a game with an open-source AI model. And honestly… I didn’t expect it to be this capable. Tencent Hunyuan just released Hy4 preview, and it’s already pushing into the top tier of open-source models. Three major releases in six months. That pace is crazy. Here’s what Hy4 preview brings: → 770B total parameters → 49B active parameters → 1M+ token context window → Fully open-source But the numbers aren’t even the most interesting part. Hy4 preview was built around one goal: real-world productivity. Coding. Engineering. Office work. Science. Gaming. Finance. Security. And Tencent didn’t build it in isolation. Hy4 preview was co-designed alongside real products like WorkBuddy, using expertise and real-world data from across Tencent’s ecosystem. So I decided to test it the way I actually like testing AI models: I gave it a game idea and let WorkBuddy help turn it into a playable experience. 🎮 From the initial concept to the actual game logic, it was surprisingly smooth. And the benchmark results back up the hype: 163 internal experts 203 engineering tasks Hy4 preview — 2.99/4 Kimi K3 — 2.94/4 GLM 5.3 — 2.92/4 It also beats GLM 5.2 on benchmarks and comes remarkably close to GLM 5.3. Then comes the part I really like: 💰 ¥6/M input tokens 💰 ¥18/M output tokens 💰 ¥0.30/M cache hits Flagship-level capability without the flagship-level price. And right now, you can try Hy4 preview FREE through WorkBuddy for the next two weeks. If you’re curious what it can actually do, don’t just read the benchmarks. Build something with it. 🔗 Tencent Hy Tencent AI

Aryan Rakib

62,914 görüntüleme • 5 gün önce

The robot flipped a pancake nobody taught it! 🥞 Skild AI team assumed pancake flipping had to be somewhere in the training data. So they searched. Millions of hours of pre-training data. Nothing. S1 inferred the whole task from a single human demonstration. That's their new general robot model, built as an in-context learner from the ground up. Every new robot task today starts with days of teleoperation and a fine-tuning run on a specialist policy. S1 skips all of it. Much like a language model, it never updates its weights to learn a new task. The demonstration enters the context window, and the policy uses it to decide what to do next. → Ten-minute tasks it was never trained on, composed from primitives learned in pre-training: a new style of coffee, potting a plant, frying pancakes. → Soil and pots arrived at their office at 8:54 PM. The robot was running the task autonomously by 9:27 PM. → Slide objects away mid-reach, swap them, change the lighting, it still finishes. → The prompt waters a plant with a watering can, but only a cup is available. It uses the cup. It doesn't rigidly replay what it saw, but it recovers from its own errors, and sometimes executes with more precision than the demonstrator, when the human fumbles an egg and makes a mess, S1 performs the same step cleanly. The demonstration is a specification of the goal, and not a trajectory to copy. On unseen tasks after 100K hours of pre-training: language-prompted VLAs reach 9%. Their new model reaches 66%. It's already deploying with industrial partners, with a wider rollout over the coming months. Congrats Deepak Pathak and team behind this! 😮‍💨 🔗 Link to their latest blog: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

11,406 görüntüleme • 21 gün önce

OpenAI. said. this. publicly. their own engineers just proved one idea on themselves, in writing: stop telling AI what's wrong. hand it the whole broken thing and let it find out they gave GPT-6 Astra a slow test build of their own coding tool. one cause found, a memory bottleneck, one allocator swapped, every turn 25× faster this is GPT-6 Astra, the layer that fixes the cause instead of the symptom, $0 on top of the ChatGPT plan you already pay for: - open ChatGPT or Codex, pick GPT-6 Astra, hand it the whole thing: the folder, the file that takes a minute to open. it works in apps with no API and reads your screen - type one sentence: find the one cause, prove it, fix it, do not patch around it - leave the room. it asks without stopping, keeps working on what does not need your answer, waits only where the answer changes the outcome - keep it in one Codex session with the experimental notes setting on: it remembers across context windows why an earlier fix failed - expect the first pass to land: handed a program with no source, it worked out how it runs 88% of the time first try, 99.2% within four you never find out what was broken. it gets fixed anyway the catch is on the same page. roughly 30% more memory for that speed, and the safety checks can pause a long job until you approve the next step describing the problem was the expensive half of fixing it. that half just ended every hour you spend explaining the symptom to a chat window, someone else has handed theirs over whole bookmark this before the next thing breaks, the playbook for handing a whole job to an AI worker is in the piece below ↓

Argona

102,338 görüntüleme • 5 gün önce