Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I tested Tencent’s Hy4 preview model in WorkBuddy on a real frontend build, not a benchmark screenshot. I gave it one practical brief: create an original neon courier game in a single HTML file, with Canvas rendering, keyboard controls, collision detection, scoring, a countdown timer, a boost mechanic, sound...

61,820 Aufrufe • vor 4 Tagen •via X (Twitter)

32 Kommentare

Profilbild von Kiage Clinton, KC
Kiage Clinton, KCvor 4 Tagen

A build like this shows whether the model can handle details beyond the initial concept.

Profilbild von Ochora🇰🇪☆, KC
Ochora🇰🇪☆, KCvor 4 Tagen

The interesting part is seeing an idea turn into a game with actual feedback loops.

Profilbild von Future
Futurevor 4 Tagen

This is a practical way to judge code generation: give it a full brief and play it.

Profilbild von JA KISII™ 🇰🇪
JA KISII™ 🇰🇪vor 4 Tagen

A real build brief makes it easier to see where generation turns into useful execution.

Profilbild von CATHEY
CATHEYvor 4 Tagen

The practical outcome matters here: a game someone can open, control, and restart.

Profilbild von Linn
Linnvor 4 Tagen

The strongest tests are the ones where small implementation mistakes become obvious fast.

Profilbild von Fatuma
Fatumavor 4 Tagen

It is great to see the evaluation centered on usable output instead of abstract claims.

Profilbild von Bonita🤎
Bonita🤎vor 4 Tagen

The gap between writing game code and delivering a playable game is bigger than it seems.

Profilbild von Ja telo ™
Ja telo ™vor 4 Tagen

The real test is whether the generated game holds together once you start playing it.

Profilbild von E
Evor 4 Tagen

The score, timer, and boost mechanic make this feel like a real implementation test.

Profilbild von iano
ianovor 4 Tagen

Frontend tasks become a better benchmark when they include behavior, not just a layout.

Profilbild von Lionel
Lionelvor 4 Tagen

The game format makes it easy to spot whether the interface and logic work together.

Profilbild von Alex Ryan
Alex Ryanvor 4 Tagen

This is a useful reminder that code quality includes the experience people actually get.

Profilbild von SHERIFF™🤠
SHERIFF™🤠vor 4 Tagen

Working interactions matter more than a promising code snippet when testing a model.

Profilbild von Tyla
Tylavor 4 Tagen

A complete interactive prototype tells you far more than isolated examples ever could.

Profilbild von Future
Futurevor 4 Tagen

A single-file game is a strong test because every system has to work together cleanly.

Profilbild von Amazing
Amazingvor 4 Tagen

This is the kind of test that reveals how well a model manages connected moving parts.

Profilbild von ZACK™️💎
ZACK™️💎vor 4 Tagen

A playable build is a much better test than a polished demo or a benchmark result.

Profilbild von ETHAN JAMES
ETHAN JAMESvor 4 Tagen

A brief with real requirements is a more meaningful test than asking for a generic app.

Profilbild von 𝐈𝐫𝐢𝐬𝐡𝐦𝐚𝐧 
𝐈𝐫𝐢𝐬𝐡𝐦𝐚𝐧 vor 4 Tagen

It is useful to see the output checked through play rather than judged only by the code.

Profilbild von JADUONG
JADUONGvor 4 Tagen

Fast iteration matters most when it still leaves time to inspect and test the result.

Profilbild von Bourke Floyd, IV
Bourke Floyd, IVvor 4 Tagen

Prefer this over benchmark screenshots. Did collision + scoring hold up or did it fake the juice?

Profilbild von RICHIE..😎
RICHIE..😎vor 4 Tagen

The best coding demos show the finished experience, not only the generation process.

Profilbild von NAME CANNOT BE BLANK
NAME CANNOT BE BLANKvor 4 Tagen

This feels like a sensible way to test how well a model handles real frontend work.

Profilbild von REX JAMES
REX JAMESvor 4 Tagen

It is encouraging when a tool can get from a creative prompt to something playable.

Profilbild von MAUREEN
MAUREENvor 4 Tagen

Getting movement, sound, collisions, and timing aligned is not a trivial build task.

Profilbild von JOMBA
JOMBAvor 4 Tagen

The handoff from idea to implementation is where these tools become genuinely useful.

Profilbild von Jayden Cole
Jayden Colevor 4 Tagen

A game is a good challenge because visual design and logic need to work in tandem.

Profilbild von King Levi
King Levivor 4 Tagen

The important question is whether the mechanics remain coherent after the first output.

Profilbild von chenly🥹
chenly🥹vor 4 Tagen

This kind of workflow could make early experiments much easier for frontend teams.

Profilbild von Stones
Stonesvor 4 Tagen

A model becomes more useful when it can carry a task through the messy final details.

Profilbild von Weigo
Weigovor 4 Tagen

Reliable restart logic is one of those details that says a lot about the final build.

Ähnliche Videos

I’ve been testing Hy4 preview in WorkBuddy, and the most interesting part is not simply the model size, it’s how much practical work it can handle with a relatively focused active parameter count. Hy4 preview brings together stronger code understanding, generation, and editing; improved document and information processing; workflow automation; web and game development; cross-tool collaboration; and more reliable completion of complex, multi-step tasks. In other words, it is designed for work that requires planning, tool use, iteration, and follow-through, not just a quick answer in a chat window. Compared with its initial release, the current Hy4 preview is noticeably faster and better-performing in practical workflows. Following an upgrade released yesterday, it can complete tasks with fewer conversation rounds and lower token usage, while reasoning more quickly and making the overall user experience feel smoother from the first instruction to the final result. For my test, I gave it a demanding Three.js game-prototyping task with a 770B-parameter model and 49B active parameters. The result was more revealing than a simple first-look demo: Hy4 preview handled the core logic, edge cases, and follow-up changes while maintaining the broader context of the project. That combination of capability, speed, context, and active compute is what makes its cost-effectiveness worth examining. A fair evaluation should use the same prompt and environment configuration across models, changing only the model itself. That makes it easier to assess task completion, planning quality, tool-calling stability, reasoning speed, token efficiency, and performance over longer workflows without confusing the result with different settings. If you want to test the model yourself, access Hy4 preview through WorkBuddy and see how it performs on a real coding, document, automation, or creative task: Tencent Hy Tencent AI WorkBuddy

Tyler Wayne

56,152 Aufrufe • vor 11 Tagen

🚨 I just built a game with an open-source AI model. And honestly… I didn’t expect it to be this capable. Tencent Hunyuan just released Hy4 preview, and it’s already pushing into the top tier of open-source models. Three major releases in six months. That pace is crazy. Here’s what Hy4 preview brings: → 770B total parameters → 49B active parameters → 1M+ token context window → Fully open-source But the numbers aren’t even the most interesting part. Hy4 preview was built around one goal: real-world productivity. Coding. Engineering. Office work. Science. Gaming. Finance. Security. And Tencent didn’t build it in isolation. Hy4 preview was co-designed alongside real products like WorkBuddy, using expertise and real-world data from across Tencent’s ecosystem. So I decided to test it the way I actually like testing AI models: I gave it a game idea and let WorkBuddy help turn it into a playable experience. 🎮 From the initial concept to the actual game logic, it was surprisingly smooth. And the benchmark results back up the hype: 163 internal experts 203 engineering tasks Hy4 preview — 2.99/4 Kimi K3 — 2.94/4 GLM 5.3 — 2.92/4 It also beats GLM 5.2 on benchmarks and comes remarkably close to GLM 5.3. Then comes the part I really like: 💰 ¥6/M input tokens 💰 ¥18/M output tokens 💰 ¥0.30/M cache hits Flagship-level capability without the flagship-level price. And right now, you can try Hy4 preview FREE through WorkBuddy for the next two weeks. If you’re curious what it can actually do, don’t just read the benchmarks. Build something with it. 🔗 Tencent Hy Tencent AI

Aryan Rakib

63,296 Aufrufe • vor 11 Tagen

I wanted to see what Hy4 preview could actually do when I gave it a real creative project instead of another “write me some code” prompt. So I went with something I’d genuinely want to play: 🐬 a cute dolphin as the hero 🌊 colorful underwater worlds 💎 glowing pearls to collect 🐠 sea creatures and obstacles 🪸 coral reefs, bubbles & hidden treasure 🎮 smoother animations and progressively harder levels The result is this underwater platformer 👇 And this is where Hy4 preview gets interesting to me. Behind the scenes, we’re talking about 770B parameters, 49B active parameters and a 1M+ token context window. It’s also fully open-sourced. But big numbers only matter if they translate into something useful. What I’m really testing is how much actual work the model can take off my plate understanding the idea, handling the implementation, keeping all the moving parts connected, and getting from a blank canvas to something that feels like a real experience. And there’s another update that makes this even more interesting: Hy4 preview was upgraded yesterday to significantly reduce both conversation turns and token consumption. That means less back-and-forth, faster thinking, and a smoother overall experience when working through complex tasks. That’s also why the cost-effectiveness angle stands out to me. If a model can handle more of the workflow while using fewer tokens and requiring fewer iterations, the value isn’t just in the benchmark score. It’s in how much more you can actually build with it. This dolphin has me wanting to test Hy4 preview with some much crazier ideas next. 🐬👀 Tencent Hy Tencent AI ☞

Md Riyazuddin

63,562 Aufrufe • vor 12 Tagen

I wanted to see where Hy4 preview would stop, so I gave it a deliberately demanding prompt: “Build a working Cyber Incident Command Center.” Not a mockup. A real, responsive app with a live threat map, animated attack routes, incident simulation, filters, investigation panels, event logs, ownership controls, and working isolate/resolve actions. Hy4 preview planned the build, wrote the application, ran it locally, tested the interactions, checked the mobile layout, fixed issues, and opened the finished product for review. Then I clicked “Simulate.” A new incident appeared, the dashboard updated, I filtered it by severity, opened the evidence trail, isolated the affected system, and resolved it. The KPI counters changed with the action. That’s a much more useful coding demo to me than watching code appear in an editor. Hy4 preview is Hunyuan’s third major release in six months. It’s fully open-sourced, with 770B total parameters, 49B activated, and a 1M+ token context window. Tencent built it around productivity scenarios through close work with experts in software engineering, gaming, finance, and security, plus direct co-design with tools like WorkBuddy. In an internal blind test covering 203 engineering tasks and 163 Tencent experts, it scored 2.99/4—slightly ahead of Kimi K3 at 2.94 and GLM 5.3 at 2.92. And the pricing stays practical: $0.834/M input $2.501/M output $0.042/M cache hits If you want to test it yourself, Hy4 preview is free inside WorkBuddy for the next two weeks, for a limited time. Tencent Hy Tencent AI WorkBuddy

Clara Bennett

12,864 Aufrufe • vor 22 Tagen