Alexa Benchmark's banner
Alexa Benchmark's profile picture

Alexa Benchmark

@Allexa_AI29,205 subscribers

AI & Tech Educator | Vibe Coding 🫶 | Making AI simple, useful & exciting for everyone | @ iaproactiv partner

Shorts

I spend my days explaining to teams why a 770-billion-parameter open-weight model will never fit into their infrastructure. This week, I asked it to code a complete mobile game from a single prompt. The concept is one everyone has probably played before. A hole moving through an open-plan office, viewed from above, swallowing everything in its path. You start tiny, only able to swallow pens and cups. You grow, moving on to keyboards and plants, then chairs and printers, then desks and vending machines. Eventually, you swallow the entire meeting room. 60 seconds on the clock. One HTML file, zero external libraries, zero errors on launch. The video shows the generation and then the actual gameplay. What broke is more instructive than what worked. The structure came out right on the first try: fixed-timestep loop, spatial grid for collisions, tier system, spring camera. The balancing and rendering, not so much. The first version scored 70 points in 13 seconds with a tier threshold at 500, and drew colorful circles and triangles instead of furniture. I had to give it numbers and exact recipes. Speed: 640 pixels per second. Radii: 30, 58, 96, 150, 225. Tier thresholds: 150, 550, 1500, 3500. And for every object, a pixel-perfect drawing recipe. Once I gave it that, it followed the instructions exactly. A model that doesn't execute its own code won't tune itself. But it will execute, down to the exact numbers, what you tell it to build. The model is Hy4 preview, released by Tencent Hunyuan on August 28. 770 billion parameters in total, but 49 billion active per token. And that second number is what determines your serving bill. Native 1M context. Apache 2.0 license. vLLM and SGLang supported from day one, with an official FP8 checkpoint. Text-only preview. The part that matters for deployment is the compression. Tencent describes it as seven times smaller with almost no loss. GGUF builds use mixed per-layer quantization, where calibration data determines the bit width layer by layer. Some layers go as low as 1.31 bits, while others go up to 2.06 bits, averaging 2.38 bits per weight. The model drops from 1.5 TB in BF16 to 213.66 GiB while, according to their measurements, staying in the same performance range on real-world tasks. On this build, they report 204 tokens/s in prefill and 20 tokens/s in decoding, measured on an 8-GPU node. Their numbers, not mine. Their blind evaluation scores 2.99 out of 4 across 203 engineering tasks rated by 163 experts. Ahead of Kimi K3 at 2.94 and GLM-5.3 at 2.92. Their numbers too. My run went through the official hosted studio, not a local build, so I’m not claiming to have benchmarked the compressed GGUF myself. Two honest caveats. None of these builds run on standard llama.cpp. The hyv4 architecture isn't upstream yet, so patches are required. And 214 GiB of resident weights is still a server, not your laptop. This is a preview, and Tencent explicitly asks users to break it and report what fails. So here's my contribution.

I spend my days explaining to teams why a 770-billion-parameter open-weight model will never fit into their infrastructure. This week, I asked it to code a complete mobile game from a single prompt. The concept is one everyone has probably played before. A hole moving through an open-plan office, viewed from above, swallowing everything in its path. You start tiny, only able to swallow pens and cups. You grow, moving on to keyboards and plants, then chairs and printers, then desks and vending machines. Eventually, you swallow the entire meeting room. 60 seconds on the clock. One HTML file, zero external libraries, zero errors on launch. The video shows the generation and then the actual gameplay. What broke is more instructive than what worked. The structure came out right on the first try: fixed-timestep loop, spatial grid for collisions, tier system, spring camera. The balancing and rendering, not so much. The first version scored 70 points in 13 seconds with a tier threshold at 500, and drew colorful circles and triangles instead of furniture. I had to give it numbers and exact recipes. Speed: 640 pixels per second. Radii: 30, 58, 96, 150, 225. Tier thresholds: 150, 550, 1500, 3500. And for every object, a pixel-perfect drawing recipe. Once I gave it that, it followed the instructions exactly. A model that doesn't execute its own code won't tune itself. But it will execute, down to the exact numbers, what you tell it to build. The model is Hy4 preview, released by Tencent Hunyuan on August 28. 770 billion parameters in total, but 49 billion active per token. And that second number is what determines your serving bill. Native 1M context. Apache 2.0 license. vLLM and SGLang supported from day one, with an official FP8 checkpoint. Text-only preview. The part that matters for deployment is the compression. Tencent describes it as seven times smaller with almost no loss. GGUF builds use mixed per-layer quantization, where calibration data determines the bit width layer by layer. Some layers go as low as 1.31 bits, while others go up to 2.06 bits, averaging 2.38 bits per weight. The model drops from 1.5 TB in BF16 to 213.66 GiB while, according to their measurements, staying in the same performance range on real-world tasks. On this build, they report 204 tokens/s in prefill and 20 tokens/s in decoding, measured on an 8-GPU node. Their numbers, not mine. Their blind evaluation scores 2.99 out of 4 across 203 engineering tasks rated by 163 experts. Ahead of Kimi K3 at 2.94 and GLM-5.3 at 2.92. Their numbers too. My run went through the official hosted studio, not a local build, so I’m not claiming to have benchmarked the compressed GGUF myself. Two honest caveats. None of these builds run on standard llama.cpp. The hyv4 architecture isn't upstream yet, so patches are required. And 214 GiB of resident weights is still a server, not your laptop. This is a preview, and Tencent explicitly asks users to break it and report what fails. So here's my contribution.

16,315 Aufrufe

Videos

Keine weiteren Inhalte verfügbar