
Future Coded
@future_coded • 26,131 subscribers
AI enthusiast / tech lover always exploring the future of innovation/ for Promotion 👉[email protected] 📧 https://t.co/jfA688J5ZY
Shorts
Videos

hy4 preview changes the question from “can a lab run 770B?” to “can your team stand it up?” tencent hunyuan released an open w8 text model: 𝟕𝟕𝟎𝐁 𝐭𝐨𝐭𝐚𝐥, 𝟒𝟗𝐁 𝐚𝐜𝐭𝐢𝐯𝐞 𝐩𝐞𝐫 𝐭𝐨𝐤𝐞𝐧, Apache 2.0 and native 𝟏𝐦 𝐜𝐨𝐧𝐭𝐞𝐱𝐭. Day 0 paths exist for vLLM and SGLang, plus an official FP8 checkpoint. That already matters. What matters more is the compressed route: mixed GGUF takes the weights from 𝐚𝐩𝐩𝐫𝐨𝐱 𝟏.𝟓𝐓𝐁 𝐭𝐨 𝟐𝟏𝟒𝐆𝐢𝐁 while reported accuracy stays close to BF16 small score movement not a different model. Say it cleanly. this is mixed per layer quant not a blanket 1 bit model. Sensitive layers keep more precision others go lower. Pair every size claim with that quality claim. A 214GiB file that still behaves on coding, agents and long documents is a deployment path. A 214GiB file that does not is just storage. Reality check, because this is a preview: GGUF needs the 𝐩𝐚𝐭𝐜𝐡𝐞𝐝 𝐥𝐥𝐚𝐦𝐚.𝐜𝐩𝐩 path for hyv4. Offload is still offload. Nobody should sell this as a laptop native 770B. Measure tokens/s on your own infra. Compare a real task BF16 or FP8 vs the mixed GGUF and treat vendor charts and your run as separate evidence. The story is ownership, download the weights, serve them and keep lowering the bar without throwing away the 49B-active capability that made the 770B class worth hosting.
Future Coded32,855 views • 24 days ago
No more content to load