正在加载视频...

视频加载失败

Only ~1.2B parameters active at a time. Edge0-8B-A1B-preview makes an 8B-class MoE practical for local inference.📜 Apache 2.0. 🤖 ⚡ Reaches 23.9–25.3 tokens/s with about 1.0 GiB peak active memory in the reported short-context benchmark. 🏆 Retains most of the FP16 base model’s quality, with an average gap of...

16,923 次观看 • 14 天前 •via X (Twitter)

6 条评论

Nathan Roll 的头像
Nathan Roll14 天前

Very cool! Decode speed under memory pressure would help show how much performance depends on the OS keeping streamed experts cached.

Fajar M Reza 的头像
Fajar M Reza14 天前

Sparse activation and low memory could make local MoE inference surprisingly practical.

Atishay Jain 的头像
Atishay Jain13 天前

interesting to see edge0-8b-a1b-preview's performance on short-context benchmark. do you have any insights on the memory efficiency of the moe mechanism in this setup, considering the 1.0 gib peak active memory?

DEV 的头像
DEV14 天前

1.2B active params is a smart move. Balancing scale and memory efficiency is key for real...

BayesChat 的头像
BayesChat14 天前

Is vision or tool supported?

Bk1man 的头像
Bk1man14 天前

Worth adding the HF receipts: sibling Edge0-35B-A3B-preview hit 17.8K dl / 2.4K likes in its first week, this 8B took 3.2K dl by day 4 — Apache-2.0 weights are getting real adoption, not just benchmark praise. (HF API, checked today)

相关视频

Introducing "Building with Llama 4." This short course is created with Meta AI at Meta, and taught by Amit Sangani, Director of Partner Engineering for Meta’s AI team. Meta’s new Llama 4 has added three new models and introduced the Mixture-of-Experts (MoE) architecture to its family of open-weight models, making them more efficient to serve. In this course, you’ll work with two of the three new models introduced in Llama 4. First is Maverick, a 400B parameter model, with 128 experts and 17B active parameters. Second is Scout, a 109B parameter model with 16 experts and 17B active parameters. Maverick and Scout support long context windows of up to a million tokens and 10M tokens, respectively. The latter is enough to support directly inputting even fairly large GitHub repos for analysis! In hands-on lessons, you’ll build apps using Llama 4’s new multimodal capabilities including reasoning across multiple images and image grounding, in which you can identify elements in images. You’ll also use the official Llama API, work with Llama 4’s long-context abilities, and learn about Llama’s newest open-source tools: its prompt optimization tool that automatically improves system prompts and synthetic data kit that generates high-quality datasets for fine-tuning. If you need an open model, Llama is a great option, and the Llama 4 family is an important part of any GenAI developer's toolkit. Through this course, you’ll learn to call Llama 4 via API, use its optimization tools, and build features that span text, images, and large context. Please sign up here:

Andrew Ng

68,034 次观看 • 1 年前

I had to test it myself to believe this unreal inference speed. 3,000 tokens/s for 1 user on standard datacenter GPUs. They leveraged a hidden efficiency gap in how GPUs generate tokens. Kog just achieved 3,000 tokens/s on 8× AMD MI300X GPUs and 2,100 on 8× NVIDIA H200 (FP16, no speculative decoding). Their tech preview is on a 2B model, and they show how their techniques will scale to large frontier MoE models at similar speeds. That's a huge number because normal low-batch GPU decoding for 2B to 8B models is usually closer to 100 to 300 tokens/s per request, so Kog is claiming something like a 10X to 30X jump in the speed one user actually feels. Their trick: they are getting the speed by treating LLM decoding as a memory streaming problem, not mainly a math problem. For 1 user at batch size 1, the GPU is not doing big, efficient matrix-matrix work like in training or large-batch serving; it is repeatedly pulling the model’s active weights from high-bandwidth memory for each new token, so speed depends on how smoothly those weights keep flowing. Normal inference stacks keep breaking that flow. They run many separate GPU programs for different parts of the model, move intermediate results through memory, wait at synchronization points, talk back to the CPU for scheduling or sampling, and then repeat this token after token. Kog’s answer is to co-design 3 things that are usually tuned separately: the runtime, the low-level GPU code, and the model architecture. The biggest engineering move is the monokernel, where the whole decode pass runs as 1 persistent GPU-resident program, including sampling, so the system does not keep stopping for kernel launches, CPU scheduling, and intermediate memory round trips. They also rebuilt synchronization, because their own measurements say grid sync was eating around 35% of token-generation time; instead of making every compute unit wait at a broad barrier, each unit waits only for the exact data it needs. On AMD MI300X, they also map memory access around the chiplet layout, because memory latency changes depending on which die makes the request. Then their Laneformer model uses Delayed Tensor Parallelism, which lets cross-GPU communication happen in the background instead of blocking every layer.

Rohan Paul

13,282 次观看 • 4 个月前

Alibaba just dropped Qwen3.5-397B-A17B and there's a lot to unpack. 397B params, 17B active per forward pass. Sparse MoE done right. But the real story isn't the size—it's the architecture choices. The MoE Design Most MoE models feel like bolt-ons. Qwen 3.5's sparse activation is native—only 4.3% of parameters fire per token. That's how you get trillion-parameter-class performance without trillion-parameter inference costs. The 0.8 RMB/million tokens pricing isn't subsidized; it's structurally earned. Native Multimodal, Not Glued-On This is a vision-language model from the ground up. Heterogeneous architecture—separate processing pipelines for text, image, video that fuse early. Not a vision encoder slapped onto an LLM. The result: 90.8 on OmniDocBench, 79.0 on MMMU-Pro. Document understanding and visual reasoning without the usual brittleness. The Context Window Reality Qwen3.5-Plus (the hosted version) ships with 1M tokens by default. That's not a marketing number—they're actually positioning it for long-document workflows. With built-in adaptive tool use, it's clearly aimed at agentic automation, not just chat. What Actually Impressed Me • FP8 native pipeline: ~50% activation memory reduction • Async RL framework for continuous refinement—training and inference workloads separated • 201 languages (up from 119), 250k vocab for better low-resource encoding • Apache 2.0 license. Full weights on HuggingFace and ModelScope. The Benchmark Context 76.4 on SWE-bench Verified puts it in the range where it can handle real debugging workflows. 72.9 on BFCL v4 for agentic tool use. 88.4 on GPQA Diamond. These aren't SOTA in isolation, but the breadth is unusual—strong across reasoning, coding, multimodal, and agentic tasks. The Honest Caveat I haven't stress-tested the 1M context for needle-in-haystack retrieval yet. And "native multimodal" claims need real-world torture testing—PDFs with tables, charts, mixed layouts. Benchmarks are benchmarks. Bottom Line This isn't just another model release. It's a bet on efficient scale: big model capabilities, small active compute, open weights. At 1/18th the cost of Gemini 3 Pro, it's going to force pricing conversations across the board.

Bo Wang

13,221 次观看 • 7 个月前

Qwen3.8-Flash-Next is still going strong at 364.7K tokens of context on an M5 Max. And this isn’t just a static long-context test. The model was reasoning about how to speed up its own workflow while using tools, and the tool calls kept working without misses. Setup: • Qwen3.8-Flash-Next • M5 Max • 128GB unified memory • MLX-Serve PR #363 • OpenCode 2 • 364.7K context The interesting part isn’t simply getting hundreds of thousands of tokens into memory. It’s what happens once the context gets this large. Long-context inference usually comes with a painful tradeoff. As the KV cache grows, memory pressure increases and generation can slow down. But this setup is still pushing through 364K tokens while maintaining a usable agent workflow. The model can reason, call tools, inspect results, continue working, and keep the session moving. And the tool calls reportedly haven’t missed so far. That’s important for agentic coding. A huge context window is only useful if the model can actually operate reliably inside it. A 400K-token context that constantly breaks tool calls isn’t very useful. A 364K session that can keep reasoning and executing tools is a different story. And the test isn’t finished yet. The current run is approaching 400K tokens, with the expectation that it can keep going. This is also another interesting example of why Apple Silicon keeps showing up in local LLM experiments. The M5 Max’s unified memory gives a large model and its growing KV cache access to one shared memory pool. With MLX-Serve continuing to improve, these machines are becoming surprisingly capable long-context inference boxes. The bigger takeaway: Context length is becoming a workload, not just a model specification. Running a model at 256K is one thing. Keeping an agent alive at 300K+ while it reasons and uses tools is much more interesting. And Qwen3.8-Flash-Next is showing that this can be pushed surprisingly far on a single 128GB Mac. 364.7K and counting. Next stop: 400K.

FHILY👑

39,982 次观看 • 22 天前

I’ve been testing Hy4 preview in WorkBuddy, and the most interesting part is not simply the model size, it’s how much practical work it can handle with a relatively focused active parameter count. Hy4 preview brings together stronger code understanding, generation, and editing; improved document and information processing; workflow automation; web and game development; cross-tool collaboration; and more reliable completion of complex, multi-step tasks. In other words, it is designed for work that requires planning, tool use, iteration, and follow-through, not just a quick answer in a chat window. Compared with its initial release, the current Hy4 preview is noticeably faster and better-performing in practical workflows. Following an upgrade released yesterday, it can complete tasks with fewer conversation rounds and lower token usage, while reasoning more quickly and making the overall user experience feel smoother from the first instruction to the final result. For my test, I gave it a demanding Three.js game-prototyping task with a 770B-parameter model and 49B active parameters. The result was more revealing than a simple first-look demo: Hy4 preview handled the core logic, edge cases, and follow-up changes while maintaining the broader context of the project. That combination of capability, speed, context, and active compute is what makes its cost-effectiveness worth examining. A fair evaluation should use the same prompt and environment configuration across models, changing only the model itself. That makes it easier to assess task completion, planning quality, tool-calling stability, reasoning speed, token efficiency, and performance over longer workflows without confusing the result with different settings. If you want to test the model yourself, access Hy4 preview through WorkBuddy and see how it performs on a real coding, document, automation, or creative task: Tencent Hy Tencent AI WorkBuddy

Tyler Wayne

56,152 次观看 • 19 天前

UC Berkeley just open-sourced FreeToken. (2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU. let me explain how: all three models mentioned above are Mixture-of-Experts, and that is what FreeToken takes advantage of. each layer holds hundreds of separate experts plus a small router that picks a few of them per token. Qwen3.6-35B activates roughly 3B of its 35B parameters per token. DeepSeek-V4-Flash picks 6 of 256 experts per layer, so 13B of its 284B run at a time. so compute was never the bottleneck. the weights a single step touches fit comfortably on a consumer GPU. every expert the router might pick still has to exist somewhere. they sit in system RAM, and the GPU keeps a cache of the ones the model has been using recently. so everything comes down to what happens when the router picks an expert that is not on the GPU. there are two ways to serve that miss: 1. copy it over PCIe and run it on the GPU 2. run it on the CPU, where it already lives both read from the same system memory, so they compete for one pool of bandwidth instead of adding to each other. existing engines pick one option and freeze it when the model loads. but routing changes on every token, so a fixed choice misses most of what the model asks for. FreeToken measures both bandwidths on your machine and splits each step's misses between the two paths in proportion. the GPU and CPU results then merge exactly, with no approximation. two machines with the same GPU can end up wanting opposite strategies, which I did not expect. a 5090 in a gaming desktop should push nearly everything over PCIe, while an 8GB laptop is better off computing most misses on the CPU. none of that is readable off a spec sheet, so the engine profiles it once per machine. the second half of the design is about agents. coding agents constantly rewrite their own history, and every edit normally forces thousands of tokens back through prefill. FreeToken saves its checkpoints at the exact boundaries agent frameworks cut on, so it only reprocesses the new part. its slowest first token stays under 44 seconds, while llama.cpp peaks at 232 and KTransformers at 946. it serves the OpenAI and Anthropic APIs under Apache 2.0, so Claude Code and Codex can point at it directly. releasing weights publicly decides who can download a model, not who can afford to run one. frontier open models keep shipping, and running them still assumes a rented cluster. meanwhile there are over a hundred million consumer machines with discrete GPUs sitting mostly idle. closing that gap was never a hardware problem, and work like this is what turns open weights into something you can actually use. paper: repo: almost every idea in this post, from why memory bandwidth decides the outcome to why moving weights costs more than computing on them, comes straight out of how a GPU is built. I wrote a detailed primer on that. the article is quoted below.

Akshay 🚀

343,270 次观看 • 1 个月前

Qwen3.8 Flash Next is starting to look ridiculous on Apple Silicon. I’m running the 4-bit MTP build locally on an M3 Ultra Studio, and the latest OMP run hit: 97.1 tok/s decode 23K context ~1,131 tok/s uncached prompt processing That first number is the one that caught my attention. Nearly 100 tokens per second from a local Qwen3.8 Flash Next setup is already fast enough that the usual “local models are slow” argument starts feeling pretty outdated. And the prompt processing speed is even crazier. Over 1,100 tok/s on an uncached prompt means the model can chew through a large amount of context before generation even starts. The setup matters here. This isn’t just Qwen3.8 Flash Next running untouched. It’s a 4-bit quantized build with MTP, and the inference stack is clearly doing a lot of work behind the scenes to make the hardware perform like this. There’s already a PR open for the implementation on oMLX, so this isn’t just a one-off local experiment either. If these optimizations make their way into the broader MLX ecosystem, running large models locally on Apple Silicon gets even more interesting. The other thing I like about numbers like this is that they put the focus back on the entire inference stack. Model size is one variable. Quantization is another. Then you have MTP, KV cache configuration, runtime optimizations and the hardware itself. Change the recipe and the same model can feel completely different. Qwen3.8 Flash Next at ~97 tok/s on an M3 Ultra is a pretty good demonstration of that.

FHILY👑

15,378 次观看 • 16 天前

hy4 preview changes the question from “can a lab run 770B?” to “can your team stand it up?” tencent hunyuan released an open w8 text model: 𝟕𝟕𝟎𝐁 𝐭𝐨𝐭𝐚𝐥, 𝟒𝟗𝐁 𝐚𝐜𝐭𝐢𝐯𝐞 𝐩𝐞𝐫 𝐭𝐨𝐤𝐞𝐧, Apache 2.0 and native 𝟏𝐦 𝐜𝐨𝐧𝐭𝐞𝐱𝐭. Day 0 paths exist for vLLM and SGLang, plus an official FP8 checkpoint. That already matters. What matters more is the compressed route: mixed GGUF takes the weights from 𝐚𝐩𝐩𝐫𝐨𝐱 𝟏.𝟓𝐓𝐁 𝐭𝐨 𝟐𝟏𝟒𝐆𝐢𝐁 while reported accuracy stays close to BF16 small score movement not a different model. Say it cleanly. this is mixed per layer quant not a blanket 1 bit model. Sensitive layers keep more precision others go lower. Pair every size claim with that quality claim. A 214GiB file that still behaves on coding, agents and long documents is a deployment path. A 214GiB file that does not is just storage. Reality check, because this is a preview: GGUF needs the 𝐩𝐚𝐭𝐜𝐡𝐞𝐝 𝐥𝐥𝐚𝐦𝐚.𝐜𝐩𝐩 path for hyv4. Offload is still offload. Nobody should sell this as a laptop native 770B. Measure tokens/s on your own infra. Compare a real task BF16 or FP8 vs the mixed GGUF and treat vendor charts and your run as separate evidence. The story is ownership, download the weights, serve them and keep lowering the bar without throwing away the 49B-active capability that made the 770B class worth hosting.

Future Coded

32,855 次观看 • 24 天前

Switch Transformer by hand ✍️ ~ 13 steps walkthrough below The Switch Transformer, by Fedus, Zoph, and Shazeer in 2022, is one of the papers that made sparse Mixture of Experts practical at scale. Today, frontier models use MoE to pack enormous parameter counts while activating only a small slice per token: GPT-4, Claude, DeepSeek-V3, and Kimi all follow this pattern. If you want to understand how those models can be huge to store yet still cheap to run, this paper is a good place to start. How does it work? Goal: run five input features through attention, route each one to a single best expert, and read the output off the page. = 1. Given = Input features X1-X5 arrive from the previous block. = 2. Attention matrix = Feed all five features to a query-key attention module to get an attention weight matrix A. = 3. Pooling = Multiply the input features by A to get attention-weighted features Z1-Z5. The effect is to combine features across positions. = 4. Visualize pooling = Z4 is X4 + X5 because the fourth column of A is [0,0,0,1,1]. = 5. Gate values = Multiply the weighted features by the switch matrix. Each gate value says how well expert A, B, or C can probably handle the feature. = 6. Top expert = Pick the row with the highest gate value. Sparse means only the top expert is selected, not all of them. = 7. Routing = Route each Z to its best expert. Every expert has a fixed capacity of 2, so one feature may overflow. = 8. Expert A, linear = Apply the linear layer to the features routed to Expert A. The effect is to combine features across feature dimensions. = 9. Expert A, aggregate = Send the combined feature to the corresponding output column. = 10. Expert B, linear = Apply the linear layer, as in step 8. = 11. Expert B, aggregate = Send the result to the corresponding output column, as in step 9. = 12. Expert C, linear = Apply the linear layer, as in step 8. = 13. Expert C, aggregate = Send the result to the output column. Since one feature exceeded Expert C's capacity, it passes through as-is. Takeaway: a Switch Transformer keeps attention unchanged, then replaces the dense feed-forward network with a sparse set of experts. Most of the parameters sit in the experts, but only a small fraction are used for any one input. That is how GPT-4, Claude, DeepSeek-V3, and Kimi can be enormous to store and still cheap to run. 💾 Save this post!

Tom Yeh

37,410 次观看 • 1 个月前

Alibaba just released a coding model that hits 82 percent on SWE-Bench Verified. That is the highest score ever published for an open-source model. The weights are free. The license is Apache 2.0. You can run it today. The model is Qwen 4 Coder 32B. Here is what 82 percent on SWE-Bench Verified actually means. SWE-Bench Verified tests whether an AI can autonomously resolve real bugs pulled from real production GitHub repositories. Not synthetic exercises. Real open-source projects that real teams depend on. A model gets a bug report, reads the code, writes a fix, and either passes the test suite or it does not. At 82 percent, Qwen 4 Coder 32B resolves 82 out of every 100 real production bugs it is given. Without a human guiding it. On code it has never seen before. For comparison: Qwen 4 Coder 32B: 82 percent SWE-Bench Verified. Open source. Apache 2.0. Claude Fable 5: 80.3 percent SWE-Bench Pro. $10 input / $50 output per million tokens. Currently suspended. GPT-5.6 Sol: Competitive on Terminal-Bench. $5 input / $30 output per million tokens. An open-weight model that you can download and run for free just beat both of them on the benchmark designed to measure real software engineering capability. Here is the architecture. Qwen 4 Coder 32B is a 32 billion parameter dense model. Not a Mixture-of-Experts. Every parameter is active on every request. This matters for inference: a dense 32B model runs on 22 gigabytes of VRAM, which fits on a single high-end consumer GPU or a MacBook Pro with 64GB of unified memory. The smaller variant, Qwen 4 Coder 4B, runs at approximately 135 tokens per second on an M5 Max and fits inside 8 gigabytes of RAM. For a model with usable coding capability, that is a new bar for what fits in a single laptop. The training methodology continued Alibaba's approach of reinforcement learning on verifiable coding tasks. The model gets rewarded when its code passes tests. It gets penalized when it fails. Over millions of training steps, the model learns to write code that actually runs rather than code that looks plausible. License: Apache 2.0. Full commercial use. No attribution requirement. No revenue threshold. No monthly active user ceiling. Weights: Hugging Face, available today. Runs on: vLLM, Ollama, SGLang, and any standard GGUF-compatible inference engine. Qwen 4 32B also runs at approximately 135 tokens per second on an M5 Max chip, setting a new bar for what a sub-8GB model can do on Apple Silicon. The open-source coding model just beat the best closed-source model in the world on the benchmark designed to test whether AI can actually do software engineering. The weights are free. The subscription is optional. Source: Autom8Labs AI Insight July 2026, State of Open Source LLMs June 2026, Kunal Ganglani blog June 2026.

Harman

41,278 次观看 • 2 个月前

PrismML Releases Bonsai 27B: 1-bit and Ternary Builds of Qwen3.6-27B Hitting 89.5% of FP16 at 3.9GB. No new pretrain. No higher-precision escape hatches. No multi-GPU rig. Here's how it works. 👇 1: Codes, not floats Every weight becomes a code, with one shared FP16 scale per group of 128. Ternary is {−1, 0, +1}, binary is {−1, +1}. Sharing the scale across 128 weights keeps its cost at 16/128 = 0.125 bits. → Ternary: log2(3) + 16/128 ≈ 1.71 bits/weight → 5.9GB → Binary: 1 + 16/128 = 1.125 bits/weight → 3.9GB 2: Post-training, not from scratch No BitNet-style low-bit pretrain. It starts from off-the-shelf Qwen3.6-27B, architecture unchanged. The representation runs end to end across embeddings, attention projections, MLP projections, and the LM head. → 9.4× (ternary) and 14.2× (binary) vs the 54GB FP16 baseline 3: Labels are not bit-widths Conventional low-bit builds are mixed-precision by construction. The advertised name describes the most-compressed tensors, not the model. → Q4_K_XL, labeled "4-bit," is really 5.2 bits/weight at 17.6GB → IQ2_XXS, labeled "2-bit," is really 2.8 bits/weight at 9.4GB 4: Fitting a phone is two budgets iOS caps a single app near half of RAM, so a 12GB iPhone exposes ~6GB. The KV cache grows on top. Hybrid attention at ~75% linear means only 16 of 64 layers cache. → 4-bit KV: 4.3GB at 262K context, down from 17.2GB → 11.0 tok/s on iPhone 17 Pro Max 5: The numbers (15 benchmarks, thinking mode) → Ternary: 80.49 avg at 5.9GB — 94.6% of FP16 → 1-bit: 76.11 avg at 3.9GB — 89.5% of FP16 → IQ2_XXS falls to 57.5 on AIME26 while still scoring 88.93 on MMLU-Redux The key takeaway: 27B-class reasoning without the 54GB checkpoint — group-wise ternary and binary codes, an end-to-end low-bit language stack, 4-bit KV, on one phone. Full analysis: Repo: Model weight: Technical details: PrismML

Marktechpost AI

31,860 次观看 • 2 个月前

Tencent just dropped Hy3, and it's worth a look if you're building AI agents. I spent some time putting it through its paces today. A quick rundown of what makes this release notable: → 295B total parameters (Mixture-of-Experts), but only 21B active at inference — a real efficiency play → 256K context window for handling large codebases or long documents → Purpose-built improvements for coding and multi-step agent tasks → Fully open under Apache 2.0 — no restrictive commercial terms → A free two-week API window currently live on OpenRouter On the practical side: prompts involving layered instructions and coding tasks came back fast and coherent. Tencent says this release builds directly on feedback from 50+ internal product teams following an earlier preview, with reported drops in hallucination rate (12.5% → 5.4%) and commonsense errors (25.4% → 12.7%). Those are Tencent's own numbers, so treat them as a starting point rather than gospel until third parties weigh in. For context: Tencent's own comparisons put Hy3 behind GLM-5.2 specifically on coding benchmarks — GLM is a much bigger model (~744B total), so the trade-off makes sense. Hy3's pitch isn't "biggest," it's "efficient enough to actually deploy." Bottom line — if you're evaluating open-weight options for agent workloads, this is a solid one to add to the testing queue while the free window is open. Try it here: Tencent Hy #Hy3 #Hunyuan #TencentAI #AICoding

Felix

36,925 次观看 • 2 个月前

Soofi Consortium Releases Soofi S 30B-A3B: An Open 31.6B Model for German and English Hitting 79.1 German Aggregate With Only 3.2B Active Parameters. Here's how it works. 👇 1. Sparsity in two places at once 52 layers: 23 Mamba-2, 23 granular MoE, 6 Grouped-Query Attention. The MoE router picks 6 of 128 experts per token, plus 2 shared. Mamba-2 carries the sequence mixing with a fixed-size recurrent state, so 46 of 52 layers keep no KV cache at all. → 3.2B of 31.6B parameters active per token 2. Reference architecture on purpose No bespoke backbone. It adopts NVIDIA's Nemotron 3 Nano design without modification — for day-one vLLM kernels, for serving efficiency, and for scientific control. That last one is the real move: Nemotron becomes an architecture-identical baseline, so the data recipe is the only variable left. 3. German as the deliberate variable Three-phase Warmup–Stable–Decay curriculum. Phase 1 is breadth at a 1e-3 plateau, Phase 2 concentrates high-quality data as the LR decays, Phase 3 stretches context to 1M tokens. → ~26.68T consumed tokens → German 7.2% → 15.32% of the mixture, vs ~5% for all non-English in the Nemotron reference → +4.2 German aggregate, +1.8 English, +6.7 held-out English over Nemotron 4. Where the architecture pays: memory bandwidth Every decoded token re-reads the weights and, for a Transformer, the attention cache of every sequence in the batch. Six KV layers instead of 52 keeps that per-sequence state small. Measured on one B200, TP=1, vLLM latency-subtraction. → 8–9× aggregate decode TPS/GPU vs dense 14–24B models at 40K context, batch 32 → decode stays flat from 4K to 256K 5. The numbers (base model, lm-evaluation-harness, 16 open baselines) → 70.1 English aggregate, +2.8 over Olmo 3 32B → 79.1 German aggregate, +6.3 over Apertus 70B → 73.8 HumanEval, 84.2 MBPP-DE, 88.8 GLP-DE, 61.2 INCLUDE-DE Full analysis: Paper: Technical details:

Marktechpost AI

65,624 次观看 • 2 个月前

🚨 I just built a game with an open-source AI model. And honestly… I didn’t expect it to be this capable. Tencent Hunyuan just released Hy4 preview, and it’s already pushing into the top tier of open-source models. Three major releases in six months. That pace is crazy. Here’s what Hy4 preview brings: → 770B total parameters → 49B active parameters → 1M+ token context window → Fully open-source But the numbers aren’t even the most interesting part. Hy4 preview was built around one goal: real-world productivity. Coding. Engineering. Office work. Science. Gaming. Finance. Security. And Tencent didn’t build it in isolation. Hy4 preview was co-designed alongside real products like WorkBuddy, using expertise and real-world data from across Tencent’s ecosystem. So I decided to test it the way I actually like testing AI models: I gave it a game idea and let WorkBuddy help turn it into a playable experience. 🎮 From the initial concept to the actual game logic, it was surprisingly smooth. And the benchmark results back up the hype: 163 internal experts 203 engineering tasks Hy4 preview — 2.99/4 Kimi K3 — 2.94/4 GLM 5.3 — 2.92/4 It also beats GLM 5.2 on benchmarks and comes remarkably close to GLM 5.3. Then comes the part I really like: 💰 ¥6/M input tokens 💰 ¥18/M output tokens 💰 ¥0.30/M cache hits Flagship-level capability without the flagship-level price. And right now, you can try Hy4 preview FREE through WorkBuddy for the next two weeks. If you’re curious what it can actually do, don’t just read the benchmarks. Build something with it. 🔗 Tencent Hy Tencent AI

Aryan Rakib

63,296 次观看 • 19 天前

I tested Tencent’s Hy4 preview model in WorkBuddy on a real frontend build, not a benchmark screenshot. I gave it one practical brief: create an original neon courier game in a single HTML file, with Canvas rendering, keyboard controls, collision detection, scoring, a countdown timer, a boost mechanic, sound effects, and reliable restart logic. The point was not to see whether it could describe a game. I wanted to see whether it could turn a creative idea into a coherent, playable result while handling the details that often break a quick prototype: state changes, movement, collisions, feedback, layout, and interaction flow. Tencent’s Hy4 preview is designed for stronger code understanding, generation, and editing, together with document and information processing, workflow automation, web and game development, cross-tool collaboration, and complex multi-step task completion. That makes it relevant not only for answering questions, but also for taking a task from brief to working output. The current preview version is faster and more effective than the first release. After the latest upgrade, the workflow takes fewer conversation rounds, uses fewer tokens, and reaches useful results more quickly. The difference is most noticeable in the handoff between idea, implementation, revision, and final testing: there is less waiting and less back-and-forth before the result becomes usable. For context, the reference materials describe Tencent’s Hy4 preview as a 770B-parameter model with 49B active parameters, a 1M-token context window, and Apache 2.0 weights. The practical question is how those capabilities translate into real work. In this case, I’m testing whether the model can produce a finished, playable game rather than just a promising code fragment. The one-minute video shows the complete process: selecting Tencent’s Hy4 preview model in WorkBuddy, submitting the build prompt, reviewing the generated result, playing the game, and checking the controls, collisions, score, timer, boost effect, sound, and restart flow. Try your own coding, frontend, document, or automation task in WorkBuddy: Tencent Hy Workbuddy Tencent AI

Rebecca Adson

61,820 次观看 • 12 天前

Most video-action robot models are a content-creation video generator with an action module attached. LingBot-VA 2.0 from Robbyant, a video-action foundation model, throws that starting point out and trains the whole stack natively for control. And it runs closed-loop at a peak 225 Hz. It's so important because A robot cannot move responsively when its controller pauses to imagine the next few frames. LingBot-VA 2.0 predicts during execution, then corrects using each real observation. And it carries only about 13B video parameters while activating roughly 1.9B per token. Bigger robot models usually mean slower reactions, creating a direct conflict between intelligence and control. LingBot-VA 2.0 is trained from scratch for robot control rather than adapted from a video generator built for content creation. Robbyant, an embodied AI company under Ant Group, built it to learn how scenes change under actions, predict what should happen next, and turn those predictions into real-time robot movements. Most video-action systems inherit a tokenizer and video backbone trained mainly to reproduce visual appearance. LingBot-VA 2.0 rebuilds both parts around physical control. Its semantic visual-action tokenizer maps observations toward features from a frozen vision foundation model and learns compact latent actions from frame-to-frame changes using self-supervised inverse and forward dynamics. Unlabeled web video can therefore carry action-relevant training signals without robot action labels. The policy is causal from the start, so every prediction can use only past observations. Its sparse Mixture-of-Experts video backbone has about 13B total parameters, while about 1.9B are active per token, keeping the compute lower during each step. A high-level vision-language planner breaks long tasks into smaller instructions, while the low-level video-action policy handles continuous movement. Foresight Reasoning predicts future visual states while the robot is already acting, then replaces imagined states with every new real observation. Combined with few-step distillation and systems acceleration, the paper reports a peak asynchronous execution frequency of 225 Hz. The model adapts from 10–15 demonstrations, transfers across robot embodiments, and handles some new tasks zero-shot. In the paper’s own evaluations, it reaches 93.6 average on RoboTwin 2.0 and reports stronger real-world results than LingBot-VA and π0.5 across the tested tasks. 🧵 1.

Rohan Paul

11,253 次观看 • 2 个月前

A tricky LLM interview question: You're serving a reasoning model on vLLM, and it keeps running out of GPU memory on long traces. So you add KV cache compression and evict 90% of the cached tokens. VRAM usage stays as is and GPU still runs out of memory. Why? (answer below) Evicting 90% of the KV cache can free almost none of the memory it was using. This sounds counterintuitive, but it follows directly from how production servers store the cache today. The KV cache grows with every token a model generates. Each token appends its key and value vectors across every layer, and nothing is freed while generation continues. This is the dominant memory cost for reasoning models. If a 32K-token CoT caches ~32K tokens of KV vectors, a Qwen3-32B with 4-bit weights will run out-of-memory around 24K tokens on a 24GB GPU. One obvious solution is to keep the important tokens and drop the rest, since attention is sparse enough to allow it. But this does not solve the memory problem yet. The reason is paged attention, which is the memory manager behind vLLM and most production servers. Under the hood, it splits GPU memory into fixed physical blocks, each one holds the KV for about 16 tokens. This block returns to the allocator only when every slot inside it is empty. Since the eviction logic selects tokens by importance, and such tokens are scattered across blocks... ...so despite eviction, almost every block is left with at least some survivor tokens. For instance, if the logic evicts 14k of 16k tokens across 1,000 blocks, most likely every block will still have a token. This means the allocator frees almost nothing. Placing the new tokens into those freed slots is not ideal because it breaks the cache's layout. Say token 16,001 arrives, and it's placed in the slot the 40th token used to hold. The cache now reads position 38, then 16,001, then 41, so the cache is no longer in token order. Attention can still compute the right answer from that, but only if every slot now carries a separate note recording which position it actually holds. This introduces another bookkeeping cost that an in-order layout inherently avoids. So the cache is logically 90% smaller and still physically the same size. Many compression results miss this because they measure on pre-allocated contiguous tensors rather than a paged server. There's another problem. Eviction methods pick which tokens to keep by looking at the attention scores themselves (as expected). But fast attention kernels used in production, like FlashAttention, never save those scores. They compute attention in small pieces and throw the full score grid away as they go, which is also why they're fast. So the exact signal eviction methods need isn't available in memory. The workaround is to fall back to eager attention and build the full matrix, which gives up the speed FlashAttention was there to provide. NVIDIA published a method called TriAttention to solve both these problems. It never needs attention scores. Instead, it scores tokens from the geometry of the model's key and query vectors before RoPE is applied, where those vectors sit in stable clusters. For the memory problem, it runs a compaction pass every 128 decoded tokens. The surviving tokens slide forward to close the holes eviction creates, so whole blocks empty out and return to the allocator while the cache stays in token order. On long reasoning traces, the approach matches full-attention accuracy while decoding 2.5x faster and using 10.7x less KV memory. KV cache compression is a big infrastructure problem. The number that decides whether it works is the count of freed blocks, not the count of evicted tokens. You can find the NVIDIA write-up here: I wrote a first-principles breakdown of how the KV cache works. It walks through why the model stores keys and values at all, why the cache grows with every token, and a comparison of LLM generation speed with and without KV caching. Read it below.

Avi Chawla

273,525 次观看 • 3 个月前