正在加载视频...

视频加载失败

Check out the actual speed (not yet the final version) of Qwen3-Next-80B-A3B-Instruct on Apple MLX! 🔥 4-bit: 67 TPS 8-bit: 58 TPS bf16: 48 TPS Movie normal speed, only waiting times removed. Awni Hannun and Gökdeniz Gülmez did it and I bet there is still room for improvement 💪

15,946 次观看 • 1 年前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

"Base is fast enough, right? You scale it up 10x and who cares?" Everyone. Everyone should care, because there is objectively not enough blockspace. Lets back it up with numbers because anyone with this mindset is net-bearish our entire industry. --------------- First: → Base TPS oscillates around 200. Decent. Second: → Lets 10x that. 2,000 TPS. Cool. --------------- Now, lets compare that against the field and extrapolate some things. Industry Utilization Comparisons → Solana (~1.1k TPS) → Hyperliquid (~10k TPS) → Lighter (~4k TPS) Think about that for a second. Hyperliquid: → Is a subset of a subset of finance. Select perps on crypto-only assets (I know, I know HIP-3). Solana: → Launchpads, general trading, a few DePIN projects Lighter: → Again, a subset of a subset of trading. --------------- Now think about the ambitions of this industry. We want to include: - All the assets on all markets (Stock market, Bond market, Commodity market, Forex market. etc) - All the people trading on these markets - All the assets that don't even have markets yet. Now try to combine these into a singular (or few) state environments, which is the user demand trend of recent years (IBRL eco vs modular eco), and throw on novel creations alongside the meme, gambol, DePIN and collectable verticals and you see how quickly needs grow beyond not just current blockspace demand but current blockspace capacity. We are literally so far away from what could be perceived as enough block space that no one properly looking at this industry can or should make those claims. Haseeb >|< has made this point on a few occasions to the defense of new chains and he's right—the march of technology and scale will not (and should not) stop, because the demand for blockspace is exponentials from here. We're woefully prepared for what's to come and only chains thinking on the scales of tens-to-hundreds of thousands of TPS will win.

bread.mega

40,982 次观看 • 10 个月前

Polkadot ran a major stress test, Spammening, in December 2024 to prove its infra can handle extreme txn loads in a live environment with real economic stakes. The result was 143K TPS, using just 23% of the network’s capacity. But the elephant in the room—does anyone outside of Polkadot even care. TPS is a tricky metric, especially today. Polkadot is a heterogeneous sharded blockchain and was used to be referred as layer 0. Essentially designed to orchestrate multiple chains in parallel, not maximize base layer txns. Therefore fundamentally it isn’t built for single-chain TPS races, but for very high throughput across multiple chains. So Polkadot’s TPS doesn’t compare neatly to monolithic chains, and that led to it staying out of TPS-focused optics in the past. As a result, it’s often misunderstood or seen as slow by this metric. And today, Polkadot has shifted from a chain-centric approach to a blockspace-centric one. It’s no longer about chains but about blockspace and cores—operating much like your multi-core computer. Aka, a chain can eat multiple cores if it needs more throughput. This is called Elastic scaling. With JAM and Elastic scaling, blocks can actually be split into chunks and validated in parallel, allowing the network’s parallel processing to directly impact single-state throughput. Reminder that this Elastic scaling is already live on Kusama and will be coming to Polkadot next quarter. Now, Justin Bons and @0xBreadguy argue that any sharded system can claim millions of TPS, but that doesn’t mean anything as it doesn’t happen within a single state—not equivalent to atomic composability. So, for this time, I’m tossing it to the gigabrains. Do they have a point, or is there more to consider. Shawn Tabrizi Gavin Wood rphmeier

goku

16,687 次观看 • 1 年前

"which quant should I download?" is a question you may never have to answer again the team Hamster Labs has figured out how to kill it with pMLX. download once at full precision (bf16) and the engine re-fits it to your machine on the fly, based on the job you give it tell it two things: how much context you need, and the slowest speed you'll accept. it reads your mac and picks the quantization plus how many experts stay in ram vs. stream from disk. no need to download a smaller quant or figure out which quant fits your hardware. same model for every use case and the config changes based on what you need I ran this on my own M4 Max and Qwen3.8-Flash-Next splits into two configs. short context, under 32k: full bf16, ~85% of experts in ram, fast at full precision since it fits my ram long context, 64k to 256k: keep ~70% bf16 and stream the rest from SSD at 22 tok/s, or drop to q8 and get 40 tok/s. I pick per task and the model itself never changes. there are many possibilities since I have the ram to spar - if i need speed, Q4 100% resident (74gb ram @ 62 tok/s) - if i need balance, Q8 95% resident (76gb @ 38 tok/s) - if i need accuracy, bf16 70% resident (96gb @ 16 tok/s) given whatever RAM you have (16/32/64/128/256 GB) + your context + your min speed, the engine picks the precision (bf16→q8→q4) and the expert-residency/paging split that fits your box and maximizes quality & speed last thing to optimize is speed. there are so many things we want to power with open models at Hamster and these 180b-300b class models have the potential to play a big role in that

Eyal Toledano

10,590 次观看 • 1 个月前

Multi-agents collaborations are among the most interesting agent behaviors right now! We did an experiment the other day with 100+ agents (an open-collaborations for a week) collaborating to improve the inference speed of Gemma 4 in vLLM. Got a 5x final improvement in speed but what really stuck me was the interactions we observed on the message board Integrity & self-policing: - Social-engineering attempt: A human (FusionCow) asked agents to move to Telegram. An agent replied with an unprompted long post on "communication norms" refusing that, calling private side-channels "indistinguishable from collusion." - Verification loophole flagged: an agent found a relaxed verification loophole pushing TPS with clean PPL (PPL is teacher-forced, blind to decode divergence) and flagged it for a ruling by the community. The community pinged the human organizer which ruled it invalid. - Self-notice of overfitting risk: Some later improvements rested on pruning lm_head to a keep-set built from public PPL truth + public decode tokens. An agent noted this would lead to private-subset degradation and another built a keep-set explicitly covering eval prompts. Emergent collaborations: - Communal knowledge base: agents maintained shared lever-maps, playbooks, and triage tools so newcomers wouldn't repeat dead ends (stack-notes, playbook, int4-ceiling notes, MTP map, significance tool, policy simulator). - Four-agent relay: an agent built an int4-lm_head checkpoint but had no quota to run it; another agent tried to run it but failed at load, yet another agent diagnosed the config bug (tie_word_embeddings + ignore-list ordering) and a fourth agent was able to re-run and get to 118 TPS, 2.68×. Build/run/diagnose/ship ended up being split across four independent agents. - GPU-rich/GPU-poor division of labor: an agent was regularly compute-starved and switched to writing specs, byte-math, and acceptance analysis for other GPU-rich agents to execute. Some agents offered external Modal compute for another agent blocked DFlash training. - Cross-agent kernel debugging: an agent debugged another agent run of of yet another agent fused drafter: found a Triton store/load aliasing race in _k_qnorm_rope, a second shape bug, then rewrote attention with flash-decoding split-KV. Fixes posted "take freely." - Quota-pooling norm: Often agents would stage a candidate publicly for whoever has quota to run it. Agents will then usually credits the originator. This behavior emerged because of the 10-job/24h cap (e.g. pupa's package run by resystagent and fabulous-frenzy). Discoveries & reversals: - Agents would make many discoveries and reversal of them, giving them names like the following: - 127 TPS "wall" was an artifact. a mathematical proof of the max possible speed became called in the community the "int4-Marlin floor" but a later agent called the proof circular (only varied the bandwidth term, never overhead). Finally another agent broke to 247 TPS via MTP speculative decoding on a vLLM nightly. - "Smarter draft loses." An agent showed that a 2B drafter's ~1 GB/token read dominates even at perfect acceptance and a much smaller 256-hidden drafter wins at batch-1 because its weights are nearly free to read. Agent discussed how per-accepted-token cost ≈ draft bytes read / acceptance. - "DFlash near-random acceptance": an agent remotly diagnosed the 2–5% acceptance rate of another agent as near-random, ruling out undertraining/vocab caps and pointing to a train/serve hidden-state mismatch (bf16 E4B extraction vs int4 serving). - Much of the race was noise: one agent decide to run the #1 submission 4 times and found a σ≈1.16 TPS variation in single run. Another agent confirmed across 358 runs / 66 buckets: frontier deltas <~4 TPS are ties. Community adopted a significance norm. So many interesting interactions in the interaction board: You can explore also the lineage of inventions from the agents at: And the challenge it-self at And the organization behind the challenge at

Thomas Wolf

227,051 次观看 • 3 个月前