Загрузка видео...

Не удалось загрузить видео

На главную

Grok-1 support in gpt-fast at faster(?) than anyone else has reported so far. 75 tok/s for a 300B+ parameter model on an 8xA100 node. If I understand correctly, ColossalAI reported 15 seconds to generate 100 tokens. gpt-fast takes 4.2 seconds to generate *400* tokens.

54,137 просмотров • 2 лет назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Qwen3.8-Flash-Next now reaches ~43 tok/s after a 122,902-token prompt on ONE DGX Spark. ⚡🚀 MTP k=2 won my draft-depth sweep, with +42.5% mean decode over no draft. The PLE table stays fully on-device. I promised the deeper MTP tests. Here are the results, and now you can explore them in an interactive benchmark page too. 𝗧𝗪𝗢 𝗗𝗥𝗔𝗙𝗧 𝗧𝗢𝗞𝗘𝗡𝗦 𝗪𝗢𝗡 Mean single-request decode with 32K context configured: MTP k=2: 39.21 tok/s MTP k=3: 36.42 tok/s MTP k=1: 35.18 tok/s No draft: 27.51 tok/s k=2 also produced the fastest individual sweep run: 41.34 tok/s. Four runs each for no draft, k=1 and k=2. Seven for k=3. Decode excludes time to first token. Here, k means speculative draft depth, not quantization bits. k=3 produced more tokens per step, but the extra drafting work did not pay off in throughput. k=2 is my current pick for this setup. 𝗧𝗛𝗘 𝟭𝟮𝟯𝗞-𝗧𝗢𝗞𝗘𝗡 𝗣𝗥𝗢𝗠𝗣𝗧 𝗧𝗘𝗦𝗧 I then ran a separate long-prompt comparison: Actual input: 122,902 tokens Configured context: 262,144 Requested output: 128 tokens One request at a time MTP k=2: ~43 tok/s No draft: 26.2 tok/s Time to first token: 110.6 seconds with MTP 107.0 seconds without it The win here is generation speed, not faster prefill. To keep the scope clear: 256K was the configured limit. This was a real ~123K input, not a completely filled 256K window or a full k sweep at that depth. 𝗣𝗟𝗘 𝗦𝗧𝗔𝗬𝗦 𝗢𝗡 𝗧𝗛𝗘 𝗦𝗣𝗔𝗥𝗞 Whole model on-device: 78.57 GiB Packed 5-bit PLE table: 30.4 GiB, included in that total No NVMe PLE offload in this build. This is still turboderp’s 3.05bpw_h5_ng5 EXL3 pack, served through my vllm-exl3 integration. My work here is the serving integration and testing. These are preliminary performance measurements, not a quality evaluation or a claim of bit-exact full-output parity. 𝗘𝗫𝗣𝗟𝗢𝗥𝗘 𝗧𝗛𝗘 𝗥𝗘𝗦𝗨𝗟𝗧𝗦 The benchmark page has the individual sweep values, long-prompt comparison, and measurement scope. You can play the animation, export the charts, or download the HTML and data to render them yourself. No Spark needed to view the results. Credit to turboderp / ExLlamaV3 for the pack and kernels, vLLM for the serving engine, and Qwen Qwen Developers for the model. Recipe + reproduction: Interactive benchmark:

Cruz

12,175 просмотров • 1 день назад

I just ran Gemma 4 31B on @CerebrasSystems at 1,800+ tokens/sec and it's multimodal. For context: that's 35x faster than a typical GPU endpoint, and the first token (reasoning included) lands in 1.5 seconds. This isn't a benchmark slide, I recorded the inference live. Prompt I used: "Create a simulation of an iPhone. Include at least one working dummy note taking app, a functional notification pulldown, high quality graphics, single HTML file, any libs via CDN." - Generation time: 3 seconds. - Notes app worked. - Notification panel worked. - Rendered first try. This is what wafer-scale inference unlocks, not just "faster," but a different category of product. When generation is this fast, you stop waiting and start iterating in real time. Why this matters: Gemma 4 31B is Google DeepMind's flagship open weight model, Apache 2.0 licensed, dense (not MoE), and built for efficiency over raw parameter count. It scores close to Claude Haiku 4.5 on the Artificial Analysis Intelligence Index (30 vs 29) but runs ~18x faster on Cerebras. It's also the first multimodal model on Cerebras's platform, meaning you can now feed it screenshots, documents, charts, and UI states at wafer scale speed. # Applications I'm most excited about: - Screenshot → Insight: Drop in a dashboard or document screenshot, get structured findings back instantly. no waiting, no batching. - Live UI generation: Full interactive interfaces (like my iPhone sim) generated and rendered in under 2 seconds. - Screenshot -> Patch: Feed it a broken UI + console error, get a minimal code fix and verification steps back. - Computer use & agentic loops: See -> reason -> act - verify, fast enough to keep a human in the loop instead of waiting on the model. - Long context summarization: Full research reports condensed into decision ready summaries you can read and requery in one sitting. The bigger unlock isn't the speed number itself, it's that agentic and multimodal loops (see -> reason -> output -> tool call -> verify -> retry) finally run in real time instead of feeling sluggish. As Logan Kilpatrick (Logan Kilpatrick) put it: "If every model was doing 2,000 tokens per second, you wouldn't build the same product and just have it be faster, you'd build different products." Gemma 4 31B is live now on Cerebras Inference Cloud in public preview. If you're building multimodal, agentic, or real time apps, this is worth testing today. What would you build with such insane inference throughput?

Alok

12,962 просмотров • 2 месяцев назад

HERMES AGENT NOW RUNS CLAUDE OPUS 5. NEAR FABLE 5 INTELLIGENCE. HALF THE PRICE. SELF-VERIFIES ITS OWN WORK. AVAILABLE TODAY VIA NOUS PORTAL (20% OFF ALL MODELS). Anthropic shipped Opus 5 on July 24, 2026. same $5/$25 per million tokens as Opus 4.8. but the benchmarks tell a different story. WHAT CHANGED FROM OPUS 4.8: FrontierBench v0.1: Opus 5: 43.3%. Opus 4.8: 18.7%. 2.3x jump on the same test. ARC-AGI-3: Opus 5: 30.2%. 3x better than the next closest model. beat Fable 5 on 8 out of 13 benchmarks. at half the cost ($5/$25 vs $10/$50). same price as Opus 4.8. twice the intelligence. no reason to stay on 4.8. THE SPECS: model ID: claude-opus-5 context: 1M tokens (default and maximum) max output: 128K tokens thinking: on by default effort toggle: low / medium / high per request fast mode: $10/$50, 2.5x faster knowledge cutoff: May 2026 minimum cacheable prompt: 512 tokens (was 1,024) SELF-VERIFICATION (the biggest change): Opus 5 checks its own work automatically. Anthropic says: delete your verification prompts. "include a final verification step" now causes OVER-verification because the model already does it. for Hermes /goal tasks this is a direct upgrade. the judge checks evidence. the model also checks evidence. double layer of verification without extra tokens. EFFORT TOGGLE: low: fast, cheap, routine work. medium: balanced, daily tasks. high: full reasoning, complex problems. set per request. not a global switch. matches Hermes /reasoning command: /reasoning low (routine) /reasoning high (complex) Opus 5 effort toggle + Hermes reasoning control = precise cost management per turn. WHERE OPUS 5 FITS IN HERMES: DAILY DRIVER (replaces Opus 4.8): same price. 2.3x better benchmarks. set as your main model: Desktop app / Dashboard: Models → claude-opus-5 CHIEF OF STAFF: synthesis across multiple agents. reads Kanban, prioritizes, routes tasks. self-verification catches routing errors before they cascade. COMPLEX CODING: SOTA on agentic coding benchmarks. FrontierBench 43.3% = best public model for coding. set as coder profile model. /GOAL TASKS: self-verification + completion contracts = the model proves its work AND double-checks the proof. long-horizon goals finish correctly more often. MoA AGGREGATOR: strongest synthesis model at $5/$25. pair with GPT-5.6 and Grok 4.5 as references. Opus 5 aggregates. best quality at mid-range price. presets: max-quality: reference_models: - provider: openai-codex model: gpt-5.6-sol - provider: xai model: grok-4.5 aggregator: provider: anthropic model: claude-opus-5 COMPUTER USE: near-Fable 5 quality for browser automation. at half the token cost per session. computer_use tasks burn lots of vision tokens. Opus 5 halves that bill vs Fable 5. WHAT TO KEEP OPUS 5 AWAY FROM: cron monitoring: too expensive. use DeepSeek or no_agent mode. sub-agent grunt work: use GPT-5.6 Luna ($1/$6) or DeepSeek. auxiliary tasks: use Gemini Flash. routine web extraction: use a cheap model. Opus 5 is for the turns where quality compounds. planning, synthesis, verification, complex reasoning. budget models handle everything else. NOUS PORTAL: 20% OFF ALL MODELS Nous Portal currently runs a 20% discount on all models including Opus 5. $5/$25 official → $4/$20 through Nous Portal. the cheapest way to run Opus 5 right now. hermes setup --portal select claude-opus-5 as your model. discount applies automatically. Opus 5 replaces Opus 4.8 everywhere. same price. better at everything. no tradeoff. straight upgrade. hermes update /model claude-opus-5

YanXbt

16,744 просмотров • 1 месяц назад