Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Cerebras Code just got an UPGRADE. It's now powered by GLM 4.6 Pro Plans ($50): 300k ▶️ 1M TPM @ 24M Tokens/day Max Plans ($200): 400k ▶️ 1.5M TPM @ 120M Tokens/day Fastest GLM provider on the planet at 1000 tokens/s and at 131K context. Get yours before we...

177,904 Aufrufe • vor 9 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Qwen 3.8 27B Q4_K_M - 90 tokens/sec on a single NVIDIA RTX 4090 (24 GB VRAM) with Dflash2! (MTP 60 tps -> 90 tps Dflash2!!!!) Local AI moves so fast (literally!) it’s terrifying. Z lab just dropped DFlash 2 for Qwen 3.8 27b and Muse Glimmer. I patched llama.cpp (PR #27342) and paired it with Unsloth’s Qwen 3.8 27B UD-Q4_K_XL quant. The result? Lossless 90 tokens/s decode. My last post highlighted native MTP hitting 60 t/s at 130,000 context. But DFlash 2 just completely shattered that ceiling. By using parallel block diffusion drafting (predicting whole blocks of tokens in a single pass using dynamic convolutions), DFlash achieves a massive 5.39 token acceptance rate. THE ALPHA TWEAK: `n-max 7` eats too much VRAM for draft states. But if you drop the draft limit to `--spec-draft-n-max 4`, you slash the VRAM overhead and actually increase the throughput. Here is the new 24GB VRAM Physics Matrix (DFlash 2 @ n-max 4): - 30k Context: 1,725 t/s prefill | 87.05 t/s decode | 22.2 GB VRAM - 80k Context: 1,789 t/s prefill | 84.20 t/s decode | 23.3 GB VRAM - 110k Context: 1,767 t/s prefill | 83.35 t/s decode | 23.96 GB VRAM (110k context at 83+ tokens a second sitting exactly on the 24GB hardware limit is absolute wizardry). How to compile the PR today: git clone cd llama.cpp git fetch origin pull/27342/head:pr-27342 git switch pr-27342 cmake -B build -DGGML_CUDA=ON && cmake --build build -j Llama.cpp flags for Dflash (110k Context Ceiling): ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q4_K_M.gguf --spec-type draft-dflash --spec-draft-n-max 4 -c 110000 -ngl 99 --port 8080 -ctv q4_0 -ctk q4_0 The fact that the open source community is shipping block diffusion drafters so quickly that run entirely locally on a gaming GPU is unbelievable. If you own a single RTX 3090 or 4090, it is officially time to upgrade to qwen 3.8 27b with dflash 2 and cancel your API subscriptions and let local silicon eat the cloud. This model beats GPT 5.6 Terra, GLM 5.2 DeepSeek V4 Pro, Muse Spark 1.2 and Claude Opus 4.8 on the artificial analysis agentic index (details in the replies) Hugging Face GGUF links (Base + DFlash2) and the full visual VRAM scaling and Dflash2 vs MTP graphs are also in the replies below. are you sticking to native MTP for the 130k context, or sacrificing 20k context to redline your decode speed? How many tokens/sec are you pushing on your current local rig?

Alok

102,967 Aufrufe • vor 8 Tagen

New open-source agent harness just landed! I got early access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.

elvis

11,303 Aufrufe • vor 8 Tagen

Do you want to own part of a AAA game? I know, you hear it all the time. “Triple A game”, you go to play it, it’s crap. This is different, and it’s only possible with Sonic (Sonic) speed, transaction cost, and of-course FeeM. A game that includes talent from Kojima, Ubisoft, EA Sports, Gameloft & more with advisors from NVIDIA. A game that you’ll be able to play on mobile, desktop, and then Xbox and PlayStation (yes really)! YES! A PRETTY BIG DEAL! Before I tell you about the sale, let me at least tell you about this game (being a massive gamer nerd, this excited me), so…. Introducing Animera (Search for Animera): • Fast-paced skill-based PvP in the Nubera galaxy • Compete in real-time space battles for real rewards It will be powered with $STRIKE: • Compete2Earn: win matches, earn tokens • Play2Burn: 5% of $STRIKE used in matches gets burned Oh, and with 8.75% of all game revenue will be used to buy & burn $SWPx, so the SwapX (SwapX) community owns a real stake in this AAA title. Absolutely insane. > Now let me tell you about its beta run quickly: • 16K+ beta signups • 500+ players added weekly • 7.5K+ matches already played • Launching to 500K+ mobile users via Nomina Games > How can you own a piece of Animera? June 5th at 2pm EDT the sale will go live on SwapX, it will go in three phases each lasting 12 hours or until sold out: PHASE 1️⃣: xNFT Holders Early access with exclusive perks and bonuses. These are for xNFT holders only you can get these here on paintswap PHASE 2️⃣ Whitelisted Communities These will be whitelisted from Creo Engine, SFA AGC, derp, and GOGLZ | SONIC 🥽💥. PHASE 3️⃣ Public Round Any remaining allocation will open to the public - only if Phases 1 & 2 don’t sell out. > What is the raise? Token Price & Allocation: • Token: $STRIKE • Currency: USDC • Total tokens for sale: 101.75M Unlock structure: • 50% unlocked at TGE • Remaining 50% claimable in 30 days • Raise cap: Max $100,000 per user, capped at $10,000 per xNFT • Purchase window priority: xNFT holders get early access (see above)! Transparency is key: Why I love working with the team is because transparency is crucial, so I’m going to tell you about its tokenomics, seed, and fully diluted valuation here: Token Symbol: STRIKE Total Supply: 370,000,000 Initial FDV: $1.48M Total Raise: $950,160 Total Initial Unlock: 112,947,501 STRIKE Initial Market Cap (excluding liquidity): $303,790 Token Allocation: • Seed Round: 59.2M tokens (16% allocation), with a 1-month cliff and linear vesting over 9 months. • Private Round: 94.35M tokens (25.5% allocation), with a 1-month cliff and 6-month vesting period. • Crowdsale: 10.75M tokens (2.91% allocation), unlocked 50% at TGE. • xNFT Holders: 10M tokens (2.7% allocation), with a 1-month cliff. • Liquidity: 37M tokens (10% allocation), with no lock or vesting. • Team: 18.5M tokens (5% allocation), with a 6-month cliff and 12-month vesting. • Rewards: 28.6M tokens (8% allocation), vested over 18 months. • Product Growth: 19.6M tokens (5.3% allocation), vested over 24 months. Token Offering: • Seed Round: Priced at $0.0033 per token, raising $195,360 by selling 59.2M tokens. 10% unlocks at TGE, with a 1-month cliff and 9-month vesting. The initial market cap from seed unlock is $234,127. • Private Round: Priced at $0.0037 per token, raising $349,095 for 94.35M tokens. 15% unlocks at TGE, with a 1-month cliff and 6-month vesting. Initial market cap contribution is $262,508. • Crowdsale: Priced at $0.0040 per token, raising $407,000 by selling 10.75M tokens. 50% unlocks at TGE, with no cliff or vesting. Adds $283,790 to the initial market cap. It’s important you had the full information at hand so you can decide whether or not you’d like to participate. I will be, because it’s a low FDV and it looks great. This is not financial advice, I’m helping the team out. Below is real gameplay: Further details: 👇

hoeem

21,634 Aufrufe • vor 1 Jahr

🔥 Congratulations on the Launch of the Die Last Playtest & Upcoming Airdrop! 🔥 We are thrilled to announce the official start of the Playtest! Moreover—we've prepared a bunch of surprises just for you. Alongside the Playtest, we're hosting the first-ever in-game token AirDrop 💎. Your mission: dig and craft Lost Objects, remnants of a forgotten civilization. To participate in the upcoming token AirDrop, you need to find a Spy Radio to craft an NFT, your key to this exciting journey. Anyone can find in-game NFTs, which are the key to owning wallets with different amounts of tokens needed to participate in the AirdDrop: the more tokens you have, the more you'll get 💸. The luckiest players can obtain an NFT with a full-filled wallet, which means you can participate in the AirDrop at no cost! Didn't find an NFT with a 100% discount? No worries, keep digging! Every NFT you find boosts the number of tokens you receive during the AirDrop. For instance, if you have 100 tokens in your wallet and craft 6 NFTs, you’ll receive 80 tokens in the AirDrop. The more NFTs you have, the more tokens you’ll get. ✅ To participate, add #DieLast to your Steam wishlist and apply for the Playtest (approval of your application may take up to 24 hours or be approved instantly): ✅ To grab your NFT once it's dug up, visit our website: ❗️ Important note: We will approve applications for the Playtest on a limited basis to balance the load on the network. Once we are confident that the network is functioning as intended, we will accept all applications at once. We expect this to occur within one day and will keep you updated. 💥 Our dedicated followers know about our quest campaign with Bitget Wallet 🩵 and DeGame. Join us and complete quests to earn rewards totaling $3000+ USDT 👉 Dive into the adventure and unearth hidden treasures. Welcome to the Die Last Planet! 🌎 #crypto #taptap #clicker #game

Die Last

13,396 Aufrufe • vor 2 Jahren

HERMES AGENT NOW RUNS CLAUDE OPUS 5. NEAR FABLE 5 INTELLIGENCE. HALF THE PRICE. SELF-VERIFIES ITS OWN WORK. AVAILABLE TODAY VIA NOUS PORTAL (20% OFF ALL MODELS). Anthropic shipped Opus 5 on July 24, 2026. same $5/$25 per million tokens as Opus 4.8. but the benchmarks tell a different story. WHAT CHANGED FROM OPUS 4.8: FrontierBench v0.1: Opus 5: 43.3%. Opus 4.8: 18.7%. 2.3x jump on the same test. ARC-AGI-3: Opus 5: 30.2%. 3x better than the next closest model. beat Fable 5 on 8 out of 13 benchmarks. at half the cost ($5/$25 vs $10/$50). same price as Opus 4.8. twice the intelligence. no reason to stay on 4.8. THE SPECS: model ID: claude-opus-5 context: 1M tokens (default and maximum) max output: 128K tokens thinking: on by default effort toggle: low / medium / high per request fast mode: $10/$50, 2.5x faster knowledge cutoff: May 2026 minimum cacheable prompt: 512 tokens (was 1,024) SELF-VERIFICATION (the biggest change): Opus 5 checks its own work automatically. Anthropic says: delete your verification prompts. "include a final verification step" now causes OVER-verification because the model already does it. for Hermes /goal tasks this is a direct upgrade. the judge checks evidence. the model also checks evidence. double layer of verification without extra tokens. EFFORT TOGGLE: low: fast, cheap, routine work. medium: balanced, daily tasks. high: full reasoning, complex problems. set per request. not a global switch. matches Hermes /reasoning command: /reasoning low (routine) /reasoning high (complex) Opus 5 effort toggle + Hermes reasoning control = precise cost management per turn. WHERE OPUS 5 FITS IN HERMES: DAILY DRIVER (replaces Opus 4.8): same price. 2.3x better benchmarks. set as your main model: Desktop app / Dashboard: Models → claude-opus-5 CHIEF OF STAFF: synthesis across multiple agents. reads Kanban, prioritizes, routes tasks. self-verification catches routing errors before they cascade. COMPLEX CODING: SOTA on agentic coding benchmarks. FrontierBench 43.3% = best public model for coding. set as coder profile model. /GOAL TASKS: self-verification + completion contracts = the model proves its work AND double-checks the proof. long-horizon goals finish correctly more often. MoA AGGREGATOR: strongest synthesis model at $5/$25. pair with GPT-5.6 and Grok 4.5 as references. Opus 5 aggregates. best quality at mid-range price. presets: max-quality: reference_models: - provider: openai-codex model: gpt-5.6-sol - provider: xai model: grok-4.5 aggregator: provider: anthropic model: claude-opus-5 COMPUTER USE: near-Fable 5 quality for browser automation. at half the token cost per session. computer_use tasks burn lots of vision tokens. Opus 5 halves that bill vs Fable 5. WHAT TO KEEP OPUS 5 AWAY FROM: cron monitoring: too expensive. use DeepSeek or no_agent mode. sub-agent grunt work: use GPT-5.6 Luna ($1/$6) or DeepSeek. auxiliary tasks: use Gemini Flash. routine web extraction: use a cheap model. Opus 5 is for the turns where quality compounds. planning, synthesis, verification, complex reasoning. budget models handle everything else. NOUS PORTAL: 20% OFF ALL MODELS Nous Portal currently runs a 20% discount on all models including Opus 5. $5/$25 official → $4/$20 through Nous Portal. the cheapest way to run Opus 5 right now. hermes setup --portal select claude-opus-5 as your model. discount applies automatically. Opus 5 replaces Opus 4.8 everywhere. same price. better at everything. no tradeoff. straight upgrade. hermes update /model claude-opus-5

YanXbt

16,744 Aufrufe • vor 1 Monat

This Chinese developer launched Llama 70B locally on a MacBook on a plane and for a full 11 hours without internet ran client projects. He was sitting by the window on a transatlantic flight with a MacBook Pro M4 with 64 GB of memory. WiFi on board cost $25 for the flight. He declined. No cloud API, no connection to Anthropic or OpenAI servers, no internet at all. Just a local Llama 3.3 70B on bf16 and his own orchestrator script. The model runs through llama.cpp. Generation speed, 71 tokens per second. Context around 60,000 tokens. Memory usage, 48.6 GiB out of 64. Battery at takeoff, 3 hours 21 minutes. And he gave the orchestrator this system prompt before takeoff: "You are an offline orchestrator running on a single MacBook. There is no network. The only resources you have are local files in /Users/dev/work, the Llama 70B inference server at localhost:8080, and a battery budget of 3 hours 21 minutes. Process the queue at /Users/dev/work/queue.jsonl (one client task per line). For each task: draft → run local evals → save artefact to /Users/dev/work/done/. Save context checkpoints every 12 tasks so you can resume after a battery swap. Stop only on empty queue or when battery drops below 5%." So the system knows exactly what resources it is running on. It knows it has no connection to the outside world for the next 11 hours. It knows it has finite memory and a finite battery. It knows the human will not intervene until the plane lands. The system runs in 1 loop. Takes a task from the queue, runs it through inference, saves the artifact, writes a checkpoint. Task after task, just like that. And only when the battery drops below 5% does the orchestrator automatically pause, waits for the laptop to switch to the backup power bank, and continues from the last checkpoint. Here is what the system actually writes in his log during the flight: "saved context checkpoint 8 of 12 (pos_min = 488, pos_max = 50118, size = 62.813 MiB)" "restored context checkpoint (pos_min = 488, pos_max = 50118)" "prompt processing progress: n_tokens = 50 / 60 818" "task 37016 done | tps = 71 s tokens text → /Users/dev/work/done/proposal_westside.md" Outside the window, clouds, blue sky, and no WiFi. On the tray, 1 MacBook, an open terminal on 2 screens, and an inference server on localhost. From what I have observed, this is the cleanest offline AI workflow I have seen in the past year: 11 hours of flight, $0 for WiFi, and the entire client queue closed before landing.

Blaze

1,841,161 Aufrufe • vor 3 Monaten

GROK SURGES TO THE FRONT OF THE GENAI RACE AS GROWTH SKYROCKETS Grok is absolutely amazing, continuing to stun with incredible results. Web traffic jumped nearly 15% month-over-month in November 2025, the fastest growth in the generative AI industry, proving Elon’s xAI project isn’t just competing, it’s taking real market share from ChatGPT. Grok hit around 234.4 million visits in November, up from 204 million in October. That’s a staggering 1,300% year-over-year surge, pushing it past Perplexity and Claude in user growth and cementing it as the world’s #2 chatbot by market share. The breakout came with Grok 4.1’s release in mid-November. The update debuted at number 1 on LMSYS Arena, that’s the global benchmark where AI models are ranked through blind human evaluations. Grok’s new “Thinking Mode” scored 1483 Elo, a 31-point lead over every open model, beating Gemini 2.5 Pro and Claude 3. The upgrade also cut hallucinations (false answers) by two-thirds, expanded its context window to 2 million tokens, about 1.5 million words of memory, and dominated top reasoning and coding tests like, graduate-level logic, and the emotional intelligence aspect. U.S. traffic surged to 51.5 million visits, boosted by X integration and Grok’s unfiltered style. At just $0.20 per million input tokens, versus GPT-5.1’s $1.25—it’s winning with both speed and affordability. The momentum isn’t hype, it’s lift-off. If growth holds, Grok could reach 500 million users by mid-2026, forcing every rival to redefine what “intelligent” really means. To truthful AI winning! Source: X Freeze, NextBigFuture, CometApi, Langcopilot

Mario Nawfal

33,823 Aufrufe • vor 8 Monaten

I just ran Gemma 4 31B on @CerebrasSystems at 1,800+ tokens/sec and it's multimodal. For context: that's 35x faster than a typical GPU endpoint, and the first token (reasoning included) lands in 1.5 seconds. This isn't a benchmark slide, I recorded the inference live. Prompt I used: "Create a simulation of an iPhone. Include at least one working dummy note taking app, a functional notification pulldown, high quality graphics, single HTML file, any libs via CDN." - Generation time: 3 seconds. - Notes app worked. - Notification panel worked. - Rendered first try. This is what wafer-scale inference unlocks, not just "faster," but a different category of product. When generation is this fast, you stop waiting and start iterating in real time. Why this matters: Gemma 4 31B is Google DeepMind's flagship open weight model, Apache 2.0 licensed, dense (not MoE), and built for efficiency over raw parameter count. It scores close to Claude Haiku 4.5 on the Artificial Analysis Intelligence Index (30 vs 29) but runs ~18x faster on Cerebras. It's also the first multimodal model on Cerebras's platform, meaning you can now feed it screenshots, documents, charts, and UI states at wafer scale speed. # Applications I'm most excited about: - Screenshot → Insight: Drop in a dashboard or document screenshot, get structured findings back instantly. no waiting, no batching. - Live UI generation: Full interactive interfaces (like my iPhone sim) generated and rendered in under 2 seconds. - Screenshot -> Patch: Feed it a broken UI + console error, get a minimal code fix and verification steps back. - Computer use & agentic loops: See -> reason -> act - verify, fast enough to keep a human in the loop instead of waiting on the model. - Long context summarization: Full research reports condensed into decision ready summaries you can read and requery in one sitting. The bigger unlock isn't the speed number itself, it's that agentic and multimodal loops (see -> reason -> output -> tool call -> verify -> retry) finally run in real time instead of feeling sluggish. As Logan Kilpatrick (Logan Kilpatrick) put it: "If every model was doing 2,000 tokens per second, you wouldn't build the same product and just have it be faster, you'd build different products." Gemma 4 31B is live now on Cerebras Inference Cloud in public preview. If you're building multimodal, agentic, or real time apps, this is worth testing today. What would you build with such insane inference throughput?

Alok

12,962 Aufrufe • vor 1 Monat

I just got Gemma 4 26B A4B MoE model running fully locally with Hermes agent on an 8GB RTX 4060 and it's now backtesting trading strategies end to end, no hand holding. If you’re a trader or work on Wall Street, you don’t want to miss this. Yes. fully automated. No cloud. No APIs beyond market data. # Here's what I did: Setup: - Model: Gemma 4 26B-A4B QAT (MoE), Q4_K_XL Unsloth's quant (link in the comments) - Inference: llama.cpp (turboquant fork by Tom Turney link in the comments) - Hardware: RTX 4060, 8GB VRAM + 16GB RAM only (with 50 other chrome tabs open) - Context: 64K llama.cpp turboquant flags: -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -c 64000 --cache-type-k q8_0 --cache-type-v turbo3 --port 8080 turboquant helps achieve high prefill and decode throughput for interactive sessions. throughput with Hermes agent: decode: 25+ tokens/sec prefill: 250+ tokens/sec # Then I gave the agent one task: Backtest a strategy: - Buy when RSI crosses above 30 - Sell at +2% profit or -1% stoploss - No overlapping positions - Use Google stock via yfinance - Generate a full HTML report with candlestick charts + signals What happened next was wild. It didn't just write code, it ran the entire workflow itself: Audited the environment (pip list, dependency check) Hit a ModuleNotFoundError, multiple Python installs were conflicting Ran where python to map every interpreter on the system Manually selected the correct Python 3.13 path and re ran the script Wrote a clean statevmachine backtester (strict no overlapping trades logic) Patched a yfinance MultiIndex quirk that would've crashed the script Built Plotly candlestick + RSI charts with buy/sell markers Calculated win rate, PnL, and summary stats Exported a polished single file HTML report. check the report at the end of the video or in the comments. Biggest takeaway: local LLMs aren't just "chat assistants" anymore. They debug their own environment, write production code, and ship a finished deliverable on consumer hardware, for $0 in API costs. If you're still calling local models "toys," you're already behind. This is just the beginning. Hermes agent just surpassed 1 trillion tokens in a single day on OpenRouter. Think about the scale of total token generation happening right now. Disclaimer: This is not financial advice. Consult a professional before making any trading decisions.

Alok

105,094 Aufrufe • vor 2 Monaten