
Cerebras
@cerebras • 73,912 subscribers
The world's fastest AI inference and training. Try the latest open models at: https://t.co/jREGhLI2nj
Shorts
Videos

Previewing Ultrafast mode for OpenAI's GPT 5.6 Sol, powered by Cerebras. GPT-5.6 Sol Ultrafast generates responses at up to 750 tokens per second. That's the full, GPT-5.6-Sol model -- up to 14× faster than the same model on Standard processing. It speedran Humanity’s Last Exam in 11h 11m, nearly 7× faster than Claude Fable 5 with comparable accuracy.
Cerebras1,284,390 Aufrufe • vor 20 Tagen

Ultrafast mode for GPT-5.6 Sol is now in limited preview, powered by Cerebras. We gave OpenAI's GPT-5.6 Sol the same prompt on Ultrafast and Standard: build a financial terminal-style dashboard for analysts. Ultrafast: 1 min 50 seconds Standard: 12 min 20 seconds Same result, nearly 7x faster.
Cerebras144,511 Aufrufe • vor 20 Tagen

Right now, when you send a query to an LLM, it gets decrypted on the server. The LLM sees your data in plain text. Prof. Ajay Joshi (BU, CipherSonic AI ) on fully homomorphic encryption, which may be key for the future of AI privacy: how we can compute on data without ever decrypting it. The catch: it's a brutally memory-bound workload. Exactly the bottleneck wafer-scale was built to solve.
Cerebras293,561 Aufrufe • vor 2 Monaten

GLM 4.7 is one of the strongest open-source coding models available—but most developers aren't prompting it correctly. We put together 10 rules to help you get the most out of it: - Front-load instructions (it has a strong recency bias) - Use firm language: "must" and "strictly" > soft suggestions - Break complex tasks into smaller steps - Disable reasoning for simple tasks, enable it for hard ones - Use critic agents for code review, QA, and validation - Pair it with a frontier model for the hardest 10% of workloads - and more… GLM 4.7 hits 96% on Tau² Bench and 86% on GPQA Diamond. At 1,500 tokens/sec on Cerebras, it's 20x faster than closed-source alternatives on GPUs.
Cerebras633,658 Aufrufe • vor 7 Monaten

OpenAI Codex-Spark powered by Cerebras You can now just build things faster—at 1,000 tokens/s.
Cerebras287,784 Aufrufe • vor 6 Monaten

Introducing the Cerebras Ambassador Program 🌏 We're looking for AI community builders to build your local ecosystem around the world's fastest inference. In the last few months, our community has hosted Cafe Compute Meetups in Berlin, Mexico City, Mumbai, SF, Seattle, and NYC. We want your city to be next! Application link in comments.
Cerebras102,053 Aufrufe • vor 2 Monaten

Some of our top customers are still choosing Llama 3.1 8B. For a while, we jumped to whatever hottest, latest model was taking up our twitter feed. 🙈 But as we are quickly realizing, to create a SOTA product, you need a model that fits your exact use case. Here’s what our customers tell us: > a lot of the legwork is actually around prompting > there’s an art to selecting and combining multiple models > benchmarks only show part of the picture. you have to understand the unique quirks of each model. Especially as model releases become more and more frequent, we need a clear way to evaluate new models. We have to break free of the naive trend to migrate to the ‘latest and greatest’. And you can easily achieve this using tools like Cerebras and Braintrust to swap models safely (without breaking production).
Cerebras346,446 Aufrufe • vor 8 Monaten

Cerebras Code: 20x faster than Claude, 1x the price Today we are launching two monthly coding plans: ➡️Cerebras Code Pro: $50/m – for indie developers ➡️Cerebras Code Max: $200/m – for power users with 5x rate limits Both plans get: Qwen3-Coder at 2,000 tokens/s, 131K context, and no weekly limits. Sign up now:
Cerebras461,380 Aufrufe • vor 1 Jahr

Let's talk about MoE: 🔶 How many experts should you use? 🔶 How does dynamic routing actually behave in production? 🔶 How do you debug a model that won’t train? 🔶 What does 8x7B actually mean for memory and compute? 🔶 What hardware optimizations matter for sparse models? Mixture of Experts (MoE) is changing how the biggest AI models are built — but it’s still hard to get right. That's why we are launching a new MoE 101 series, led by Daria Soboleva to bridge the gap between theory and practice. Dive in to our MoE guide:
Cerebras345,732 Aufrufe • vor 1 Jahr

🎁 We're giving away 5 Windsurf plans ($250 credit each)! Try SWE-1.6 — Cognition’s latest fast and intelligent agentic coding model, powered by Cerebras. In a side-by-side with Claude, the speed difference is clear. More iterations, faster fixes, better code. 💬Comment why you want access to enter. Five winners will be selected at random within 48 hours.
Cerebras105,709 Aufrufe • vor 3 Monaten