Grok-1 support in gpt-fast at faster(?) than anyone else... has reported so far. 75 tok/s for a 300B+ parameter model on an 8xA100 node. If I understand correctly, ColossalAI reported 15 seconds to generate 100 tokens. gpt-fast takes 4.2 seconds to generate *400* tokens.show more

Horace He
54,137 просмотров • 2 лет назад
HappyHorse 1.0 is now on OpenArt 🐎 It's the... #1 ranked text-to-video model on Artificial Analysis right now. Generate up to 15 seconds 1080p video with synchronised audio in a single pass and support for 7 languages.show more

OpenArt
2,310,185 просмотров • 4 месяцев назад
China is literally on 🔥 Baidu from China has... launched ERNIE 4.5 and ERNIE X1 and it’s freaking cheap . Here is everything you need to know. ERNIE 4.5 - Native multimodal and Outperforms GPT 4.5 in multiple benchmarks at just 1% of GPT 4.5 price - OpenAI GPT 4.5 – Input: $75 / 1M tokens, Output: $150 / 1M tokens; - ERNIE 4.5 – Input: $0.55 / 1M tokens, Output: $2.20 / 1M tokens ERNIE X1 - A deep thinking reasoning model with multimodal capabilities on par with DeepSeek R1 at only half the price See it in action and check out the pricing details👇 📹 source : yiyan[.]baidu[.]com 1/6 ERNIE 4.5 is a multimodal which can take Audio files as well.show more

AshutoshShrivastava
324,324 просмотров • 1 год назад
Today we’re open-sourcing Stable Audio Open Small, a 341M-parameter... text-to-audio model optimized to run entirely on Arm CPUs. This means 99% of smartphones can now generate music-production samples in seconds, right on-device with no internet required. Built for fast, on-the-go creation, it turns your next quick idea into up to 11 seconds of audio. Generate drum loops, foley, riffs, and textures right where you are. No cords 🔌 just chords 🎹 You can learn more here:show more

Stability AI
94,796 просмотров • 1 год назад
GROK 4.1 FAST JUST DID THE IMPOSSIBLE AGAIN: 334... BILLION TOKENS IN A SINGLE DAY The OpenRouter leaderboard isn’t just broken. It’s been nuked from orbit. Yesterday: 253B tokens Today: 334B tokens That’s an 81B token jump in 24 hours. The closest competitor needs almost an entire week to hit what Grok 4.1 Fast does before lunch. This is the free model. The one anyone can use with zero paywall. People aren’t testing Grok anymore. They’ve moved in permanently. The throne isn’t up for grabs. It’s already been welded shut with Grok sitting on it. Bow down. The king just went super Saiyan. Source: X Freezeshow more

Mario Nawfal
42,979 просмотров • 9 месяцев назад
I topped up $5 on an API aggregator ToAPIs... Then I found out GPT Image 2 costs only around $0.015 per image. If you do a lot of testing or batch-generate commercial AI images, that difference adds up fast. I think I just found the secret to generating more, testing more, and spending less. And it’s not just one model. With the same key, you can access 50+ models for image, video, and text, including GPT Image 2, Gemini Omni, Seedance 2.0, Kling AI 3.0, grok-video-1.5-preview, and more. Some models are priced up to 80% lower than official platforms. Just top up and test what you need: Made on ToAPIs with GPT Image 2 + Seedance 2.0show more

Shami
22,991 просмотров • 2 месяцев назад
Liquid's LFM2.5-8B-A1B smashed OpenAI's gpt-oss-20b on tool calling We... ran both locally on a MacBook Pro M5 Max, 64GB, and gave each the same trip-planning request that only completes if the model fires all 7 tool calls - weather for 3 cities, two currency conversions, an email and a reminder Outputs: LFM2.5-8B-A1B: 4.8 GB RAM usage, 7/7 tool-calls, 266 tok/s, 6.9s OpenAI gpt-oss-20b: 11 GB RAM usage, 3/7 tool-calls, 146 tok/s, 15.0s The 8B used less than half the RAM and still fired all 7 calls, while the 20B silently dropped more than half of its own. It also ran ~2x faster, wrapping the full agentic request in 6.9s against 15s. That's what 38T training tokens buy: a 1B-active MoE that nails the agentic tool calls a model 2.5x its active size keeps droppingshow more

atomic.chat
90,063 просмотров • 2 месяцев назад
(1/n) 🚀 With FastVideo, you can now generate a... 5-second video in 5 seconds on a single H200 GPU! Introducing FastWan series, a family of fast video generation models trained via a new recipe we term as “sparse distillation”, to speed up video denoising time by 70X! 🖥️ Live demo: (Thanks to @gmicloud for the support!) 🔗 Blog: 🔓 We fully open-source our models, code, and data with Apache-2.0 licensesshow more

Hao AI Lab
78,660 просмотров • 1 год назад
Holy moly: GLM-5.3 got much better in cybersecurity since... our pre-release evaluation with Z.ai. It now matches GPT-5.6-Sol on our cybersecurity benchmark at 0.4x the cost 🤯 - At pass@1: it went from 60.4% to 65.6% CVEs rediscovered, crushing every other open model on one-shot tasks - At pass@3: it did 75% -> 78.1%, matching GPT-5.6-Sol - Its precision remained stable, reporting fewer false positives than DeepSeek models The performance increase comes from a behavioral change: the new version is more persistent. It tends to run longer, and had a ~43% reasoning tokens increase. But the performance upgrade is worth that additional cost. 1/3 🧵show more

pilvar (Philippe Dourassov)
29,276 просмотров • 6 дней назад
I built an ios app that lets you setup... an open claw agent from your phone in less than 30 seconds. My mom has an open claw assistant now. No api keys, no vps, no setup, no mac mini. You just log in and its spins up your own ai agent. Most people don't understand how open claw is different from chat gpt. I am building one tap skills so that anyone can outsource their work and organize their life. People spend $2500 a month on tokens. I am adding cost efficient models with one tap switching so you don't waste your money. Most of these simple setup claw projects are completely unsafe, I have passed apple review and am working with world class security professionals to keep your data safe.show more

Max Blade
26,095 просмотров • 6 месяцев назад
✨ I revived my first AI startup from 6... years ago with Claude Code [ 💡 ] Back then it used GPT-3 (this was 2 years before ChatGPT existed!) to generate new startup ideas which then people can vote on And the best startup ideas rise to the top! Back then I made it because people complained they didn't have any ideas to build a startup This week I moved it to its own VPS and installed Claude Code and told it to fix everything, the DB had become big and there was stupid write operations on every page load that it made it very slow Claude Code is excellent at fixing all those small bugs from old projects and quickly fixing them As Garry Tan says "boil the oceans" as in before I'd not have the time to fix these kinds of projects, it wouldn't be worth it, I mean IdeasAI doesn't even make money, but now it takes me an hour to do this and it works again! I also upgraded GPT-3 to xAI's Grok 4.2 for new startup ideasshow more

@levelsio
180,039 просмотров • 5 месяцев назад
GPT-5.6 Sol is unbelievably good at creating and editing... videos. It can do motion design, product demos, and animations like this one I made by simply giving it a screen recording. GPT 5.6 has the best design taste and significantly outperforms Fable, which relies heavily on repetitive design patterns. To help you experiment with video editing on it, we just launched a collection of 100 ready-to-use skills that show what’s possible and help you get started with video editing using GPT-5.6. These skills can create anything from motion graphics launch videos for your product to a 3B1B-style science explainer video. You can also use them to edit existing videos: add captions, generate motion graphics, create voiceovers, redesign visual styles, translate into new languages, and much more. If you want access to the full library, comment “VIDEO SKILLS” and I’ll share it with you. (You'll have to follow me so I can DM you.)show more

Akash Anand
515,129 просмотров • 1 месяц назад
A peanut-sized Chinese model just dethroned Gemini at reading... documents. GLM-OCR is a 0.9B parameter vision-language model. It scores 94.62 on OmniDocBench V1.5, ranking #1 overall. For context, it outperforms models 100x its size. 100% open-source. It works in two stages. 1. A layout engine detects every region in a document. 2. Each region gets read in parallel. The model predicts multiple tokens per step instead of one. That's what makes it so fast at small size. It handles things most OCR tools struggle with: > Complex tables and nested layouts > Handwritten text and stamps > Math formulas and code blocks > Mixed image-and-text documents You can run it locally through Ollama. It fits on edge devices with limited compute. Every expensive OCR API just got a free competitor.show more

AlphaSignal
92,071 просмотров • 4 месяцев назад
A peanut-sized Chinese model just dethroned Gemini at reading... documents. GLM-OCR is a 0.9B parameter vision-language model. It scores 94.62 on OmniDocBench V1.5, ranking #1 overall. For context, it outperforms models 100x its size. 100% open-source. It works in two stages. 1. A layout engine detects every region in a document. 2. Each region gets read in parallel. The model predicts multiple tokens per step instead of one. That's what makes it so fast at small size. It handles things most OCR tools struggle with: > Complex tables and nested layouts > Handwritten text and stamps > Math formulas and code blocks > Mixed image-and-text documents You can run it locally through Ollama. It fits on edge devices with limited compute. Every expensive OCR API just got a free competitor.show more

Jafar Najafov
13,630 просмотров • 4 месяцев назад
Perplexity Computer in 60 seconds: 1. It's a cloud-based... AI employee that runs tasks in the background. 2. 19 models working together. Claude for reasoning, GPT-5.2 for research, Grok for speed tasks. You don't pick. It routes automatically. 3. 400+ connectors. Gmail, Slack, Notion, Salesforce, HubSpot. One click to enable each. 4. Credits, not tokens. Simple tasks cost ~30. Complex builds cost 1,000+. Vague prompts waste them. Specific prompts save them. 5. Spaces = persistent project folders. Upload context once, every task inherits it. 6. Scheduled tasks run on autopilot. "Every Monday, prep my calendar." Set it and forget it. The PRD hack alone (in the article) will save you hundreds in credits. Full breakdown in the article below.show more

Corey Ganim
106,105 просмотров • 4 месяцев назад
I just ran Gemma 4 31B on @CerebrasSystems at... 1,800+ tokens/sec and it's multimodal. For context: that's 35x faster than a typical GPU endpoint, and the first token (reasoning included) lands in 1.5 seconds. This isn't a benchmark slide, I recorded the inference live. Prompt I used: "Create a simulation of an iPhone. Include at least one working dummy note taking app, a functional notification pulldown, high quality graphics, single HTML file, any libs via CDN." - Generation time: 3 seconds. - Notes app worked. - Notification panel worked. - Rendered first try. This is what wafer-scale inference unlocks, not just "faster," but a different category of product. When generation is this fast, you stop waiting and start iterating in real time. Why this matters: Gemma 4 31B is Google DeepMind's flagship open weight model, Apache 2.0 licensed, dense (not MoE), and built for efficiency over raw parameter count. It scores close to Claude Haiku 4.5 on the Artificial Analysis Intelligence Index (30 vs 29) but runs ~18x faster on Cerebras. It's also the first multimodal model on Cerebras's platform, meaning you can now feed it screenshots, documents, charts, and UI states at wafer scale speed. # Applications I'm most excited about: - Screenshot → Insight: Drop in a dashboard or document screenshot, get structured findings back instantly. no waiting, no batching. - Live UI generation: Full interactive interfaces (like my iPhone sim) generated and rendered in under 2 seconds. - Screenshot -> Patch: Feed it a broken UI + console error, get a minimal code fix and verification steps back. - Computer use & agentic loops: See -> reason -> act - verify, fast enough to keep a human in the loop instead of waiting on the model. - Long context summarization: Full research reports condensed into decision ready summaries you can read and requery in one sitting. The bigger unlock isn't the speed number itself, it's that agentic and multimodal loops (see -> reason -> output -> tool call -> verify -> retry) finally run in real time instead of feeling sluggish. As Logan Kilpatrick (Logan Kilpatrick) put it: "If every model was doing 2,000 tokens per second, you wouldn't build the same product and just have it be faster, you'd build different products." Gemma 4 31B is live now on Cerebras Inference Cloud in public preview. If you're building multimodal, agentic, or real time apps, this is worth testing today. What would you build with such insane inference throughput?show more

Alok
12,962 просмотров • 1 месяц назад
Suppose two cars are traveling on an expressway at... 100 km/h with a 50-meter gap. If the front car slams on its brakes, it takes about 3.5 seconds and nearly 50 meters to stop. If the driver behind does nothing, a collision occurs in exactly 3.5 seconds. This gives the rear driver a 3.5-second window to avoid a crash. For an alert driver, 1 second is reaction time, leaving 2.5 seconds to brake and stop safely. Now, consider the same 50-meter gap, but the front car is reversing at 10 km/h. For the car approaching from behind at 100 km/h, the closing speed skyrockets to 110 km/h, slashing the time before collision to just 1.6 seconds. In reality, it is even worse: because drivers never expect a vehicle to reverse on an expressway, their brains freeze. Reaction time jumps from 1 second to 2, 3, or more. At high speeds, drivers often crash before even realizing they need to brake. That's why reversing on an expressway is hundred times more dangerous than a sudden braking. A maneuver that feels slow to the reversing driver creates an unexpectedly violent closing speed for those behind, leaving virtually no time to react. In the video, a family of seven missed their exit on the Dehradun expressway to Haridwar. They reversed, were struck from behind, and four of them died.show more

THE SKIN DOCTOR
547,082 просмотров • 1 месяц назад
HERMES AGENT NOW RUNS CLAUDE OPUS 5. NEAR FABLE... 5 INTELLIGENCE. HALF THE PRICE. SELF-VERIFIES ITS OWN WORK. AVAILABLE TODAY VIA NOUS PORTAL (20% OFF ALL MODELS). Anthropic shipped Opus 5 on July 24, 2026. same $5/$25 per million tokens as Opus 4.8. but the benchmarks tell a different story. WHAT CHANGED FROM OPUS 4.8: FrontierBench v0.1: Opus 5: 43.3%. Opus 4.8: 18.7%. 2.3x jump on the same test. ARC-AGI-3: Opus 5: 30.2%. 3x better than the next closest model. beat Fable 5 on 8 out of 13 benchmarks. at half the cost ($5/$25 vs $10/$50). same price as Opus 4.8. twice the intelligence. no reason to stay on 4.8. THE SPECS: model ID: claude-opus-5 context: 1M tokens (default and maximum) max output: 128K tokens thinking: on by default effort toggle: low / medium / high per request fast mode: $10/$50, 2.5x faster knowledge cutoff: May 2026 minimum cacheable prompt: 512 tokens (was 1,024) SELF-VERIFICATION (the biggest change): Opus 5 checks its own work automatically. Anthropic says: delete your verification prompts. "include a final verification step" now causes OVER-verification because the model already does it. for Hermes /goal tasks this is a direct upgrade. the judge checks evidence. the model also checks evidence. double layer of verification without extra tokens. EFFORT TOGGLE: low: fast, cheap, routine work. medium: balanced, daily tasks. high: full reasoning, complex problems. set per request. not a global switch. matches Hermes /reasoning command: /reasoning low (routine) /reasoning high (complex) Opus 5 effort toggle + Hermes reasoning control = precise cost management per turn. WHERE OPUS 5 FITS IN HERMES: DAILY DRIVER (replaces Opus 4.8): same price. 2.3x better benchmarks. set as your main model: Desktop app / Dashboard: Models → claude-opus-5 CHIEF OF STAFF: synthesis across multiple agents. reads Kanban, prioritizes, routes tasks. self-verification catches routing errors before they cascade. COMPLEX CODING: SOTA on agentic coding benchmarks. FrontierBench 43.3% = best public model for coding. set as coder profile model. /GOAL TASKS: self-verification + completion contracts = the model proves its work AND double-checks the proof. long-horizon goals finish correctly more often. MoA AGGREGATOR: strongest synthesis model at $5/$25. pair with GPT-5.6 and Grok 4.5 as references. Opus 5 aggregates. best quality at mid-range price. presets: max-quality: reference_models: - provider: openai-codex model: gpt-5.6-sol - provider: xai model: grok-4.5 aggregator: provider: anthropic model: claude-opus-5 COMPUTER USE: near-Fable 5 quality for browser automation. at half the token cost per session. computer_use tasks burn lots of vision tokens. Opus 5 halves that bill vs Fable 5. WHAT TO KEEP OPUS 5 AWAY FROM: cron monitoring: too expensive. use DeepSeek or no_agent mode. sub-agent grunt work: use GPT-5.6 Luna ($1/$6) or DeepSeek. auxiliary tasks: use Gemini Flash. routine web extraction: use a cheap model. Opus 5 is for the turns where quality compounds. planning, synthesis, verification, complex reasoning. budget models handle everything else. NOUS PORTAL: 20% OFF ALL MODELS Nous Portal currently runs a 20% discount on all models including Opus 5. $5/$25 official → $4/$20 through Nous Portal. the cheapest way to run Opus 5 right now. hermes setup --portal select claude-opus-5 as your model. discount applies automatically. Opus 5 replaces Opus 4.8 everywhere. same price. better at everything. no tradeoff. straight upgrade. hermes update /model claude-opus-5show more

YanXbt
16,744 просмотров • 1 месяц назад
Seedance 2.5 is now on CapCut. We’ve been testing... Seedance 2.5 on CapCut, and what impressed us most is how much more control it gives creators. Being able to generate and edit everything in one place makes the workflow feel much faster and more natural. The timestamp-based controls, support for up to 50 references, and video generation of up to 90 seconds are especially useful. We also noticed better multilingual performance, plus viewport rendering and green screen options for more advanced workflows. Definitely a strong update for anyone creating AI video content. If you found this useful, leave a like and retweet it so more creators can discover the update! 🔁 Web: App: #CapCut #Seedance25 #CapCutAI #CapCutDidThatshow more

DothAI
75,937 просмотров • 27 дней назад
Here is my first Seedance 2.0 generation. Also, everyone... wants to have fun so stop gatekeeping this stuff, just go here and generate. Login with a google account. If you solve a puzzle, you're logged in because even if it asks you to verify a number just refresh the page with the link below and you'll be logged in. So far I've been able to generate 15 seconds completely for free. This thing is no joke, it's actually SUPER good, I had to try it myself and not just believe what was posted on X. This is the king of AI epic fight scenes now, I don't think anything else even comes close. Sound is also EXCELLENT! China is really cooking! Holy shit! Prompt: An intense fight scene between a masked ronin with a huge sword and a massive creature during a violent thunder storm. The earth is shattering as the ronin fights the colossal monster, slicing it's chest and eventually defeats it. The scene is chaotic with handheld motion and camera shake.show more

Travis Davids
79,984 просмотров • 6 месяцев назад
a moonshot engineer leaked the benchmark anthropic, openai and... xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article belowshow more

starmex
30,987 просмотров • 5 дней назад