Llama 4 seems to actually be a poor model... for coding. Scout (109B) and Maverick (402B) underperform 4o, Gemini Flash, Grok 3, DeepSeek V3 and Sonnet 3.5/7 on the Kscores benchmark which tests on coding tasks like this. ELO-maxxing on LMarena doesn’t create the best models.show more

Deedy
252,791 views • 1 year ago
Llama 3.1 Nemotron 70B is the latest model from... NVIDIA, released only a few hours ago. Initial testing shows the model outperforms GPT-4o and Sonnet 3.5 on several benchmarks. Try it on Akash Chat for free:show more

Akash Network
38,472 views • 1 year ago
Just tried Baidu’s new ERNIE X1.1 and wow… this... is actually impressive 🤯 I’ve been testing it on complex coding problems and math reasoning tasks, and the jump from earlier models is night and day. What stood out to me: ▪️ Much stronger at factual accuracy and instruction following ▪️ Handles agent-like tasks, tool use, and multi-step reasoning smoothly ▪️ Far fewer hallucinations compared to other reasoning-heavy models ▪️ Solid performance across creative writing, Q&A, coding, and math Benchmarks put ERNIE X1.1 ahead of DeepSeek R1-0528 and on par with GPT-5 & Gemini 2.5 Pro. If you’re into AI or reasoning models, it's definitely worth checking out: This doesn’t feel like just another release — it feels like a new standard for reasoning AI.show more

Parul Gautam
68,406 views • 10 months ago
LayerAI Works with DeepSeek to Supercharge the LayerAI Product... Suite 🌐 LayerAI integrates DeepSeek by leveraging several of its features and capabilities, focusing on areas where DeepSeek's strengths align with LayerAI's objectives. 👉 Learn more about our work with DeepSeek: - Model Deployment: DeepSeek-V3 & Coder-V2 for superior code generation, debugging, and language tasks. - DeepSeek-Coder-V2: Coding in 338 languages with a 128K context for complex structures. - Efficient Inference: Enhanced performance with DeepSeek's MLA & MoE tech. - IDE Assist: Real-time coding help, error catch, and auto-complete. - GitHub Sync: Better collab with seamless version control and reviews. This integration elevates our AI tools, making coding, learning, and teamwork smarter. More:show more

LayerAI | AI2Earn
47,893 views • 1 year ago
Salute to the Qwen team 🫡 We tested Qwen... 3.7-Max, Gemini 3.5 Flash, GPT-5.5, and Claude Opus 4.7. The biggest shock came from Qwen. In less than a month (3.6 Max dropped April 20), Qwen went from the worst multimodal output on our sakura tree test, barely keeping up with Gemini, GPT, and Claude , to matching Gemini 3.5 Flash frame for frame on this soccer test, and outperforming GPT-5.5 and Claude Opus 4.7. It rendered a perfectly proportioned soccer player and the most lifelike ball in the entire test. Remarkable spatial reasoning. Also: Gemini 3.5 Flash is now faster than GPT-5.5, which used to be the fastest in our past tests.show more

GMI Cloud
75,785 views • 2 months ago
you can run claude code inside antigravity completely Free... with zero credit card and no rate limits 😳 use openrouter’s free models + antigravity. no anthropic bill. no paid api keys. takes 10 minutes to set up. what you get during this setup: - full claude code agent experience - strong coding models (including deepseek-r1, qwen2.5-coder, llama-4, grok-4 free tier) - antigravity’s clean workspace and sandbox - unlimited usage (as long as you stay on free models) - easy model swapping - zero cost full setup guide (100% free): step 1: install antigravity -go to and install it -create a new workspace step 2: install claude code - inside antigravity, install the claude code extension from the marketplace - open the built-in terminal step 3: create openrouter free account -go to - sign up with google (no card needed) - go to keys and create a new api key step 4: set the environment variables -in antigravity terminal run: export ANTHROPIC_API_KEY=sk-or-xxx export OPENROUTER_API_KEY=sk-or-xxx step 5: launch claude code with free model -run this command: claude-code --model deepseek/deepseek-r1:free or try: qwen/qwen2.5-coder:free if you already have antigravity? skip straight to step 2. after 10 minutes you’ll have a full agentic coding setup running for free. this is currently one of the cheapest ways to run serious coding agents in 2026. bookmark this before they limit the free models.show more

painn
32,057 views • 2 months ago
1/ Gemini 2.5 is here, and it’s our most... intelligent AI model ever. Our first 2.5 model, Gemini 2.5 Pro Experimental is a state-of-the-art thinking model, leading in a wide range of benchmarks – with impressive improvements in enhanced reasoning and coding and now #1 on Arena by a significant margin. With a model this intelligent, we wanted to get it to people as quickly as possible. Find it on Google AI Studio and in the Google Gemini for Gemini Advanced users now – and in Vertex in the coming weeks. This is the start of a new era of thinking models – and we can’t wait to see where things go from here.show more

Sundar Pichai
864,455 views • 1 year ago
Grok 4.5 barely had time to enjoy first place.... Grok 4.5 Medium just took the top spot on LaurenBench with 56.9%, beating Claude Sonnet 5, Claude Opus 5, GLM 5.2 and GPT 5.6 on real world agent tasks. And Elon says Grok 4.6 arrives next week. Looks like Grok 4.5 won’t be in the spotlight for long. xAI Grok / Writer: Annette, Grok Imagine Designer: Jannéshow more

Mario Nawfal
39,492 views • 4 days ago
New Hunyuan Hy3 hits Gemini 3.5 quality on physics... for 35x cheaper! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: - A bowling ball knocking down the pins - An air hockey rally that ends in a goal - A pool break scattering the rack Outputs: Hunyuan Hy3: 29,757 tokens, $0.006 Gemini 3.5: 23,300 tokens, $0.21 GLM-5.2: 25,454 tokens, $0.07 DeepSeek-V4: 50,600 tokens, $0.009 Tencent's Hy3 matched Gemini across all three: clean collisions, the puck bounced true, the pins scattered like a real strike, the rack broke with real momentum, nothing clipped or floated. GLM is genuinely strong on pure coding tasks, but the moment the job steps outside clean code it gives way. DeepSeek was the letdown, it burned the most tokens of anyone (50k, almost 2x Hy3) and still turned in the weakest scenesshow more

atomic.chat
96,930 views • 29 days ago
There are so many models that it can get... confusing so today we launched an auto-router - you get the best answers from any of the top models. At the same time, we have AI experts and model enthusiasts who will enjoy access to 10 new models we just added including: ✅ GPT-4.1 ✅ o3 & o4 Mini ✅ Grok 3 ✅ Llama 4 Maverick & Scout We're also sunsetting some older models that just can't keep up anymore.show more

Richard Socher
24,731 views • 1 year ago
🚨Gemini 3.6 Flash is trash I tested it on... a 3D Golden Gate Bridge, and the results were awful. • I had to re-prompt it three times because it repeatedly ignored the instructions. • First attempt, instead of creating the requested .html file, it first tried to build the experience inside the Gemini app using simulations. • Then second attempt it started placing images from the web into the chat rather than actually producing the file. • Even after getting it to complete the task, the final output was dramatically worse than Gemini 3.1 Pro, which is 5 months old and now not even a top 10 model on leaderboards. This feels like a regression from Gemini 3.5 Flash and honestly, it is one of the weakest models I have tested in the past few months. Has anyone else tested Gemini 3.6 Flash yet, and are you seeing the same thing?show more

Lumina
72,113 views • 14 days ago
Moonshot AI is casually giving developers free daily access... to Kimi K3 😳 no subscription no upfront payment just sign in and start using one of the largest open AI models available what you get for $0: - Kimi K3 with 2.8T parameters - 1M token context window - strong coding and reasoning performance - native vision capabilities - free daily credits that refresh automatically why this is worth checking: > access a frontier model without paying API fees > long context for large codebases and documents > works on web, desktop, mobile, and CLI getting started takes less than 2 minutes: 1. go to 2. create a free account 3. Kimi K3 is available as the default model 4. start chatting or coding with your daily free credits bonus: Moonshot Together lets you invite friends for a chance to earn 3, 7, 15, 30, or even 365 days of Kimi Membership through its rewards program benchmark highlights: > 2.8T parameter MoE model > 1M context window > strong performance across coding, browsing, and reasoning benchmarks important: free credits reset daily, rate limits apply on the free tier, and the open-weight release is expected on July 27 A simple way to try one of the latest frontier AI models without paying for API accessshow more

K2S
22,890 views • 17 days ago
Elon says Grok Code is on track to outperform... Claude at coding - by June this year Progress is expected to bring it close to Claude by April, roughly equal by May, and potentially ahead once Colossus 2 is fully running "By then, like a self-driving car that drives perfectly, it will be hard to tell the difference between the leading coding models, as they will so rarely get anything wrong"show more

X Freeze
373,992 views • 5 months ago
Claude 3 Opus, Sonnet, and Haiku are now available... as a base for you to build prompt bots with on Poe! These powerful multimodal models are the best in the world for many use cases and support text and image input. We’re excited to see what you create with them!show more

Poe
17,040 views • 2 years ago
🚨 do you understand what is happening?! OpenAI and... Anthropic are putting the best models behind approval lists. Meanwhile GLM 5.2 is running locally on a Mac Studio, coding in a loop 24/7. If someone can turn off your AI, you are renting your intelligence layer. Local models are getting so good now that you can create gameplays like this.👇show more

Vadim
14,609 views • 1 month ago
Big win for open-source LLMs! DeepSeek V4 Pro holds... the top open-weights score on SWE-bench Verified, in the GPT-5.5 range. GLM 5.2 leads the open-weight intelligence index and sits near the closed frontier on long-horizon coding. But this leaderboard number is a weak proxy for real performance. It comes from one task set, run through one harness, served at one precision. The same weights can even score differently across providers, since many hosts quantize activations to fp8 and drift the model off its reference weights. Real performance is determined based on whether a model can read a repo, make coordinated edits across files, run the tests, and recover when one breaks. By that measure, the top open models hold up, but only inside the right harness. The teams that actually put DeepSeek V4 into production pipelines as a frontier substitute got there through the harness they built around the model, not by picking a stronger model. If you want to see this in practice, Cline (64k+ stars) has actually built that harness around open models, tuned so they run at production quality. And it's tuned so that these LLMs can run at production quality, with plan and act modes, checkpoints, and terminal feedback. ClinePass is the new access layer on top of it. It runs a curated set of those models inside Cline, narrowed to the ones tested for coding-agent use, with 2 to 5x the standard rate limits and no separate provider accounts, keys, or billing to track. The video below shows the setup, and I worked with the team to put this together. It runs alongside custom keys and local models as well, not in place of them.show more

Avi Chawla
44,124 views • 1 month ago
GROK’S GOT 1.15 TRILLION REASONS TO GLOAT #1 across... major leaderboards – a takeover powered by pure scale, speed, and developer love. * 1.15T tokens daily on OpenRouter – 29% of total traffic * Grok 4.1 tops LMSYS Elo, outthinking Claude 4.5 and GPT-4o * 70% share in programming tasks, wrecking Big Tech’s best * New CryptoBench leader in price calls, DeFi risks, and on-chain intel This is what frontier AI leadership looks like in real traffic and real usage! Source: Nextbigfutureshow more

Mario Nawfal
15,817 views • 7 months ago
subagents are just recursive agents where you can apply... different prompts + models depending on the task. since they’re just a primitive, Cursor cli can actually spawn subagents by calling cursor-agent in headless mode via shell commands. that’s what makes the cli so nice. you can extend it, experiment, and have a lot of fun exploring orchestration patterns. here’s one way to do it w. dynamic model selection: 1. create a subagents.mdc rule 2. drop in: ``` --- alwaysApply: true --- ALWAYS spawn subagents by running `cursor-agent -p [task] --output-format=text --force --model [model]` in the terminal. Each subagent should return a summary of the changes it made. Subagents should be used for ALL tasks You can adopt a fan-out pattern where you spawn subagents to perform parallel isolated tasks, and then fan-in the results. Use the following models: - `--model gpt-5` for reasoning, researching, and planning - `--model sonnet-4` for implementation ``` 3. start cursor cli and try it out you can also adjust the rule to be more explicit when it should use subagents, when not to, which models when etc.show more

eric zakariasson
57,554 views • 11 months ago
"The best organizations won't manage token budgets manually. Instead,... they'll rely on orchestration layers that automatically route workloads to the right models based on performance, cost, and use case. Keeping track of which model is best for coding, finance, research, or support will become too complex for humans to manage directly, so intelligent routing will become a core part of enterprise AI infrastructure." Aravind Srinivas How will the best orgs of the future manage token budgeting Michael Mignano Nikesh Arora Ryan Petersen mmurphshow more

Harry Stebbings
29,644 views • 1 month ago