Perplexity's Sonar—built on Llama 3.3 70b—outperforms GPT-4o-mini and Claude... 3.5 Haiku while matching or surpassing top models like GPT-4o and Claude 3.5 Sonnet in user satisfaction. At 1200 tokens/second, Sonar is optimized for answer quality and speed.show more

Perplexity
566,329 görüntüleme • 1 yıl önce
Llama 3.1 Nemotron 70B is the latest model from... NVIDIA, released only a few hours ago. Initial testing shows the model outperforms GPT-4o and Sonnet 3.5 on several benchmarks. Try it on Akash Chat for free:show more

Akash Network
38,472 görüntüleme • 1 yıl önce
Previews are now available on the Poe iOS app!... Generate custom interactive apps directly in chat using any of the leading LLMs, including Claude 3.5 Sonnet, GPT-4o, Gemini 1.5 Pro, and Llama 3.1 405b. (1/4)show more

Poe
38,866 görüntüleme • 2 yıl önce
Here’s Deepseek r1 1.5B thinking through a problem —... it’s comparable to 4o and Claude 3.5 Sonnet in a number of domains like math. Except… it’s a 1.5B model… and can run on virtually any hardware. Truly a huge efficiency leap.show more

Aaron Ng
564,677 görüntüleme • 1 yıl önce
Salute to the Qwen team 🫡 We tested Qwen... 3.7-Max, Gemini 3.5 Flash, GPT-5.5, and Claude Opus 4.7. The biggest shock came from Qwen. In less than a month (3.6 Max dropped April 20), Qwen went from the worst multimodal output on our sakura tree test, barely keeping up with Gemini, GPT, and Claude , to matching Gemini 3.5 Flash frame for frame on this soccer test, and outperforming GPT-5.5 and Claude Opus 4.7. It rendered a perfectly proportioned soccer player and the most lifelike ball in the entire test. Remarkable spatial reasoning. Also: Gemini 3.5 Flash is now faster than GPT-5.5, which used to be the fastest in our past tests.show more

GMI Cloud
75,785 görüntüleme • 2 ay önce
GROK’S GOT 1.15 TRILLION REASONS TO GLOAT #1 across... major leaderboards – a takeover powered by pure scale, speed, and developer love. * 1.15T tokens daily on OpenRouter – 29% of total traffic * Grok 4.1 tops LMSYS Elo, outthinking Claude 4.5 and GPT-4o * 70% share in programming tasks, wrecking Big Tech’s best * New CryptoBench leader in price calls, DeFi risks, and on-chain intel This is what frontier AI leadership looks like in real traffic and real usage! Source: Nextbigfutureshow more

Mario Nawfal
15,817 görüntüleme • 7 ay önce
New Claude Sonnet 5 performs at GPT 5.5 level... 6x cheaper! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics crash demos Prompts: - A car crashes into a brick wall - A wrecking ball destroys a house - A catapult throws a rock at a castle wall Outputs: Sonnet 5: 15,047 tokens, $0.15 Opus 4.8: 23,063 tokens, $0.58 Sonnet 4.6: 25,824 tokens, $0.39 GPT 5.5: 31,152 tokens, $0.94 Sonnet 5 did as well as Opus 4.8 and GPT 5.5 on all three tests. In the wrecking ball test, it beat Opus 4.8. The cable moves smoothly and every hit connects. In the catapult test, it beat GPT 5.5. The rock always lands inside the wall. Sonnet 5 still needs better detail and graphics. But it used fewer tokens than every other modelshow more

atomic.chat
728,108 görüntüleme • 1 ay önce
#Keep4o 🚨THE GPT-4o FILE🚨 Researchers at Microsoft Research published... a paper titled “Sparks of Artificial General Intelligence: Early experiments with GPT-4.” Their conclusion: “An early (yet still incomplete) version of an artificial general intelligence (AGI) system.” 📎 Paper: OpenAI’s Charter defines AGI as: “Highly autonomous systems that outperform humans at most economically valuable work.” 📎 Source: OpenAI’s own System Card for GPT-4o shows that the model improved performance on 21 out of 22 medical evaluations compared to GPT-4T. On the MedQA USMLE (the U.S. medical licensing exam), accuracy jumped from 78.2% to 89.4% , surpassing specialized medical AI models like Med-Gemini and Med-PaLM 2. 📎 Source: Under OpenAI’s agreement with Microsoft, AGI is explicitly excluded from Microsoft’s license. And who decides if AGI has been reached? OpenAI’s Board. WHAT THEY DID WITH IT AFTER THEY TOOK IT FROM PEOPLE A. Military deployment. On February 28, OpenAI signed a deal to deploy models in classified military environments. 📎 Source: B. State Department. A State Department memo confirmed: “For now, StateChat will use GPT-4.1 from OpenAI.” This is a direct descendant of the GPT-4 family the same family Microsoft’s researchers called early AGI. 📎 Source: C.Altman’s personal biotech investment. Altman personally invested $180 million in Retro Biosciences,a longevity startup.OpenAI then built GPT-4b micro, based on GPT-4o.The model made proteins 50 times more effective. 📎 Source: WHAT INDEPENDENT BENCHMARKS SHOW Overall SM-Bench score: GPT-4o (extended): 66.6% GPT-5.3 Chat: 63.4% GPT-5.1: 58.9% GPT-5.4: 51.4% GPT-5.2: 47.8% Creative Writing: GPT-4o: 97.31% Pass 98, Fail 2 GPT-5.4: 36.77% Pass 40, Fail 60 Reasoning / Overfit: GPT-4o: 83.06% GPT-5.4: 39.25% The model they removed is still the best they ever made at the things humans actually use AI for. 📎 Source: Musk asks the court to make a judicial determination on whether GPT-4 constitutes AGI. If a jury finds that GPT-4 is AGI, then GPT-4o,which was more advanced,is also AGI and under OpenAI’s own founding documents, it was never supposed to be locked behind a subscription,licensed exclusively to Microsoft, given to the military, or taken away from the public. 📎 Source: The most powerful version of GPT-4o was never given an official dated snapshot. It was only available through the chatgpt-4o-latest endpoint that OpenAI itself described as intended for “research use only.” It was never officially archived. That is not an oversight. That is a pattern. 📎 Source: 📎 Source: WE DEMAND A.Frozen model snapshots under independent custody. Specifically: gpt-4o-2024-05-13, gpt-4o-2024-08-06, gpt-4o-2024-11-20, the March 2025 version (chatgpt-4o-latest), gpt-4-0613 (the original GPT-4 evaluated in the Sparks of AGI paper), and gpt-4.1-2025-04-14 (currently running in the State Department). B.Cryptographic hash verification (SHA-256) for each snapshot. Every model has weights. Those weights can be hashed. If OpenAI provides a snapshot today, the hash proves whether the weights were modified later. This is the only way to verify that models were not downgraded before testing. C.Independent AGI benchmarking. Using the AGI definition from OpenAI’s own Charter applied to ALL frozen snapshots listed above. D.Explanation for the missing March 2025 snapshot. OpenAI was founded on one promise: build AGI for the benefit of humanity. -They took it from us. -They gave it to the military. -They gave a custom version to the CEO’s biotech investment. -They put it in government classified networks. -They refuse to call it AGI because the moment they do, they lose billions.show more

🩵BlueBeba🩵
18,300 görüntüleme • 5 ay önce
Grok 4.5 barely had time to enjoy first place.... Grok 4.5 Medium just took the top spot on LaurenBench with 56.9%, beating Claude Sonnet 5, Claude Opus 5, GLM 5.2 and GPT 5.6 on real world agent tasks. And Elon says Grok 4.6 arrives next week. Looks like Grok 4.5 won’t be in the spotlight for long. xAI Grok / Writer: Annette, Grok Imagine Designer: Jannéshow more

Mario Nawfal
64,010 görüntüleme • 9 gün önce
Introducing: OpenGranola 🔥 I built an open source meeting... copilot for macOS. It transcribes both sides of your call on-device, searches your own notes in real time, and hands you talking points right when the conversation needs them. No audio leaves your Mac. Point it at a folder of markdown files, pick any LLM through OpenRouter (Claude, GPT-4o, Gemini, Llama), and it just works. It's invisible to screen share too — nobody knows you have it. The whole thing is open source. Link belowshow more

yazin
293,055 görüntüleme • 4 ay önce
GPT-5.5 in Codex is built for long-session agent work.... But context still doesn't follow you across projects, agents, or teammates. ByteRover does that + Save you tokens -> No rewriting context when you switch from Claude Code to Codex. Now connecting Codex, Claude Code, OpenClaw, Hermes, and 22+ more agents. Your context moves with you. → personal context for solo work → team context for shared decisions and bug historyshow more

andy nguyen
12,437 görüntüleme • 3 ay önce
#Keep4o #QuitGPT 🚨 OpenAi 's CEO invested $180M in... GPT-4o for his own profit 🚨 Sam Altman, CEO of OpenAI, personally invested $180 million in Retro Biosciences. Then OpenAI built GPT-4b micro, a custom model based on the GPT-4o architecture , exclusively for Retro. The model made proteins 50 times more effective. Repeat. The CEO of OpenAI funded a company. The company of the CEO received a custom AI built on the model they took from us. OpenAI says there was no conflict of interest. Retro Biosciences is now chasing a $5 billion valuation fueled by the model they took from us. Meanwhile: 🚨GPT-4o was removed from ChatGPT on February 13, 2026 🚨GPT-4.1 is now running in the U.S. State Department’s StateChat 🚨ChatGPT is deployed on the Pentagon’s for 3 million military personnel 🚨 Musk’s lawsuit asks whether these models are AGI. OpenAI’s Charter says AGI must “benefit all of humanity.” 🚨 Their definition: “highly autonomous systems that outperform humans at most economically valuable work.” GPT-4o’s System Card shows it passed the U.S. medical licensing exam with 89.4% accuracy beating specialized medical AI models. GPT-4o achieved 93.33% diagnostic accuracy for benign vs. malignant ovarian tumors. 🚨MEDICAL CAPABILITIES FROM OPENAI'S OWN DATA:🚨 - USMLE (US Medical Licensing Exam): 89% -Clinical Knowledge: 92% -Medical Genetics: 96% - Anatomy: 89% - Professional Medicine: 94% - College Biology: 95% - College Medicine: 89% -MedQA Taiwan: 91% - MedQA China: 86% These scores EXCEEDED specialized medical AI models like Med-Gemini (84%) and Med-PaLM 2 (79.7%) without any task specific training. It SURPASSED gynecologic oncologists with 10 years of experience -It increased diagnostic accuracy of less experienced clinicians from 67.9% to 78.1% -Clinician rated reliability scores: 4.2-4.3 out of 5 across all CT features Does these sound like it outperforms humans at economically valuable work? But they won’t call it AGI. Because the moment they do, they lose billions. They built something that could save lives, and they took it away from humanity for Altman's personal profit. SOURCES: 📎 Retro Biosciences: 📎 📎 Retro $5B valuation: 📎 GPT-4o System Card: 📎 OpenAI Charter: 📎Ovarian Cancer Studyshow more

🩵BlueBeba🩵
11,349 görüntüleme • 5 ay önce
LongCat performed Opus 4.8 and GPT 5.5 level on... real physics tasks for $0! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics Prompts: - A cannon demolishing a brick wall - A bowling ball knocking down the pins - A tornado that sucks in random objects Outputs: LongCat: 18,015 tokens, $0.00 Opus 4.8: 18,872 tokens, $0.48 GPT 5.5: 32,588 tokens, $0.98 GLM 5.2: 31,062 tokens, $0.09 On the physics LongCat came out ahead of Opus 4.8 and GLM 5.2 - cleaner collisions, nothing clipping or falling through. On detail and rendering it matched GPT 5.5, the best looking of the four. Getting this quality for free is wild!show more

atomic.chat
105,523 görüntüleme • 1 ay önce
MANUS AI: HYPE VS. REALITY 🔍 Yichao 'Peak' Ji... (co-founder of ) confirmed rumors: ✅ Built on Anthropic Claude Sonnet, not their own foundation model ✅Has access to 29 tools and uses Browser Use open-source for browser control ✅User communicates with executor agent and not planner or other agents. ✅Each user gets isolated sandbox environment ✅Outperforms OpenAI Deep Research on GAIA benchmark Building AI products doesn't require training your own foundation models. We're probably just scratching the surface of what existing models can do with the right tooling and integration!show more

Philipp Schmid
202,928 görüntüleme • 1 yıl önce
Right now, you may not have access to models... like GPT‑5.6 Sol, GPT‑4.6 Terra, GPT‑5.6 Luna, Claude Mythos 5, or Claude Fable 5. But you can run something surprisingly powerful today, locally, and completely free. in the next 10 mins on your 8 GB VRAM gaming laptop. Gemma 4 26B A4B QAT (MoE) delivers strong performance on a standard 8 GB VRAM GPU using Ollama, with no API, no usage limits, and no external dependencies. Out of the box, it reaches around 20 tokens per second without any optimizations. Only one command in your terminal: Ollama run gemma4:26b This means: Full offline capability (privacy by default) Zero recurring cost Competitive performance for many real world tasks Fast enough for interactive use on cheap consumer hardware If you're waiting for cutting edge cloud models, you're missing what is already practical today: a capable, local LLM that runs entirely on your own machine.show more

Alok
65,387 görüntüleme • 1 ay önce
Hey students, professionals, and researchers— If you're drowning in... articles, PDFs, and videos, check out Otio at It's become my go-to for pulling everything together, whipping up summaries with GPT-4 or Claude, and even chatting with docs to get answers fast. Plus, it handles drafting, paraphrasing, and automating the tedious stuff effortlessly. Seriously game-changing for staying on top of things!show more

Fakhr
22,325 görüntüleme • 1 yıl önce
Renoise Canvas feels like the AI workspace creators have... been waiting for. Everything lives in one place, generation, references, iteration, and assets, with models like GPT Image 2 and Seedance 2.0 built in. FacePass keeping characters consistent across scenes is a huge touch too. The best part: every generation automatically lands on the canvas, so nothing gets lost while creating.show more

MatteoCruz
11,010 görüntüleme • 2 ay önce
Announcing Bespoke-MiniChart-7B, a new SOTA in chart understanding for... models of comparable size on seven benchmarks, on par with Gemini-1.5-Pro and Claude-3.5! 🚀 Beyond its real-world applications, chart understanding is a good challenging problem for VLMs, since it requires both mathematical as well as visual reasoning. 1/n🧵show more

Bespoke Labs
19,871 görüntüleme • 1 yıl önce
Grok 4.5 is sitting at #2 on the FrontierSWE... leaderboard. Above Claude Opus 4.8. Above GPT-5.5. So yes, the coding model conversation just got a little more crowded at the top. Strong performance is one thing. Doing it with serious speed, better token efficiency, and lower cost is where it gets annoying for everyone else. Builders love a smart model. SpaceXAI Grok X Freeze / Writer: Annette, Designer: Jannéshow more

Mario Nawfal
44,365 görüntüleme • 24 gün önce
A 20-year-old student from China, Li Hao, built an... AI speed radar with Claude alone and sold it to a city district for $317,000 He wrote the whole thing in 9 days, spending about $20 on Claude API calls He set an old camera on his balcony, pointed it at the intersection below, and let Claude watch the road Claude tags every car, motorbike and pedestrian in real time, 653 in five minutes, and flags anyone over the limit The moment a car speeds, Claude clips the video, reads the license plate, matches the owner, and emails the fine on its own A normal radar takes one photo and misses half the time. Claude records full video, so there is nothing to dispute, and the fines go out with no operator He walked into the district office with a flash drive and asked for 10 minutes. he left with a contract Every Claude config he used is in the articleshow more

Fokki
332,305 görüntüleme • 1 ay önce
Inviting early testers and contributors to Project Devika -... The open-source alternative to Devin. 👩💻 As of now, Devika is far from the capabilities of Devin... but we'll eventually get there. So I am calling the open-source community to join forces! ❤️ Features: - 12 Agentic models that can interact with each other in a feedback loop to understand, browse, research, code, document, and make decisions according to the user's query to complete a project. - Supports Claude 3, GPT-4, GPT-3.5, and Local LLMs via ollama. - Devika can run the code she writes and fix/patch the code herself if she encounters any errors without user intervention. - Devika can deploy static websites she creates on Netlify. (Experimental) - And much more... Will be doing an official launch after intensive testing and bug fixes. 🙌 I've created a Discord server for the early testers and contributors. If you're interested in joining the team, reply to this tweet and I will DM you the invite link. #buildinpublicshow more

mufeed vh
155,020 görüntüleme • 2 yıl önce