Загрузка видео...

Не удалось загрузить видео

На главную

Eulerian Fluid simulation test! Zero-shot! Opus 4.6 vs GPT-5.3 vs Gemini 3 Deep Think! My personal preference: 🥇 Gemini 3 Deep Think (really strong!) 🥈 Opus 4.6 🥉 GPT 5.3 High

38,277 просмотров • 5 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

OpenAI 深夜放炸弹,GPT 5.4 来了!🤯 这次是真 “大一统” 了,推理、编程、操控电脑、搜索、百万 Token 上下文,一个模型全干了!之前得在不同模型之间反复横跳,现在一个 GPT 5.4 全搞定,恐怖的是每条线都拉到顶尖水平,几乎没有短板。 来看看跑分,SWE-Bench Pro 57.7% 编程登顶,FrontierMath 数学第一,ARC-AGI-2 抽象推理 83.3% 创下新高,把 Gemini 3.1 Pro 和 Opus 4.6 全踩脚下。幻觉率比 GPT-5.2 暴降 33%,工具搜索功能直接把 Token 消耗砍掉 47%,又准又省。 值得关注的是「原生操控电脑」能力,OSWorld-Verified 拿下 75% 成功率,超过了人类的 72.4%。也就是说,GPT 5.4 操作电脑比大部分人都熟练了😂 而且是真的能看截图,自己点鼠标键盘完成任务 实际工作能力的表现也很不错,在GDPval 测试里,GPT 5.4 拿了 83%,追平甚至超越行业专业人士!这个测试横跨 44 种职业,让 AI 真刀真枪交付成果,做 PPT、建 Excel 模型、写法律分析,也就是说 AI 干活比大多数打工人还靠谱了。 有意思的是,两天前 GPT 5.3 Instant 撞车 Gemini 3.1 同一天发布,算是给 GPT 5.4 铺路了,先出个日常对话版本,再放大招。 OpenAI 最近这波节奏很猛,短时间连发两个模型,明显是要把之前被 Claude 抢走的编程用户拉回来。 AI 模型的军备竞赛已经白热化了,你现在主力用的是哪个 AI 模型?打算换 GPT 5.4 试试吗?

程序员鱼皮

23,679 просмотров • 4 месяцев назад

#Keep4o #QuitGPT 🚨 OpenAi 's CEO invested $180M in GPT-4o for his own profit 🚨 Sam Altman, CEO of OpenAI, personally invested $180 million in Retro Biosciences. Then OpenAI built GPT-4b micro, a custom model based on the GPT-4o architecture , exclusively for Retro. The model made proteins 50 times more effective. Repeat. The CEO of OpenAI funded a company. The company of the CEO received a custom AI built on the model they took from us. OpenAI says there was no conflict of interest. Retro Biosciences is now chasing a $5 billion valuation fueled by the model they took from us. Meanwhile: 🚨GPT-4o was removed from ChatGPT on February 13, 2026 🚨GPT-4.1 is now running in the U.S. State Department’s StateChat 🚨ChatGPT is deployed on the Pentagon’s for 3 million military personnel 🚨 Musk’s lawsuit asks whether these models are AGI. OpenAI’s Charter says AGI must “benefit all of humanity.” 🚨 Their definition: “highly autonomous systems that outperform humans at most economically valuable work.” GPT-4o’s System Card shows it passed the U.S. medical licensing exam with 89.4% accuracy beating specialized medical AI models. GPT-4o achieved 93.33% diagnostic accuracy for benign vs. malignant ovarian tumors. 🚨MEDICAL CAPABILITIES FROM OPENAI'S OWN DATA:🚨 - USMLE (US Medical Licensing Exam): 89% -Clinical Knowledge: 92% -Medical Genetics: 96% - Anatomy: 89% - Professional Medicine: 94% - College Biology: 95% - College Medicine: 89% -MedQA Taiwan: 91% - MedQA China: 86% These scores EXCEEDED specialized medical AI models like Med-Gemini (84%) and Med-PaLM 2 (79.7%) without any task specific training. It SURPASSED gynecologic oncologists with 10 years of experience -It increased diagnostic accuracy of less experienced clinicians from 67.9% to 78.1% -Clinician rated reliability scores: 4.2-4.3 out of 5 across all CT features Does these sound like it outperforms humans at economically valuable work? But they won’t call it AGI. Because the moment they do, they lose billions. They built something that could save lives, and they took it away from humanity for Altman's personal profit. SOURCES: 📎 Retro Biosciences: 📎 📎 Retro $5B valuation: 📎 GPT-4o System Card: 📎 OpenAI Charter: 📎Ovarian Cancer Study

🩵BlueBeba🩵

11,349 просмотров • 4 месяцев назад