Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Your OpenAI and Anthropic weekly digest (Week 16, 2026) OpenAI announced Cloudflare Agent Cloud partnership with OpenAI frontier models including GPT-5.4, Agents SDK update with native sandbox execution and model-native harness, Codex update with background computer use, image generation, memory, automations, and 90+ new plugins, introduced GPT-Rosalind life sciences...

11,057 görüntüleme • 3 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

The week in OpenAI and Anthropic news (Week 19, 2026) OpenAI rolled out GPT-5.5 Instant as the new ChatGPT default model with memory sources and ChatGPT for Excel and Google Sheets globally, launched three new realtime voice models in the API (GPT-Realtime-2, GPT-Realtime-Translate, GPT-Realtime-Whisper) and an OpenAI CLI, introduced Trusted Contact safety feature, GPT-5.5-Cyber for defenders, B2B Signals report, ChatGPT Futures Class of 2026, EMEA youth safety blueprint, privacy in model training explainer, expanded ads pilot, published engineering posts on low-latency voice, MRC supercomputer networking, running Codex safely internally, and investigating accidental chain-of-thought grading during reinforcement learning, plus discovered ChatGPT Personal Wiki and dropped a goblin-themed merch line that sold out Anthropic hosted Code with Claude developer conference in San Francisco, announced a new enterprise AI services company with Blackstone, Hellman & Friedman, and Goldman Sachs, signed a SpaceX compute partnership and raised Claude Code and API usage limits, made Claude for Excel, PowerPoint, and Word generally available with Claude for Outlook in beta, launched Workload Identity Federation, financial services agent templates, dreaming, outcomes, and multiagent orchestration in Managed Agents, shipped 60+ Claude Code reliability fixes, published research on agentic misalignment training, sandbagging mitigation, model spec midtraining, and Natural Language Autoencoders, donated Petri to Meridian Labs, introduced The Anthropic Institute research agenda, plus discovered Orbit proactive assistant for Cowork and /radio command in Claude Code, and more

Tibor Blaho

12,602 görüntüleme • 2 ay önce

OpenAI and Anthropic week in review (Week 15, 2026) OpenAI published industrial policy ideas for the Intelligence Age, announced Safety Fellowship pilot program, acquired Cirrus Labs for agent infrastructure, introduced new $100 Pro tier with 5x Codex usage over Plus as the existing $200 Pro tier remains the highest option, published child safety blueprint, announced over $100 million in OpenAI Foundation grants for Alzheimer's research, shared enterprise AI update with enterprise now over 40% of revenue, released GPT-5.3 Instant Mini as new fallback model, announced older Codex model retirements, added Outlook shared mailboxes and calendars support in ChatGPT, paused UK Stargate project over energy costs and regulation, disclosed macOS app signing security incident from Axios supply chain compromise, and launched Prism Paper Review, plus spotted ImageGen 2 A/B testing in ChatGPT and European Commission plans to designate ChatGPT as very large online search engine under the Digital Services Act Anthropic expanded partnership with Google and Broadcom for multiple gigawatts of next-generation compute with run rate surpassing $30B and over 1,000 enterprise customers now spending $1M+ annually, officially announced previously leaked Claude Mythos Preview through Project Glasswing finding thousands of zero-day vulnerabilities with $100M in usage credits, launched Managed Agents in public beta and published engineering blog on decoupling agent architecture, published trustworthy agents framework update, made Claude Cowork generally available on all paid plans with enterprise controls, introduced the advisor tool for pairing Opus with Sonnet or Haiku, and launched Claude for Word beta, and more

Tibor Blaho

12,121 görüntüleme • 3 ay önce

OpenAI and Anthropic news roundup (Week 18, 2026) OpenAI open-sourced Symphony spec for Codex orchestration, published "Our principles" post, announced amended Microsoft partnership, achieved FedRAMP Moderate authorization, posted commitment to community safety, brought OpenAI models, Codex, and Managed Agents to AWS, shared cybersecurity action plan, posted Stargate compute infrastructure update, published "Where the goblins came from" post, posted Auto-review write-up for Codex, introduced Advanced Account Security, announced DevDay 2026, repositioned Codex as personal assistant for everyday work, launched Codex setup import, added Codex pets, announced GPT-5.5 party for next week, shared GPT-5.5 one-week launch metrics, and rolled out 360 worlds in ChatGPT Images on web, plus discovered Custom dictionary feature in development, ChatGPT search EU recipient numbers, confirmed new model selector in composer, and updated privacy policy with marketing cookies on by default for free users Anthropic opened Sydney office with new General Manager, launched Claude for Creative Work with new connectors, Claude Code can now send push notifications to your phone, published Introspection Adapters research, BioMysteryBench evaluation, "How people ask Claude for personal guidance" study, launched Claude Security in public beta, and preparing for Code with Claude developer conference next week, plus discovered "Cardinal" stats feature in development and internal red teaming for Claude Jupiter V1 P, and more

Tibor Blaho

12,254 görüntüleme • 3 ay önce

OpenAI and Anthropic this week: GPT-Red, Fable 5 plan changes, and free Claude for teachers (Week 29, 2026) OpenAI introduced GPT-Red, an internal automated red teamer trained through adversarial self-play to find prompt injection vulnerabilities at scale Training against it made GPT-5.6 their most robust model against prompt injections to date And there's a hidden "GPT-RED // Invader Patrol" game in the article OpenAI also published GPT-Live usage limits, brought ChatGPT back to WhatsApp in the European Economic Area, rolled out a new unified search in ChatGPT across chats, projects, images, and documents, raised the custom instructions limit from 1,500 to 5,000 characters, and updated the ChatGPT desktop app with a clearer Chat and Work layout, unified Recents, Projects, and cloud sync ChatGPT Finances got Apple Card and Savings support On the publishing side, OpenAI shared articles on managing AI investments in the agentic era, why teens deserve access to safe AI, state and federal AI safety action, and Sarah Friar's useful intelligence per dollar scorecard Beyond the official channels, Bloomberg reported OpenAI's first device will be a movable screenless smart speaker, and The New York Times reported a Kalshi partnership showing World Cup odds in ChatGPT search Work Louder launched the Codex Micro keypad built for Codex, OpenAI merch is back in a new Supply Co. shop, and I spotted a new private equity and investment management community in the works Anthropic made Claude Fable 5 standard in all Max and Team Premium plans at 50% of limits starting July 20, with a one-time $100 credit for Pro and Team Standard, after extending Fable 5 access on paid plans through July 19 Claude Code weekly limits stay 50% higher through August 19 Anthropic also introduced Claude for Teachers with free premium Claude access for verified K-12 educators in the US, committed 10 million Canadian dollars to Canadian AI research, and published a Canada Economic Index country brief Artifacts in Claude Code now support public sharing, multiplayer editing, and MCP connectors, plus creation via Claude Tag Claude Code got /code-review effort levels up to ultra, and HIPAA configuration is now self-serve for Claude organizations On the research side, Anthropic published work on Claude's values across models and languages, and four new agentic misalignment case studies

Tibor Blaho

13,137 görüntüleme • 16 gün önce

OpenAI just admitted Anthropic is KILLING their business. Their own applications chief told employees it was a "code red." Said Anthropic was a "wake-up call." Then admitted OpenAI had been "spreading efforts across too many apps" and it was "slowing them down." This is an internal confession. Here's why Anthropic is eating up OpenAI: 12 months ago, OpenAI owned 50% of all enterprise AI spending. Today it's just 27%. Anthropic went from nearly ZERO to winning 70% of every first-time enterprise AI deal. Seven out of ten companies buying AI tools for the first time are choosing Claude over ChatGPT. A year ago, one in 25 businesses on Ramp paid for Anthropic. Today it's one in four. OpenAI just had its biggest single-month adoption decline ever recorded. And Anthropic literally charges MORE than OpenAI for roughly the same performance. And businesses are STILL choosing them. In enterprise software, that never happens. The cheaper product usually wins. But Claude became something OpenAI never figured out how to be: Cool. Celebrities publicly switched to Claude. Senators are tweeting about using it. Engineers are shipping entire products with Claude Code in hours that used to take weeks. It started to became an identity signal. Like blue bubble vs green bubble in iMessage. Choosing Claude says something about you now. Meanwhile OpenAI went the opposite direction: They took the Pentagon contract that Anthropic refused. Greg Brockman donated $25 million to fund wars. ChatGPT uninstalls jumped 295% in a single day. Reddit posts saying "Cancel and Delete ChatGPT" got 30,000 upvotes. Anthropic said no to mass surveillance and autonomous weapons. Got blacklisted by the Pentagon. Trump called them a "Radical Left AI company." And their downloads went to #1 on the App Store the next day. Turns out refusing to build weapons is good marketing. But the real damage isn't consumer downloads. It's the MONEY. Claude Code hit $2.5 billion in annual revenue in six months. OpenAI's competing product Codex just barely crossed $1 billion. And Anthropic literally cannot meet demand. They're turning away paying customers because they don't have enough compute to serve them. A company REJECTING revenue because it's growing too fast. While OpenAI scrambles to consolidate. Last week OpenAI announced they're merging ChatGPT, Codex, and their browser into one "superapp." But what this really means: "We launched too many products, none of them worked well enough alone, so now we're cramming everything together and hoping it sticks." And remember their video tool Sora? Launched standalone. Hit #1 on the App Store. Usage flatlined within weeks. Now they're forced to shut it down. Their browser Atlas? Still hasn't launched publicly. Their IPO? Polymarket odds dropped from 55% to 35%. OpenAI has 900 million users. Anthropic has maybe 10 million daily actives. But here's the thing... OpenAI won the consumer war. ChatGPT is where your mom asks about recipes and your cousin makes memes. Anthropic won the war that actually MATTERS. The developers. The engineers. The enterprises writing 7 figure checks. OpenAI built the biggest chatbot on Earth. Anthropic built the tool that companies can't stop paying for. This is Yahoo vs Google all over again. Yahoo had the users. Google had the product. And we all know how that ended. OpenAI has 12 months to prove the superapp works, land the IPO, and stop the enterprise bleeding. If they can't, the most valuable startup in history becomes the most cautionary tale in tech. 900 million users don't mean anything if the people who actually pay are walking out the door. What do you think?

Ricardo

35,020 görüntüleme • 4 ay önce

In the future, you’ll be able to accomplish a goal by just giving Claude an outcome and a budget. That’s the direction Anthropic is building in with its new Managed Agents features, announced at this week’s Code with Claude developer event. The basic idea: Claude, wrapped in a computer in the cloud, that you can spin up, scale, and manage as needed. Anthropic is taking on the infrastructure that kills most agent products, and making sure that it scales to meet the needs of agents running 24/7. On this week’s AI & I from Every 📧, I talk with Angela Jiang (Angela Jiang), head of product for the Claude platform, and Katelyn Lesse (Katelyn Lesse), head of engineering for the Claude platform, about what Anthropic is building and what it takes to make agents reliable in production. We get into: - Why the "build a generic harness, hot-swap any model behind it" playbook is already outdated. Angela points to eval data on Memory where the same task across different harnesses performed drastically differently. - The infrastructure wall every team hits in production—and why Katelyn thinks “my sandbox died and took the agent with it” is the real reason internal agents don't ship. - Why Anthropic is so bullish on using file systems and skills within Claude, including Angela's argument that those early design choices can compound for years. This is a must-watch for anyone trying to take an agent past the demo and into production. Watch below! Timestamps: How the Claude platform evolved from API to agents: 00:01:48 The primitives that make up Claude Managed Agents: 00:04:09 Why the harness and the model are becoming a single unit: 00:10:37 The infrastructure wall that kills most agent projects in production: 00:18:49 Why team agents need a different shape than individual productivity tools: 00:24:49 How Anthropic's legal team uses an agent to review marketing copy: 00:26:36 Using multi-agent orchestration for advisor strategies, adversarial pairs, and swarms: 00:34:24 How to measure agent success with outcome and budget as the end state: 00:35:50 What the platform looks like a year from now, when Claude writes its own harness: 00:39:11

Dan Shipper 📧

66,339 görüntüleme • 2 ay önce

New short course: Collaborative Writing and Coding with OpenAI Canvas! Explore new ways to write and code with OpenAI Canvas, a user-friendly interface that allows you to brainstorm, draft, and refine text and code in collaboration with ChatGPT. In the short course, created with OpenAI, and taught by , a research lead at OpenAI, you’ll learn to use Canvas to enhance your workflows. Canvas lets you go beyond simple chat interactions. It provides a side-by-side workspace where you and ChatGPT can edit and refine text or code collaboratively. This makes brainstorming, drafting, and iterating as you write feel more natural and effective. As the first major update to ChatGPT’s visual interface since its launch in 2022, Canvas gives a new, innovative approach to collaboration with AI. For instance, after writing the first version of your code, Canvas can review it and give suggestions for improvement. It can also help with debugging by adding logging, identifying problems to fix, and writing comments. In addition, you'll also learn what it takes to train the model for an interface like Canvas. In this video-only short course, you’ll: - Learn how to ask for in-line feedback and control the iteration of your work by directly editing selected areas of your text or code from the model’s output. - Learn how to access quick automation tools in a shortcut menu that allows you to modify your writing tone and length, enhance your code, and restore previous versions of your work. - Learn how to use Canvas as a research assistant tool with an example of asking the model to reason through the screenshot of a plot to write a research report, in which you can ask questions within the created report. - Ask the model to write Python code to replicate the graph seen on a screenshot image. - Go behind the scenes of how you can create a video game, such as Space Battleship, from scratch, edit it, and display it in one self-contained HTML file. - Get a real-world application example of creating a SQL database from the image of its architecture. - Understand the model training and design processes that power Canvas! Please sign up here:

Andrew Ng

128,180 görüntüleme • 1 yıl önce

Thanksgiving-week treat: an epic conversation on Frontier AI with Lukasz Kaiser -co-author of “Attention Is All You Need” (Transformers) and leading research scientist at OpenAI working on GPT-5.1-era reasoning models. 00:00 – Cold open and intro 01:29 – “AI slowdown” vs a wild week of new frontier models 08:03 – Low-hanging fruit, infra, RL training and better data 11:39 – What is a reasoning model, in plain language 17:02 – Chain-of-thought and training the thinking process with RL 21:39 – Łukasz’s path: from logic and France to Google and Kurzweil 24:20 – Inside the Transformer story and what “attention” really means 28:42 – From Google Brain to OpenAI: culture, scale and GPUs 32:49 – What’s next for pre-training, GPUs and distillation 37:29 – Can we still understand these models? Circuits, sparsity and black boxes 39:42 – GPT-4 → GPT-5 → GPT-5.1: what actually changed 42:40 – Post-training, safety and teaching GPT-5.1 different tones 46:16 – How long should GPT-5.1 think? Reasoning tokens and jagged abilities 47:43 – The five-year-old’s dot puzzle that still breaks frontier models 52:22 – Generalization, child-like learning and whether reasoning is enough 53:48 – Beyond Transformers: ARC, LeCun’s ideas and multimodal bottlenecks 56:10 – GPT-5.1 Codex Max, long-running agents and compaction 1:00:06 – Will foundation models eat most apps? The translation analogy and trust 1:02:34 – What still needs to be solved, and where AI might go next

Matt Turck

168,007 görüntüleme • 8 ay önce

HERMES AGENT SUPPORTS 300+ MODELS. PICKING THE RIGHT ONE PER TASK IS THE DIFFERENCE BETWEEN $5/MONTH AND $50. STARTING OUT: Claude Sonnet 4.6. official recommendation from Nous Research. "the model this project was built and tested with." strong reasoning. reliable tool calling. mid-range pricing. PREMIUM TIER: Claude Opus 4.8. best coding benchmarks available. self-correcting reasoning. catches its own mistakes. 1M context. use for demanding tasks where quality matters. GPT-5.5. #1 Chatbot Arena. #1 GPQA Diamond reasoning (94.1%). #1 creative writing. 2M context. handles entire codebases in one pass. Grok 4.30. the only frontier model with live X firehose access. real-time social data, breaking news, market sentiment. connects via Grok OAuth. no separate API key. Grok-Composer-2.5-Fast (v0.17.0). Cursor's coding model. 200K context. available through your Grok subscription via OAuth. no extra cost if you already pay for Grok. MID-RANGE TIER: Claude Sonnet 4.6. best balance of quality and cost for daily use. strongest prose and tool calling in this tier. Gemini 2.5 Pro. Google Search grounding built in. cites sources. verifies claims. pulls current data. 2M context. best for research-heavy workflows. GPT-4.1. reliable tool calling. solid general reasoning. good middle ground when you need OpenAI compatibility. BUDGET TIER: Claude Haiku 4.5. fastest Anthropic model. cheapest paid Claude option. strong at classification, routing, simple queries. use for auxiliary tasks: compression, vision, web extraction, approval scoring. DeepSeek V4. best cost-to-quality ratio in the market. 90% cache discount on repeated context. use for sub-agents and bulk parallel work. DeepSeek V4 Flash. cheapest paid model worth using. 1M context. MIT license. self-hostable. use for cron jobs, monitoring, routine searches. MiniMax M3. Nous Research and MiniMax collaborating on optimization. 1M context via lightning attention. 59% SWE-Bench Pro. beats several premium models on coding. one of the most-used models inside Hermes. FREE / LOCAL: Qwen 3.5 27B via Ollama. 16GB VRAM. reliable tool calling. best free local model for Hermes as of mid-2026. Qwen 3 8B. 8GB VRAM. fits a $7 VPS. handles routine tasks at zero API cost. Llama 4 Maverick. best open-weight tool calling. 1M context. needs more VRAM but strongest local option. HOW TO ASSIGN MODELS: main model: Desktop app / Dashboard → Models → switch sub-agent model: set in Desktop app, Dashboard, or config.yaml: delegation: model: "deepseek/deepseek-v4" auxiliary models (compression, vision, web extract): Desktop app / Dashboard → Models → Auxiliary Haiku 4.5 or Gemini Flash work well here. saves significantly when your main model is premium. per-profile: each Hermes profile gets its own model. Scout on DeepSeek. Analyst on Sonnet. Briefer on budget model. Coder on Opus. per-cron-job: pin a specific model to any cron job. morning brief on Haiku. deep research on Sonnet. monitoring on DeepSeek Flash. each job uses only the model it needs. per-session: /model deepseek/deepseek-v4-flash hot-swap mid-conversation. no restart needed. FALLBACK CHAINS: if your primary model is unavailable, Hermes automatically switches to the next provider. rate limit or server error = next model in the chain. no failed runs. no manual intervention. set in Desktop app, Dashboard, or config.yaml: fallback_providers: - openrouter - nous - codex PROVIDER PATHS: OPENROUTER: 300+ models under one API key. pay per token. most flexible. NOUS PORTAL: 300+ models + Tool Gateway (web search, image gen, TTS, browser). one OAuth. one subscription. 10% off token-billed providers. CHATGPT SUB: GPT-5.5 + Grok via OAuth. included tokens with $20 subscription. OLLAMA: free. local. private. zero API cost. your hardware only. mix providers across profiles and tasks. Scout on OpenRouter. Analyst on Nous Portal. Coder on ChatGPT sub. Monitor on Ollama. THE RULE: premium for work that needs deep reasoning. mid-range for daily driver tasks. budget for volume and background work. free for monitoring and routine jobs. pricing changes fast. check openrouter ai for current rates before committing. Which is your favourite model and for what task? full 15 levels breakdown in the article 👇

YanXbt

17,138 görüntüleme • 1 ay önce

OpenAI just announced API access to o1 (advanced reasoning model) yesterday. I'm delighted to announce today a new short course, Reasoning with o1, built with OpenAI, and taught by Colin Jarvis, Head of AI Solutions at OpenAI, to show you how to use this effectively! Unlike previous language models which generate output directly, o1 “thinks before it responds,” and generates many reasoning tokens before returning a more thoughtful and accurate response. It is great at complex reasoning -- including planning for agentic workflows, coding, and domain-specific reasoning in STEM fields like law. But how you should use it is quite different from other LLMs. I think o1 will be a game changer for many AI applications; and in this course, you'll learn how to use it effectively. In detail, you’ll: - Learn to recognize what tasks o1 is suited for, and when to use a smaller model, or combine o1 with a smaller model - Understand the new principles of prompting reasoning models: Be simple and direct; no explicit chain-of-thought required; use structure; show rather than tell - Implement multi-step orchestration in which o1 plans, and hands tasks over to gpt-4o-mini to execute specific steps; this illustrates a design pattern to optimize intelligence (accuracy) and cost - Use o1 for a coding task to build a new application, edit existing code, and test performance by running a coding competition between o1-mini and GPT 4o - Use o1 for image understanding and learn how it performs better with a "hierarchy of reasoning," in which it incurs the latency and cost upfront, preprocessing the image and indexing it with rich details so it can be used for Q&A later - Learn a technique called meta-prompting, in which you use o1 to improve your prompts. Using a customer support evaluation set, you'll iteratively use o1 to modify a prompt to improve performance You'll also learn about how OpenAI used reinforcement learning to produce a model that uses "test-time compute" to improve performance. I think you'll find this course enjoyable and valuable. Please sign up for it here:

Andrew Ng

357,661 görüntüleme • 1 yıl önce