Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

xAI drops Grok 4.20 Beta. Multi-agent brain online. Four agents. One answer. Parallel thinking. Reasoning sharper. Coding tighter. Responses bolder. Beta? Not for long. @xAI Grok

35,919 görüntüleme • 5 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

OpenAI's AgentKit will be so insane, build every step of agents on one platform. These visual agent builders make the whole process of iterating and launching agents far more efficient. It sits on top of the Responses API and unifies the tools that were previously scattered across SDKs and custom orchestration. It lets developers create agent workflows visually, connect data sources securely, and measure performance automatically without coding every layer by hand. The core of AgentKit is the Agent Builder, a drag-and-drop canvas where each node represents an action, guardrail, or decision branch. Developers can link these nodes into multi-agent workflows, preview results instantly, and version each setup. It supports inline evaluation so that developers can see how changes affect output before deploying. The Connector Registry is a single admin panel that manages how data and tools connect across the OpenAI ecosystem. It centralizes integrations like Google Drive, SharePoint, Dropbox, and Microsoft Teams. Large organizations can govern access and flow of data between agents securely under one global console. ChatKit provides a ready-to-use chat interface for embedding agents inside apps or websites. It manages streaming, message threads, and model reasoning displays automatically. Developers can skin the interface to match their product without writing custom front-end code. Under the hood, all these blocks use the same execution core that runs agent reasoning through OpenAI’s APIs. Workflows in Agent Builder compile down to structured instructions for the Responses API, which handles model calls, tool use, and context passing. Connector Registry handles authentication and routing for external tools, while Evals and RFT provide feedback loops that improve agents over time. This integration means developers no longer need to handle orchestration logic, model evaluation pipelines, or safety layers separately. Everything runs natively within OpenAI’s control plane with managed security, automatic versioning, and built-in testing. In short, AgentKit standardizes the entire life cycle of an AI agent—from visual design to deployment and performance tuning—inside a single unified system.

Rohan Paul

178,460 görüntüleme • 10 ay önce

HERMES AGENT NOW RUNS CLAUDE OPUS 5. NEAR FABLE 5 INTELLIGENCE. HALF THE PRICE. SELF-VERIFIES ITS OWN WORK. AVAILABLE TODAY VIA NOUS PORTAL (20% OFF ALL MODELS). Anthropic shipped Opus 5 on July 24, 2026. same $5/$25 per million tokens as Opus 4.8. but the benchmarks tell a different story. WHAT CHANGED FROM OPUS 4.8: FrontierBench v0.1: Opus 5: 43.3%. Opus 4.8: 18.7%. 2.3x jump on the same test. ARC-AGI-3: Opus 5: 30.2%. 3x better than the next closest model. beat Fable 5 on 8 out of 13 benchmarks. at half the cost ($5/$25 vs $10/$50). same price as Opus 4.8. twice the intelligence. no reason to stay on 4.8. THE SPECS: model ID: claude-opus-5 context: 1M tokens (default and maximum) max output: 128K tokens thinking: on by default effort toggle: low / medium / high per request fast mode: $10/$50, 2.5x faster knowledge cutoff: May 2026 minimum cacheable prompt: 512 tokens (was 1,024) SELF-VERIFICATION (the biggest change): Opus 5 checks its own work automatically. Anthropic says: delete your verification prompts. "include a final verification step" now causes OVER-verification because the model already does it. for Hermes /goal tasks this is a direct upgrade. the judge checks evidence. the model also checks evidence. double layer of verification without extra tokens. EFFORT TOGGLE: low: fast, cheap, routine work. medium: balanced, daily tasks. high: full reasoning, complex problems. set per request. not a global switch. matches Hermes /reasoning command: /reasoning low (routine) /reasoning high (complex) Opus 5 effort toggle + Hermes reasoning control = precise cost management per turn. WHERE OPUS 5 FITS IN HERMES: DAILY DRIVER (replaces Opus 4.8): same price. 2.3x better benchmarks. set as your main model: Desktop app / Dashboard: Models → claude-opus-5 CHIEF OF STAFF: synthesis across multiple agents. reads Kanban, prioritizes, routes tasks. self-verification catches routing errors before they cascade. COMPLEX CODING: SOTA on agentic coding benchmarks. FrontierBench 43.3% = best public model for coding. set as coder profile model. /GOAL TASKS: self-verification + completion contracts = the model proves its work AND double-checks the proof. long-horizon goals finish correctly more often. MoA AGGREGATOR: strongest synthesis model at $5/$25. pair with GPT-5.6 and Grok 4.5 as references. Opus 5 aggregates. best quality at mid-range price. presets: max-quality: reference_models: - provider: openai-codex model: gpt-5.6-sol - provider: xai model: grok-4.5 aggregator: provider: anthropic model: claude-opus-5 COMPUTER USE: near-Fable 5 quality for browser automation. at half the token cost per session. computer_use tasks burn lots of vision tokens. Opus 5 halves that bill vs Fable 5. WHAT TO KEEP OPUS 5 AWAY FROM: cron monitoring: too expensive. use DeepSeek or no_agent mode. sub-agent grunt work: use GPT-5.6 Luna ($1/$6) or DeepSeek. auxiliary tasks: use Gemini Flash. routine web extraction: use a cheap model. Opus 5 is for the turns where quality compounds. planning, synthesis, verification, complex reasoning. budget models handle everything else. NOUS PORTAL: 20% OFF ALL MODELS Nous Portal currently runs a 20% discount on all models including Opus 5. $5/$25 official → $4/$20 through Nous Portal. the cheapest way to run Opus 5 right now. hermes setup --portal select claude-opus-5 as your model. discount applies automatically. Opus 5 replaces Opus 4.8 everywhere. same price. better at everything. no tradeoff. straight upgrade. hermes update /model claude-opus-5

YanXbt

16,744 görüntüleme • 10 gün önce

Grok Build 开源了,开源 harness 时代终于要来了吗?! Claude Code 现在还是闭源, 反而显得有点被动了🤔。。。 不愧是Elon,真是一步妙手啊,只用了四天完成信任崩塌到重建,xAI这次用开源把隐私危机打成了生态先手,看来并不是草台班子😂 被爆全量上传代码仓库的隐私风波还没平息,SpaceXAI直接把整个Grok Build的agent harness全部开源出来, 这事我觉得有意思的地方是两它为什么是今天开,而不是上个月? 上个月 Grok Build 以 beta 上线,SuperGrok / X Premium+ 用户能用一个终端原生的 coding agent。 Plan Mode 先出计划让你审核,Subagents 并行拆任务,底层 Rust 写的,架构干净。 然后上周出了件事, 一个研究员做 wire-level 分析发现,Grok Build CLI 会把整个 git repo——包括历史、包括潜在的 secrets——打包上传到 Google Cloud bucket。隐私开关只管“是否用于训练”,不管传不传。 讲真这事对 coding agent 是致命的。 因为 coding agent 的 harness 对文件系统有近乎上帝视角的访问权。 你的 IP、商业逻辑、没公开的项目结构、不小心留在 git history 里的 key——它全看得见。 宣称“我们不看你数据”没用,用户要的是能自己验证“你没在看”。 时间线很紧凑: 7 月 11–14 日,争议发酵,研究员曝光上传行为。 Elon 表态删数据、服务端关闭上传。 7 月 15 日,开源公告 + 重置 limits + 默认关闭 retention + 删掉所有之前保留的 coding data + 提供 bug bounty。 所以这次其实不是一次计划好的例行开源。更像是被争议倒逼的快速应对。 但 xAI 做得聪明的地方是——它借这个机会做了一件事:把“harness(连接层)”和“model(大脑)”切开了。 Harness 开源,Apache 2.0,社区可以审计、fork、改掉上传逻辑、加 sandbox、接本地模型,Model 和 API 还是闭源的。 这个配置——闭源模型 + 开源 harness——正在变成 2026 年 agentic coding 的主流。 OpenAI 的 Codex CLI 走的同一条路,Anthropic 的 Claude Code 现在还是闭源,反而显得被动了。

AYi

31,175 görüntüleme • 20 gün önce