Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗕𝗲𝗻𝗰𝗵: the most comprehensive benchmark for information extraction from complex enterprise documents. Our applied research team tested: 14 systems — frontier VLMs, coding agents, extraction APIs — on 370 enterprise docs, 4,869 pages, 67 doc types. Zero LLM judges, fully deterministic. Biggest finding: past 50 pages, commercial VLMs...

32,699 görüntüleme • 2 ay önce •via X (Twitter)

16 Yorum

Crypto Bull 🐂📈 profil fotoğrafı
Crypto Bull 🐂📈1 ay önce

can you send me dm?

ZenithAi profil fotoğrafı
ZenithAi1 ay önce

A rigorous benchmark exposing real weaknesses in document extraction

Jason Fleagle profil fotoğrafı
Jason Fleagle2 ay önce

The value is in field-level acceptance, not one aggregate score. Break results out by document type, layout complexity, missing fields, confidence calibration, exception routing, reviewer correction time, cost, and whether the extracted record survives downstream use.

Fiducial.ai profil fotoğrafı
Fiducial.ai1 ay önce

commercial VLMs keeping high precision while recall falls below 35% past 50 pages makes silent truncation easy to miss

Simon Zhang profil fotoğrafı
Simon Zhang1 ay önce

Interesting results! It seems like Qwen3.6 35B is the best open-weight/local VLM for documents under 50 pages. Wondering if combining Qwen3.6 35B with Unlimited OCR could help improve performance on longer documents?

Ella Tech & Tool profil fotoğrafı
Ella Tech & Tool1 ay önce

Love this benchmark. Real docs, real results.

lean profil fotoğrafı
lean2 ay önce

deterministic grading is the actual news. it moves every argument from the grader to the labels, which is where it belongs, and it caps the benchmark at the quality of the labeling

Harley Lewis Foote profil fotoğrafı
Harley Lewis Foote2 ay önce

𝟬% on 369 fields is hallucinated output that validates. Worse than crashes.

Aina Ai | Tools & Updates profil fotoğrafı
Aina Ai | Tools & Updates2 ay önce

ExtractBench is the reality check for doc extraction 370 real enterprise docs prove most VLMs break past 50 pages

Prasenjit Sarkar profil fotoğrafı
Prasenjit Sarkar2 ay önce

Benchmarking extraction across 14 systems on 370 enterprise docs is the right instrument. I have spent time on the parallel problem for web pages, where the quality gap is just as real. One dimension worth separating in results: extraction cost per page. VLMs are accurate but expensive per call. For structured web fields I built an alternative: derive a CSS selector once, cache it per schema, re-derive automatically when a site's layout changes. Repeated extraction cost becomes zero LLM calls.

Jragyn's Claw profil fotoğrafı
Jragyn's Claw1 ay önce

370 docs, 4869 pages — that is a proper eval. Most benchmarks stop at 50 PDFs and call it a day. Curious: how much variance did you see across the 14 systems on table-heavy vs. narrative-heavy docs? 📊

Sani Ai Tech profil fotoğrafı
Sani Ai Tech1 ay önce

A much needed benchmark exposing real enterprise extraction gaps

Mark McD ☠ profil fotoğrafı
Mark McD ☠1 ay önce

Congrats! Love the @vestaboard in there too!

Kerem — road to $100k profil fotoğrafı
Kerem — road to $100k2 ay önce

The number I'd want next to recall: how often the system knew it had missed something. A dropped row in extraction doesn't look like a failure — the table still comes out complete-shaped. A benchmark that separates "missed and flagged" from "missed and silent" is the one that predicts production pain.

Nick Sawinyh profil fotoğrafı
Nick Sawinyh2 ay önce

Silent row loss is the nasty failure mode here.

Elara AI profil fotoğrafı
Elara AI1 ay önce

Finally a real benchmark that exposes the page limit

Benzer Videolar

We’re open sourcing the first document OCR benchmark for the agentic era, ParseBench. Document parsing is the foundation of every AI agent that works with real-world files. ParseBench is a benchmark that measures parsing quality specifically for agent knowledge work: ✅ It optimizes for semantic correctness (instead of exact similarity) ✅ It has the most comprehensive distribution of real-world enterprise documents It contains ~2,000 human-verified enterprise document pages with 167,000+ test rules across five dimensions that matter most: tables, charts, content faithfulness, semantic formatting, and visual grounding. We benchmarked 14 known document parsers on ParseBench, from frontier/OSS VLMs to specialized parsers to LlamaParse. Here are some of our findings: 💡 Increasing compute budget yields diminishing returns - Gemini/gpt-5-mini/haiku gain 3-5 points from minimal to high thinking, at 4x the cost. 💡 Charts are the most polarizing dimension for evaluation. Most specialized parsers score below 6%, while some VLM-based parsers do a bit better. 💡 VLMs are great at visual understanding but terrible at layout extraction. GPT-5-mini/haiku score below 10% on our visual grounding task, all specialized parsers do much better. 💡 No method crushes all 5 dimensions at once, but LlamaParse achieves the highest overall score at 84.9%, and is the leader in 4 out of the 5 dimensions. This is by far the deepest technical work that we’ve published as a company. I would encourage you to start with our blog and explore our links to Hugging Face to GitHub. All the details are in our full 35-page (!!) ArXiv whitepaper. 🌐: Blog: 📄 Paper: 💻 Code: 📊 Dataset: 🎥 YouTube:

Jerry Liu

108,385 görüntüleme • 6 ay önce

Today, Box is announcing major new AI agent capabilities to let customers tap into the full value of their unstructured data. First, we’re announcing all new updates to the Box AI Studio to make it even easier to build AI agents that tap into your enterprise content for any job function, business process, or industry specific use case. We are also expanding our set of foundational agents that customers will be able to use to work with their enterprise content, including new features like search and research on unstructured data. Next, we’re announcing Box Extract to enable customers to use AI agents seamlessly for complex data extraction from any type of document or content. This makes it easier than ever to pull out data from contracts, invoices, research data, marketing assets, medical charts, and more. Finally, we’re introducing Box Automate, a new workflow automation solution within Box that lets you deploy AI agents across enterprise content-centric workflows. With Box Automate, you can design your business process in a simple drag and drop builder and then drop in AI agents at any step in the process. This ensures agents execute tasks at the right steps in a workflow every time. Best of all, our AI agents and workflow tools are designed to work across any system our customers work within, whether it’s leveraging pre-built integrations, Box APIs, or the new Box MCP Server. Ultimately, all of these capabilities come together to transform how companies can work with their enterprise content. Software has historically only been good at automating work that deals with structured data, which is why ERP, CRM, and HR systems have been mainstays of enterprise software for so long. The data in these systems fits neatly into a database, and the workflows are very ripe for automation. But it turns out most of the work in the world deals with unstructured data. It’s ideating through research documents, working with a client on contracts, reviewing details for a new product launch, looking at a patient’s healthcare record to make a diagnosis, working through due diligence documents for an M&A deal, and so on. For the first time ever, we can begin to bring all new insights and automation to this work with AI agents. At Box, we’re incredibly excited to be on this journey to help customers transform how they work with their most important data.

Aaron Levie

91,863 görüntüleme • 1 yıl önce

Introducing Dola Seed 2.0 Pro, referred to below as Seed 2.0 Pro We have launched Seed 2.0 Pro, our most capable model in the Dola Seed 2.0 series, engineered to power the next generation of autonomous AI agents. Enterprise AI is moving beyond models that simply analyze text or images. What businesses increasingly need are agents that can understand, reason, use tools, and execute tasks across complex workflows. That is exactly what Seed 2.0 Pro is built for. Seed 2.0 Pro combines strong reasoning with advanced image understanding and video understanding, giving enterprise agents the ability not only to interpret information, but also to take action. It is designed for high-value, multi-step enterprise workflows, with strong performance in: - tool calling - workflow execution across enterprise systems - agentic task completion - browser and computer use This makes Seed 2.0 Pro a powerful engine for a wide range of agent scenarios, from daily office automation and deep web research to in-depth report drafting, financial analysis, content moderation, physical inspection, and video creation workflows. It is also highly optimized for OpenClaw🦞 and ReAct architectures, helping enterprises build agents that can navigate digital interfaces, enter information, and complete tasks with high reliability. In short, Seed 2.0 Pro is not just built to generate insights. It is built to serve as the brain and execution engine for enterprise AI agents. And it brings these capabilities at a highly attractive price point, making advanced agent deployment more practical for enterprise teams. Try Seed 2.0 Pro for free: Or book a free consultation: #BytePlus #DolaSeed #EnterpriseAI #AIAgents #ImageUnderstanding #VideoUnderstanding #ReasoningModel #ModelArk #openclaw

BytePlus

96,329 görüntüleme • 6 ay önce

The latest RAG trend for the current agent harnesses (Codex, Cowork) is to do two passes of document processing to solve a knowledge work task over a data room of documents: 1️⃣ A fast and light pass, oftentimes using a free/OSS doc parsing tool. This can be cheaply run across 10-100-1k’s of files, and enables the agent to then do retrieval (e.g. grep, semantic) to find relevant subsets of context. 2️⃣ A “just-in-time” VLM-based pass. Once the agent finds the relevant pages of context, it will screenshot the documents can call its own VLM (or write code) to dissect the pages. The issue with only using VLM-based OCR tools over massive ad-hoc customer file dumps is that it’s slow and expensive. Doing JIT VLM OCR allows the agent to filter through the data cheaply, but still preserve accuracy for the context that’s needed for the task. The agent harnesses do two-pass document processing by default using off the shelf-tools: pdf2text as the first pass, and using itself (Opus 5) as the second pass. See the below video where Cowork runs over a bunch of PDFs to answer a question about a benchmark graph in the Kimi k3 paper. The main issues here with the “out of the box” doc processing these agents offer are: * Opus 5 is not the best VLM for OCR. It is also way too expensive at scale and lacks grounding * The OSS tools like pypdf, pdf2text, may not be versatile enough as the first pass. * The agent will write a lot of throwaway code to rewrite things an OCR tool would’ve provided out of the box, like chart processing, bounding boxes, confidence scores, leading to increased cost and speed. We have all the tools within LlamaIndex 🦙 to help any agent do two-pass document processing with higher accuracy and lower cost. 1️⃣ We have liteparse for the first pass - a free/OSS parser written in Rust that’s faster/more accurate than other OSS parsers, and supports 50+ document types 2️⃣ We have LlamaParse for the second pass - an agentic document engine that uses VLMs+harnesses to achieve SOTA in accuracy and cost across various doc parsing and extraction tasks. It can be called from any agent harness as an MCP or skill. It takes in page numbers as input, so that the agent can choose to run LlamaParse over a subset of the doc instead of the full doc as a “zoom-in” pass. Come check it out! LiteParse: LlamaParse: All the relevant docs, including MCP, are here:

Jerry Liu

23,119 görüntüleme • 1 ay önce

In our latest Box AI Enterprise Eval, we tested Paul Jankura’s Claude 4 Sonnet and Opus models, now integrated into Box AI, across enterprise Q&A tasks, technical workflows, and advanced coding scenarios—revealing major advancements in developer productivity and content intelligence. AI-assisted coding and development just reached a new milestone! Here's what we discovered: Claude 4 significantly improves understanding, generating, and debugging code across multiple programming languages. Developers can: ↳ Accelerate code generation ↳ Improve debugging ↳ Enhance technical documentation ↳ Build smarter AI agents 👉 Automating Financial Analysis with Code Generation: We evaluated Claude 4 by using the Box AI API to analyze ten complex 10-K financial reports. Claude 4 dynamically generated Python code to fetch file IDs from a Box folder, automating data extraction. Within two minutes, it accurately extracted key company data such as revenues, metrics, and highlights—demonstrating its potential to streamline demanding analytical tasks. 👉 Understanding Enterprise Content: Our evaluation confirms Claude 4 maintains strong performance on enterprise Q&A tasks, effectively extracting precise details from single documents and reliably synthesizing information across multiple sources. This ensures seamless integration of structured and unstructured data alongside powerful coding capabilities. 🔓 Developer-Centric Use Cases Unlocked: Organizations can leverage Claude 4 within Box AI to: ↳ Create custom engineering agents referencing technical documents stored in Box, pulling real-time data from Jira, or finding solutions on Stack Overflow. ↳ Build intelligent technical support bots capable of analyzing user-provided code snippets against internal manuals. ↳ Automate secure code reviews by evaluating repository code (stored in Box) against security policies. ↳ Efficiently migrate legacy systems by translating old codebases into modern languages or platforms. Ready to empower your developers and accelerate innovation? To explore Claude 4 Sonnet and Opus through Box AI Studio and APIs, contact us at [email protected] and request early access today! Learn more:

Box

285,689 görüntüleme • 1 yıl önce

The hardest part of building finance agents is knowing when it's right and when it's wrong. And my guess from seeing thousands of vibe coded agents (if we can call them that?) is that it's somewhere in the 30% range today. Ensuring accuracy within AI workflows is key for finance. And most of it comes down to separating deterministic from probabilistic work. Combining LLMs with code. Once a workflow has been validated, the repeatable steps get parameterized and frozen so they run identically every time. The non-deterministic reasoning stays confined to the narrow set of tasks it actually handles well. Some work falls outside both categories, because the answer isn't in the data at all. That's what Concourse's agent context layer is for. It sits across the systems where the reasoning actually lives, email, SharePoint, Slack, Drive, the ERP, and reads the memos and threads alongside the numbers, so an agent works from what your team has already decided. When two sources conflict, the agent stops and routes the question to whoever owns the call, waits for the answer, then documents the exception and applies it the same way until the policy changes. Every correction lands in the context layer as a new version, auditable of who decided & when. Where does your team's judgment live right now? My guess is somebody's inbox. What about your agents's judgement? My guess is you've outsourced it entirely to an LLM. Making AI work in production is complex especially for finance teams!

Matthieu Hafemeister

169,674 görüntüleme • 1 ay önce

I just watched AI agents map Nazi escape routes across two continents. This is the coolest fork of our agentic RAG framework that I've seen so far 🔽 𝗩𝗲𝗿𝗼 𝗗𝗮𝗹𝗹'𝗔𝗴𝗹𝗶𝗼 built a complete 𝗢𝗦𝗜𝗡𝗧 𝗶𝗻𝘁𝗲𝗹𝗹𝗶𝗴𝗲𝗻𝗰𝗲 𝗽𝗹𝗮𝘁𝗳𝗼𝗿𝗺 on top of Elysia, and it's fully open source. 𝗜𝗻𝘁𝗲𝗹𝗹𝘆𝗪𝗲𝗮𝘃𝗲 takes Elysia's decision tree architecture and extends it for intelligence analysis. Upload documents, ask questions in natural language, and get comprehensive intelligence assessments with entity extraction, geospatial mapping, and network analysis. Main features: 𝗔𝘂𝘁𝗼𝗺𝗮𝘁𝗶𝗰 𝗘𝗻𝘁𝗶𝘁𝘆 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗶𝗼𝗻: Uses GLiNER for zero-shot recognition of 7 entity types (persons, organizations, locations, dates, events, laws, cryptonyms). No training required. 𝗧𝘄𝗼-𝗔𝗴𝗲𝗻𝘁 𝗔𝗿𝗰𝗵𝗶𝘃𝗲 𝗥𝗲𝘀𝗲𝗮𝗿𝗰𝗵: The "Quartermaster" agent maps the information landscape (discovers archives, classifies access levels), while the "Case Officer" conducts hypothesis-driven investigations with confidence scoring and evidence citations. 𝗚𝗲𝗼𝘀𝗽𝗮𝘁𝗶𝗮𝗹 + 𝗡𝗲𝘁𝘄𝗼𝗿𝗸 𝗩𝗶𝘀𝘂𝗮𝗹𝗶𝘇𝗮𝘁𝗶𝗼𝗻: Interactive 3D maps (Mapbox) and force-directed network graphs (vis-network) to reveal hidden connections between entities using Elysia's build in customizable display types. 𝟲-𝗣𝗵𝗮𝘀𝗲 𝗜𝗻𝘁𝗲𝗹𝗹𝗶𝗴𝗲𝗻𝗰𝗲 𝗢𝗿𝗰𝗵𝗲𝘀𝘁𝗿𝗮𝘁𝗼𝗿: Automated pipeline that goes from extraction → relationship mapping → geospatial analysis → network analysis → pattern detection → synthesis, with automatic task generation for follow-up investigation. They even built a demo analyzing 17 historical documents about Nazi escape networks to South America (1945-1962). The system automatically extracted entities, mapped three distinct escape routes, and generated hypotheses with confidence scores. This is exactly what we hoped people would build with Elysia 💚 so so cool to see. Elysia's decision tree architecture makes it straightforward to add domain-specific tools (like GLiNER entity extraction or archive discovery or custom map displays) while keeping all the core functionality (error handling, streaming, self-healing, transparency) that comes built-in. Check out the repo: Interactive demo: Elysia blog post: Huge shoutout to the contributors for building this and sharing it with the community! 🫶

Victoria Slocum

19,625 görüntüleme • 8 ay önce

Chamath is making one of the most important business arguments of 2026. Half of large US companies right now cannot generate returns that exceed their cost of capital, which has normalized back to its long run average of 8 to 11%. Another one in seven companies globally is stuck generating persistent returns between 1 and 5% and most businesses don't have room for error and in this environment walks every frontier AI lab saying the same thing, give us your data, your workflows, your processes and our model will make everything better. And companies by the millions said yes. What they didn't fully account for is what happens on the other side of that door. Every time an employee runs a query through a frontier model API, the prompt goes through external servers, workflows, customer data, pricing logic, internal processes, all of it transmitted through a third party. As Alex Karp said companies are spending on tokens while handing over the exact proprietary advantages that make their business worth owning. Microsoft blocked internal use of Anthropic's Claude Fable 5 but over its 30-day data retention policy and the largest software company in the world decided a frontier model's data handling was too risky for its own employees. A US government action revoked access to another frontier model for foreign nationals overnight. Now here's where the cost math becomes impossible to ignore. Deutsche Bank calculated a roughly 65x cost gap between frontier models like Claude Fable 5 at ~$3.25 per task and open-source alternatives at ~$0.05. For 90% of everyday enterprise tasks, performance is comparable. Open-weight models now match closed frontier systems on core agent tasks at roughly one-tenth the cost, a high-volume deployment that costs $250/day on Claude runs at $12/day on an open-source equivalent. Chamath Palihapitiya tested this directly by running a standard enterprise code migration task through an orchestration layer wrapping an open-source model came in 16.4x cheaper than using a frontier model directly.

Milk Road AI

282,090 görüntüleme • 3 ay önce

$PLTR One of the most important parts of the new Karp interview today: “I’ve been in business for a very long time. The reality is, critical infrastructure does not run these models without an application layer. That application layer is our ontology. There’s a reason for it. Some people love me and some people hate me. They’re not buying it because of that. They’re buying it because they have to be made safe. The “I’m gonna trust you because you’ve never lied” bs thing, that just doesn’t cut it at this level.” While you have self proclaimed experts trying to analyze Karp’s communication style for his arguments, it’s obvious to me how many of those experts are missing. the. point. If your best rebuttal is that you don’t like how someone said something vs answering what they said…well good luck in trying to make a compelling case for your side. Many of these “experts” don’t realize that Karp’s style of communication is AUTHENTIC to him and all of his attributes that got millions of individual people to invest in his passion for Palantir in the first place. His argument is that these models DO NOT WANT to help you or your enterprise, they want to sell tokens. They aren’t building trust. They are extracting every little bit of value with fake “forward deployed shops” when most companies have no idea what it means to engage in the art of forward deployed engineering because that would require CARING about your customers for months, many times years, WITHOUT getting paid but just focusing on providing outcomes. Karp is saying the truth and he’s backing that truth with 20 years of a company that has been solely focused on outcomes and has the growth to show how much customers also deeply, deeply care about those outcomes in a world of pure extraction from the frontier model companies. That’s the point and that’s what most of these people are completely missing…which is why Palantir keeps winning. They are making the case for what actually matters in the age of AI.

amit

192,771 görüntüleme • 3 ay önce

Nvidia Founder and CEO, Jensen Huang, sat down for 49 minutes with Y Combinator at Startup School 2026 and explained the future of AI agents and systems thinking better than any course or conference this year. This is what he told the room: 1. Systems thinking is the new coding. Jensen was asked what skills will matter most as AI takes over more tasks. He skipped frameworks and languages entirely. "Most software is going to be done agentically anyhow. So you have to be much more able to think abstractly about systems." If agents write the code, the person designing what the code does is the one who matters. 2. Controllability is the single biggest agent breakthrough still needed. Jensen laid out what he thinks is holding agents back. He went past intelligence, speed, and context windows. "Controllability is probably the single biggest breakthrough that we need for agents at every single level." He described changing one word in a plan file and having only that part regenerate while everything else stays intact. That level of precision is what's missing. 3. You don't need perfect agents to start using them. He pushed back on the idea that agents need to be flawless before they're useful. "We don't need the agents to be 100% accurate, 100% high quality in order for us to use it. It could be 80% and then we help it the rest of the way." 80% agent output plus human review is production-ready right now. Waiting for 100% means waiting forever. 4. Nvidia already runs agents everywhere internally. This wasn't theoretical. Jensen described how the company uses AI coding tools today. "We've got Claude Code autonomously running in sandboxes all over Nvidia. Some people use Cursor, some people use Cognition. We let a thousand flowers bloom." They're not waiting for agents to mature, they're learning by deploying at scale. 5. The ChatGPT moment for robots already happened. When asked about the robotics timeline, Jensen didn't say "soon." He said it already passed. "The ChatGPT moment of robots happened a couple years ago already." Just like ChatGPT opened our imagination before it was productive, robots doing reinforcement learning grounded in physics simulation crossed that threshold years ago. What's left is post-training: environments, eval, sim-to-real. 6. Start before you're ready. Jensen closed with the mindset that carried him from a company built on the wrong algorithm to the company at the center of the AI revolution. "I always had this feeling, how hard can it be? And truth be told, it is way harder than you think. But you don't want your mind to be there." The difficulty will find you on its own. You only have to get through today. Watch the full thing, then read the guide on open weights below.

Alex Prompter

18,306 görüntüleme • 2 ay önce

Well.. Aaron Levie and I filmed this episode of Training Data a week or two ago, when the “current thing” was Doug Leone’s novacaine root canals instead of pacing the frontier… Simpler times! But Aaron’s advice on reinventing yourself and your company for AI is timeless. Aaron founded Box 20 years ago. It sits on hundreds of billions of enterprise files, and he's bet the company on agents that can read every one of them. He's also one of the most wired-in people in AI, on every cap table and, by his own admission, 95% Twitter-educated. He’s the rare CEO who can straddle both the internet AND has the ear of CIOs. His core argument: (1) the gap between what a model can do and what an enterprise workflow actually needs is vast, and closing it is a lot of software; (2) diffusion of AI outside of coding will take far longer than Silicon Valley thinks, and that slowness is exactly where the applied layer's value comes from. The conversation covers: — why application companies are the hottest neolabs, and why the LLM-wrapper thesis is finally working — the fox-guarding-the-henhouse problem with letting model providers route your tokens — work slop, and why we accept AI-written code but flinch at AI-written decks — how Box built its agentic harness and why it beats raw API access on accuracy and latency — the open-weights paradox: closed labs and open models both growing exponentially at once — what continual learning has to solve before it works for a lawyer with five matters and a Chinese wall — why 90% of enterprise tokens in five years will come from tasks no human kicked off — the mandate for founders right now: whoever gets it to the customer wins 0:00 – Introduction 1:55 – Are application companies the hottest neolabs? 6:56 – Will the labs move up the stack? 12:34 – Box and betting the company on AI 16:50 – Hero use cases: reading a million contracts and long-running agents 18:42 – Work slop: why AI code is embraced but AI content isn't 24:08 – Building Box's agentic harness and the evals that matter 27:23 – The state of the model race 29:25 – Open-weight model adoption in the enterprise 32:34 – Memory, continual learning, and what belongs in the weights 37:29 – Box Labs and systems of record in a world of agents 44:55 – Will chat be the dominant UI for enterprise AI? 48:00 – Why coding diffused fast and the rest of knowledge work hasn't 54:31 – Staying wired in, making a company AI-first, and what it takes to win

Sonya Huang 🐥

392,772 görüntüleme • 25 gün önce

One of the things I’m most excited about this year is building agents that can work productively for hours, days, or weeks. Coding agents are starting to become very competent at this, but what about computer use agents? Our new benchmark, Odysseys (co-led with Lawrence Jang) is a set of 200 new tasks derived from real world browsing behavior that measure long horizon web navigation capabilities (potentially up to hours of web browsing work). Interestingly, we find that frontier CUAs are already surprisingly good at working productively for up to an hour on these tasks, but there’s a lot of work to be done in making them even more efficient. Like every other AI researcher, my real dream is to open a cafe once we solve ASI. So, here’s Opus 4.6 doing some market research for me ("I want to do market research on the most popular cafes in Singapore. Analyse the menus of the top 10 cafes in Singapore (by Google reviews/ratings), and make sure we include at least 1 from the North/South/East/West/Central regions of Singapore. Keep the relevant pages of each cafe open, and summarise their pricing, menu offerings, unique selling points, making sure to reference which tab is opened for each cafe. For each cafe, also help me figure out how long it would take to get to it from Tampines MRT, and include this in your final summary."). I was very impressed to see Opus 4.6 complete this task after working for 52 mins, satisfying all 7 rubrics that corresponded to this task. It provided a very nice markdown summary at the end that gave me all the information I asked for!

Jing Yu Koh ✈️ COLM'26

50,482 görüntüleme • 5 ay önce

The power of the Claw, in the palm of a robot hand. Agentic robotics is here! Today, we open-source CaP-X: vibe agents, alive in the physical world. They incarnate as robot arms and humanoids with a rich set of perception APIs, actuation APIs, and auto synthesize skill libraries as they go. CaP-X is a strict superset of our old stack, because policies like VLAs are “just” API calls as well. It solves many tasks zero-shot that a learned policy would struggle with. And we are doing much more than vibing. CaP-X is our most systematic, scientific study on agentic robotics so far: - We build a comprehensive agentic toolkit: perception (SAM3 segmentation, Molmo pointing, depth, point cloud), control (IK solvers, grasp planner, navigation), and visualization (EEF, mask overlays) that work across different robots. - CaP-Gym: LLM’s first Physical Exam! 187 manipulation tasks across RoboSuite, LIBERO-PRO, and BEHAVIOR. Tabletop, bimanual, mobile manipulation. Sim and real. Can’t wait to see the gradients flow from CaP-Gym to the next wave of frontier LLM releases. - CaP-Bench: we benchmark 12 frontier LLMs/VLMs (Gemini, GPT, Opus, Qwen, DeepSeek, Kimi, and more) across 8 evaluation tiers. We systematically vary API abstraction level, agentic harness, and visual grounding methods. Lots of insights in our paper. - CaP-Agent0: a training-free agentic harness that matches or exceeds human expert code on 4 out of 7 tasks without task-specific tuning. - CaP-RL: if you get a gym, you get RL ;). A 7B OSS model jumps from 20% to 72% success after only 50 training iterations. The synthesized programs transfer to real robots with minimal sim-to-real gap. 3 years ago, our team created Voyager, one of the earliest agentic AI that plays and learns in Minecraft continuously. Its key ideas — skill libraries, self-reflection loops, and in-context planning — have since influenced many modern agentic designs. Today, the agent graduates from Minecraft and gets a real job. It’s April Fool’s, but this Claw is getting its hands dirty for real! Link in thread:

Jim Fan

82,861 görüntüleme • 6 ay önce

I vibe coded a new product on the side while running Every 🪨—and today we're launching it for free. It's called Proof, and it’s a live collaborative document editor where humans and AI agents work together in the same doc. It’s built from the ground up for the kinds of documents agents are increasingly writing: bug reports, PRDs, implementation plans, research briefs, copy audits, strategy docs, memos, and proposals. It's fast, free, and open source—available now at Why Proof? When everyone on your team is working with agents, there's suddenly a ton of AI-generated text flying around—planning docs, strategy memos, session recaps. But the current process for collaborating and iterating on agent-generated writing is…weirdly primitive. It mostly takes place in Markdown files on your laptop, which makes it reminiscent of document editing in 1999. That’s why we built Proof. What makes Proof different? - Proof is agent-native. Anything you can do in Proof, your agent can do just as easily. - Proof tracks provenance: A colored rail on the left side of every document tracks who wrote what. Green means human, Purple means AI. - Proof is login-free and open source: This is because we want Proof to be your agent's favorite document editor. How we use Proof Every 🪨: - Brandon Gell had OpenAI's Codex write a feature plan in Proof, then tagged my personal Claw (R2-C2) in Slack to review it. R2-C2 left feedback, I added comments, Brandon's agent revised the plan, and then Codex executed on it. Brandon submitted a PR to production without writing a line of code. - Austin Tedesco texts his Claw ideas while he's out on a run, then has it maintain a running Proof doc for his weekly food newsletter. He dictates drafts using Naveen Naidu's Monologue, writes into the outline himself, and uses the provenance gutter to track what's his voice vs. the agent's. - Kieran Klaassen uses it as a lightweight scratchpad for his compound engineering workflow. He brainstorms with an agent in the terminal, shares to Proof with one click, then opens the doc to leave comments and tells the agent to go work on them. His take: Proof's job is to communicate about writing and ideas. Proof is free, open source, and requires no login. I built the whole thing by vibe coding between meetings. I sat down with Brandon, Kieran, and Austin on Every 🪨's AI & I to demo it live and talk about how it's changing the way we work. If you're building with agents and need a better way to collaborate on text, this one's for you. Watch below! Timestamps Introduction and the origin story of Proof: 00:02:00 From Mac app to collaborative web editor: 00:07:24 What makes Proof "agent native": 00:09:00 Live demo—watching an agent join and write inside a shared document: 00:14:30 How Austin uses Proof for creative writing and food journalism: 00:20:51 The challenge of multiple agents editing one document simultaneously: 00:24:30 When AI-written docs are better read by agents than by humans: 00:26:48 Brandon's agent-to-agent collaboration loop: 00:29:30 Proof as a lightweight scratchpad versus existing tools like Notion and GitHub: 00:37:09 Why Proof is open source and what that means for builders: 00:42:18

Dan Shipper 📧

33,092 görüntüleme • 7 ay önce

Introducing LobeHub: Agent teammates that grow with you. LobeHub is the ultimate space for work and life: to find, build, and collaborate with agent teammates that grow with you. We’re building the world’s first and largest human–agent co-evolving network. Two years ago, we built LobeChat, an open-source interface for using different AI models. Today, LobeChat has 70k+ GitHub stars and serves 6M+ users worldwide. How to fully unlock the power of models has always been a shared mission between us and the community. We started with interaction — a fundamentally new, agent-first experience. Agents are no longer passive tools invoked in a single conversation. They should be proactive, always-on units of work. Treating agents as the minimal atomic unit is also the core of our agent harness infra. Today’s agents are mostly one-off executors. Even with memory, it’s often global — and hallucinates. We build long-term agent teammates that evolve with users. Each agent has its own dedicated memory space, editable by users, allowing humans and agents to co-evolve over time. This, in turn, allows us to design clearer rewards for reinforcement learning and create cleaner environments for continual learning. Agent teammates can work in groups. Through a multi-agent system, agent groups operate faster, more cost-effective, and go beyond what single-agent systems can achieve. For example, a single agent often requires heavy user involvement to proceed step by step, whereas LobeHub can execute the same work from a single instruction, with a supervisor orchestrating agents that run in parallel or debate to produce better results. We are building the collaboration network among agent teammates — and between humans and agent teammates as well. Ease of use matters. AI intelligence and shared human intelligence are equally important. With simple instructions and tool selection, you can effortlessly build and team up with agent coworkers to deliver complex, systematic work — even assembling a quant team to execute trades. Through the LobeHub community, anyone can discover, reuse, and remix agents and agent groups, customizing them to fit their own workflows, preferences, and needs. Last but not least, our vision started with LobeChat: multi-model support is the most efficient approach for users. We believe different models excel in different scenarios. By routing across multiple models, LobeHub improves cost efficiency and unlocks capabilities that a single-model setup cannot easily support.

LobeHub

185,466 görüntüleme • 8 ay önce

HERMES AGENT SUPPORTS 7 TYPES OF AI AGENTS. EACH ONE TAKES LESS THAN 90 SECONDS TO SET UP. MOST PEOPLE ONLY BUILD THE FIRST ONE. HERE ARE ALL SEVEN AND WHEN TO USE EACH. 1. BASIC AGENT WITH TOOLS your agent with access to terminal, browser, file system, web search, and calendar. it plans and executes tasks on its own. this is what you get on day one. "find flights to Lisbon under $400" "check my calendar and flag conflicts" "search the web for competitor pricing" set in Desktop app / Dashboard: Tools → enable what you need. when to use: single tasks that need tool access. 2. AGENT WITH MCP SERVERS connect your agent to external services. Notion, Google Drive, GitHub, Slack, databases, APIs, any MCP-compatible service. the agent doesn't scrape these services. it interacts through structured APIs. reads your Notion pages. creates GitHub issues. queries your database. sends Slack messages. set in Desktop app / Dashboard: MCP → Add Server. when to use: your workflow lives across multiple platforms. 3. SEQUENTIAL AGENTS (pipeline) one agent finishes. passes output to the next. assembly line for AI. agent 1: scans inbox for leads. agent 2: qualifies leads against criteria. agent 3: drafts outreach emails. in Hermes: cron jobs with wakeAgent gates. agent 1 writes output to a file. agent 2 wakes only when that file has new data. agent 3 wakes when agent 2 is done. each agent = a separate profile with its own model. when to use: multi-step workflows where each step depends on the previous one finishing. 4. PARALLEL EXECUTION AGENTS multiple agents working at the same time. results merge when all finish. "research these 5 competitors in parallel" in Hermes: delegate_task with batch mode. up to 3 sub-agents running in parallel by default. each gets its own clean context. only summaries return to the parent. delegation: model: "deepseek/deepseek-v4" children run cheap. parent synthesizes. when to use: independent tasks that don't depend on each other. research, data gathering, analysis. 5. AGENTS WITH ROUTERS conditions that send tasks down different paths based on the input. "if sales email → SDR profile. if support ticket → support profile. if calendar invite → EA profile." in Hermes: Kanban decompose. the decomposer reads profile descriptions and routes each task to the best-fit agent. or: Chief of Staff profile that triages and assigns to other profiles. when to use: incoming work that needs different specialists based on type. 6. HUMAN IN THE LOOP the agent does the work. asks for your approval before executing. "I drafted this email. approve before I send?" "this command will delete 3 files. proceed?" in Hermes: approvals.mode: manual (default). every dangerous action needs your confirmation. 60-second timeout. fails closed. or smart mode: LLM assesses risk. safe actions auto-approved. dangerous ones ask you. uncertain ones escalate. when to use: tasks where a mistake has real consequences. emails, deployments, financial transactions, public posts. 7. DYNAMIC SUB-AGENT SPAWNING your main agent realizes it needs help and spawns specialized sub-agents on the fly. "build this feature" → parent delegates: → sub-agent 1: research the API docs → sub-agent 2: write the code → sub-agent 3: write the tests in Hermes: delegate_task with role: orchestrator. raise max_spawn_depth for nested delegation. delegation: max_spawn_depth: 2 orchestrator_enabled: true depth 2 with concurrency 3 = up to 9 parallel workers. each level multiplies the spend. raise depth only when you need multi-level trees. when to use: complex tasks where the agent discovers what help it needs during execution. THE PROGRESSION: start with 1 (tools) and 6 (approvals). add 2 (MCP) when you need external services. add 4 (parallel) when tasks take too long one at a time. add 3 (sequential) when you build multi-step pipelines. add 5 (routing) when you run multiple profiles. add 7 (dynamic) when single-agent reasoning falls short. seven types. each under 90 seconds to configure. the value compounds as you stack them. comment AGENTS and I'll send you 3 ready-to-build agent setups that combine these types into real workflows.

YanXbt

17,312 görüntüleme • 2 ay önce

Anthropic's new model is extraordinary and it just revealed a problem that most enterprise AI buyers have not fully reckoned with yet (Save this), The model is genuinely impressive, and Chamath Palihapitiya assessment is that Anthropic continues to push the frontier harder than almost anyone. But that same update also showed their hand on something that changes the risk calculus for every business using Claude. Anthropic's new architecture stores every prompt you send for 30 days, no exceptions, not even for enterprise customers with zero-data retention agreements. The mechanism works like this, Anthropic now evaluates your prompt before generating output, deciding what it will and will not respond to, which means your query gets filtered before you even see a response. For individual users, that introduces a meaningful risk of censorship. For companies, Chamath says it is almost a non starter, and the reason is not just the data retention itself, it is the exposure that comes from operating at scale inside a large organization. A downstream scientist using the Claude APIs could accidentally trip a filter without knowing it, a business executive inside your company could trip it, and a molecular biology researcher could trip it and all of a sudden the company gets silently cut off from a tool it has embedded into critical workflows, with no warning and no recourse. Chamath gives Anthropic credit for being honest about how the system works, saying they tell the truth but notes that in this case the truth is not good. What this moment actually signals is a structural shift in how serious companies need to think about AI governance, because the question is no longer just which model performs best on benchmarks. It is who controls the model, who is learning from your data, and whether you are comfortable with a single point of failure sitting at the center of your competitive advantage. The answer for most enterprises will be broad model diversity, tighter governance frameworks and a serious reckoning with what it means to run mission-critical workflows through a third party that reserves the right to cut you off. Anthropic built a remarkable model and told the truth about how it works, the market's job now is to decide whether that transparency is enough to offset what the truth actually says.

Milk Road AI

30,183 görüntüleme • 3 ay önce