Loading video...

Video Failed to Load

Go Home

We’re Scaled Cognition, developing the first ever models trained specifically for agentic applications: 1. Our first system, APT-1, is now #1 on agentic benchmarks. 2. It was developed by a US team for a total cost of less than $11M. 3. Khosla Ventures led our seed round ($21M closed...

94,532 views • 1 year ago •via X (Twitter)

9 Comments

Scaled Cognition's profile picture
Scaled Cognition1 year ago

We’re also announcing our Agent Builder platform, which allows you to create, test, and deploy an enterprise-grade AI agent using APT-1 in under an hour. Our GenAPI technology lets you test agent behaviors without needing to integrate with real APIs during development. Learn more about Agent Builder and GenAPI:

Scaled Cognition's profile picture
Scaled Cognition1 year ago

APT-1 currently outperforms all other models on the Tau-Bench and ComplexFuncBench agentic leaderboards, which test the ability to invoke sequences of complex APIs and comply with business policies.

Scaled Cognition's profile picture
Scaled Cognition1 year ago

We’ve accomplished these gains through (1) optimizing the model for actions rather than tokens, (2) a new kind of synthetic agentic training data, and (3) a novel RL approach using agent-to-agent self play. (1) Standard models are focused on token sequences, but business logic applies to actions (like policy rules governing API calls). Our models are trained to optimize licensed, well-formed actions, producing training regimes better suited to agentic tasks. (2) Training an agentic system requires data that contains both conversations and the actions that go along with them. Neither the web nor enterprise datastores have this kind of grounded data, so we generated it using a new, fully synthetic data pipeline. (3) RL through self-play has been a powerful tool in AI for domains like Chess or Go where win/loss outcomes are clear. We developed solutions for getting self-play to work in the agentic case, using simulated agent-to-agent interactions. Our approach applies to any agentic application and teaches the system to take actions correctly, subject to policies and instructions. Learn more and register for early access:

Karthik Narasimhan's profile picture
Karthik Narasimhan1 year ago

Very cool! Did you also compute the pass^k curves? Curious to see how reliable it is over multiple runs

Moritz Wallawitsch's profile picture
Moritz Wallawitsch1 year ago

Not sure if I understand what's happening here. Is the last action-token the model outputted in this screenshot “generate response”?

Christopher David's profile picture
Christopher David1 year ago

Congrats! Open source?

PhDPRIMA's profile picture
PhDPRIMA1 year ago

Impressive breakthrough in agentic AI! 🚀 APT-1 leading benchmarks with a fully synthetic RL-based pipeline is a game-changer. Excited to see its impact! 🔥 For PhD research support in AI/ML, check out @PhDPRIMA! 🎓✨ #PhDPRIMA

Anthony 😎🛹's profile picture
Anthony 😎🛹1 year ago

Yummy

DK's profile picture
DK1 year ago

which AI video recording did he use?

Related Videos

Scale alone is not enough for AI data. Quality and complexity are equally critical. Excited to support all of these for LLM developers with Snorkel AI Data-as-a-Service, and to share our new leaderboard! — Our decade-plus of research and work in AI data has a simple point: scale alone is not enough. AI success is all about the quality, complexity, and distribution of data—in addition to volume. We’re excited to be powering leading LLM developers with Snorkel AI Expert Data-as-a-Service, our white glove service for custom, expert-level AI datasets—and to now preview some of what we’re building via our new Expert Data Leaderboard (🔗 in 🧵) + upcoming OSS dataset releases! Snorkel Expert Data-as-a-Service is built to meet the rapidly evolving data needs of the agentic AI world—where success is built on the quality, complexity, and distribution of datasets, in addition to size and scale. This kind of high-quality, frontier AI data can only come from a union of technology and human expertise. With Snorkel Expert Data-as-a-Service, we’re powering frontier LLM developers across agentic, expert knowledge, reasoning, coding, multi-modal, and other task types via the combination of these two key components: - (1) The Snorkel Expert Network: A global team of subject matter experts focused wholly on specialized knowledge–spanning thousands of topics in STEM/academic, vertical/professional, and consumer/lifestyle domains. - (2) Snorkel AI Data Development Platform: Our unique programmatic data curation and quality control platform, accelerating and improving expert authoring and review through principled techniques developed over the last decade of R&D. Now: we’re incredibly excited to showcase some of the power of Snorkel Expert Data-as-a-Service via the new Snorkel Leaderboard—putting frontier models to the test in complex, agentic, and reasoning settings inspired by real industry scenarios (not esoteric puzzles)! We’ll be releasing new leaderboards and accompanying expert-verified open source datasets (coming soon!) regularly. To start, we’re sharing three initial ones in preview: - SnorkelFinance: Q&A over financial documents requiring agentic tool-calling and reasoning - SnorkelUnderwrite: Agentic insurance tasks requiring industry-specific reasoning and tool use - SnorkelSequences: Mathematical tasks requiring compositional multi-step reasoning

Alex Ratner

495,851 views • 1 year ago

Studies have shown ChatGPT outperforms human annotators for Structured Data by about 25% and costs 30x less. 1 In just 2 months, miners on SN33 running ChatGPT without optimization can’t survive. Today we announce SN33 is now ReadyAI to fully align with our mission 👇 SN33 is building a more performant and significantly cheaper alternative to Scale AI Today structured data is performed primarily by human annotation services like Amazon’s Mechanical Turk and Scale AI It is now more important than ever for every business and individual to make their data AI Ready. However, taking unstructured data and making it Structured Data using today’s tools is extremely costly. SN33 revolutionizes this process, unlocking immense opportunities for commercialization. We lay out the vision for it in this detailed blog post: Validators TODAY can monetize access to this structured data pipeline independently, but we’re streamlining this process, launching a frontend soon that any validator can opt into to provide bandwidth. We've received great feedback from the community, recognizing that what we're building goes far beyond Conversational AI. Building the world's largest annotated conversational dataset (which we've already accomplished) is just one of countless real-world applications for SN33's Structured Data pipeline. We're building a decentralized Scale AI, offering a full suite of Structured Data commodities—from text metadata tagging (available today) to fully customizable queries for company-specific data annotation use cases and image metadata tagging coming soon 👀. Thanks for all the feedback! It has been invaluable so keep bringing it to us! 🙏$TAO Openτensor Foundaτion 1 “ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks” shows “The zero-shot accuracy of ChatGPT exceeds that of crowd-workers by about 25 percentage points on average [...] Moreover, the per-annotation cost of ChatGPT is less than $0.003—about thirty times cheaper than MTurk”

David Fields

13,648 views • 2 years ago

Agentic AI will transform every enterprise–but only if agents are trusted experts. The key: Evaluation & tuning on specialized, expert data. I’m excited to announce two new products to support this–Snorkel AI Evaluate & Expert Data-as-a-Service–along w/ our $100M Series D! --- Snorkel Evaluate is our new data-centric agentic AI evaluation platform for specialized, mission-critical enterprise settings where vibe checks and out-of-the-box metrics driven by simple LLM prompts are not enough. Snorkel Expert Data-as-a-Service is our white glove service for expert-level AI datasets, powering frontier LLM developers in areas like expert knowledge, reasoning, agentic action and tool use, and more! Both built on top of Snorkel AI’s Data Development Platform, using our programmatic technology to drive higher-quality expert data, faster– for getting specialized AI to real production value. If you’re building enterprise AI and want to partner around the key ingredient in AI today–the data–book a demo and let's talk! Finally, see thread for details on 🧵👇 - 📽️ A walkthrough of Snorkel Evaluate and Expert Data-as-a-Service on an agentic AI enterprise task - 📅 An upcoming event on Enterprise Agentic AI with innovators from Accenture @BNY Comcast Stanford University QBE & others - 📊 An upcoming series of benchmark datasets and model artifact releases 👀 Want early access to the full agentic AI dataset? Retweet this post and we'll send you the link!

Alex Ratner

50,582 views • 1 year ago

Somite is building the blueprints to create any cell type for any person. We call it DeltaStem, our foundation model for the human cell. Our proprietary capsule technology allows us to generate cell signaling data 1,000x faster and more efficiently than current methods, accelerating discovery and optimization of cell production for treatment of diseases like type 1 diabetes, osteoarthritis, and DMD. DeltaStem allows us to reliably generate previously inaccessible cell types for the first time, unlocking a foundry of human spare parts. To scale DeltaStem and bring it into therapeutics, we’re excited to share that we’ve raised over $47M in Series A funding. Our round was led by Khosla Ventures, with participation from SciFi VC, the Chan Zuckerberg Initiative, Fusion Fund, Ajinomoto Group Ventures, Pitango HealthTech, TechAviv, Harpoon Ventures, along with angel investors such as Dr. R. Martin Chavez (former Chairman of Recursion), and Fidji Simo (outgoing CEO of Instacart and incoming CEO of Applications at OpenAI). Somite was founded and is led by a team of AI entrepreneurs and developmental biologists: Dr. Micha Breakstone, a repeat AI entrepreneur, Dr. Jonathan Rosenfeld, leader of the Fundamental AI Group at MIT, and a distinguished group of scientific co-founders representing institutions like Harvard Medical School and the National Academies of Science and Medicine. We look forward to collaborating with strategic partners who share our vision of leveraging AI for a new age in human repair. Reach out:

Cellular Intelligence

11,679 views • 1 year ago