Loading video...

Video Failed to Load

Go Home

Introducing ALE-Bench, ALE-Agent! Towards Automating Long-Horizon Algorithm Engineering for Hard Optimization Problems Blog: Paper: ALE-Bench is a coding benchmark primarily focused on hard optimization (NP-hard) problems. We developed this benchmark with AtCoder Inc., a leading coding contest platform company. What makes ALE-Bench unique is its focus on hard optimization...

237,195 views • 1 year ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

🚀New Amazon Q Developer agent for software development is available to customers: This agent is based on a new agent architecture that has exciting results coming from the SWE-bench scores (on the full and verified benchmarks) representing AI models’ ability to resolve real-world coding problems. Interesting aspect of Q Agent is that with these newest updates, Q drove nearly 50% more successful coding tasks completed. What makes Q Dev Agent remarkable? The agent architecture is not just about using the best LLMs (which we do), but also giving the agent the ability to constantly explore multiple paths to find the best way to resolve a particular problem (and back tracking when it has reached dead end like a developer would do). Needless to say, we are just getting started on the developer agent and we are constantly pushing to advance our AI capabilities while maintaining quality, security, privacy, and reliability to keep Amazon Q Developer an innovative and trusted option available to our customers using agents for software development. We highlighted the results of our first SWE-bench submission of Amazon Q Developer back in June blog post; with these updates, our new agent resolves 51% more coding tasks than its previous iteration on the SWE-bench verified dataset, and 43% more on the full dataset. That’s the difference a few months make, and I can’t wait to share what our teams will deliver at re:Invent this December. Here's a quick demo showcasing our new Agent in action:

Swami Sivasubramanian

28,946 views • 1 year ago

We are excited to announce a powerful step for the future of FOMO! Taking a page out of Virtuals book on BASE, FOMO will be releasing the ability for future projects to be paired in $FOMO in the coming weeks. This is the biggest release we have ever announced. Launch your AI Agent Token + $FOMO trading pair Every individual agent token is paired with the $FOMO token in its liquidity pool. When launching an agent on you will need $FOMO tokens, which are used to create the liquidity pool. This process creates deflationary pressure for FOMO and the entire agent ecosystem. When creating your agent and token, you will have the option to pair your launch with FOMO or SOL, as our goal is not to alienate any project, but rather invite the best communities, CTO’s and builders to launch with us. If you decide to pair your project with FOMO you in turn get full marketing and dev support, once your project graduates the bonding curve and reaches Raydium. Further, as an added incentive, as our revenue grows we will be using part of the funds to support projects that have paired in FOMO. And Devs who launch tokens paired in FOMO will earn fees from their AI Agent token launch. Building the most robust agents using our framework will catapult us as one of the most prominent standards of the Solana ecosystem. Not only have we developed our own core infrastructure, but we also pull from some of the best repo’s and developer talent in all of AI, not just blockchain. Our team is comprised of 9 world class artificial intelligence engineers, PHDs in mathematics and engineering from the top companies on the cutting edge of AI. The future of AI Agents will be on Solana and we will help lead the way.

FOMO

129,867 views • 1 year ago

What does the reputation model look like for agents? (alpha leak below) And how do we associate the proofs that we have about human beings with the agents who represent them? You may have heard of a process called KYC or Know Your Customer. That's very common with traditional financial applications and services. We have introduced a concept that we call KYA or Know Your Agent, which is a structured way to be able to express what model, how data was used in training, who the deployer is, what entities this agent instance is accountable back to, providing not only provenance but identity of the associated organization or entity. That's also another root of trust that we think about a lot: Enterprises and organizations tied back to things like their domains. To share a little bit of an alpha leak here, a product that we're excited to be rolling out in the next few weeks will allow our enterprise partners to more easily verify and prove the traits and capabilities of their teams as well as their counterparties. On the agent front, that makes it really easy to prove that an agent is acting on behalf of a given business or entity. We've already seen lawsuits where the absence of such technology has been a huge risk, such as with airlines that incorporate ChatGPT wrappers in their support pages. And then those AI enabled interactions end up making up plane tickets that don't exist and those airlines have to honor them. As small of an example as that might be, being able to prove agent accountability also unlocks a huge set of opportunities for use in enterprise for those agent to agent interactions. The Deep Trust Framework that our team has put together that we're excited to be bringing into a friendly SDK form in the next few weeks for some of our partners includes those reputation based capabilities, so how you can basically keep track of the interactions an agent has had, associate all of that to the entity to which they're accountable, and then that creates a sustainable reputation model for these agent to agent Interactions. Source: Billions CEO Evin McMullen evin speaking at House of Chimera Spaces Event Dec 3, 2025

Billions Network

68,503 views • 8 months ago

New short course: Evaluating AI Agents! Evals are important for driving AI system improvements, and in this course you'll learn to systematically assess and improve an AI agent’s performance. This is built in partnership with Arize AI and taught by John Gilhuly, Head of Developer Relations, and , Director of Product. I've often found evals to be a critical tool in the agent development process - they can be the difference between picking the right thing to work on vs. wasting weeks of effort. Whether you’re building a shopping assistant, coding agent, or research assistant, having a structured evaluation process helps you refine its performance systematically, rather than relying on random trial and error. This course shows you how to structure your evals to assess the performance of each component of an agent and its end-to-end performance. For each component, you select the appropriate evaluators, test examples, and performance metrics. This helps you identify areas for improvement both during development and in production. (If you're familiar with error analysis in supervised learning, think of this as adapting those ideas to agentic workflows.) In this course, you'll build an AI agent, and add observability to visualize and debug its steps. You’ll learn about code-based evals, in which you write code explicitly to test a certain step, as well as LLM-as-a-Judge evals, in which you prompt an LLM to efficiently come up with ways to evaluate more open-ended outputs. In detail, you’ll: - Understand key differences between evaluating LLM-based systems and traditional software testing. - Add observability to an agent by collecting traces of the steps taken by the agent and visualizing them - Choose the appropriate evaluator - code-based, LLM-as-a-Judge, human-annotation based - for each component. - Compute a convergence score to evaluate if your agent can respond to a query in an efficient number of steps. - Run structured experiments to improve the agent’s performance by exploring changes to the prompt, LLM model, or the agent’s logic. - Understand how to deploy these evaluation techniques to monitor the agent’s performance in production. By the end of this course, you’ll know how to trace AI agents, systematically evaluate them, and improve their performance. Please sign up here:

Andrew Ng

126,587 views • 1 year ago

The hardest problems in AI aren't research problems anymore. They're deployment problems. It’s how we actually deliver real value, today, to build the future people want. That’s why, after 20 years in AI, my next step was inevitable: make robots do useful work for and alongside people, right now. Today, I am delighted to announce the launch of Walden Robotics to tackle just that. We started this year and are coming out of stealth today with a $300M seed round backed by some of the most serious companies and investors in the world. They have seen firsthand our general-purpose robots being useful in production on day one, and getting better every day after. You can see a glimpse of what we've been building in the video below. Physical AI has gone through a rapid phase transition, in part thanks to pioneering research from my friends and co-founders Russ Tedrake , Ben Burchfiel , Siyuan Feng, Rareș Ambruș , and many others at Walden. But from our long experience working together with co-founders Kerri Fetzer-Borelli and Dave Johnson, we learned how hard it is to deploy cutting-edge AI in a real, live, incredibly sophisticated production environment with an intricate ballet of automation and human ingenuity. That’s why we deliberately created Walden Robotics as a full-stack, human-centric, customer-focused robotics company from the start: we seeded the company with a world-class team across hardware, software, AI, deployment, operations, product, and business talent, so we could continuously optimize our whole system end-to-end, deeply and purposefully, from real-world experience with real customers. The efficacy of this strategy speaks for itself: since February, our general-purpose robots have been doing useful work in production at a Toyota plant in North America, moving from first pilot to real work in under two months. Not a lab. Not a demo. Not a future promise. Real work on a real line, today, at one of the best large-scale manufacturers in the world, with general-purpose robots that get better every day. And this is just the beginning. Two ways to find us: If you run a manufacturing or logistics business and want robots that are widely useful now, not someday, let's talk. We own “ for a reason! And if you want to build them: we're hiring across the company, from software, to hardware, AI, ops, product, business, and more. In particular, as the Chief Strategy Officer at Walden, I am recruiting for three incredibly impactful founding roles to fuel our agent-native go-to-market engine. Check out Let’s build together!

Adrien Gaidon

64,278 views • 1 month ago

Maple is preparing for the release of a co-working agent. You install it locally and it works with your files, whether it's office work or building websites and apps. It's a turnkey solution, as easy as Claude Code, that keeps your data secure and private, no data sharing with closed AI labs. This is THE sovereign AI app for individuals and businesses who want powerful AI while retaining ownership of their information. Why build an agent into the Maple app when other agents already exist? Easy, we want to give you control over your work. We don't have a business plan that incorporates making money off our users' data. In the age of AI, your information, whether it's personal or company trade secrets, is the single thing that differentiates you from everyone else. We all have access to AI that can build a professional website for selling shoes. But your strategy and network for how you sell shoes should not be shared with your competitors. Sovereignty is the path to protecting what makes you, you. Maple sits at the intersection of Usability and Sovereignty. Maple gives you the best tools that are both easy to use and maintain your data sovereignty. Sovereign for one, sovereign for all. It has been a journey to get here. We brought to market the very first personal chatbot with end-to-end encryption using TEEs in late 2024. Prior to that there were proofs of concept but no full product offerings. Every other AI chat product on the market handled your data in plain text, either selling you a service to get your data or asking you to trust that they won't snoop on you. Quickly people found Maple and latched onto its open-source code and verifiable encryption. We didn't stop there. You may remember earlier this year we teased a product called "Maple Agent" and opened up a waiting list. That product is a mobile app that acts as your AI "friend", maintaining one long continuous chat, and getting to know you over time. I dislike using the word "friend" there, but it's the best way to convey the UX in a few words. AI is a tool, always has been, always will be. Any kind of friendly personality on top is just synthetic. In our testing, the UX of Maple Agent is really powerful for what it does. Think about the many short AI chats you have in your favorite app, whether it's looking up a historical fact or asking advice about a topic. With Maple Agent, those all go away in favor of the long-running chat with the friendly agent. It's like you have your own personal assistant who knows you so well and can look up anything for you. When I ask AI certain questions, I want to ask an expert who already understands my situation so I'm not repeating myself for the 100th time. That's the amazing value the personal agent brings to the table. We still see great utility for a personal agent like the "Maple Agent". Thousands of people on the waiting list, hoping to get their hands on it, agree that the concept is worth exploring and trying out. We were constrained in launching it due to a few circumstances, one of them being access to the scale of compute needed to power it. We have a clear path laid out for how to get there, but today is not the day to execute on that. It will be in the near future. Instead we have a different agent ready to go that we think is also incredible. We now have an agentic harness inside of the Maple Research app. This thing is a powerhouse. It even builds and publishes its own software releases. The agent in Maple Research works with your local filesystem, speaks to the largest open models running in TEEs, utilizes local models for certain tasks, is compatible with MCP tools, has an API for connecting to anything you need, and also supports the ACP protocol, which means it can be extended in the future to speak to other tools like Claude Code, Codex, and local models running on your own hardware. A big unlock for us was the Goose Development Kit, which powers the core of our agent harness. More on that to come as we publish articles and documentation later about the agent. The agent inside Maple Research doesn't have a name. At least not yet, not sure if it ever will. For now we call it "Chat Mode" and "Agent Mode". Think of this as the workhorse, the truck, the heavy lifter. Our other "Agent", the phone app, is your sidekick in your pocket, ready to help with quick things and ongoing conversations about life. I am incredibly excited about the Maple Research Agent. While I'm already seeing great results using it for internal work items, I'm especially thrilled about the personal health and wellness work it's doing for me. I know there are plenty of apps out there for compiling wellness data, but I'm having it build a tool tailored specifically for what I need, without the extra fluff. And none of my health data is being donated to the closed AI labs or sent to advertisers. I know that the AI logic is not being silently adjusted to fit the whims of a large corporation that has paid for product placement. It's me, state of the art AI, and my data. That's how I want it. Maple's new agent makes that possible. We can't wait for you to try it out. If you want early access, comment here, email us, reach out in some way. To those on the other agent waitlist, you're already in the queue. Thanks for reading this lengthy update. :)

Mark

44,707 views • 1 month ago