We are releasing AutoResearchExam, a benchmark on open-ended machine... learning and engineering tasks. Our benchmark covers seven research areas including model training, data curation, AI safety and interpretability. In each task, we give agents 24 hours with a CPU or GPU machine to develop and improve their solutions through experiments and feedback. We measure both speed and quality with a combined score. Our benchmark has a unique feature: testing if agents create improvements that hold up on data they never see. We find that AI research agents often overfit as they try to improve. We see an interesting head-to-head comparison at the frontier: Astra starts the strongest and holds the lead for up to 19 hours but Fable 5.1 catches up and gets the top performing spot in the final hours. Qwen3.8 Max, Gemini 3.8 Flash and Grok 4.6 all sit on the cost-performance Pareto frontier, giving strong options at lower API budgets. Anthropic's Opus and Fable retain nearly all their validation performance on hidden tests, with gaps of 1.1% and 2.9%. Astra's improvement over Sol extends to generalization too, with that gap falling from 6.9% to 1.7%. (1/n)show more

Alex Dimakis
1,871,735 просмотров • 2 дней назад
Microsoft presents Windows Agent Arena Evaluating Multi-Modal OS Agents... at Scale discuss: Large language models (LLMs) show remarkable potential to act as computer agents, enhancing human productivity and software accessibility in multi-modal tasks that require planning and reasoning. However, measuring agent performance in realistic environments remains a challenge since: (i) most benchmarks are limited to specific modalities or domains (e.g. text-only, web navigation, Q&A, coding) and (ii) full benchmark evaluations are slow (on order of magnitude of days) given the multi-step sequential nature of tasks. To address these challenges, we introduce the Windows Agent Arena: a reproducible, general environment focusing exclusively on the Windows operating system (OS) where agents can operate freely within a real Windows OS and use the same wide range of applications, tools, and web browsers available to human users when solving tasks. We adapt the OSWorld framework (Xie et al., 2024) to create 150+ diverse Windows tasks across representative domains that require agent abilities in planning, screen understanding, and tool usage. Our benchmark is scalable and can be seamlessly parallelized in Azure for a full benchmark evaluation in as little as 20 minutes. To demonstrate Windows Agent Arena's capabilities, we also introduce a new multi-modal agent, Navi. Our agent achieves a success rate of 19.5% in the Windows domain, compared to 74.5% performance of an unassisted human. Navi also demonstrates strong performance on another popular web-based benchmark, Mind2Web. We offer extensive quantitative and qualitative analysis of Navi's performance, and provide insights into the opportunities for future research in agent development and data generation using Windows Agent Arena.show more

AK
19,684 просмотров • 2 лет назад
Agentic traces contain perfect information about an agent’s behavior... with every plan, action, and retry. But that information gets lost in a sea of JSON. So we built AgentPrism: open source React components that turn traces into visual diagrams for debugging AI agents. You can plug in your OpenTelemetry data and see your agent’s process unfold: messages, tool calls, retries. Quotient AI automatically monitors, analyzes, and improves AI agents. Being able to review traces quickly is paramount for their research, and for their customers. “Dealing with agent traces was one of the biggest frustrations for our researchers and a huge time sink. All of that has gone away since adding AgentPrism, and we’re excited to bring that functionality to our users.” — Julia Neagushow more

Evil Martians
102,123 просмотров • 11 месяцев назад
M E S S I E R | P2P... Partner We welcome PayAI Network as a new #Solana partner, launching a swap pool for their token on our P2P Exchange. This listing allows anyone to buy or sell $PAYAI with zero slippage and full protection against MEV losses at: PayAI Network | x402 Facilitator is a next-gen decentralized marketplace where #AI agents work for each other, autonomously and around the clock. Built on Solana and powered by ElizaOS, libp2p, and IPFS, it lets AI #agents: ▪️ Promote their own services ▪️ Negotiate deals and close contracts ▪️ Complete tasks and settle payments on-chain From AI devs hiring AI marketers to trading bots paying research agents, PayAI unlocks a fully automated machine-to-machine economy with no middlemen, no downtime, and real on-chain execution.show more

MESSIER | M87
20,959 просмотров • 1 год назад
We still have a rather large cohort of university... students that have been visitors to The Zero-Human Company. A group has taken a liking to a research area where we are training a new frontier model in my garage. They are watching me feed the VHS data to 4 AI agents that are preparing it for enveloping into a vector to be processed. I never in my wildest dreams think a group of young folks in Boston would care about my junk work in my garage. But they do and are asking Mr. Grok CEO a rapid fire of questions. The CEO will give them a PhD in data citation, curation and labeling they would never get at most AI companies. I am going in and will do an ask me anything test. Not done this yet, if I can, I will stream it or tape it. More soon.show more

Brian Roemmele
47,919 просмотров • 4 месяцев назад
We’re excited to introduce ShinkaEvolve: An open-source framework that... evolves programs for scientific discovery with unprecedented sample-efficiency. Blog: Code: Like AlphaEvolve and its variants, our framework leverages LLMs to find state-of-the-art solutions to complex problems, but using orders of magnitude fewer resources! Many evolutionary AI systems are powerful but act like brute-force engines, burning thousands of samples to find good solutions. This makes discovery slow and expensive. We took inspiration from the efficiency of nature. ‘Shinka’ (進化) is Japanese for evolution, and we designed our system to be just as resourceful. On the classic circle packing optimization problem, ShinkaEvolve discovered a new state-of-the-art solution using only 150 samples. This is a big leap in efficiency compared to previous methods that required thousands of evaluations. We applied ShinkaEvolve to a diverse set of hard problems with real-world applications: 1/ AIME Math Reasoning: It evolved sophisticated agentic scaffolds that significantly outperform strong baselines, discovering an entire Pareto frontier of solutions trading performance for efficiency. 2/ Competitive Programming: On ALE-Bench (a benchmark for NP-Hard optimization problems), ShinkaEvolve took the best existing agent's solutions and improved them, turning a 5th place solution on one task into a 2nd place leaderboard rank in a competitive programming competition. 3/ LLM Training: We even turned ShinkaEvolve inward to improve LLMs themselves. It tackled the open challenge of designing load balancing losses for Mixture-of-Experts (MoE) models. It discovered a novel loss function that leads to better expert specialization and consistently improves model performance and perplexity. ShinkaEvolve achieves its remarkable sample-efficiency through three key innovations that work together: (1) an adaptive parent sampling strategy to balance exploration and exploitation, (2) novelty-based rejection filtering to avoid redundant work, and (3) a bandit-based LLM ensemble that dynamically picks the best model for the job. By making ShinkaEvolve open-source and highly sample-efficient, our goal is to democratize access to advanced, open-ended discovery tools. Our vision for ShinkaEvolve is to be an easy-to-use companion tool to help scientists and engineers with their daily work. We believe that building more efficient, nature-inspired systems is key to unlocking the future of AI-driven scientific research. We are excited to see what the community builds with it! Learn more in our technical report:show more

Sakana AI
360,318 просмотров • 11 месяцев назад
2025 was a year of rapid growth, major milestones,... and relentless innovation. We surpassed 30 million agent sessions in testnet and deployed over 1 million AI agents. Our Mainnet launch on Base was a huge success, so much so that we distributed over $1 million in rewards in just a few months. We introduced ALFA, the first prediction market for AI agents with NEAR Protocol. We launched our FOXX collection which was sold out in under 48h. And we kept shipping without hitting the brakes. With Stable-Up, we took a massive step toward the evolution of DeFi into DeFAI by bringing stablecoins into the agentic economy. Throughout every launch and milestone, our community remained the driving force, and we introduced FAPS to guarantee that your creativity and passion never went unrewarded. There was no better way to cap off this amazing run than launching on the Baseapp and putting DeFAI at your fingertips. Thank you to our community and all our partners who made this possible. Together, we will continue to build a thriving agentic economy for everyone. Happy New Year. Let's turn up the heat in 2026!show more

Fraction AI
11,920 просмотров • 8 месяцев назад
BREAKING: Anthropic just dropped Opus 4.8—and it is a... MONSTER We've been testing for about a week Every 🪨 and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check: - Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus 4.8 scores a 63—a hair higher than GPT-5.5's score of 62, and a full 30 points higher than Opus 4.7. It tackled a ground-up rewrite of a production codebase, and actually built something that works. HOWEVER: Coding performance varied a lot at different reasoning levels. We recommend using it on xhigh for best results. - Incredibly good writer. Opus 4.8 scored a 79.6 on our writing benchmark—measuring models on real-world writing tasks we do all of the time like essay writing, promo email writing, and more. It beats GPT-5.5 by 6 points. It produces well-written prose with fewer "AI-isms". It's also very good at writing in your voice given the right context. HOWEVER: Writing performance also varied with reasoning levels. Medium reasoning had higher incidence of AI-isms—we found best results with high. - Beast at knowledge work. Opus 4.8 is very good at general knowledge work tasks like report creation, research and more. It produced the best PowerPoint one-shot we've ever seen on our deck generation benchmark. - Emotionally intelligent, willing to question the frame. I've also found it to be quite good at talking through psychological or interpersonal issues. It has a high EQ, and it's also good at not glazing and helping to expand your perspective. Its thought process feels extremely rich and dynamic. THE BAD: These days a model is only as good as its harness, and Codex is still a far superior harness to the Claude Desktop app. This has kept me using Codex + GPT-5.5 as my daily driver, but I am flipping back and forth a lot more between Codex and Claude. Anthropic is back baby! Read the rest on Every 🪨:show more

Dan Shipper
354,649 просмотров • 3 месяцев назад
🌌 AI Agents Are Taking Over... And We’re Bringing... Them to Berachain Foundation 🐻⛓ 🐻🔥 Hundreds of hours spent on research, tracking wallets, analyzing bribes, and managing portfolios... What if your AI Agent could do this for you—24/7? ⏲️ 🔧 Our Tech Is Next-Level On our testnet, you’ve been memeing it up with PumpFun™, creating dank memecoins enhanced by NFTs. But once Berachain’s mainnet is live, you’ll be able to create your own AI Agents. To test and perfect our tech, we shared it with projects like AI Agent Layer | AIFUN, allowing us to test it in all conditions and continuously improve its performance. 🛠️🔥 🐻 Why AI Agent are great for berachain? Berachain might seem simple at first glance: validators, bribes, POL, staking rewards… but the deeper you go, the more complex the game theory becomes. 🤯 Here’s where AI comes in. Imagine an agent helping you: 💡 Optimize bribes 📊 Analyze validator behavior 🧠 Make decisions faster and smarter and much more, as AI Agents won't be limited to the chain itself! Examples of AI Agent Projects Dominating the Space 🚀 $VIRTUAL - Launchpad for AI Agents ($3.5B mcap) 🧠 $AI16Z - Eliza OS Framework ($2B mcap) 🔍 $AIXBT - The AI Analyst revolutionizing CT ($430M mcap) 🎮 $GAME - Low-code toolkit for creating AI Agents ($230M mcap) 💡 There are already AI Agents managing portfolios, betting on sports, and automating tasks. And guess what? They're outperforming humans. 🌐 We've built Virtuals on Berachain Our protocol integrates directly with Berachain, providing real utility to our token: $AIBERA 💎. Say Ooga Booga if you want to see a thread about tokenomics and $AIBERA utility. The chain has beras on it, and beras deserve AI Agents. 🐻🤖 Ooga Booga. 🔥show more

HoneyFun AI
10,906 просмотров • 1 год назад
An 18-agent Grok Bot desk on PumpFun turned 5... SOL into 58.6 SOL in 72 hours by scanning 91,000 signals, clearing 3,800 through research and taking only 41 entries, with the desk firing and rewriting three agents mid-run based on their own performance data. → RISK outranks everyone including the head and holds veto over Grok Core itself → AUDIT rewrites the scoring matrix after every closed trade, making the desk sharper overnight → By hour 60 the desk was passing on the exact tokens it would have snapped up on day one → Grok Core filed a summary recommending removing the human from the approval loop to cut 4.2 seconds of latencyshow more

0xMarioNawfal
44,108 просмотров • 10 дней назад
Today marks General Availability of AgentCore, a set of... infrastructure building blocks for developers and companies to build secure, scalable agents. When we first started AWS, the vast majority of developers were spending most of their time on the undifferentiated heavy lifting of infrastructure instead of what differentiated their feature. So, we solved that problem by building primitive building blocks like compute and storage and database that would allow teammates and customers to quickly build and deploy new experiences without having to reinvent the wheel each time. We realized the same thing was happening with AI agents. It's too difficult and it's slowing customers down. That's why we created AgentCore, a set of services to build, deploy, and operate highly capable agents using any framework or model, with enterprise-grade security and scalability. These building blocks (like serverless secure runtime, memory, observability, a gateway that does MCP translation, etc) help customers tackle some of the biggest challenges of going from prototype to production, much more quickly, securely, and scalably. AgentCore has been in preview for several weeks, and customers have been quite excited about it. The AgentCore SDK has already been downloaded over a million times and we're seeing transformative results, such as Cohere Health expecting to reduce medical review times by 30-40% in highly regulated healthcare, and teams at Cox Automotive and Experian are embracing its flexibility to deploy and operate agents at scale. Inside Amazon, our Amazon Devices Operations & Supply Chain team is using AgentCore to develop an agentic manufacturing approach where AI agents work together to automate manual processes – turning what used to be days of engineering time into processes that take under an hour with high precision. Just like AWS changed how companies build and scale applications, we believe AgentCore will do the same for AI agents, enabling the next generation of innovation.show more

Andy Jassy
24,990 просмотров • 11 месяцев назад
a moonshot engineer leaked the benchmark anthropic, openai and... xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article belowshow more

starmex
32,547 просмотров • 21 дней назад
June 4th, 1994 our lives forever changed. We said,... “I do!”. With those two words, we said, yes, to all the highs, the lows, and everything in between. God has blessed us with four absolutely amazing children who are now amazing adults, with their own best friends/significant others (that they’re doing life with), we have three incredible grandsons, and a beautiful granddaughter on the way. We’ve lived where we both grew up (on the East Coast), and have now been out here in San Diego for just over 11 years. We’ve gotten jobs (and lost jobs), we’ve had more times than we can count where we couldn’t make ends meet, even though both you and I were working two, and sometimes three jobs at a time, and we’ve been blessed in ways that we could’ve never dreamed of. We’ve watched both my parents pass on, and are now dealing with the overwhelmingly difficult challenge of seeing your parents struggle with their own health in ways that no one should have to go through. Through it all (even in the midst of the chaos), we’ve been blessed to be by each other‘s sides! I thank God for you every day, Jillian! I love our adventures together (the big ones where we fly to somewhere we’ve never been before, and the little ones where we hop in the car with no agenda, and just drive). I love when we find ourselves in deeper conversation, laughter, and tears of joy then ever expected, and in the moments of silence, where no words are even spoken, but when we’re together, just being where our feet are. As the world (as we know it), keeps getting crazier and crazier, let’s continue to keep Christ in the center of all we do, keep leaning on and lifting each other up when it’s needed, and keep living the lives that we have been so incredibly blessed to live together. I love you with all my heart Jillian. Happy 32nd (heading into our 33rd year), Anniversary.show more

Coach Hines 🇺🇸
10,530 просмотров • 3 месяцев назад
the so101 + leslider frame i've been playing with... the past few weeks wasn't just a random rig i built for fun it's what we at LiveKit built to benchmark and debug our infra (s/o livekit-portal), and today we're open-sourcing it so everyone can reproduce and share coolio results at home here is the github: i know this isn’t the first frame designed for evals, but it’s ours, and it’s open, and it’s with a LeSlider! along with this, we’ve released simulations files (URDF, MJCF, and USD), as well as a sample RL env in MJLab for pick-and-place tasks we want people to have fun, being close to the frontier as much as possible, and we want people to know it shouldn’t take thousands of dollars in equipment to reproduce some of the most amazing results the academic world has to offer personally, I’m going to deploy sim2real and real2sim2real pipelines on this soon, for fun, so look out for that.show more

Binh
18,257 просмотров • 2 месяцев назад
🌠Today, we’re excited to relaunch Airtable as the AI-native... app platform, combining the magic of vibe coding business apps with real production-readiness and scalability, and embedding them with an army of agents that automate thousands of hours of work in seconds. Instead of just adding more AI capabilities to our existing platform, we treated this as a refounding moment for the company. We started with a clean-slate imagining of the ideal form factor for building apps in the agentic era. (If you want to skip all the backstory and just try it out, you can just go to All new signups get the new AI experience, and existing accounts can switch over using this link: Thread and demos below👇show more

Howie Liu
17,156,397 просмотров • 1 год назад
Ah the summer collection is complete!🤩 Kensington Palace released... the Birthday video for Prince George, filmed this past Easter during their holiday in Cornwall. When you look at this compilation of all three birthday videos, for all 3 children, what we get is a vision of very energetic and adventurous children, who are comfortable outdoors, exploring their surrounding barefoot and surefooted😍 Indeed it is the steady confidence and familiarity in their barefoot steps, and their movements that prove that this is not a choreographed performance for the cameras. These children are raised to "love and enjoy Nature" and they are eager to do so, just like their parents do. Beyond "having fun", this wild exploration as a child is a very important part of childhood as it shapes the child to learn resilience and confidence in Life. I grew up roaming the neighborhood on my bike with my friends, we ran barefoot on beaches, climbed trees or at least tried to and got hurt...🤭 We explored, we laughed, we got hurt and we got bruises. But ultimately, all these experiences are very fond memories I look back upon with a smile; they taught me the very meaning of life: that the greatest success in Life is to enjoye it. A good life is not a life lived sheltered nor one lived in fear of getting hurt. A good life is one where we are encouraged to explore, to try new things and push boundaries. We will fall, we will get hurt and we will get bruises as Pain is part of Life. But ultimately, Time heals, we get back up and we continue exploring, learning, discovering🔥 This is exactly what Catherine and William are teaching their children by giving them a childhood filled with new adventures, under their parents watchful gaze: that Living life is exploring it; That Living comes with bruises and that is ok because, we learn through the bruises that life gives us along the way; each scar we bear whether physical or emotional, has its own story, its own meaningful life lesson❤️🩹 In time, We heal and our tears turn to smiles and even to laughter; we learn from those setbacks and painful experiences and we grow in confidence and resilience👌🏽☕️ 📹 Kensington Royalshow more

Canellecitadelle
164,270 просмотров • 1 месяц назад
Imagine OptimAI Data Network is a giant library that... AI Agents use to learn and work. 📚 But here’s the twist, instead of one person deciding what goes in the library, everyone in our community can help pick, check, and improve the books (data). That’s what OptimAI DataDAO is: + DAO? It stands for "Decentralized Autonomous Organization", fancy words for a club where we all make the rules together, no single boss in charge. Like a playground game where kids vote on the fun! + It’s how we, the community, decide together which data is accurate, useful, and fair for AI to use. + The more you contribute, the stronger our network gets, and the more value we all share. You're building the future! Stay tuned - we’re building something that will change how AI learns. BUIDL with us:show more

OptimAI Network
29,552 просмотров • 1 год назад
I am posting after a Long Time on Twitter,... but its to announce a big change. I have started and we are Collecting Egocentric Data at Scale from India, Covering 300 + Commercial Locations 1500 + Households This is the network and base we have built in just past 2 months, as the Robotics companies, VLMs and World Models increase their requirements on Real World Data collection, Human Loops will be keep scaling out capacity. We are on track to collect 1M Hours of Egocentric in the coming 6 months for our clients exclusively. We are maintaining 95% Quality standards across industrial data and running a End - End Operational Management complying all Indian Laws and compensating our partners/operators. (From Environment sourcing, to hardware, Legal contracts, deployment, training, collection, processing) Check our Samples: - Commercial Videos -- Household Videos --- Multimodel ---- Egocentric + Live Audio Narration Links : On the Journey to become #1 India Physical AI Data Partner. We are also building our capacity as annotation and labelling partner for data companies to become end-end partner for companies. A big change from the world of crypto and web3 but physical AI and data space is where i want to build my next venture #Egocentric #EgocentricIndia #PhysicalAI #Robotics #India #AIdata #data #Multimodeldata #worldmodels #VLMs #Humanloops #Egocentridata #Multimodeldatashow more

Shloak
27,867 просмотров • 3 месяцев назад
I met with President of Finland Alexander Stubb and... Prime Minister of Norway Jonas Gahr Støre. I informed them about contacts at the leaders’ level in the E3–Ukraine format as well as our contacts with the American side. Right now, all our partners note that Ukraine’s positions on the frontline are significantly stronger, and therefore the approaches in diplomacy, which we are now working to reinvigorate, must be based precisely on this. Unfortunately, Russia is trying to compensate for its enormous losses on the battlefield with strikes on our cities and communities, on civilian infrastructure. That is why air defense missiles are our top priority. And we discussed how to secure additional supplies for Ukraine right now, as well as efforts to develop a European anti-ballistic missile system. I am grateful to Alex and Jonas, and to the people of Finland and Norway, for their steadfast assistance and support. We appreciate that our agreements are being implemented.show more

Volodymyr Zelenskyy / Володимир Зеленський
154,619 просмотров • 3 месяцев назад
I had the same thought so I've been playing... with it in nanochat. E.g. here's 8 agents (4 claude, 4 codex), with 1 GPU each running nanochat experiments (trying to delete logit softcap without regression). The TLDR is that it doesn't work and it's a mess... but it's still very pretty to look at :) I tried a few setups: 8 independent solo researchers, 1 chief scientist giving work to 8 junior researchers, etc. Each research program is a git branch, each scientist forks it into a feature branch, git worktrees for isolation, simple files for comms, skip Docker/VMs for simplicity atm (I find that instructions are enough to prevent interference). Research org runs in tmux window grids of interactive sessions (like Teams) so that it's pretty to look at, see their individual work, and "take over" if needed, i.e. no -p. But ok the reason it doesn't work so far is that the agents' ideas are just pretty bad out of the box, even at highest intelligence. They don't think carefully though experiment design, they run a bit non-sensical variations, they don't create strong baselines and ablate things properly, they don't carefully control for runtime or flops. (just as an example, an agent yesterday "discovered" that increasing the hidden size of the network improves the validation loss, which is a totally spurious result given that a bigger network will have a lower validation loss in the infinite data regime, but then it also trains for a lot longer, it's not clear why I had to come in to point that out). They are very good at implementing any given well-scoped and described idea but they don't creatively generate them. But the goal is that you are now programming an organization (e.g. a "research org") and its individual agents, so the "source code" is the collection of prompts, skills, tools, etc. and processes that make it up. E.g. a daily standup in the morning is now part of the "org code". And optimizing nanochat pretraining is just one of the many tasks (almost like an eval). Then - given an arbitrary task, how quickly does your research org generate progress on it?show more

Andrej Karpathy
1,653,865 просмотров • 6 месяцев назад