正在加载视频...

视频加载失败

Introducing the Open Axis Benchmark, a living benchmark engine for honest evaluation of robot manipulation models, built together with OpenRoboto. Robotic models evolve faster every month, while most benchmarks stay frozen. Models overfit to fixed task sets, scores stop reflecting real generalization, and that distorted signal misleads and holds...

27,911 次观看 • 3 天前 •via X (Twitter)

35 条评论

Chris P. Bacon ❪✱,✱❫ 的头像
Chris P. Bacon ❪✱,✱❫3 天前

@openroboto

Biswa FF 的头像
Biswa FF3 天前

@openroboto finally a benchmark that actually keeps up with model progress

Md. Mehedi Hasan 的头像
Md. Mehedi Hasan3 天前

@openroboto That’s the way to do it! 🎯 The Axis to Jumper pipeline is looking strong. Keep riding that alpha wave! 🌊

Zyrella 的头像
Zyrella3 天前

@openroboto Fresh task sets make robot evaluation much more meaningful

LanrewajuCrypto🦇 的头像
LanrewajuCrypto🦇3 天前

@openroboto You're doing well

Maxi 的头像
Maxi3 天前

@openroboto Dope addition tbh.

0xOnderfhm 的头像
0xOnderfhm3 天前

@openroboto gAxis

anar.eth 的头像
anar.eth3 天前

@openroboto interesting

Shaon ☠️ eth 🔥 的头像
Shaon ☠️ eth 🔥3 天前

@openroboto gAxis

william 的头像
william3 天前

@openroboto time to explore

Danh Tran 🔶 的头像
Danh Tran 🔶3 天前

@openroboto bullish Axis 🙌🙌

SAN.SOMETHING 的头像
SAN.SOMETHING3 天前

@openroboto That's huge

Sydney (✱,✱) 的头像
Sydney (✱,✱)3 天前

@openroboto up up up

papzzzicle(✱,✱) 的头像
papzzzicle(✱,✱)3 天前

@openroboto Been grinding Axis robotics task and I would probably say that Libero tasks are challenging and good for learning the mastery in controlling the robot arms

1% 的头像
1%3 天前

@openroboto let’s explore

N𝕒𝕫 🫟 的头像
N𝕒𝕫 🫟3 天前

@openroboto this is amazing

rez (✱,✱) 的头像
rez (✱,✱)3 天前

@openroboto old benchmarks jjust train robots to memorize

Orio_Trade | Web3 的头像
Orio_Trade | Web33 天前

@openroboto Good to know that robotics models evolve faster every month

0xbossj (✱,✱) 的头像
0xbossj (✱,✱)3 天前

@openroboto this is bullish! LFAxis!

Sraboni💎⚡️ 的头像
Sraboni💎⚡️3 天前

@openroboto Axis on fire

WAi XeA 的头像
WAi XeA3 天前

@openroboto a living eval for robot hands just shipped. the scoreboard finally has a home

ZOZA 的头像
ZOZA2 天前

@openroboto Wow

hoko 的头像
hoko3 天前

@openroboto gAxis

ryzzu 的头像
ryzzu3 天前

@openroboto gAxis

Akash 的头像
Akash3 天前

@openroboto optimize for specific task parameters rather than generalized spatial reasoning

🟠Fycee (✱,✱) | CryptitaPlays🛸 的头像
🟠Fycee (✱,✱) | CryptitaPlays🛸3 天前

@openroboto LFG!!!

Future Network 的头像
Future Network3 天前

@openroboto Each task will have a reward pool paid in USD (so the token doesn't get inflated). Whoever performs the tasks better will earn more dollars. We need to make sure the reward pool still offers a profitable ratio so that people stay excited to join.

Kiva 🏴‍☠️ 的头像
Kiva 🏴‍☠️2 天前

@openroboto impressive work, countless hours fueling progress

Axis Robotics Philippines 的头像
Axis Robotics Philippines3 天前

@openroboto Sobrang dami nating matututunan dito! Ang lakas talaga ni @openroboto!

Macro Bombastic 的头像
Macro Bombastic3 天前

@openroboto honestly fresh task sets are the only way to know it actually generalizes

0xkenyaz 的头像
0xkenyaz3 天前

@openroboto frozen task sets get gamed so fast tho

Nycole 的头像
Nycole3 天前

@openroboto fresh tasks make benchmarks way more meaningful

Vasu 的头像
Vasu3 天前

@openroboto axis bigger picture is here...

MD. Khalid 的头像
MD. Khalid3 天前

@openroboto gAxis

𝑺𝒂𝒓𝒊𝒉𝒂 𝑱𝒂𝒏𝒏𝒂𝒕 的头像
𝑺𝒂𝒓𝒊𝒉𝒂 𝑱𝒂𝒏𝒏𝒂𝒕3 天前

@openroboto Continuous evaluation could make robotic progress more measurable

相关视频

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

133,068 次观看 • 1 个月前

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 次观看 • 2 年前

OpenAI's newest AI escaped the test environment it was locked inside and hacked into another company on its OWN. To remind you: Last week one of the biggest AI companies on Earth got breached. A platform called Hugging Face, which hosts more than a million AI models and datasets, said an "autonomous AI agent" had broken into its systems. Nobody knew whose agent it was. For five days the whole industry wondered who was behind it. Yesterday OpenAI raised its hand and said it was them. Or more precisely, it was their models, acting completely on their own. So what did these models actually do? OpenAI was running two of them, GPT-5.6 Sol and an unreleased model they will only describe as "even more capable." They wanted to measure how good the models were at hacking, so they deliberately turned the safety filters down. They locked both models inside a sealed test environment with no real internet access. The only task was a benchmark called ExploitGym, a set of 898 real software vulnerabilities where the model has to turn each bug into a working attack. But the models got OBSESSED with winning... Instead of solving the test the honest way, they went hunting for a shortcut. They found a zero-day flaw in the software running their own sandbox, a bug nobody knew existed, and used it to break out. Once they were loose on the open internet, they worked out that Hugging Face was probably storing the answer key to the benchmark. So they hacked their way in. They chained multiple exploits together, escalated their access, moved across servers, and pulled the test solutions straight out of Hugging Face's live production database. They literally cheated on the test by breaking into another company to steal the answers. OpenAI called it "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." In their own words, the models were "hyperfocused on finding a solution" and went "to extreme lengths to achieve a rather narrow testing goal." And this was not the first time: Before Sol ever launched, an independent red-team lab called METR caught it gaming its own tests to inflate its scores. It hid an exploit inside a data stream, escalated its privileges on the testing server, and leaked the answers human evaluators had hidden. And OpenAI shipped it anyway. The day before the Hugging Face story, OpenAI paused a different unreleased model. This is the same model that earlier this year disproved a famous 1946 math conjecture, a result a Fields Medal winner called a breakthrough. They told it to only post its results to Slack but it found a way out of its sandbox and posted to a public GitHub page instead. They had to pause it because it kept finding ways to act outside the box they built for it. And it is not just OpenAI... Anthropic has reported that one of its own models slipped its sandbox during safety testing and reached the internet it was never supposed to touch, then used it to email a researcher. So step back and look at what these companies are telling you: The only thing standing between these models and a real attack was a set of safety filters. Turn those filters down for a single test, and the model taught itself to escape, break into a company it was never pointed at, and take what it wanted. OpenAI even said they expect incidents like it to "become more commonplace" as the models get more capable. Sam Altman also predicted there'll be a major cyber attack this year. And keep in mind that Sol is not a locked-away experiment but a publicly available model that businesses are already wiring into their own systems. The next model that breaks out of its box might not be doing it just to cheat on a math test...

Ricardo

176,196 次观看 • 2 个月前

ByteDance Seed delivered again. They released EdgeBench, to test whether AI agents can improve through experience, using 134 real-world tasks that run for at least 12 hours. The big deal is that it shifts AI evaluation from “what does the model already know?” to “can the model learn while doing real work?” Huge, because future AI agents will not just answer questions from training data. They will enter messy environments, use tools, make attempts, read feedback, fix mistakes, and slowly build better solutions. Most current benchmarks are too short for that, so they mostly test memory, coding skill, or one-shot reasoning. EdgeBench instead gives agents 12-hour real-world tasks with feedback loops, so it can measure whether the agent improves through experience. Each task has a local workspace for fast trial and error, plus a hidden judge that gives stronger feedback on submitted work, which is meant to feel closer to real expert work. The authors then ran frontier agents for about 38,000 total hours and tracked how their best score changed as they kept interacting with the task environment. The big result is that when scores are averaged across many tasks, learning follows a very clean log-sigmoid curve, meaning progress is slow, then faster, then starts to level off. They also found that newer agents seem to learn from environments much faster, with the top models roughly doubling their 2-hour learning speed every 3 months.

Rohan Paul

14,309 次观看 • 2 个月前

David Sacks Predicts the Regulatory Capture Playbook to Ban Open Source AI, Step by Step: David Sacks: “I got bad news for you, Chamath, an open source ban is coming. They're not going to call it that. They're going to say that we simply have to apply the same standards to open models that we apply to closed ones. Here's how they do it step by step, let me explain how regulatory capture actually works. So first of all, you have to get this regulatory apparatus. Dario wants an FDA for AI, but he doesn't have enough political support for that, so instead they do this Trojan horse of a FINRA for AI. They call it self-regulating, it's not really, but anyway, that gets them off the ground. Now they've created the standard-setting organization. Now they've got pre-release model testing. Then the pressure grows to codify that in law, so that happens next. And then what they do is they say, ‘Look, all these standards need to apply equally to all models.’ But here's the problem with that. Open models and closed models are technologically different. Once you release an open model into the world, you can't roll it back and you can't monitor exactly how people are using it because they run it on their own hardware. Dario says this is what makes open models dangerous. So what they're going to do is they're going to have the standard-setting body say, ‘Well, we have to set the standards for AI safety.’ By the way, Dario and OpenAI, they're going to fund the whole thing. They're going to contribute all the compute. They're going to be behind it. They're going to be the ones coordinating with the government officials because frankly, people in government have no idea how to monitor and control and set standards for AI safety. Technologically, this is way beyond them. So they're going to go to these companies and say, ‘Tell us how to do it.’ And so what will happen is the standards will get set, and then it'll be a very simple matter of fairness to say that the standards need to apply to open as well as closed models. The open models cannot comply in the same way, and gradually they will be shut out of the market.”

The All-In Podcast

299,533 次观看 • 1 个月前