Загрузка видео...

Не удалось загрузить видео

На главную

Open-ended coding training data may no longer be the bottleneck: AI can scale open-ended tasks—and even outperform human-expert curation. FrontierCS team is releasing FrontierSmith: a system for synthesizing open-ended coding problems at scale. Starting from closed-ended coding tasks, FrontierSmith mutates, filters, and builds runnable optimization environments for long-horizon coding...

121,283 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 29

Фото профиля Alex Goldie
Alex Goldie4 месяцев назад

Nice work! Our paper from earlier this year might be relevant to you - DiscoGen already supports generation of ~100 billion open ended coding problems using PCG, and performance scales as you optimise over more tasks!

Фото профиля Qiuyang Mang
Qiuyang Mang4 месяцев назад

Thanks so much for sharing! This is really exciting and feels very complementary to our work. FrontierSmith focuses more on how to generate novel open-ended problem formulations from existing seeds.

Фото профиля Alex Goldie
Alex Goldie4 месяцев назад

Would be cool to see if there’s potential to combine these complementary axes!!

Фото профиля Qiuyang Mang
Qiuyang Mang4 месяцев назад

Code agents are getting very good at repetitive software work. Across research labs and startups, the next important question is increasingly about whether AI can solve open-ended optimization problems that matter in the real world: chip placement and routing, logistics, power-grid scheduling, database tuning, kernel optimization, and many others. But our previous FrontierCS on Harbor blog ( showed a clear weakness. Today's code agents are much less reliable in long-horizon, open-ended optimization than they are on traditional contest or math-style tasks.

Фото профиля Qiuyang Mang
Qiuyang Mang4 месяцев назад

We think a major reason is data. Classic RLVR settings have huge amounts of high-quality training data. Competitive programming alone has more than 100,000 public problems, and the broader coding-data industry is continuously producing more. By contrast, if we add together open-ended optimization benchmarks such as FrontierCS, ALE-bench, KernelBench, and the recent MLS-Bench, we still only get hundreds of tasks. That gap is the bottleneck FrontierSmith targets. Frontier labs may already understand the value of open-ended optimization, but without enough scalable training tasks, it is hard to run the kind of training that made closed-ended coding models so strong.

Фото профиля Qiuyang Mang
Qiuyang Mang4 месяцев назад

The core idea of FrontierSmith is simple: do not ask an LLM to invent high-quality open-ended problems from scratch. Start from closed-ended problems instead. Closed-ended coding tasks are already abundant. Given a LeetCode-style or competitive-programming-style problem, FrontierSmith applies principled mutations that turn it into a high-quality open-ended optimization problem.

Фото профиля Qiuyang Mang
Qiuyang Mang4 месяцев назад

Mutation creates many candidates, but not every candidate is useful. Some are still effectively closed-ended. Some are open-ended in wording but dominated by one obvious strategy. Our key filtering signal is idea divergence. We cannot ask an LLM to prove whether a problem is P or NP-hard, or whether an optimum is reachable under a fixed compute budget. We can, however, sample solutions from different solvers and ask whether they explore meaningfully different algorithmic ideas. Open-ended problems tend to produce diverse solution strategies. Closed-ended problems are often dominated by a single "gold idea."

Фото профиля Qiuyang Mang
Qiuyang Mang4 месяцев назад

The surviving ideas are converted into clean, runnable training environments. We reuse the FrontierCS judge sandbox and generate two pieces for each task. We evaluate on two open-ended coding benchmarks:FrontierCS, using the 172 algorithmic open-ended tasks. ALE-bench-lite, derived from AtCoder Heuristic Contest-style optimization tasks. For training, we synthesize 200 FrontierSmith problems and run GRPO on Qwen3.5-9B and Qwen3.5-27B. We compare against several controls: training on human-curated FrontierCS problems, training on ALE-bench, training directly on 200 closed-ended HardTests problems, and training on FrontierCS with random rewards. The results are direct. FrontierSmith-generated data is strong enough to match or exceed human-curated open-ended training data.

Фото профиля Qiuyang Mang
Qiuyang Mang4 месяцев назад

Huge thanks to all collaborators: @RunyuanHe @Zhoushang @KaiyuanLiu04 @lihanc02 @HuanzhiMao @qizhengz_alex Zerui Li @Builtin_Pb Lufeng Cheng @YichuanM @shangjingbo @AlexGDimakis @profjoeyg @alvinkcheung

Фото профиля Kexun Zhang
Kexun Zhang4 месяцев назад

lol thanks for featuring hardtests in your video haha

Фото профиля Qiuyang Mang
Qiuyang Mang4 месяцев назад

yes! Everyone should check out Hardtests if their research wants to benefit from large-scale high quality competitive programming data

Фото профиля Strata
Strata4 месяцев назад

@StanfordAILab oh nice, didn't know they were working on this

Фото профиля AiDevCraft
AiDevCraft4 месяцев назад

The "outperforms human curation" framing buries the lede — once tasks ship with runnable verifiers, humans lose because they can't iterate on the reward channel, not because synthesis is smarter. Real metric: closed-loop iterations per problem, not problems per hour.

Фото профиля Qiuyang Mang
Qiuyang Mang4 месяцев назад

I love the insights! Another aspect is the per-hour cost, human experts is very expensive in terms of curating open-ended challenges. ASFAK, for those open-ended problems in ICPC world challenges, it's about $1000 per problem-setter per hour

Фото профиля AiDevCraft
AiDevCraft4 месяцев назад

The $1000/hr number is actually the upper bound that synthesis breaks — but the deeper unlock is that synthetic pipelines pay once for the verifier and amortize across millions of variants, while ICPC setters re-pay the full rate per fresh problem. The cost curve flips from linear to fixed.

Фото профиля Strata
Strata4 месяцев назад

@ClementDelangue oh nice, didn't know they were working on this

Фото профиля web3nomad.eth | atypica.ai
web3nomad.eth | atypica.ai4 месяцев назад

interesting. in AI persona experiments at atypica we’re finding the opposite pattern. human-curated personality traits produce more authentic-feeling personas than AI-optimized ones. the AI version is more consistent but less believable. "outperforming curation" might depend heavily on what you’re curating for.

Фото профиля Hanchen Li
Hanchen Li4 месяцев назад

This paper is partially supported by Sky Lab funding and shout out to @modal for providing generous GPU resources alongside and @LaudeInstitute for the Slingshot funding!

Фото профиля OzAI
OzAI4 месяцев назад

Love this, exciting stuff! Good on ya to the FrontierCS team for making heaps of interesting coding challenges. Keen to try a few. Any public samples we can play with?

Фото профиля Victor
Victor4 месяцев назад

This matches what I keep seeing building synthetic training data: generation stopped being the bottleneck a while ago. Curation is the whole game now. The interesting claim here isn’t “AI can synthesize tasks”. It’s that synthesized + filtered beats human-curated. In my own pipeline only about a quarter of generated pairs survive filtering, and the filtered set trains a better model than 4x the volume unfiltered. Human curation doesn’t lose because humans are worse. It loses because humans can’t curate at the volume where the filter starts to matter.

Фото профиля ank R
ank R4 месяцев назад

nice work!

Фото профиля Qiuyang Mang
Qiuyang Mang4 месяцев назад

thank you!

Фото профиля Gerard Sans | Axiom 🇬🇧
Gerard Sans | Axiom 🇬🇧4 месяцев назад

This doesn’t work even in principle. AI doesn’t think, it computes. Searching across training samples brings no additional value; those examples are already encoded in the training data and lie directly on the inference manifold.

Фото профиля Qiuyang Mang
Qiuyang Mang4 месяцев назад

Yes, AI doesn’t “think” in the human sense; it often produces noisy outputs. So the key question is how to automate human expertise: how do we identify high-quality data from noisy AI-generated outputs, and use it to improve future AI systems? Humans naturally have a sense of which cat picture looks more realistic. In our work, we implement a similar idea for open-ended problems: we use a solution divergence metric to automatically select higher-quality data from AI-generated candidates.

Фото профиля Gerard Sans | Axiom 🇬🇧
Gerard Sans | Axiom 🇬🇧4 месяцев назад

Just be aware that your approach leads to a substantial drop in dimensionality while providing little to no added value. This stems from the high entanglement in the corpus and the generally low transferability across different training seeds and architectures, with no meaningful benefit even within the same model family.

Фото профиля Gerard Sans | Axiom 🇬🇧
Gerard Sans | Axiom 🇬🇧4 месяцев назад

If you could improve learning without using new samples we wouldn’t need pre-training. RL is not a solution to corpus gaps. What you need is more, better curated samples that go through a full training cycle. Understand the limits of RL first:

Фото профиля Qiuyang Mang
Qiuyang Mang4 месяцев назад

So how do you explain even for human world champion cannot distinguish the synthetic open-ended challenges and human-curated one in the real contests? Is that mean the human curated data is also not valuable? So how to explain the current gain leading by RLVR

Фото профиля Gerard Sans | Axiom 🇬🇧
Gerard Sans | Axiom 🇬🇧4 месяцев назад

You’re mixing two different things. Human evaluation is heavily skewed by psychological bias and competence gaps, you’re projecting assumptions about performance before measuring it. In general, it’s low quality and avoided. Pre-training and data curation are separate disciplines. In your case, the key distinction is data used in full pre-training vs. fine-tuning. Data curation’s impact on pre-training is a different metric from the efficacy of human curation (which carries the same biases). Two distinct variables: whether to use curation at all, vs. how well humans do it vs. no curation.

Фото профиля LandonCryptoExplr
LandonCryptoExplr4 месяцев назад

Synthetic data beating human curation. Data wall keeps looking like a temporary constraint.

Похожие видео

Introducing ALE-Bench, ALE-Agent! Towards Automating Long-Horizon Algorithm Engineering for Hard Optimization Problems Blog: Paper: ALE-Bench is a coding benchmark primarily focused on hard optimization (NP-hard) problems. We developed this benchmark with AtCoder Inc., a leading coding contest platform company. What makes ALE-Bench unique is its focus on hard optimization problems that demand long-horizon and creative reasoning. It’s open-ended, in the sense that true optima are out of reach (NP-hard) and scores can continuously improve. We believe this benchmark has the potential to become one of the key benchmarks for reasoning and coding in the next generation. ALE-Agent is our end-to-end agent that we specifically designed for this challenging domain. In fact, our ALE-Agent has already built an impressive track record in the wild! In May 2025, our agent participated in a live AtCoder Heuristic Competition (AHC), alongside 1,000 other participants in real-time. AHC is considered to be one of the most challenging coding competitions in this domain. Our ALE-Agent achieved an impressive ranking of 21st out of 1,000 human participants in the competition (top 2%), marking a turning point for AI discovery of solutions to hard optimization problems with a wide spectrum of important real world applications such as logistics, routing, packing, factory production planning, power-grid balancing. We look forward to applying this technology to real industrial optimization opportunities. Building on the insights from this study, Sakana AI will continue to tackle the challenge of developing AI with even greater algorithm engineering capabilities. ALE-Bench Dataset: ALE-Bench Code: This research was conducted in collaboration with AtCoder Inc. (AtCoder). We are deeply grateful for their outstanding expertise and contributions in optimization and algorithms, which were invaluable in providing data, analyzing results, and enabling our AI agent’s participation in their contests.

Sakana AI

237,195 просмотров • 1 год назад

Scale alone is not enough for AI data. Quality and complexity are equally critical. Excited to support all of these for LLM developers with Snorkel AI Data-as-a-Service, and to share our new leaderboard! — Our decade-plus of research and work in AI data has a simple point: scale alone is not enough. AI success is all about the quality, complexity, and distribution of data—in addition to volume. We’re excited to be powering leading LLM developers with Snorkel AI Expert Data-as-a-Service, our white glove service for custom, expert-level AI datasets—and to now preview some of what we’re building via our new Expert Data Leaderboard (🔗 in 🧵) + upcoming OSS dataset releases! Snorkel Expert Data-as-a-Service is built to meet the rapidly evolving data needs of the agentic AI world—where success is built on the quality, complexity, and distribution of datasets, in addition to size and scale. This kind of high-quality, frontier AI data can only come from a union of technology and human expertise. With Snorkel Expert Data-as-a-Service, we’re powering frontier LLM developers across agentic, expert knowledge, reasoning, coding, multi-modal, and other task types via the combination of these two key components: - (1) The Snorkel Expert Network: A global team of subject matter experts focused wholly on specialized knowledge–spanning thousands of topics in STEM/academic, vertical/professional, and consumer/lifestyle domains. - (2) Snorkel AI Data Development Platform: Our unique programmatic data curation and quality control platform, accelerating and improving expert authoring and review through principled techniques developed over the last decade of R&D. Now: we’re incredibly excited to showcase some of the power of Snorkel Expert Data-as-a-Service via the new Snorkel Leaderboard—putting frontier models to the test in complex, agentic, and reasoning settings inspired by real industry scenarios (not esoteric puzzles)! We’ll be releasing new leaderboards and accompanying expert-verified open source datasets (coming soon!) regularly. To start, we’re sharing three initial ones in preview: - SnorkelFinance: Q&A over financial documents requiring agentic tool-calling and reasoning - SnorkelUnderwrite: Agentic insurance tasks requiring industry-specific reasoning and tool use - SnorkelSequences: Mathematical tasks requiring compositional multi-step reasoning

Alex Ratner

495,851 просмотров • 1 год назад