Загрузка видео...
Не удалось загрузить видео
Open-ended coding training data may no longer be the bottleneck: AI can scale open-ended tasks—and even outperform human-expert curation. FrontierCS team is releasing FrontierSmith: a system for synthesizing open-ended coding problems at scale. Starting from closed-ended coding tasks, FrontierSmith mutates, filters, and builds runnable optimization environments for long-horizon coding... show more
121,283 просмотров • 4 месяцев назад •via X (Twitter)
Комментарии: 29

Nice work! Our paper from earlier this year might be relevant to you - DiscoGen already supports generation of ~100 billion open ended coding problems using PCG, and performance scales as you optimise over more tasks!

Thanks so much for sharing! This is really exciting and feels very complementary to our work. FrontierSmith focuses more on how to generate novel open-ended problem formulations from existing seeds.

Would be cool to see if there’s potential to combine these complementary axes!!

Code agents are getting very good at repetitive software work. Across research labs and startups, the next important question is increasingly about whether AI can solve open-ended optimization problems that matter in the real world: chip placement and routing, logistics, power-grid scheduling, database tuning, kernel optimization, and many others. But our previous FrontierCS on Harbor blog ( showed a clear weakness. Today's code agents are much less reliable in long-horizon, open-ended optimization than they are on traditional contest or math-style tasks.

We think a major reason is data. Classic RLVR settings have huge amounts of high-quality training data. Competitive programming alone has more than 100,000 public problems, and the broader coding-data industry is continuously producing more. By contrast, if we add together open-ended optimization benchmarks such as FrontierCS, ALE-bench, KernelBench, and the recent MLS-Bench, we still only get hundreds of tasks. That gap is the bottleneck FrontierSmith targets. Frontier labs may already understand the value of open-ended optimization, but without enough scalable training tasks, it is hard to run the kind of training that made closed-ended coding models so strong.

The core idea of FrontierSmith is simple: do not ask an LLM to invent high-quality open-ended problems from scratch. Start from closed-ended problems instead. Closed-ended coding tasks are already abundant. Given a LeetCode-style or competitive-programming-style problem, FrontierSmith applies principled mutations that turn it into a high-quality open-ended optimization problem.

Mutation creates many candidates, but not every candidate is useful. Some are still effectively closed-ended. Some are open-ended in wording but dominated by one obvious strategy. Our key filtering signal is idea divergence. We cannot ask an LLM to prove whether a problem is P or NP-hard, or whether an optimum is reachable under a fixed compute budget. We can, however, sample solutions from different solvers and ask whether they explore meaningfully different algorithmic ideas. Open-ended problems tend to produce diverse solution strategies. Closed-ended problems are often dominated by a single "gold idea."

The surviving ideas are converted into clean, runnable training environments. We reuse the FrontierCS judge sandbox and generate two pieces for each task. We evaluate on two open-ended coding benchmarks:FrontierCS, using the 172 algorithmic open-ended tasks. ALE-bench-lite, derived from AtCoder Heuristic Contest-style optimization tasks. For training, we synthesize 200 FrontierSmith problems and run GRPO on Qwen3.5-9B and Qwen3.5-27B. We compare against several controls: training on human-curated FrontierCS problems, training on ALE-bench, training directly on 200 closed-ended HardTests problems, and training on FrontierCS with random rewards. The results are direct. FrontierSmith-generated data is strong enough to match or exceed human-curated open-ended training data.

Huge thanks to all collaborators: @RunyuanHe @Zhoushang @KaiyuanLiu04 @lihanc02 @HuanzhiMao @qizhengz_alex Zerui Li @Builtin_Pb Lufeng Cheng @YichuanM @shangjingbo @AlexGDimakis @profjoeyg @alvinkcheung

lol thanks for featuring hardtests in your video haha

yes! Everyone should check out Hardtests if their research wants to benefit from large-scale high quality competitive programming data

@StanfordAILab oh nice, didn't know they were working on this

The "outperforms human curation" framing buries the lede — once tasks ship with runnable verifiers, humans lose because they can't iterate on the reward channel, not because synthesis is smarter. Real metric: closed-loop iterations per problem, not problems per hour.

I love the insights! Another aspect is the per-hour cost, human experts is very expensive in terms of curating open-ended challenges. ASFAK, for those open-ended problems in ICPC world challenges, it's about $1000 per problem-setter per hour

The $1000/hr number is actually the upper bound that synthesis breaks — but the deeper unlock is that synthetic pipelines pay once for the verifier and amortize across millions of variants, while ICPC setters re-pay the full rate per fresh problem. The cost curve flips from linear to fixed.

@ClementDelangue oh nice, didn't know they were working on this

interesting. in AI persona experiments at atypica we’re finding the opposite pattern. human-curated personality traits produce more authentic-feeling personas than AI-optimized ones. the AI version is more consistent but less believable. "outperforming curation" might depend heavily on what you’re curating for.

This paper is partially supported by Sky Lab funding and shout out to @modal for providing generous GPU resources alongside and @LaudeInstitute for the Slingshot funding!

Love this, exciting stuff! Good on ya to the FrontierCS team for making heaps of interesting coding challenges. Keen to try a few. Any public samples we can play with?

This matches what I keep seeing building synthetic training data: generation stopped being the bottleneck a while ago. Curation is the whole game now. The interesting claim here isn’t “AI can synthesize tasks”. It’s that synthesized + filtered beats human-curated. In my own pipeline only about a quarter of generated pairs survive filtering, and the filtered set trains a better model than 4x the volume unfiltered. Human curation doesn’t lose because humans are worse. It loses because humans can’t curate at the volume where the filter starts to matter.

nice work!

thank you!

This doesn’t work even in principle. AI doesn’t think, it computes. Searching across training samples brings no additional value; those examples are already encoded in the training data and lie directly on the inference manifold.

Yes, AI doesn’t “think” in the human sense; it often produces noisy outputs. So the key question is how to automate human expertise: how do we identify high-quality data from noisy AI-generated outputs, and use it to improve future AI systems? Humans naturally have a sense of which cat picture looks more realistic. In our work, we implement a similar idea for open-ended problems: we use a solution divergence metric to automatically select higher-quality data from AI-generated candidates.

Just be aware that your approach leads to a substantial drop in dimensionality while providing little to no added value. This stems from the high entanglement in the corpus and the generally low transferability across different training seeds and architectures, with no meaningful benefit even within the same model family.

If you could improve learning without using new samples we wouldn’t need pre-training. RL is not a solution to corpus gaps. What you need is more, better curated samples that go through a full training cycle. Understand the limits of RL first:

So how do you explain even for human world champion cannot distinguish the synthetic open-ended challenges and human-curated one in the real contests? Is that mean the human curated data is also not valuable? So how to explain the current gain leading by RLVR

You’re mixing two different things. Human evaluation is heavily skewed by psychological bias and competence gaps, you’re projecting assumptions about performance before measuring it. In general, it’s low quality and avoided. Pre-training and data curation are separate disciplines. In your case, the key distinction is data used in full pre-training vs. fine-tuning. Data curation’s impact on pre-training is a different metric from the efficacy of human curation (which carries the same biases). Two distinct variables: whether to use curation at all, vs. how well humans do it vs. no curation.

Synthetic data beating human curation. Data wall keeps looking like a temporary constraint.
