Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Today we’re releasing Ramp SWE-Bench: a private, production-grounded coding benchmark created from real engineering problems we've faced at Ramp.

185,981 Aufrufe • vor 1 Monat •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Introducing ALE-Bench, ALE-Agent! Towards Automating Long-Horizon Algorithm Engineering for Hard Optimization Problems Blog: Paper: ALE-Bench is a coding benchmark primarily focused on hard optimization (NP-hard) problems. We developed this benchmark with AtCoder Inc., a leading coding contest platform company. What makes ALE-Bench unique is its focus on hard optimization problems that demand long-horizon and creative reasoning. It’s open-ended, in the sense that true optima are out of reach (NP-hard) and scores can continuously improve. We believe this benchmark has the potential to become one of the key benchmarks for reasoning and coding in the next generation. ALE-Agent is our end-to-end agent that we specifically designed for this challenging domain. In fact, our ALE-Agent has already built an impressive track record in the wild! In May 2025, our agent participated in a live AtCoder Heuristic Competition (AHC), alongside 1,000 other participants in real-time. AHC is considered to be one of the most challenging coding competitions in this domain. Our ALE-Agent achieved an impressive ranking of 21st out of 1,000 human participants in the competition (top 2%), marking a turning point for AI discovery of solutions to hard optimization problems with a wide spectrum of important real world applications such as logistics, routing, packing, factory production planning, power-grid balancing. We look forward to applying this technology to real industrial optimization opportunities. Building on the insights from this study, Sakana AI will continue to tackle the challenge of developing AI with even greater algorithm engineering capabilities. ALE-Bench Dataset: ALE-Bench Code: This research was conducted in collaboration with AtCoder Inc. (AtCoder). We are deeply grateful for their outstanding expertise and contributions in optimization and algorithms, which were invaluable in providing data, analyzing results, and enabling our AI agent’s participation in their contests.

Sakana AI

237,195 Aufrufe • vor 1 Jahr

Alibaba just released a coding model that hits 82 percent on SWE-Bench Verified. That is the highest score ever published for an open-source model. The weights are free. The license is Apache 2.0. You can run it today. The model is Qwen 4 Coder 32B. Here is what 82 percent on SWE-Bench Verified actually means. SWE-Bench Verified tests whether an AI can autonomously resolve real bugs pulled from real production GitHub repositories. Not synthetic exercises. Real open-source projects that real teams depend on. A model gets a bug report, reads the code, writes a fix, and either passes the test suite or it does not. At 82 percent, Qwen 4 Coder 32B resolves 82 out of every 100 real production bugs it is given. Without a human guiding it. On code it has never seen before. For comparison: Qwen 4 Coder 32B: 82 percent SWE-Bench Verified. Open source. Apache 2.0. Claude Fable 5: 80.3 percent SWE-Bench Pro. $10 input / $50 output per million tokens. Currently suspended. GPT-5.6 Sol: Competitive on Terminal-Bench. $5 input / $30 output per million tokens. An open-weight model that you can download and run for free just beat both of them on the benchmark designed to measure real software engineering capability. Here is the architecture. Qwen 4 Coder 32B is a 32 billion parameter dense model. Not a Mixture-of-Experts. Every parameter is active on every request. This matters for inference: a dense 32B model runs on 22 gigabytes of VRAM, which fits on a single high-end consumer GPU or a MacBook Pro with 64GB of unified memory. The smaller variant, Qwen 4 Coder 4B, runs at approximately 135 tokens per second on an M5 Max and fits inside 8 gigabytes of RAM. For a model with usable coding capability, that is a new bar for what fits in a single laptop. The training methodology continued Alibaba's approach of reinforcement learning on verifiable coding tasks. The model gets rewarded when its code passes tests. It gets penalized when it fails. Over millions of training steps, the model learns to write code that actually runs rather than code that looks plausible. License: Apache 2.0. Full commercial use. No attribution requirement. No revenue threshold. No monthly active user ceiling. Weights: Hugging Face, available today. Runs on: vLLM, Ollama, SGLang, and any standard GGUF-compatible inference engine. Qwen 4 32B also runs at approximately 135 tokens per second on an M5 Max chip, setting a new bar for what a sub-8GB model can do on Apple Silicon. The open-source coding model just beat the best closed-source model in the world on the benchmark designed to test whether AI can actually do software engineering. The weights are free. The subscription is optional. Source: Autom8Labs AI Insight July 2026, State of Open Source LLMs June 2026, Kunal Ganglani blog June 2026.

Harman

41,179 Aufrufe • vor 21 Tagen

It’s day 1905 at Ramp. Today, I’m thrilled to announce the launch of Ramp Travel, a new solution designed to make booking travel and managing travel expenses more intuitive, low-cost, and streamlined. Today, 1 in 5 (20%) dollars spent on Ramp cards go towards flights, hotels, and other trip-related entertainment - double the 10% of just a few years ago. Companies are hungry for travel as a means of uncovering new growth, but are hamstrung by outdated, expensive tools that employees hate, or consumer booking sites that might be cheaper and easier to use, but lack controls to enforce travel policies. Either way, hours are wasted at the end each month on manual expense reports and cumbersome reconciliations. But, these tradeoffs don’t need to exist. Companies deserve total control over their travel spend without being saddled with fees. And their employees deserve a delightful booking experience without the additional work. We’re excited to deliver on this vision with Ramp Travel. Our mission has always been to help companies spend less money and time, and this launch is a significant step towards that goal. Employees can book flights and hotels directly through Ramp, with real-time visibility into what’s in-policy and dynamically adjusted rates based on travel destinations. On the trip, Ramp AI handles the rest. On the ground, the quality of life improvement for employees is substantial. Receipts are automatically assigned to trips, and all expenses are seamlessly integrated into our platform. Put more simply, on other platforms, you have to do your expenses as you go (or 1-3 months later); on Ramp, your expenses do themselves. As the saying goes, “your margin is our opportunity.” We are not a travel company and aren’t in this to take price, we’re a savings company and are in this to deliver value, control, and a better experience. Our new partnership with priceline allows us to offer our customers access to a global selection of inventory from major airline partners and hotel affiliates at competitive rates, while maintaining the high level of control and visibility that Ramp is known for. Rather than ratcheting up the price through hidden fees, all savings are passed directly to our customers. We’re excited to see how Ramp Travel helps businesses streamline their travel processes, save money, and enhance the overall travel experience for their employees. As always, we love feedback, so try it out and let us know what you think.

Eric Glyman

336,719 Aufrufe • vor 2 Jahren