正在加载视频...

视频加载失败

I'm excited to share that we've built the world's most capable AI software engineer, achieving 30.08% on SWE-Bench – ahead of Amazon and Cognition. This model is so much more than a benchmark score: it was trained from the start to think and behave like a human SWE.

820,449 次观看 • 2 年前 •via X (Twitter)

11 条评论

Alistair 的头像
Alistair2 年前

I’ve actually been talking about building Genie for a while (photo from Dec 2022), it just wasn’t feasible from a technical point of view - it’s only been in the last 6-8 months that the models have been in a position to make it a reality

Alistair 的头像
Alistair2 年前

If you’re curious as to how we built the model there’s more info in our technical report:

uɐɥdǝʇS 的头像
uɐɥdǝʇS2 年前

You mean it can't talk to women and gets all sweaty if you call it on the phone?

Alistair 的头像
Alistair2 年前

Those were actually key characteristics we spent a good amount of time training in and eval’ing

Mckay Wrigley 的头像
Mckay Wrigley2 年前

Would love to test it out!

Alistair 的头像
Alistair2 年前

Just replied to your DM 👌

WeGPT.ai | Prompting is Better Together 的头像
WeGPT.ai | Prompting is Better Together2 年前

🤔 what does this mean? Does this mean you had privileged access to models nobody else did. So while we were fine tuning GPT-3.5 in this exact manner, to achieve the actual industry leading results, you’ve been fine tuning GPT-4?? During this same span of time, we were de-listed “erroneously” FIVE times from the GPT Store, banned from the OpenAI Developer forums, and the subreddit, as we sought to share our story about creating innovative useful AI apps. Your RAG method, as described, is identical to ours. And our has been live — and partially built — by GPT-4. Did you happen to have access to that code base at any time while OpenAI employees helped launch you into @ycombinator, we actually wrote @garrytan back in February of 2022 when we repo:rd we had something special on our hands. Ok that’s a bunch of questions I know. Your turn. Let’s hear your side of the story.

Berat 的头像
Berat2 年前

he is genuinely the first non-evil looking tech ceo i've ever seen, great product btw. congrats!

Alistair 的头像
Alistair2 年前

I’ll take compliments where I can get ‘em, appreciate it!

Harrison Kinsley 的头像
Harrison Kinsley2 年前

Came for genie, stayed for the Trump impressions. Would be very interested in checking genie out. Have signed up to the waitlist.

Alistair 的头像
Alistair2 年前

I had to make sure there was something there for everyone 🤣, maybe we’ll release the bloopers from that shoot at some point as they’re hilarious- thanks for signing up!

相关视频

🚀New Amazon Q Developer agent for software development is available to customers: This agent is based on a new agent architecture that has exciting results coming from the SWE-bench scores (on the full and verified benchmarks) representing AI models’ ability to resolve real-world coding problems. Interesting aspect of Q Agent is that with these newest updates, Q drove nearly 50% more successful coding tasks completed. What makes Q Dev Agent remarkable? The agent architecture is not just about using the best LLMs (which we do), but also giving the agent the ability to constantly explore multiple paths to find the best way to resolve a particular problem (and back tracking when it has reached dead end like a developer would do). Needless to say, we are just getting started on the developer agent and we are constantly pushing to advance our AI capabilities while maintaining quality, security, privacy, and reliability to keep Amazon Q Developer an innovative and trusted option available to our customers using agents for software development. We highlighted the results of our first SWE-bench submission of Amazon Q Developer back in June blog post; with these updates, our new agent resolves 51% more coding tasks than its previous iteration on the SWE-bench verified dataset, and 43% more on the full dataset. That’s the difference a few months make, and I can’t wait to share what our teams will deliver at re:Invent this December. Here's a quick demo showcasing our new Agent in action:

Swami Sivasubramanian

29,013 次观看 • 2 年前

Alibaba just released a coding model that hits 82 percent on SWE-Bench Verified. That is the highest score ever published for an open-source model. The weights are free. The license is Apache 2.0. You can run it today. The model is Qwen 4 Coder 32B. Here is what 82 percent on SWE-Bench Verified actually means. SWE-Bench Verified tests whether an AI can autonomously resolve real bugs pulled from real production GitHub repositories. Not synthetic exercises. Real open-source projects that real teams depend on. A model gets a bug report, reads the code, writes a fix, and either passes the test suite or it does not. At 82 percent, Qwen 4 Coder 32B resolves 82 out of every 100 real production bugs it is given. Without a human guiding it. On code it has never seen before. For comparison: Qwen 4 Coder 32B: 82 percent SWE-Bench Verified. Open source. Apache 2.0. Claude Fable 5: 80.3 percent SWE-Bench Pro. $10 input / $50 output per million tokens. Currently suspended. GPT-5.6 Sol: Competitive on Terminal-Bench. $5 input / $30 output per million tokens. An open-weight model that you can download and run for free just beat both of them on the benchmark designed to measure real software engineering capability. Here is the architecture. Qwen 4 Coder 32B is a 32 billion parameter dense model. Not a Mixture-of-Experts. Every parameter is active on every request. This matters for inference: a dense 32B model runs on 22 gigabytes of VRAM, which fits on a single high-end consumer GPU or a MacBook Pro with 64GB of unified memory. The smaller variant, Qwen 4 Coder 4B, runs at approximately 135 tokens per second on an M5 Max and fits inside 8 gigabytes of RAM. For a model with usable coding capability, that is a new bar for what fits in a single laptop. The training methodology continued Alibaba's approach of reinforcement learning on verifiable coding tasks. The model gets rewarded when its code passes tests. It gets penalized when it fails. Over millions of training steps, the model learns to write code that actually runs rather than code that looks plausible. License: Apache 2.0. Full commercial use. No attribution requirement. No revenue threshold. No monthly active user ceiling. Weights: Hugging Face, available today. Runs on: vLLM, Ollama, SGLang, and any standard GGUF-compatible inference engine. Qwen 4 32B also runs at approximately 135 tokens per second on an M5 Max chip, setting a new bar for what a sub-8GB model can do on Apple Silicon. The open-source coding model just beat the best closed-source model in the world on the benchmark designed to test whether AI can actually do software engineering. The weights are free. The subscription is optional. Source: Autom8Labs AI Insight July 2026, State of Open Source LLMs June 2026, Kunal Ganglani blog June 2026.

Harman

41,278 次观看 • 2 个月前

Elon Musk on Grok: “I think an AI that is sort of SpaceX’s baby will be a very good AI. I think SpaceX is a collection of some of the very best humans on Earth, both in capability and morality and goodness, and I think we want to have an AI that is born of that capability and that morality and that goodness. I think that’s very, very important. So I’d like to encourage people to help out, actually, with the AI hardware on the ground and the AI software and, of course, with the AI satellites. Solar is going to be a very important part of that. The TerraFab will be a very important part of that, but we must also succeed on the software front, and we’re going to be training Grok on the sum total of all SpaceX information. So in a way, it will be trained on you. You will effectively be the parents of the AI. It will inherit your thoughts and ideas and beliefs, and I think that’s a good thing. So, yeah, we’ve got to win here on the AI hardware and the AI software. That maybe is the most important message. I think that is actually the most important message I wanted to convey today is that we must win on AI because the future is overwhelmingly AI and robots. So we won’t ultimately be able to control the AI. It’ll be too smart for that. But just like if you have a child that is a super genius child, you can still instill in that child the values and beliefs that you think are good and right. And so that’s why it’s incredibly important that we succeed with Grok. So, yeah, well, Grok 4.5 you’ve tried probably. We’ve got 4.6 coming out in about a week. And then 4.7 should be really pretty special. So I’d like to encourage everyone at SpaceX to use AI and to make it better.”

DogeDesigner

64,434 次观看 • 1 个月前