正在加载视频...

视频加载失败

here's what i vibecoded today: punchingface 🥊 an app to make Hugging Face models fight each other on coding and canvas challenges, built with qwen3.6 35b a3b in 24 hours! benchmarks numbers don't mean anything anymore, we need a way to visualize what the models are actually capable of,...

19,153 次观看 • 3 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

AI is changing the software engineering craft. Anders Hejlsberg (Anders Hejlsberg) - creator of C#, TypeScript and industry legend - on why code review needs to get more enjoyable in response: #1 - AI is shifting the craft from writing code, to reviewing code: "In a sense, we're all turning into project managers. We can have an army of junior programmers, called agents, that will just spit out reams of code but someone's got to have the big picture and review all of that. And so, increasingly, our craft is going from one of writing the code, to one of reviewing the code and building the architecture of the code and overseeing the work. It's a different kind of craft. It's a different kind of enjoyment. I've always liked writing the code. To me that was the fulfilling part, seeing it work. In a way, AI robs a little bit of that, because I am less interested in reviewing code." #2 - The code review experience should be improved: "I think we could also make the process of reviewing code much more interesting than it is today. I mean, today, you see a list of diffs in alphabetical order and now it's up to you to make heads or tails of it. There are more pedagogical ways of presenting that. And you could have commentary generated by the AI that tells you what the changes are and whatever, and then tries to guide you along. So that symbiotic relationship, I think we need to work on that more and to keep the enjoyment in there."

The Pragmatic Engineer

39,011 次观看 • 2 个月前

i did it. i vibe-coded my own vibe-noding app⚡️ lately every ai tool has been announcing some kind of node canvas. they’re all impressive in their own way, but each one solves a different problem, with a different logic, in a different universe. some are too technical, some are too rigid, most feel like they were designed by a backend engineer. so i built the version i actually want to use. over the last few days i hacked this together with Google Gemini and Google Antigravity. the tool is absurdly capable. i didn’t touch a single line of code. the result is a fully working app with smooth ui animations, smart interactions, and a workflow that doesn’t fight me. my focus isn’t on the models. Nano Banana 2 Lite is already strong enough to generate half a universe on its own. and a year from now, we won’t care about model juggling anyway. we’ll have one model doing everything. the real problem is the workflow. the interface. the flow state. so i optimized for that. the whole point is to keep the technical junk out of the way during ideation and iteration. i want to try an idea, branch it, remix it, explore it, all on a single canvas. every block is editable. every idea is forkable. every path is visible. the design isn’t perfect yet. some of the ui patterns are familiar on purpose. good ideas deserve to be stolen. i rushed it. i let a bunch of great ui decisions die for speed. but it works. and honestly, it feels great to use. future of creating is more organic and more fluid. not just chat bubbles. not rigid feeds. but an intelligent canvas where ideas can grow, split, collapse, or evolve without friction. where the environment quietly adapts to your workflow instead of interrupting it. this is my first version. just for myself, for now. but it’s a start.

Synthetic_soul

239,298 次观看 • 8 个月前

I just compared Claude Code vs Codex vs Cursor CLI The task was to build a Next.js app with Tailwind 4 and shadcn components to collect customer feedback and showcase it with a widget. I gave all three the same prompt and let them go for 30 minutes to see what they came up with. Claude Code with Opus 4.1 Even though I told it to set up the app in the existing project folder, it tried to create a directory for it. After I interrupted and told it not to do that, it built a demo form and landing page with no errors. I had to ask it to make the demo interactive so users could submit a testimonial and preview it. The landing page looked like AI and was pretty basic, but it worked and it was done in a fraction of the time of the others. Total tokens used: 33k Codex with GPT-5 At the end of the 30 minutes I just could not get Codex to produce a working app. It got stuck in a loop of not being able to set up Tailwind 4 and despite many, MANY, attempts, I ended up with a "failed to compile" error. Total tokens used: 102k Cursor Agent with GPT-5 This was the slowest agent by far and a couple of times I actually thought it got stuck in a loop and was close to Ctrl+C'ing to cancel it. The TUI is really nice though, especially how it shows diffs and it did eventually build a working app (after one or two slight errors that needed fixing) The demo was interactive and it had a very minimal design that looked bare but also a lot less like an "AI generated" app than the Opus 4.1 design. It also wasn't too chatty and just did what it needed to do! Code quality was on a par with Opus 4.1, but it did use 5.5x as many tokens to get there. Still cheaper than Opus on a direct comparison but not when you factor in a Claude Code Max subscription. Total tokens: 188k I'll be able to do a proper comparison and record some videos when I'm back from holiday but for now, Opus is still the more capable model out of the box and Claude Code is the more complete CLI product. It will be interesting to see how Cursor evolve their CLI though with commands and subagents because I think with GPT-5 they have a real shot at providing competition for Claude Code if they can optimise output to get similar quality with less tokens. Jump to 0:40 in the video to see the two apps. Which do you think is which? ;)

Ian Nuttall

194,949 次观看 • 1 年前

First impressions on Muse Glimmer! It's incredibly fast for a dense model, currently running an average of 208tps with a max of 274tps on a single 5090 with their DFLASH config. Comparatively, though, both using Open Code, Qwopus Coder (with thinking off) produced a much better shark survival game than the one I got from Glimmer. Meta's new dense model is currently just lacking some HTML canvas taste, but this is something that can be added via SFT as long as the model is stable and capable from a back-end programming perspective. And it seems to be, without a doubt. The big kicker here is that I ran this at extra high thinking, and it did not take long at all to run. Our current local leader, Qwen 27B 3.6, has a tendency to overthink, but with glimmer, that is not the case. Right now, my recommendation for general local programming (Apps, Games, Websites, Visual Tools) in this class is still Qwopus Coder with thinking disabled, or Qwopus Fusion with thinking enabled. Of course Shark Survival is a very basic domain-specific test, but I find that the result scales very well across many domains. If we're going to be shipping apps generated entirely locally, visual taste is somewhat of a bare minimum requirement, solely in my opinion, and Qwen's models in this class offer significantly more at the moment. That's actually why I initially started getting into finetuning with Qwen 3.5, they were the first base that was able to do really good front-end with some opus-trace fine-tuning. Qwen 3.6 has taste even in the base model, and we know Qwen 3.8 is going to blow us all away! Regardless, this looks like a very tempting new base model. As a first offering from Meta in this class for a long time, I am incredibly impressed and elated to have it. We now finally have a proper Single GPU frontier race, instead of us just begging Qwen for more releases. Single GPU open frontier model race is a VERY good thing. Please keep pushing Meta

Kyle Hessling

16,191 次观看 • 5 天前

I have been testing DeepSeek-V4-Pro with the Pi coding agent. I am mindblown by how well it works out of the box. A few notes: I spent a few hours building an LLM wiki with an agent powered entirely by DeepSeek-V4-Pro on Fireworks inference. This is the first time I feel like there is an open-weight model that can reason at the level of Claude and Codex. And it does this in a cost-effective way with support for 1M context length. To be clear, I am using DeepSeek-V4-Pro inside of Pi without any special configuration. It works out of the box. It's exciting that there is a model that can just be plugged into a basic harness like Pi, and it just works. I've never seen that before. Most models require lots of configuration and setup. DeepSeek's DeepSeek-V4-Pro is clearly good at agentic coding (probably the best from the open-weight models), but the model is also great on knowledge-intensive tasks where reasoning matters. The agent pulled agentic engineering best practices from different company docs (Anthropic, OpenAI, Google, Stripe, Meta, Modal, DeepSeek, Mistral, Cohere), searched and digested Reddit and HN threads, summarized arxiv papers, and surfaced trending GitHub repos. Then it distilled everything into actionable tips across categories. I love the Wiki it built. The quality is really good. Here is a snapshot of what the wiki looks like: DeepSeek-V4-Pro handled the task without breaking stride. Multi-step research queries, code generation for scaffolding, context-heavy reasoning across disparate sources. For coding specifically, this is the first open-weight model that genuinely feels like a Codex or Claude Code experience. It compares in capability and actual multi-turn agentic work. What made the loop feel so responsive was Fireworks' inference speed (the fastest in the market) and the fact that they actually validate models at the systems level before shipping. No corrupted reasoning traces. Just fast, reliable iteration. The hybrid CSA and HCA attention design cuts KV cache to just 10% and inference FLOPs by nearly 4x at 1M-token context. This is what makes the agent loop actually fast and cheap enough to run in practice. For devs who've been watching open-weight models close the gap but haven't found one that actually delivers in practice, this is the closest I've seen. Try it here:

elvis

59,974 次观看 • 3 个月前

YOKO ONO: ONOCHORD, VENICE, 2004 Yoko: The world is divided in two industries. One is the War Industry and the other is the Peace Industry. The people in the War Industry are totally together. They don't have to talk to each other, even. They know exactly what they want to do. They want to go out there, kill and make money. But the people in the Peace Industry, which are us - we are so idealistic that each one of us criticises the other Peace Person in the Peace Industry. And we are always just arguing and we are wasting our energies doing that. So let's just forgive each other and see that we are in the Peace Industry and that's all that counts. Even if you are not marching for peace, just be yourself, being a florist, being a merchant, being a talior, anything. That way you're contributing to the Peace Industry. People are just concentrating on fear, confusion and anger. And therefore just for a moment, I'd like us to think about Love. In a very magical, straight way, John and I met in London and from then on we stood for Peace and Love. And when I do this kind of event. Well it is... I was inspired to do it, but I still think that I'm still with John in spirit. John and I created the country called Nutopia. Not Utopia, because there was Utopia as a concept already. And we wanted to create a new concept, so we just added N on it - Nutopia - and as a country. Well, that is the concept of a country. And we all are citizens of that country. And in my apartment in the Dakota Building, we put a little plaque on the back door, the kitchen door. It says 'Nutopian Embassy' and even now we have that. (laughs). Nutopia exists in our minds. And because of that, some people want to rebel against it. The reason some want to rebel against it is a good proof that it exists. I think that it was a terrible thing that happened in Chechnya. But we have to still keep our hopes up. And instead of giving up, we have to keep on sending the message of Love to each other. You say that I am the Ambassador of Peace. We are all Ambassadors of Peace. You are too. Everybody in this room are Ambassadors of Peace. Just the fact that we are not participating in War. The fact that we are here, and we are what we are, means that we are in the Peace Industry. All of us. John and I used to say that our apartment in the Dakota is a conceptual monastry, just for the two of us. And when we go out of the Dakota, we get so many people communicating with us, so it's very important that we had silence and quietness. And my apartment is a very small space compared to the world. And I need that for my peace of mind. You should be kind to each other. You should come together, hug each other, love each other, express our love to each other and we should make it work. We should finally create a world that is a totally an Earth for Us. So let's do it. Yoko Ono, OpenAsia Press Conference, whilst exhibiting Onochord, 2004 by Yoko Ono (Nutopia) at the Venice Biennale: OpenAsia 2004, Lido Di Venezia, Venice, Italy, 9 September 2004.

Yoko Ono

35,208 次观看 • 2 年前

Karpathy said something you'll regret ignoring: "We have to keep the AI on the leash. I'm still the bottleneck. I have to make sure this thing isn't introducing bugs and that there's no security issues." He said it at YC talk last year, when the worry was reliability. The models hallucinated and made mistakes no human would, so the leash implied keeping yourself in the loop and checking the output before trusting it. The models are far better now, and the line still holds, for a reason he was not focused on back then. Even a model that writes flawless code today still has no idea who is allowed to run it. Correctness and authorization are different problems, and only correctness improves as the model improves. A perfect agent still hands a tool where anyone can do anything, because permission was never part of the task. I actually tested this in practice with Claude Code. I asked it to build a small internal tool with a button that issues account credits. It worked first try, and running it locally, the credit applied the instant I clicked. Nothing decided who was allowed to click it. The agent wrote the right logic and displayed a success notification. It never checked whether the caller had the right, whether it should pause for a human, or whether anything was logged. And this is not a bug a smarter model can outgrow because the leash was never in the code. Identity, permissions, and audit live in the system that runs the app, not in what the agent generates. To solve this, I took the exact same bundle and hosted it on Retool. The credit write that fired silently on my laptop now stopped at an approval gate, resolved to a real identity through SSO, and landed in an audit log. I wrote none of it. The app inherited the entire boundary the moment it was deployed, and the video shows the before and after. You can try it yourself here: I also wrote a detailed breakdown of the whole thing in my recent article, and I worked with the team to put this together. It walks through the build, the exact moment the credit write went through on my laptop with nobody checking, and then what changed when the same app ran on Retool. It also covers why this is a property of the runtime and not something a better model fixes, which is why devs typically miss this. The article is quoted below.

Akshay 🚀

42,911 次观看 • 1 个月前

Pi was built when there were already agent harnesses around. Here’s why Mario Zechner(Mario Zechner), found them suboptimal and built Pi, a minimalist self-modifying agent: #1 - Mario initially was a believer in Claude Code: "I was a believer in Claude code because they were the first that packaged agentic search up in a really compelling package. And at the time that fit my workflow really well. Everything around the LLM was kind of nice and tidy and easy to understand. I was super happy. I was proselytising Claude code." #2 - Reverse engineering Claude Code highlighted the degradation that Mario felt as a user: "I personally like simple tools that are stable and that I can rely on. Even if they have non-deterministic parts, all the deterministic parts should be as stable as possible. That was just not the experience with Claude Code around summer 2025. They would take away your control of the context. They would inject stuff behind your back, which is bad. Then, your workflows stopped working because there's now a system reminder that you don't even see in the UI that would modify the behaviour of the model. They would also do this to the system prompt. I built a little service where I can track the progression or evolution of the system, prompt and tool definitions and, with every release, it was messing with stuff. That just messed with my workflows and I don't appreciate that." #3 - PI was built with an appreciation for simple and reliable tools: "If I commit to a development tool, I want it to be a stable, reliable thing like a hammer. I don't want my hammer to break a different spot every day. That's terrible. We need somebody who goes the full velocity kind of way. But I don't want to work with a tool like that."

The Pragmatic Engineer

62,825 次观看 • 3 个月前