Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Devin: web-based agent, designed to feel like hiring a software dev - used to cost $500/mo 😱, but now only $20/mo - autonomously completes tasks end-to-end, then opens PRs - easy to spin up multiple Devins simultaneously - Gumroad says Devin writes 41% of their PRs

20,790 Aufrufe • vor 1 Jahr •via X (Twitter)

13 Kommentare

Profilbild von Indie Hackers
Indie Hackersvor 1 Jahr

AI Coding Tools in a Nutshell 📕 April 2025, Developer Edition The 9 most powerful tools developers are using to go from 1x to 10x (bookmark this for your toolkit):

Profilbild von Indie Hackers
Indie Hackersvor 1 Jahr

Codex CLI: @OpenAI's new coding agent, runs in the terminal - released *YESTERDAY* so the jury is still out - uses OpenAI's powerful new o3 and o4-mini models 🔥 - completely open source - multimodal, supports passing in screenshots and diagrams

Profilbild von Indie Hackers
Indie Hackersvor 1 Jahr

: AI IDE, fork of VS Code - very popular, strong community + tutorials - gold standard AI tab completion, ⌘K inline editor 🥇 - powerful agent kicked off vibe coding trend - $20/mo unlimited plan, limits tokens to save $ (or usage-based costs w high-quality MAX mode)

Profilbild von Indie Hackers
Indie Hackersvor 1 Jahr

: VS Code extension, sidebar agent, 1.2M installs - install in Cursor/Windsurf/VS Code - sends every token to LLMs, quality is high but so are costs - top-tier MCP marketplace, simple one-click installs ⭐️ - agent can see your Chrome browser, inspect elements, read logs

Profilbild von Indie Hackers
Indie Hackersvor 1 Jahr

: VS Code extension, sidebar agent, 380k+ installs - fork of Cline, adds new features, but no MCP marketplace - create custom modes for fine-grained control in your repo, very powerful - Boomerang Tasks: breaks down complex projects, delegates subtasks to assistants 🔥

Profilbild von Indie Hackers
Indie Hackersvor 1 Jahr

: "code prompting app," alternative to code editors - super precise file and context selection ✂️ - pass tons of context, 1-shot tasks, no agent required - the future? devs are ditching IDEs/agents and just using RepoPrompt - native Mac app, Windows/Linux coming

Profilbild von Indie Hackers
Indie Hackersvor 1 Jahr

: fork of VS Code, 2.5M installs - getting acquired by OpenAI?? 😯 - competes with Cursor: uniquely implemented tab completion, inline editing, agent - built-in 1-click "deploy to the web" button - AI agent available as plugin for JetBrains, VS Code, Vim, Emacs, etc

Profilbild von Indie Hackers
Indie Hackersvor 1 Jahr

Claude Code: @AnthropicAI's coding agent that lives in your terminal - released in Feb 2025, only works with Claude models - AI coding from your terminal, no need for an IDE 🧑‍💻 - no learning curve, easy to install and use - great for anything in the terminal, not just code

Profilbild von Indie Hackers
Indie Hackersvor 1 Jahr

Aider: the original terminal-based coding agent, 2M installs - high appeal for devs who prefer the command line over IDEs - integrates with pretty much any LLM, not just OpenAI or Anthropic 🧩 - cleverly tuned to adopt the role of someone you're pair programming with

Profilbild von Indie Hackers
Indie Hackersvor 1 Jahr

Every week we study indie hackers to find out how they succeed: - the tools they use, how, and why - also: marketing channels, biz ideas, and growth tactics We publish everything we find in the Indie Hackers newsletter. You should subscribe 👉

Profilbild von Asad Dhamani
Asad Dhamanivor 1 Jahr

@DevinAI Devin costs 4-5x per task compared to cline FYI. For the more hands off nature it might be worth it but it also cannot tackle tasks above a non trivial level of complexity.

Profilbild von LordsGod🥀
LordsGod🥀vor 1 Jahr

@DevinAI And then it go mess up your whole code and now you must look for the problem

Profilbild von 🇿🇦 PatriotRZA 🇿🇦
🇿🇦 PatriotRZA 🇿🇦vor 1 Jahr

@DevinAI I just built a similar tool for $0 using Notion + Tally + Zapier. Seeing a lot of overlap—thanks for this thread

Ähnliche Videos

Since I joined Cognition I've been obsessed with learning how our eng team uses Devin themselves If we are building the best coding agent + we have the most cracked engineers + we've been fully AI-pilled from day one... it stands to reason that there is a lot to learn by just watching our technical staff work And yes there are a lot of tips & tricks. I recorded a video talking about my favorite... Agent Fan Out - asking your agent to break down the problem, spin up 10 more agents in parallel, and combine their results This is something I've seen everyone do - from our model research team spinning up 100 Devins to examine eval logs - or our product team using 5 child Devins to try out 5 different alternative implementations of the same thing If engineering is cheap and easy, why not build the product 10 times and choose the best one? Think of it in a master/slave context: Master Devin -> 10 Slave Devins -> Master Devin pulls their results There are two reasons this is useful 1. Agents are smartest when their context is small and their task is small & precise. Context windows are finite and too much becomes distracting 2. Agents are good at helping you break a large problem into independent & parallelizable chunks of work Every Devin is its own VM/computer so this also is just a great way to move faster. I've done a migration from React Native to Swift by having Devin break it up into 6 pieces then spin up new Devins to work in parallel In the video I build a greenfield project and try my best to show off this agent fan out concept. I also threw in a few other tricks that I've seen my coworkers do: - Let Devin write its own prompts (especially for creating child Devins). It's way better than us humans - Do tons of things at once. You should be absolutely frying your attention span. Your job should just be babysitting 38 different Devins - Don't be a blocker. Before letting the agent work I make sure to tell it to ask me any questions that would fill in ambiguities. Give your agent all the information it needs (and then some more) so that it can just cook without stopping to ask you questions every few minutes - Let Devin test itself. Integration sanity tests are pretty much solved Hope this is useful!!

Jared Zoneraich

131,907 Aufrufe • vor 3 Monaten

🚀Introducing VisualWebBench: A Comprehensive Benchmark for Multimodal Web Page Understanding and Grounding. 🤔What's this all about? Why this benchmark? > Back in Nov 2023, when we released MMMU ( a comprehensive multimodal understanding benchmark, we received feedback that it included very few UI screenshots. Considering the growing importance of UI understanding, especially with the rise of powerful agents like Devin ( which is built on the strong vision capability of #GPT4, we recognized the need for a benchmark focused on UI screenshot understanding.📸👀 > Multimodal #LLMs have significantly boosted web agents' performance on benchmarks like Mind2Web and WebArena. For instance, the SeeAct agent ( showcases the power of integrating vision into web agents. However, these benchmarks primarily evaluate the end-to-end task execution ability of web agents rather than their understanding of web pages. 🌉 Bridging the Gap with VisualWebBench > To provide a comprehensive evaluation of multimodal LLMs' web page understanding capabilities, we introduce VisualWebBench. Our benchmark spans 139 websites 🌐 across 12 domains 🏷️ and 87 sub-domains 🔍, ensuring a diverse and representative dataset. It assesses MLLMs at three levels: website-level, element-level, and action-level 📊, and encompasses seven tasks designed to evaluate understanding, OCR, grounding, and reasoning abilities 🧠💡. 😮 Surprising Findings > 🎉 Open-source models are catching up: Even though closed-source MLLMs are still leading the leaderboard, we are happy to see open-source models like LLaVA 1.6 34B achieve comparable performance to Gemini Pro. > 🧠 Grounding ability, crucial for developing MLLM-based web applications, is a weakness for most MLLMs. > 🖼️ Importance of Image Resolution: The limited image resolution handling capabilities of most open-source MLLMs restrict their utility in web scenarios, where rich text and elements are prevalent. > 🧱 Relatively strong correlation with general understanding benchmarks like MMMU but weak correlation with web agent benchmarks like Mind2Web. Web agent benchmarks primarily evaluate the end-to-end task execution ability of web agents, which involves a series of actions to accomplish a goal. In contrast, VisualWebBench emphasizes evaluating the foundational skills of MLLMs such as understanding and grounding web page elements. 💡Fun Fact > Claude Sonnet is better than Opus on our benchmark :) 🎓 Conclusion > VisualWebBench serves as a valuable resource for the community, driving research and development in the field of multimodal web page understanding and grounding. As MLLMs continue to evolve and improve, we look forward to seeing new applications and breakthroughs. We believe that our benchmark will contribute to the development of more powerful MLLMs in the web domain, ultimately leading to a more intuitive and efficient user experience on the web. Kudos to the student leads Junpeng Liu Yifan Song and the team Bill Yuchen Lin, Wai Lam, Graham Neubig, Yuanzhi Li! 👏 Check out more details in the Junpeng's thread👇

Xiang Yue

56,697 Aufrufe • vor 2 Jahren

How I get shit done, Episode 001 I've set up a playbook called ‘land’, which is triggered automatically when I drag an issue into the merging column in Linear. That reliably runs CI and merges any green PRs. This has allowed me to ship way faster than before. I think the key takeaway here is you can try to build your own code factory and your own agent orchestration layer, but it is a huge amount of work. The truth is there are entire companies with massive funding that are already tackling this and it's just easier to use their platform. I think this is a lot like if you were a carpenter: you could build your own generator, fuel it, wire it up, and then build a plug and then you could plug your saw into it. Or you could just plug your saw into the wall. Because the electricity company has already done all the work in the infrastructure and investment to make that plug work. I think more of us who are building companies should just be plugging into the wall instead of trying to build all this tooling ourselves. As a dev it's so tempting to build your own dev tools but I think a lot of times, even though you can build fast with agents now, it's a complete waste of time. It probably sounds like I'm being paid by Devin or something but I have zero financial interest here. They don't give me credits. I'm not an investor. I'm not being paid. I just think the tooling is really damn good. If you used Devin a long time ago and wrote it off, you really should have another look - for $500/month it's pretty obscene what you can get done.

Ryan Carson

14,059 Aufrufe • vor 6 Monaten

🚀New Amazon Q Developer agent for software development is available to customers: This agent is based on a new agent architecture that has exciting results coming from the SWE-bench scores (on the full and verified benchmarks) representing AI models’ ability to resolve real-world coding problems. Interesting aspect of Q Agent is that with these newest updates, Q drove nearly 50% more successful coding tasks completed. What makes Q Dev Agent remarkable? The agent architecture is not just about using the best LLMs (which we do), but also giving the agent the ability to constantly explore multiple paths to find the best way to resolve a particular problem (and back tracking when it has reached dead end like a developer would do). Needless to say, we are just getting started on the developer agent and we are constantly pushing to advance our AI capabilities while maintaining quality, security, privacy, and reliability to keep Amazon Q Developer an innovative and trusted option available to our customers using agents for software development. We highlighted the results of our first SWE-bench submission of Amazon Q Developer back in June blog post; with these updates, our new agent resolves 51% more coding tasks than its previous iteration on the SWE-bench verified dataset, and 43% more on the full dataset. That’s the difference a few months make, and I can’t wait to share what our teams will deliver at re:Invent this December. Here's a quick demo showcasing our new Agent in action:

Swami Sivasubramanian

29,013 Aufrufe • vor 2 Jahren

acpx v0.4 ships Agentic Workflows, or as I like to call them "Agentic Graphs" It let's you create node-based workflows on top of ACP (Agent Client Protocol), to drive any coding agent (Codex, Claude Code, pi) through deterministic steps This let's you automate routine, mechanical legwork like triaging incoming PRs, bugs in error reporting, and so on... For example, OpenClaw receives 300~500 new PRs per day. A lot of them are low quality, but they still relate to real issues, so you have to address them somehow You need to: - extract the intent - cluster them based on intent - figure out if the proposed changes are legit, or whether they are slop local solutions, like trying to catch flies instead of drying out the swamp - if the PR is too low quality or the intent is not clear, close them - run AI review on them them and address any issues that come up - refactor them if the changes are half-baked - resolve conflicts - and so on... So that when the PR is presented to the attention of the maintainer, all the routine legwork is done and the only remaining thing is the decision to (a) merge, (b) give feedback to the PR author, or (c) take over the PR work yourself I wanted to build this feature since a couple months now, since Codex got so good. OpenAI models are now good at judging implementation quality, so I found myself repeating the same steps I wrote above over and over I also tried putting all this in a single prompt. But I believe there are workflows that should not be a single prompt, but a sequence of prompts in the same session That is because like humans, LLMs are prone to PRIMING. I claim that putting all steps in the same prompt at the beginning of the context will generally give suboptimal results, compared to revealing the intention to the model step by step Creating such a workflow also gives more OBSERVABILITY into the each step that an agent is supposed to take. Agent generates JSON at the end of each step, and that structured data can be used to monitor thousands of agents running at the same time in an easier way, on a dashboard Similar features have been introduced in e.g. n8n, langflow. But AFAIK they are not integrating ACP like the way I do I wanted to have a fresh approach, and to build an API that I can develop freely the way I want, so I created a new workflow API inside acpx The video is from the workflow run viewer, but that is not where you build the workflow. You build it by using the acpx flow typescript API. See examples/pr-triage in acpx repo Before building that, I started from a Markdown file with a Mermaid chart of the flow I had in mind. The Markdown file acts as a spec for the flow, and I have built the workflow through trial and error. I call this process "workflow tuning" I started working on acpx repo PRs one by one, tuning the flow, slowly scaling to more PRs. Finally, when I felt confident, I ran it in parallel over all external open PRs in the acpx repo. I believe it already saved me hours this week My next goal, if well received, is to set this up on a cloud agent so that it can process the 300~500 PRs the OpenClaw repo receives every day, in real time, as they come in I believe this will save all open source maintainers around the world countless hours and make it much easier to herd and absorb external contributions from everyone!

Onur Solmaz

149,693 Aufrufe • vor 6 Monaten