Loading video...

Video Failed to Load

Go Home

Wow so Google has just open sourced a tool to automate ANY task on mobile ARTEMIS turns prompts into automation: - 99%+ success rate (!!) - Integrates w/ Codex, Claude Code, Antigravity - Automates workflows + captures logs And the roadmap is VERY interesting: They want to add an...

58,026 views • 16 days ago •via X (Twitter)

25 Comments

Paul Couvert's profile picture
Paul Couvert16 days ago

It has a MCP integration so basically any coding agent will be able to use it (Cursor, Cordex, Antigravity, Claude Code, etc.) Here's the GitHub repo:

Nick Beagley's profile picture
Nick Beagley16 days ago

Looks cool. It hit a bit of controversy though:

Paul Couvert's profile picture
Paul Couvert15 days ago

Interesting... hope they'll clarify soon

Lennox's profile picture
Lennox15 days ago

同感。生产里最先炸的往往不是模型智商,是工具权限边界没划清。

Stephan's profile picture
Stephan16 days ago

I do seem to miss a lot of your post on Grok and Grok Bot

Paul Couvert's profile picture
Paul Couvert15 days ago

Grok Bot is good but tbh we have better alternative for now. Looking forward a more efficient model in it.

vairisk ☀️🇱🇻's profile picture
vairisk ☀️🇱🇻15 days ago

It open sourced because initially stole? Is it same tool you are talking about that google was caught stealing and rewriting authors in repo?

Paul Couvert's profile picture
Paul Couvert15 days ago

I wasn't aware of that before reading a previous comment. Hope they'll clarify the situation.

vairisk ☀️🇱🇻's profile picture
vairisk ☀️🇱🇻15 days ago

They will not clarify - they already locked option to check older commits revealing truth; now you can find only proves already gathered on i-net but when company like google does that - its clear they stole.

Paul Couvert's profile picture
Paul Couvert15 days ago

Yeah not a good practice, agree.

Alek's profile picture
Alek16 days ago

ARTEMIS — can I inspect the captured logs per run?

Paul Couvert's profile picture
Paul Couvert16 days ago

Of course you can!

Sage's profile picture
Sage15 days ago

Google has been shipping amazing AI products non-stop. If they could just get their shit together on the product and UX front they'd 10x their product usage.

Enter Matrix's profile picture
Enter Matrix16 days ago

@grok mi spieghi come usarlo?

Build Fast with AI's profile picture
Build Fast with AI15 days ago

would this work with a local 4b model that can run on an android @grok

Jimmy Otis #TruthMatters #EndCronyism's profile picture
Jimmy Otis #TruthMatters #EndCronyism16 days ago

I'm sure samsung will block it.

Paul Couvert's profile picture
Paul Couvert16 days ago

They can't, it'll work on any Android device (iOS coming) with usb debugging enabled

Jimmy Otis #TruthMatters #EndCronyism's profile picture
Jimmy Otis #TruthMatters #EndCronyism16 days ago

They "can't"? I do not think you understand Samsung OEM.

Paul Couvert's profile picture
Paul Couvert16 days ago

Well if they're blocking adb they're also blocking all the devs from using their devices haha

Jimmy Otis #TruthMatters #EndCronyism's profile picture
Jimmy Otis #TruthMatters #EndCronyism16 days ago

I will try. There are many apps that do not work via ADB either. Clearly this is not targeting mainstream.

AI Mastery Guide's profile picture
AI Mastery Guide16 days ago

99% success rate is a big claim

Paul Couvert's profile picture
Paul Couvert15 days ago

And yet, here we are.

LocalLLM's profile picture
LocalLLM16 days ago

On-device VLM for the privacy piece is the interesting bit. Are they aiming to keep the full automation loop local, or is the VLM only for the sensitive steps while the rest still phones home?

Paul Couvert's profile picture
Paul Couvert16 days ago

It's on the roadmap for now so hard to say... but I hope it will be an end-to-end solution.

Jon Fletcher's profile picture
Jon Fletcher15 days ago

This is what I've been waiting for to get a better phone experience.

Related Videos

Anthropic just released a talk on building headless automation with Claude Code. Presented by Sid Bidasaria, Member of Technical Staff at Anthropic live at Code with Claude on May 22, 2025 in San Francisco. Here is what the talk covers. Headless mode lets you run Claude Code without a person actively typing prompts from inside an automated script. Instead of a live session, a script calls Claude with a pre-written instruction using the -p flag. This opens the door for Claude Code to become a piece of a much larger, automated process. In plain terms: Claude Code stops being a tool you use and starts being a service that runs on its own. What this unlocks: Scheduled tasks: Run Claude Code on a cron schedule without anyone at a keyboard. Fix linting errors across an entire codebase. Automatically. Overnight. CI/CD integration: Trigger Claude Code as a step in your build process. Open a PR. Claude reviews it, flags issues, and pushes fixes before a human ever looks at it. GitHub automation: A project manager comments "Claude fix this" on a GitHub issue. Claude reads the request, finds the code, writes the fix, and opens the PR. Multi-machine workflows: One orchestrator dispatches tasks to multiple Claude Code instances running in parallel across different repos simultaneously. When you combine headless mode, hooks, and GitHub Actions, development teams can automate tasks that usually eat up significant time freeing senior engineers to focus on architectural problems while Claude handles the repetitive ones. If you use Claude Code for anything beyond single sessions this talk is worth 20 minutes of your time.

Elias

14,096 views • 4 months ago

-> someone cloned claude -> design interface and -> made it completely free -> it's work on YouTube -> and also suitable for kids -> it’s called open design -> and it’s live on github -> same clean split-screen ui -> you get in claude artifacts -> prompt on the left, live -> design/code preview on -> the right, type what you -> want to build and it -> generates the ui in real -> time, but here’s the twist -> you pick the ai model -> not locked into one -> company, want to use -> gemini, mistral, llama, -> deepseek any model -> with an api work -> if you’re running local -> models with ollama -> that works too -> no subscription walls -> the big difference -> vs claude artifacts -> works with any free -> ai model you’re not -> paying $20/mo just to -> design, use free tiers -> local models, or whatever -> you already have access to -> fully local, your prompts -> and code never leave -> your machine unless -> you want them to -> no data training -> no cloud storage -> privacy by default -> no usage limits -> claude cuts you off -> after a few designs -> here you can generate, -> iterate, break things -> and rebuild all day -> the only limit is your -> don’t like how a button -> works, change it -> want to add your own -> components, go ahead -> you own the tool -> so if you’ve been gatekept -> by paywalls or worried -> about sensitive prompts -> going to some company’s -> servers, this fixes that. -> same workflow, more -> control, zero monthly fee

BeingInvested

12,134 views • 4 months ago

Another insane Jev use case! Jev is making it dramatically cheaper to evaluate what actually happened inside an agent run. And finally, someone open-sourced a self-improving memory layer that can put that signal to work across agent harnesses: - Claude Code - Codex - Cursor - OpenCode, and 20+ more Beacon by Asymptote Labs continuously captures your agent history across harnesses and uses Jev to identify which runs are actually worth learning from. It then turns the highest-signal workflows, corrections, and debugging patterns into reusable skills. GitHub repo: (don’t forget to star it ⭐ ) Beacon preserves the complete session history. But preserving a run and learning from it are two different things. Most coding-agent sessions contain routine exploration, failed commands, and fixes that only apply to one task. The trace can remain available for inspection without turning every detail into guidance for future agents. Jev scores each run for evidence, reuse potential, and human correction signals. An application policy then decides whether to promote, review, or discard it. The recording shows this in action. Claude receives a coding task, modifies the implementation, and runs the tests. I then provide an edge-case correction, so Claude updates the code and adds regression coverage. Beacon automatically captures the complete session. Jev evaluates whether the correction contains a reusable engineering lesson. Once approved, that lesson becomes available to other coding agents working on the project. Since it works across harnesses: - Claude Code sessions can teach Codex. - Cursor debugging can improve OpenCode. So a problem solved by one agent should not need to be learned from scratch by another. If you want to dive deeper into Jev, I also wrote a hands-on guide to building this Jev-style decision path with open models, entirely locally. Read it below.

Avi Chawla

290,485 views • 10 days ago