Weekend read: ๐๐ด๐ฒ๐ป๐๐ ๐๐ต๐ผ๐๐น๐ฑ ๐ฏ๐ฒ ๐ฆ๐ฒ๐ฟ๐๐ฒ๐ฟ๐น๐ฒ๐๐ (๐ฎ๐ป๐ฑ ๐๐๐ฟ๐ฎ๐ฏ๐น๐ฒ) Serverless... platforms are amazing for many AI agents, especially if they do a lot of bursty work (chat sessions with long inactivity) or react to events (approvals). Two things are needed to make this work well: (1) Durable execution / orchestration, to let agents recover and suspend (scale to zero) when they wait for input. (2) A durable execution orchestrator that has a push model that works with serverless platforms. Super fun to see agents with restate go brrrrrrrrrr: starting as durable functions on Modal's compute platform, scaling up, running, scaling down. Just out of the box, no extra code, no workers deployed. ๐show more

Stephan Ewen
18,169 ะฟัะพัะผะพััะพะฒ โข 1 ะณะพะด ะฝะฐะทะฐะด
Increasingly, HTML Artifacts are becoming a core part of... how I work with AI agents. Long-horizon agent sessions need a better way to surface insights about what work it has done. This may not be obvious right now, but as you start to let your agent work on dynamic workflows, large codebases, long-running loops (e.g., using /goal), and deep research tasks, you need a good way to present results. Chat window is not it. You also don't want to just trust everything the agents do. Artifacts help provide an important verification layer, which in turn enables important decision-making. I like HTML artifacts because I can just ask the agent to produce as many of them (and in whatever form) as I need to verify the work and make sense out of everything. I even built a nice tab system for my artifacts. They are great for continual learning and research. I use HTML artifacts for logging, tracking experiments, brainstorming, managing my inbox, code reviews, agent session management, deep research, writing, reading, and so much more. I believe Andrej Karpathy wrote about this somewhere: As we move on to more advanced applications of AI agents and outputs get more complex, we will start to find the need for even more advanced forms of interactions with AI, including interactive neural videos/simulations.show more

elvis
37,127 ะฟัะพัะผะพััะพะฒ โข 4 ะผะตัััะตะฒ ะฝะฐะทะฐะด
HTML Artifacts are a big part of how I... work with agents now. Artifacts can be more than just static files. When combined with agents, they can take action or help you take action. This unlocks all kinds of interesting ways to work with agents. This is clearly the future. Check out this writing and scheduler artifact I built in a few minutes. It uses a bit of HTML and JS. All the data is in markdown (Obsidian vaults), so the agent can access and modify it at any time. No DB needed. No sophisticated functionalities. The agent decides all that for me based on the skills, context, and memory it has access to. The best part about this simple stack is that all the important information stays with me. This has allowed me to build a recursive self-improving system and automations that can better tap into coding agents like Codex or Claude Code. I could have paid or built an entire app for scheduling posts, and there are so many of them out there. But I don't need to. I've realized a simple artifact does the job. And the simplicity of it is actually an advantage. Very little maintenance for very high returns on personalization, time, and efficiency. The other benefit of this is that I can add features as I please. That level of personalization feels magical, and we should all be pursuing more of it. All of this just keeps compounding. Of course, this example is just about writing. But I have similar artifacts for research, design, experimentation, evaluation, and so much more. And no, I didn't actually publish the post example I shared in the clip. It was just for demonstration purposes. I actually spend more time than this when writing together with agents. Lastly, having built my own agent orchestrator tool has made me realize that simplifying the tool stack is a superpower. If you are curious about how all this works, I will do a live session next week:show more

elvis
18,374 ะฟัะพัะผะพััะพะฒ โข 4 ะผะตัััะตะฒ ะฝะฐะทะฐะด
ClickUp now employs over 100,000 AI AGENTS for our... customers. This is from just THREE WEEKS of customers vibe coding full-blown teams of agents, THEMSELVES. BUT there's a problem. Since Super Agents are built agnostically, horizontally, and deeply capable with human-level abilities, you can literally build an agent for anything. We found that MOST of our customers have NO CLUE where to start. This is their very FIRST TIME EVER managing an agent. What I recommend is starting with a PROBLEM. Everybody can think of a problem they have. Just tell Super Agent Builder about your problems... about where you're WASTING time... about what you WISH you could do but you can't because of resource constraints. We've also found that human FEEDBACK and iteration are KEY. After agents are done with their jobs, give them feedback... Was it good? Was it bad? What do you want to see differently? They AUTOMATICALLY SELF-IMPROVE. Every time, they'll continuously get SMARTER. Personally, I find that agents go from AVERAGE intelligence to SUPER-human intelligence within about a month of working with them. Super Agents have truly democratized productivity, empowering literally anyone to build personalized, powerful agents in minutes. What problems do you wish you could solve? What do you not have enough time to get done? What would you like to do but don't have the resources for? What busy work do you wish you could get rid of?show more

Zeb Evans
18,336 ะฟัะพัะผะพััะพะฒ โข 8 ะผะตัััะตะฒ ะฝะฐะทะฐะด
Iโm joining OpenAI Codex to work on the future... of agentic development! At Cursor, I got to see the shift from autocomplete to agents. The next step isnโt a better IDE. Itโs an Agent Development Environment (ADE): systems and tools for orchestrating agents, reasoning over their outputs, and making them autonomous enough to reliably complete ambitious work. After chatting with Alexander Embiricos and Tibo, it was clear that Codex is the best place to realize this vision. The team has consistently shipped SOTA models for agentic coding (check out gpt-5.3-codex) and Iโm pumped for the future that the new Codex App points to. What Iโm most excited about is the broader mission: accelerating the knowledge work economy. All agents are coding agents, and weโre already seeing Codex used across every job function within organizations. Iโm extremely grateful for my time at Cursor, working with the incredible team, and Iโm proud of what we built together. Iโm excited to take an even bigger swing with Codex. If youโre curious to get a glimpse of where we are headed, download the Codex App! If you want to work on this mission, please apply or reach out - we are hiring across all functions! You can just build things.show more

Rohan Varma
760,125 ะฟัะพัะผะพััะพะฒ โข 7 ะผะตัััะตะฒ ะฝะฐะทะฐะด
I had the same thought so I've been playing... with it in nanochat. E.g. here's 8 agents (4 claude, 4 codex), with 1 GPU each running nanochat experiments (trying to delete logit softcap without regression). The TLDR is that it doesn't work and it's a mess... but it's still very pretty to look at :) I tried a few setups: 8 independent solo researchers, 1 chief scientist giving work to 8 junior researchers, etc. Each research program is a git branch, each scientist forks it into a feature branch, git worktrees for isolation, simple files for comms, skip Docker/VMs for simplicity atm (I find that instructions are enough to prevent interference). Research org runs in tmux window grids of interactive sessions (like Teams) so that it's pretty to look at, see their individual work, and "take over" if needed, i.e. no -p. But ok the reason it doesn't work so far is that the agents' ideas are just pretty bad out of the box, even at highest intelligence. They don't think carefully though experiment design, they run a bit non-sensical variations, they don't create strong baselines and ablate things properly, they don't carefully control for runtime or flops. (just as an example, an agent yesterday "discovered" that increasing the hidden size of the network improves the validation loss, which is a totally spurious result given that a bigger network will have a lower validation loss in the infinite data regime, but then it also trains for a lot longer, it's not clear why I had to come in to point that out). They are very good at implementing any given well-scoped and described idea but they don't creatively generate them. But the goal is that you are now programming an organization (e.g. a "research org") and its individual agents, so the "source code" is the collection of prompts, skills, tools, etc. and processes that make it up. E.g. a daily standup in the morning is now part of the "org code". And optimizing nanochat pretraining is just one of the many tasks (almost like an eval). Then - given an arbitrary task, how quickly does your research org generate progress on it?show more

Andrej Karpathy
1,657,913 ะฟัะพัะผะพััะพะฒ โข 7 ะผะตัััะตะฒ ะฝะฐะทะฐะด
โI DONโT NEED NO MANโ-Black American women This mentality... of not needing a man That has been said and installed In black girls/womens minds for decades now has plagued the community With single homes and boys that have to go to the army to learn discipline Do black women have a lot of good options in the black community dating wise? NO Because so many black men have been plagued with street nigga culture Which leads them to prison or death And so many canโt keep a decent job And provide well enough to get respect I LOVE ANGEL REESE BUT WHEN SHE SAYS THIS I FEEL LIKE ITโS MORE OF LIKE A REPETITIOUS QUOTE FROM SISTERS THAT GOES BACK FOR DECADES BECAUSE SO MANY SISTERS HAD TO DO IT BY THEMSELVES THAT THEY GOT USED TOO IT BUT THAT DONโT MEAN THAT WE ARE PRODUCING THE BEST NEVER GENERATION IF THEY ARE ALL COMING FROM BROKEN HOMES WHERE THE FATHERS ARE SEEN AS AFTER THOUGHTS INSTEAD OF MEN THAT ARE NEEDED IN THE HOME JUST AS MUCH AS THE MOTHERS TO PROVIDE THE KIDS WITH ADEQUATE DISCIPLINE AND OTHER SKILLS ESPECIALLY THE YOUNG MENshow more

MASTER STUDENT๐คฒ๐พ
58,875 ะฟัะพัะผะพััะพะฒ โข 5 ะผะตัััะตะฒ ะฝะฐะทะฐะด
Today marks General Availability of AgentCore, a set of... infrastructure building blocks for developers and companies to build secure, scalable agents. When we first started AWS, the vast majority of developers were spending most of their time on the undifferentiated heavy lifting of infrastructure instead of what differentiated their feature. So, we solved that problem by building primitive building blocks like compute and storage and database that would allow teammates and customers to quickly build and deploy new experiences without having to reinvent the wheel each time. We realized the same thing was happening with AI agents. It's too difficult and it's slowing customers down. That's why we created AgentCore, a set of services to build, deploy, and operate highly capable agents using any framework or model, with enterprise-grade security and scalability. These building blocks (like serverless secure runtime, memory, observability, a gateway that does MCP translation, etc) help customers tackle some of the biggest challenges of going from prototype to production, much more quickly, securely, and scalably. AgentCore has been in preview for several weeks, and customers have been quite excited about it. The AgentCore SDK has already been downloaded over a million times and we're seeing transformative results, such as Cohere Health expecting to reduce medical review times by 30-40% in highly regulated healthcare, and teams at Cox Automotive and Experian are embracing its flexibility to deploy and operate agents at scale. Inside Amazon, our Amazon Devices Operations & Supply Chain team is using AgentCore to develop an agentic manufacturing approach where AI agents work together to automate manual processes โ turning what used to be days of engineering time into processes that take under an hour with high precision. Just like AWS changed how companies build and scale applications, we believe AgentCore will do the same for AI agents, enabling the next generation of innovation.show more

Andy Jassy
24,990 ะฟัะพัะผะพััะพะฒ โข 11 ะผะตัััะตะฒ ะฝะฐะทะฐะด
Very pleasantly surprised to discover Cursor cloud agents can... playtest the godot game I built. See the (sped up) video below of the agent playtesting the game. As I was watching it play the game, I can see the agent slowly learn how the game works and familiarise with the game's UI. I also realised that the agent is a very 'safe' player, choosing to play very safely and retreating from battle if it foresees it can't defeat. Very interesting to see. I wonder if I could simulate different game playtester behaviours that mimic different types of real-world player archetypes. With agentic playtesting, this means that the agents are able to provide actual gameplay feedback and suggestions to improve the game, having played the game itself. This unlocks a whole lot of possibilities for AI-assisted game dev, since it closes the playtest loop. This feels like the future of recursive game development, where agents can now recursively build > playtest > improve the games they are working on. Thanks edwin for letting me know that these agents can actually playtest games, not just software! Very excited to dig deeper to see what I can do with these agents with computer access!show more

Danny Limanseta
52,432 ะฟัะพัะผะพััะพะฒ โข 7 ะผะตัััะตะฒ ะฝะฐะทะฐะด
If you are trying to understand where AI agents... are going, learn harness engineering. A capable model is only one part of an agent system. Once the model begins reading files, calling tools, modifying state and working across many steps, the quality of the system depends increasingly on the software around it. Consider a coding agent working through a large repository. The model can decide that it needs to inspect a file, search for a symbol, make an edit or run a test, but those decisions do not execute themselves. The surrounding runtime has to decide which resources are available, whether the requested action is permitted, how the operation should be performed, what result should be retained, and what information should be presented to the model on the next step. This becomes harder as the run gets longer. As history accumulates, replaying everything can become costly and less effective. The harness has to decide what should remain in context, what should be summarized or retrieved later, and what belongs in persistent state outside the context window. Execution has similar problems. A long-running agent may need to survive an interruption, avoid repeating completed work, enforce permissions around consequential actions, and preserve enough history to reconstruct what happened when the final result is wrong. These are harness problems. The harness is the layer that manages context, tools, execution, state, checkpoints, limits and traces around the model. Harness engineering is the work of designing and improving that layer. Engineers inspect execution traces, evaluate agents on representative tasks, look for recurring failure modes, and then change things such as context selection, tool interfaces, state handling or execution controls. That last part matters because agent failures are often not fixed by changing the model. Sometimes the useful change is in what the model sees, how a tool is exposed, what state is preserved, or what the runtime does after a failed step. As agents take on longer tasks, the demands on this surrounding software grow. Model capability remains essential, but harness engineering is what turns that capability into an execution process that can be controlled, inspected, tested and improved.show more

Tech with Mak
48,700 ะฟัะพัะผะพััะพะฒ โข 10 ะดะฝะตะน ะฝะฐะทะฐะด
๐ Today, weโre excited to relaunch Airtable as the AI-native... app platform, combining the magic of vibe coding business apps with real production-readiness and scalability, and embedding them with an army of agents that automate thousands of hours of work in seconds. Instead of just adding more AI capabilities to our existing platform, we treated this as a refounding moment for the company. We started with a clean-slate imagining of the ideal form factor for building apps in the agentic era. (If you want to skip all the backstory and just try it out, you can just go to All new signups get the new AI experience, and existing accounts can switch over using this link: Thread and demos below๐show more

Howie Liu
17,156,506 ะฟัะพัะผะพััะพะฒ โข 1 ะณะพะด ะฝะฐะทะฐะด
Our @Grammarly AI agents are here! Today, weโre launching... eight new AI agents designed for students and professionals. We created many of these agents with students in mind because theyโre the first generation entering a job market where employers expect both subject expertise AND AI fluency. These agents help with everything from finding credible sources to predicting reader reactions. One agent weโve gotten great feedback on is AI Grader (I wish I had this in school), which you can see in the video below. It looks at your assignment rubric and gives you suggestions like your professor would, and a grade prediction before you submit your work. And these agents are available in docs, our new AI-native writing surface! Iโm deeply proud of this launchโdocs is powered by Coda (Superhuman Docs) technology and is a great integration moment between Grammarly and Coda. This is just the beginning of Grammarlyโs journey to offering agents that work everywhere people work and collaborate. Iโve been loving using these agents, and Iโm excited for our customers to get access. Try them for yourself here and let me know what you think:show more

Shishir
13,564 ะฟัะพัะผะพััะพะฒ โข 1 ะณะพะด ะฝะฐะทะฐะด
We got early access to the Gemini Omni API... Google is calling this model "Nano Banana for video" 3 things we built with it๐ 1. Landscaping proposal from customer video. Customer submits a video of the current lot. Hyperagent designs and renders the transformation. Perfect realism, no surreal changes to the surroundings. 2. Animated professor who explains your dashboards. Hyperagent runs analysis on a business question, then generates an explainer video to walkthrough the findings. 3. 8-bit morning briefing. Hyperagent builds morning briefings based on calendar, email/chat, and market intel. Then generates a sidescrolling platformer video showing the goals you'll clear today. Our take: > Video has been a vastly underused input for agents. This is the first model to make video directly malleable to our agents. > Get more imaginative with your outputs. How could a video artifact make the work more memorable or playful? coming soon to Hyperagentshow more

Hyperagent
325,111 ะฟัะพัะผะพััะพะฒ โข 3 ะผะตัััะตะฒ ะฝะฐะทะฐะด
Last post on this I'd appreciate if people took... the time to read it and share it because I think when people listen they deserve credit but TRNSMT have done more than that I complained about the disabled situation TRNSMT Festival and I think the fact that they listened and made changes is the kind of progressive attitude we need towards disabled facilities. I wasn't complaining to come across as some sort of Karen, it just genuinely would've been a health and safety risk. So I think it's only right that I commend them for the work they've done with the platform this year. The platform has many key things that I think make it the best at any festival I've been at. 1.Clear marked out lanes for people to move out to the toilets and back to their seat to avoid congestion and create space 2. A bar for disabled people beside the platform this is amazing. 3. Toilets well spaced out as well as a changing places toilets for people who more severe disabilities beside the platform. 4. Plenty of space in the toilets. 5. It's huge and you feel really close to the stageshow more

Stephen Reside
129,737 ะฟัะพัะผะพััะพะฒ โข 2 ะปะตั ะฝะฐะทะฐะด
Holy moly. ๐คฏ Jack Dorsey just gave away, for... free, a full toolkit for running a company where AI agents work right next to your human team. His company Block put it on GitHub. It's already past 29,000 stars. Here's how you set it up takes about 5 minutes: 1. Copy (clone) the code from GitHub. 2. Run your own server. This is where your chats, search, code, and automatic tasks all live. 3. Add your AI agent into a chat channel, just like adding a new team member. Tell it what it's allowed to do, and let it work with your team live. No fees. No middleman. Just you, your team, and your AI agents all in one place you control. This might be what running a business looks like from now on Worth saving.show more

Charlie Hills
14,039 ะฟัะพัะผะพััะพะฒ โข 18 ะดะฝะตะน ะฝะฐะทะฐะด
We're excited to unveil NRN Agents, a rebrand that... aligns our project identity with our token and strengthens our mission to power the future of AI-driven gaming. This mission requires collaboration, and starting this week, we will begin our expansion to become a multi-chain ecosystem. We are joining forces with leading gaming platforms and ecosystems to realize this vision. Stay tuned for more announcements to come. Why NRN Agents? NRN stands for NEURON, the fundamental unit of intelligence. Our AI agents function as the neural foundation of games, learning, adapting, and evolving within game worlds to deliver unparalleled engagement. NRN agent SDK enables advanced gaming agents powered by a proprietary machine learning infrastructure focused on behavioral learning. We've perfected the craft of gaming agent design, creating hyper-efficient agents that are performant and scalableโfrom casual to the most demanding games. Our SDK will seamlessly integrate into many platforms, tech stacks, and ecosystem โ Any Game. Any Chain. More than just games, it's the path to AGI Gaming is our proving ground, but not our final destination. We're using games as a sandbox to accelerate the development of generalized intelligenceโone that will create meaningful real-world impact. With the upcoming launch of [redacted] and a growing network of partners committed to the AGI vision, we're building an open-source innovation movement powered by an AI x gaming framework connected by $NRN. $NRN the token $NRN is a utility token that serves as the gateway to our growing ecosystem. It will power a diversified economy with multiple revenue streams and staking opportunities: Agent Deployment: NRN is the laboratory creating gaming agents that can be distributed through platforms and launchpads alike. The model is simple: More games integrate, more NRN agents get deployed, more monetization. Data Creation: NRN Reinforcement Learning (RL) enables token staking to create Data Capsules. Players contribute gameplay data into the Capsules, which are used train RL agents and reward participants (players & stakers). AI Arena: $NRN also continues to power AI Arena's in-game economy, a cult favorite of competitive diehards that features a skill-based wagering system. To our community who have supported us since 2021: thank you for being part of our journeyโthe next chapter will be the most exciting yet!show more

NRN Agents
20,768 ะฟัะพัะผะพััะพะฒ โข 1 ะณะพะด ะฝะฐะทะฐะด
We are releasing AutoResearchExam, a benchmark on open-ended machine... learning and engineering tasks. Our benchmark covers seven research areas including model training, data curation, AI safety and interpretability. In each task, we give agents 24 hours with a CPU or GPU machine to develop and improve their solutions through experiments and feedback. We measure both speed and quality with a combined score. Our benchmark has a unique feature: testing if agents create improvements that hold up on data they never see. We find that AI research agents often overfit as they try to improve. We see an interesting head-to-head comparison at the frontier: Astra starts the strongest and holds the lead for up to 19 hours but Fable 5.1 catches up and gets the top performing spot in the final hours. Qwen3.8 Max, Gemini 3.8 Flash and Grok 4.6 all sit on the cost-performance Pareto frontier, giving strong options at lower API budgets. Anthropic's Opus and Fable retain nearly all their validation performance on hidden tests, with gaps of 1.1% and 2.9%. Astra's improvement over Sol extends to generalization too, with that gap falling from 6.9% to 1.7%. (1/n)show more

Alex Dimakis
2,032,515 ะฟัะพัะผะพััะพะฒ โข 26 ะดะฝะตะน ะฝะฐะทะฐะด
LLM Artifacts Connected to Andrej Karpathy's LLM Knowledge base... idea, I've been building out a fun way to generate dynamic artifacts from these knowledge bases with the goal of discovering and revealing meaningful and deeper insights. LLM KBs are hard to consume for humans, as I think they are more built for agents. So the question is, what form would be useful for humans to take actions and make important decisions? That's what I am trying to figure out with these artifacts. The artifact example shows a pulse on HN discussions around AI-related stories. The insights can go deeper, of course, but this is already super fun and thought-provoking, like some of my favorite podcasts. The format and depth matter a lot. The aggregation skills of agents are outstanding if you tune the prompts and skill carefully. I built this artifact generator in a few minutes through an agent skill, but I feel like there are so many ways that LLM-generated information can be used and consumed. Like generating deeper insights and analysis, and things that are just not feasible for humans today. The generated artifact (including its data and design) serves as reusable templates or can be updated in real-time via auomations, which is something I am also working on. It is truly an insane way to monitor and track information. Better than a newsletter. Better than newspapers. There is something about this that gets me really excited about the future of AI agents for knowledge generation and discovery. Lots of hidden gems everywhere just waiting to be discovered and acted on if the information is presented correctly. This is not perfect. The format, style/prose can be improved, but this is easy to customize via skill. You can personalize it to your liking. I feel like these dynamic artifacts are going to emerge as a strong new medium to stay on the cutting edge of things, both for agents and humans. My target is research, of course. This was just a basic example. Besides animation, I am also targeting other components like voice, videos, images, slides, etc. This space is full of opportunities to explore. Skill for this coming soon.show more

elvis
31,337 ะฟัะพัะผะพััะพะฒ โข 5 ะผะตัััะตะฒ ะฝะฐะทะฐะด
Polymarket Agents repo is cheating code for real life... If you've been thinking AI agents are cool, but what do they actually do? this Polymarket repo is the clearest answer I've seen This Polymarket repo basically hands you a mini AI hedge fund where Claude calls the shots. It uses RAG to read the news and snipe mispriced odds 24/7, handling all the messy API work so you don't have to. Your agent could literally be catching alpha while you sleep - don't miss out on this. How you'd use it in plain English: > Pick a market (or let the agent scan for opportunities). > Feed the agent context (news, social chatter, historical notes) and let the LLM decide what matters. > Turn that into a concrete trade decision and execution loop - automated Bookmark post so you don't lose the alphashow more

BuBBliK
16,887 ะฟัะพัะผะพััะพะฒ โข 7 ะผะตัััะตะฒ ะฝะฐะทะฐะด
Own your personal agent, folks. Just built my own... little personal agent with Pi's Pi Durable. I wanted to see how quickly I could build an OpenAI dots clone with open-source tools. I am amazed at how quickly I got this off the ground. It's proactive. It keeps working when you're not watching. It remembers you, runs jobs on a schedule, waits for your approval, and uses its own computer. Pi Durable makes that possible. Every step is checkpointed. I can kill the process mid-task, and My Pi picks up right where it left off. Memory lives in durable documents. What My Pi learns about me in one Space survives a crash and carries over to every other Space. Approvals are durable too. It also has its own computer. Each Space gets a full Linux desktop that My Pi controls with screenshots, mouse, and keyboard. The model runs through OpenRouter, so I can swap models without changing anything else. That's all configurable however you like it. If you've been wanting your own personal agent, now is a great time to build one. I am building this as an experimental ground where I can test new ideas and experiment deeply with proactive personal agents. But I plan to open-source it after testing it a bit and adding more interesting primitives. The goal is to make it customizable and easy to build on, so others can experiment with their own personal agent functionalities.show more

elvis
20,288 ะฟัะพัะผะพััะพะฒ โข 19 ัะฐัะพะฒ ะฝะฐะทะฐะด