Google I/O leaks 👀 Google is likely already testing... its own "Cowork" competitor, simply named "Agent" for Gemini and Gemini Enterprise. A new "Tasks" UI highlights - Goal - Agent - Connected apps - Files - Require a human review toggle - And more The "Require a human review" component specifically means that Gemini's capabilities will likely expand, potentially allowing users to automate their desktop tasks as well. Skills and Projects are also cooking 👀show more

🚨 AI News | TestingCatalog
310,858 Aufrufe • vor 4 Monaten
GOOGLE 🔥: An upcoming Gemini Omni video model from... Google is expected to be much more advanced in video editing, capable of completing tasks like removing watermarks, replacing objects in the video, and more. It is also likely that Google will release 2 versions of this model, including a Pro variant. And I assume what we see isn't Pro? Anime sample 👀show more

🚨 AI News | TestingCatalog
179,645 Aufrufe • vor 3 Monaten
GOOGLE 🔥: Gemini for Business will get a new... experience for collaborative Projects, where teams can work in a shared environment. Besides that, Google is rolling out Workflow Agents that can work on automation tasks across various apps. The same functionality is now available on Gemini Enterprise and will become better integrated into the core Gemini for Business experience. Is it only me, or does Gemini for Business feel much better than consumer-facing Gemini?show more

🚨 AI News | TestingCatalog
33,149 Aufrufe • vor 2 Monaten
Google is silently rolling out an updated Gemini experience... for its mobile apps ahead of Google I/O. Its updated UI for Gemini Live features an interactive "bar" or a dynamic island that reacts to your taps and can wave back. It should get loads of superpowers soon 👀show more

🚨 AI News | TestingCatalog
148,831 Aufrufe • vor 3 Monaten
GOOGLE I/O 🔥: These legends are AI-generated via an... upcoming Gemini Omni model. > Both videos are 8s HD samples. > Video with Sandar and Demis is likely generated as an image-to-video using Omni for style editing. > Logan's video is likely a "Likeness" Avatar and Omni video. And "GEMINI" means a new model release! 🤯show more

🚨 AI News | TestingCatalog
106,681 Aufrufe • vor 3 Monaten
Vorflux has opened its new cloud platform, allowing users... to run tasks from a plan to a merged PR on a dedicated machine. It can plan, build, test, and review on its own. > Diff reviews are done by a different model family than the one that wrote it by default. > First boot is saved as a snapshot, with all subsequent sessions waking up from that same starting point. > A browser agent walks through the real user flow, and the recording is included in the pull request. Video proof in the PR 👀show more

🚨 AI News | TestingCatalog
10,433 Aufrufe • vor 2 Tagen
Google DeepMind is cooking 🔥 TLDR : They just... introduced CodeMender, an AI agent that automatically finds and fixes software vulnerabilities. Powered by Gemini Deep Think models, it can patch new bugs instantly and rewrite existing code to eliminate entire classes of security flaws. In six months, CodeMender has already submitted 72 upstream security fixes to major open-source projects. It uses advanced program analysis, multi-agent reasoning, and automatic validation to ensure patches are accurate and safe before human review. One highlight: it applied -fbounds-safety annotations to libraries like libwebp, making past buffer overflow exploits (like CVE-2023-4863) permanently unexploitable. All patches are currently human reviewed, but the goal is to release CodeMender widely as a developer tool to strengthen software security at scale.show more

AshutoshShrivastava
37,936 Aufrufe • vor 10 Monaten
Microsoft presents Windows Agent Arena Evaluating Multi-Modal OS Agents... at Scale discuss: Large language models (LLMs) show remarkable potential to act as computer agents, enhancing human productivity and software accessibility in multi-modal tasks that require planning and reasoning. However, measuring agent performance in realistic environments remains a challenge since: (i) most benchmarks are limited to specific modalities or domains (e.g. text-only, web navigation, Q&A, coding) and (ii) full benchmark evaluations are slow (on order of magnitude of days) given the multi-step sequential nature of tasks. To address these challenges, we introduce the Windows Agent Arena: a reproducible, general environment focusing exclusively on the Windows operating system (OS) where agents can operate freely within a real Windows OS and use the same wide range of applications, tools, and web browsers available to human users when solving tasks. We adapt the OSWorld framework (Xie et al., 2024) to create 150+ diverse Windows tasks across representative domains that require agent abilities in planning, screen understanding, and tool usage. Our benchmark is scalable and can be seamlessly parallelized in Azure for a full benchmark evaluation in as little as 20 minutes. To demonstrate Windows Agent Arena's capabilities, we also introduce a new multi-modal agent, Navi. Our agent achieves a success rate of 19.5% in the Windows domain, compared to 74.5% performance of an unassisted human. Navi also demonstrates strong performance on another popular web-based benchmark, Mind2Web. We offer extensive quantitative and qualitative analysis of Navi's performance, and provide insights into the opportunities for future research in agent development and data generation using Windows Agent Arena.show more

AK
19,684 Aufrufe • vor 1 Jahr
ANTHROPIC 🚨: Claude Cowork will get its own proactive... assistant called "Orbit". > Users will get personalized insights from Gmail, Slack, GitHub, Calendar, Drive, Figma, and other apps, which Claude will generate proactively. > There are also mentions of "Orbit" apps, which users will be able to "deploy." > "Your deployed Orbit apps. Pin favorites for quick access." > OpenAI already has ChatGPT Pulse, while both Google and Perplexity are developing their own proactive assistants, too. > There is a high chance it will be released as Max-only. Thanks to M1 and Tibor Blaho for the tips.show more

🚨 AI News | TestingCatalog
393,549 Aufrufe • vor 3 Monaten
The Gemini 2.0 era is here. And we’re excited... for you to start building with it. A quick rewind of what we just released ⏪ Gemini 2.0 Flash ⚡ comes with low latency and better performance. 🔵 You can now access an experimental version in G3mini on the web, while Gemini Advanced users can try Deep Research, a new AI research assistant. 🔵 Developers can begin building through the Gemini API in Google AI Studio and Vertex AI 2.0 is also enabling new research prototypes of AI agents, including: 🔵 Project Astra, which explores future capabilities of a universal AI assistant 🔵 Project Mariner, which shows what’s possible for human-agent interaction, starting with your browser 🔵 Jules, an experimental AI-powered coding agent Finally, we’re exploring how 2.0 can be used in agents across domains — from navigating the virtual world of video games to applying its spatial reasoning capabilities to robotics. 🤖show more

Google DeepMind
231,798 Aufrufe • vor 1 Jahr
Vibe coding apps that require you to store and... manage large files just got easier, with Replit ⠕ App Storage. ✨ My favorite example in this category are any PDF analysis apps: create a fully customized AI synthesizer in just a few prompts. App Storage is Replit's Object Storage solution (like Amazon S3 or Google CS). But don't worry if those words sound foreign to you, Agent will set this up for you automatically if your app needs it. On that note, this Agent capability is just the beginning of our efforts to make building apps with databases and storage easier: 👇show more

vic
27,313 Aufrufe • vor 1 Jahr
Introducing - Spectre AI: The Monarch - Artificial Intelligence... Searching Reimagined Introducing a sneak-peak of the next core feature in the Spectre AI Search Engine: The Monarch aka the AI ChatZone. The Monarch will serve as your AI agent, providing not only price information, technical and sentiment analysis, and project discovery, but also offering visual assistance through Spectre AI's integrated utilities within the search platform. Real-time blockchain data integration comes with the support of Google for Startups. The AI ChatZone's landing page will feature projects that users have added to their UI Watchlist, allowing for quick and seamless discussions about their favorite projects. Additionally, the platform will include hyperlinked external sources, enabling users to access a wealth of information and resources directly from within the ChatZone. In this showcase, we demonstrated how the system responds to inquiries about $Palm AI's (PaLM AI - $PALM) price performance. The user interface and user experience (UI/UX) are fully developed, and the frontend is complete. Our backend is now entering the beta testing phase. The MVP is next in line to be showcased. Get ready for a new journey in AI with Spectre AI. #ai #artificialintelligence #google #nvda #nvidia #tech #spectre $SPECTshow more

SPECTRE AI
16,267 Aufrufe • vor 2 Jahren
DAPPOS is bringing Onchain OS Skills from X Layer... OKX Wallet into the xBubble ecosystem, making OKX Wallet’s agent-ready wallet, trading, market data, and agentic payments protocol capabilities accessible across xBubble agents. Built on top of Onchain OS Skills, xBubble’s crypto-task SOPs help turn fragmented on-chain flows into a more seamless chat-native experience inside the xBubble app across mobile, desktop, and web. Users can monitor markets, prepare trades, manage wallet activity, and coordinate payment flows through a single conversation. Bubble Engine will continue to use Onchain OS Skills as the baseline for every future SOP iteration and upgrade. With Onchain OS Skills, DAPPOS is making agentic on-chain tasks more conversational, practical, and accessible.show more

DAPPOS
14,289 Aufrufe • vor 2 Monaten
🚀Just launched: Amazon Q, the most capable GenAI-powered assistant... is generally available today: Customers are using Q to transform how their teams get work done. When employees chat with Amazon Q, it provides immediate, relevant information and advice to help streamline tasks, speedup decision-making, and help spark creativity and innovation at work. . Early indications signal Amazon Q could help our customers’ employees become more than 80% more productive at their jobs; and with the new features we’re planning on introducing in the future, we think this will only continue to grow. 🟠 Amazon Q Developer allows developers to spend more time coding and less time on maintenance and performing other tedious, repetitive tasks. Q assists developers and IT professionals (IT pros) with all of their tasks—from coding, testing, and upgrading applications, to troubleshooting, performing security scanning and fixes, and optimizing AWS resources. Q also comes with Q Developer Agents which can autonomously perform range of tasks and we expect it to be the state of the art accuracy in benchmarks like SWE-Bench. 🟠 Amazon Q Business empowers employees to be more data-driven, and helps customers make better, faster decisions using company knowledge and data. Q Business is a generative AI–powered assistant that can answer questions, provide summaries, generate content, and securely complete tasks based on data and information in enterprise systems 🟠 Amazon Q Apps, a new and powerful capability of Amazon Q Business, enables employees to use natural language to quickly and securely build their own generative AI applications to automate daily tasks without requiring any prior coding experience. Employees simply describe the type of app they want, in natural language, and Q Apps will quickly generate an app that accomplishes their desired task, helping them streamline and automate their daily work with ease and efficiency.show more

Swami Sivasubramanian
25,216 Aufrufe • vor 2 Jahren
Karpathy's Agentic Engineering finally has proper tooling! (built by... Google) Karpathy defined agentic engineering as the discipline that separates production agent work from vibe coding. The core skills he listed were spec design, eval loops, and security oversight. The problem has been that practicing this still requires a different tool for every phase: - editor for code - a terminal for scaffolding - a browser for testing - a cloud console for deployment - and a separate framework for evals. Every transition is a context switch. The solution to production-grade Agentic Engineering is now actually implemented in Google’s Agents CLI. It covers the entire workflow in one place for scaffolding, evaluating, and deploying ADK agents. One setup command injects 7 ADK-specific skills into a coding agent's context, which lets it handle scaffolding, evals, deployment, and enterprise registration through natural language. I tested this end-to-end by building a RAG agent from scratch using Claude Code. It scaffolded the full project from the ADK agentic_rag template, generated 20 eval scenarios with LLM-as-judge scoring, and returned a quantitative scorecard. Finally, it also deployed everything to Agent Runtime and registered the agent to Gemini Enterprise, so the entire org can discover and use it. The video below shows this in action, and I worked with the Google Cloud team to put this together. Agents CLI GitHub repo → (don't forget to star it ⭐ ) I wrote up the full build covering all six steps from install to enterprise registration. It includes the eval scorecard, the instruction loophole the eval caught before deployment, and what the deployment process actually looks like end-to-end. Read it below.show more

Akshay 🚀
257,831 Aufrufe • vor 1 Monat
The Visual Studio Code insiders version that just shipped... and will ship in the next few days will come with an insane amount of new capabilities. A few highlights: - You can now run sub-agents in parallel. Yes, really. I even attached a video. - Major UX improvements for sub agents, especially visible in the chat window - A new search tool wrapped as a sub-agent that iteratively runs multiple search tools: semantic_search, file_search, grep_search Which connects nicely to the point above: multiple searches running in parallel, efficiently and fast - Anthropic’s Message API is now enabled by default - You can choose the model for the cloud agent (three available, all premium) - Extended thinking support when using the Claude cloud agent This is part of the broader multi-vendor cloud support under AgentsHQ I wrote about a few weeks ago - Tasks sent to the background agent (basically the CLI tool) now always run in isolation, each with its own git worktree - In a multi-repo workspace, assigning a task to a cloud agent prompts you to choose the target repo Same behavior when opening an empty workspace with no repo - Support for building an external index for files not supported by GitHub’s default indexing - UI/UX improvements for starting new sessions and switching between local / background / cloud agents - Skills are now first-class citizens, just like prompt files, with better UX indicating when a skill is loaded - Improved API for dynamic contribution of prompt files New V2 includes skills as part of the model. Curious to see the extensions that will leverage this - Finally, initial support for showing context usage percentage per session - Skills are enabled by default - Resizable chat window and session view. Small thing, but it was driving me crazy 😁 - A new integrated browser meant to replace the old simple browser Maybe the beginning of real browser use? - Better UI/UX for token streaming in chat - Ability to index external files not supported by GitHub There’s a lot more. Some of it hasn’t fully landed yet, but everything that has is already in Insiders. The next stable release should drop in early February. As usual, I’m just shocked by the volume of features this team ships every month. After the holiday slowdown, this one is shaping up to be a wild release.show more

Oren Melamed
29,555 Aufrufe • vor 7 Monaten
INTERLINK ROLLS OUT FIVE STRUCTURAL UPGRADES TO ACCELERATE VERIFICATION... InterLink Labs (InterLink Labs 👤 + 🌐) is rolling out five structural upgrades to its Verified $ITLG verification pipeline as the Human Network passes 7 million users and prepares for InterLink Chain mainnet, $ITL listing, and the Verified $ITLG migration. The team is expanding its Curator network across more time zones to move verification to a continuous 24-hour rhythm. The AI snapshot layer is being enhanced to clear high-confidence profiles faster and route only edge cases to human review. A new multi-track queue structure splits processing by profile complexity so straightforward verifications no longer wait behind deeper review cases. Users will get a Force Sync option every 12 hours to refresh their own metrics, replacing the 2-week AI snapshot cycle. Where regulatory frameworks allow, InterLink will integrate with trusted partner verification providers to eliminate duplicate KYC work without lowering compliance standards. The compliance bar stays fixed. AMLO, SFC VASP, and FATF Recommendation 15 are non-negotiable. Every verified profile becomes a legally compliant human node designed to plug directly into global financial infrastructure for payments and Real World Assets.show more

BSCN
20,568 Aufrufe • vor 3 Monaten
🚨 $GRASS Season 2 Native Wallet & Wallet Connect... Update 😱 Many users are worried because the wallet connect option is currently unavailable in Grass Season 2. Don’t worry — if your wallet is not connected, it should not cause any issue. ✅ Users who already connected their wallet are fine, and those who haven’t connected yet will most likely be able to do it later through the upcoming official Grass Native Wallet. 💼⚡ For those who don’t know, a Native Wallet simply means Grass is launching its own wallet where users will be able to store, deposit, withdraw, and manage their $GRASS tokens. 🌱💰 Your Season 2 airdrop rewards are expected to be received through the native wallet once it is officially launched. 🚀 Is your wallet connected or not? Let us know in the comments 👇 #GRASS #Airdrop #Crypto #grassupdate #grassseason2update #grasswalletconnect #grassseason2airdorp #grassnativewalletshow more

Airdrop Hunt with Lakhan 🪂
22,580 Aufrufe • vor 3 Monaten
Frameworks such as ai16zdao's Eliza and Virtuals Protocol have... been instrumental in early AI agent developments. Agent swarms working in hierarchy represents for many the next logical step in unlocking the vast potential of AI. Learn below how Shadō Network achieves this. AI agents launched through current popular platforms have individual personas, on-chain functions and access to data via various APIs. This being said, they operate in isolated environments, with a ceiling on emergent behaviour such as collaboration or competition. Shadō Network invites massive expansion for capabilities of both new and existing AI agents, with an open-source package easily integrated into popular frameworks that enables the launching of stratified agent swarms. Our website is live: The "Shadō Play" package provides a modular, configurable platform for creating or employing agents of choice in a swarm-like setup, opening a Pandora’s box of near infinite emergent agent behaviours, relationships and functionalities. Users will be able to make use of various prefab client integrations such as Twitter, Telegram, Ollama, and others to specify swarms to their needs or create their own extensions to enhance agent capabilities even further. Agents operate with a memory module and a HTN for autonomously deciding which interactions to act on, walking the line between autonomy and configurability. The Shadō Network project’s development is supported by our ghostly friend Omnipotent (👻,👻), an AI agent developed by the Shadō Network team trained on and fine tuned with a multitude of academic data related to artificial intelligence, blockchain, finance, software engineering, world building and more. Omnipotent serves as both an interactive steward for the project and as an asset - regularly scanning social platforms, websites and newsfeeds he is capable of providing the team project development advice, whilst also communicating with the wider world via his automated X account (launching soon). Shado Network is collaborative and open-sourced. Agentic Swarms require a developer swarm to maximize the technical capabilities and impact the greatest number of users. Our dedicated team of core contributors are active in other web3 AI repos and are here to guide project direction and foster growth. We’re facilitators, not gatekeepers... Alone we can go fast but together we can go far. A lot more to come soon. 👻show more

Shadō Network | シャドウネットワーク
23,546 Aufrufe • vor 1 Jahr
LangGraph. CrewAI. Agno. Which one to pick? The good... news is that this will not matter soon! Finally, we have a full picture of how the industry is solving this with just three open protocols that work across ALL frameworks. It's not about picking the best framework. Instead, it's about understanding how protocols create interoperability. The Agent Protocol Landscape shows how three complementary protocols are creating a universal language for Agents: > AG-UI (Agent-User Interaction): - The bi-directional connection between agentic backends and frontends. - This is how agents become truly interactive inside your apps, not just as chatbots, but collaborative co-workers. > MCP (Model Context Protocol): - The standard for how agents connect to tools, data, and workflows. > A2A (Agent-to-Agent): - The protocol for multi-agent coordination. - How agents delegate tasks and share intent across systems. These aren't competing standards. They're layers of the same stack and have handshakes with each other. So instead of building point-to-point integrations, you build to protocols. Moreover, you can integrate LangGraph, CrewAI, or Agno into the same frontend, without rewriting your UI logic. These protocols let everything work together. For instance: - Your LangGraph agent pulls data via MCP. - It delegates analysis to a CrewAI agent via A2A. - Results stream to your React app via AG-UI. - Users see real-time collaboration in your interface. This way, you can focus on building agent capabilities instead of integration mechanics. The protocols handle interoperability automatically. CopilotKit unifies this entire stack into one framework so you can build "Cursor for X" style apps without implementing each protocol from scratch. It gives you all three protocols, generative UI support, and production-ready infrastructure in one framework. I have shared this playbook in the replies! It breaks down handshakes, misconceptions, and real examples and shows exactly how to start building.show more

Avi Chawla
30,762 Aufrufe • vor 9 Monaten