instead of generating text, jev from TypeSafe AI generates... structured output this makes it great for classification tasks like model routing, tool selection/search, and guardrails of many forms! it's also ridiculously fast and cheap compared to LLMs doing the same tasksshow more

Sydney Runkle
52,469 views • 8 days ago
Jev by TypeSafe AI is now on OpenRouter, in... beta. Jev is a System One model. Instead of generating text, it takes your app's state plus a typed question and returns a typed decision with a probability attached. There is no JSON prompting, parsing layer, and nothing to validate against.show more

OpenRouter
564,101 views • 8 days ago
Robotics keeps hitting the same wall. Single task RL... works, but... it does not scale to hundreds of tasks or new embodiments. This new paper looks like a real step toward fixing that. The team introduces MMBench, a benchmark with 200 tasks across many domains and robots, and Newt, a language conditioned world model trained online across all 200 tasks at once. The simple idea behind Newt: The model learns from demos to get the right priors It trains across many tasks through online interaction It uses language to ground the goal It adapts fast when a new task shows up What stood out to me: ✅ One model trained on 200 tasks at the same time ✅ Language conditioned control for both states and RGB ✅ Better data efficiency than strong baselines ✅ Strong open loop control ✅ Fast adaptation to new tasks and embodiments ✅ Full release of 200 checkpoints, 4000 demos, code, and benchmark This is a good push toward general control instead of one model per task. If you want the full paper: Project page: —- Weekly robotics and AI insights. Subscribe free:show more

Ilir Aliu
70,090 views • 10 months ago
BREAKING 🚨: OpenAI is actively polishing its Tasks feature... and there is a big chance we will see them announced today 👀 - Tasks Beta will allow users to schedule tasks like "send me AI news from TestingCatalog at 9 am" - These automations will be handled by a new model tool "jawbone" - There will be a new Notifications tab in settings, assumingly to control the way you will receive notifications about scheduled tasks Interestingly, the same feature is being in development for Gemini. What is the chance of seeing both of them released on the same day?show more

🚨 AI News | TestingCatalog
204,165 views • 1 year ago
Dynamic workflows are a generalization of harnesses, automations, loops,... routing, and graphs. It's the most powerful feature I have built into my agent orchestrator. Supports all kinds of patterns that leverage different agent backends (claude, codex, pi, hermes,...). It's a meta-harness approach that unlocks new forms of test-time compute. Example of use cases it supports: > LLM councils to get different perspectives from LLMs or plan more intensively > Dynamically routing tasks to different agents based on needs (e.g., cost efficiency and optimal intelligence) > Advisor/Judge + executor workflows and pretty much any complex graph-based pattern required by the task. I find it especially useful for long-running work and code reviewing. > Agent teams that talk to each other if needed for the task. I like to use this for AI editing, artifact creation, and other creative tasks. And I am sure it supports so many things that I haven't discovered yet. I got inspired by the dynamic workflow feature released by the Claude Code team. I had actually built it earlier this year but wanted to generalize it across different agent backends. I think this is going to become more popular in the coming days. I will share more of my findings soon.show more

elvis
32,897 views • 2 months ago
Boom! Grok Tasks Make It One Of The Most... POWERFUL Real-Time AI Systems In The World. — My How to Use Grok Tasks With Hidden Tools For Powerful Daily Output. Grok Tasks are customizable AI workflows that integrate a variety of tools to streamline daily activities, from research and analysis to creative planning and problem-solving. I have been using them for quite sometime and because of the vital heartbeat of news and first person data on X, it is the most powerful AI platform available. By combining Tasks with tools like web searches, X platform interactions, code execution, and media viewers, you can build efficient, automated processes. These tasks work by prompting Grok with a clear description of what you want to achieve, and Grok will intelligently call the necessary tools in sequence or parallel to deliver results. Here's a step-by-step guide to creating and using Grok Tasks: Step 1: Define Your Task Start by clearly outlining the daily activity or goal. Consider what inputs you have (e.g., a URL, a query, or an attachment) and what output you need (e.g., a summary, calculation, or visual analysis). Break it down into subtasks to identify tool needs. For example, if your task involves researching current events, note that you'll need search and browsing capabilities. Step 2: Review Available Tools Familiarize yourself with the tools Grok can access. Here's a quick overview: - Code Execution: Run Python code for calculations, data processing, or simulations using libraries like numpy, pandas, or sympy. - Browse Page: Fetch and summarize content from any website URL with custom instructions. - Web Search: Perform general internet searches, returning results with optional operators like site:. - Web Search With Snippets: Get quick, detailed excerpts from search results for fact-checking. - X Keyword Search: Advanced search for X posts using operators like from:, since:, or filter:. - X Semantic Search: Find semantically related X posts based on a query, with filters for dates or users. - X User Search: Locate X users by name or handle. - X Thread Fetch: Retrieve a full X post thread, including context like replies and parents. - View Image: Analyze an image from a URL or conversation ID. - View X Video: Extract frames and subtitles from an X-hosted video. - Search PDF Attachment: Query a PDF file for relevant pages using keyword or regex modes. - Browse PDF Attachment: View specific pages of a PDF with text and screenshots. Select tools that align with your task. Aim for a mix to handle data gathering, processing, and visualization. Step 3: Craft Your Prompt Write a detailed prompt to Grok describing the task. Include: - The overall goal. - Specific steps or subtasks. - References to tools if you want to guide the process (e.g., "Use web_search to find sources, then code_execution to analyze data"). - Any constraints, like dates or limits. Example prompt: "Create a Grok Task for my morning routine: Search recent X posts about tech news using x_keyword_search, fetch a key thread with x_thread_fetch, and summarize with browse_page on linked articles." Step 4: Submit and Interact Send your prompt to Grok. It will process the task by calling tools as needed, often in parallel for efficiency. Review the output and refine with follow-up prompts if required (e.g., "Expand on that using view_image for visuals"). Iterate to fine-tune the workflow for reuse. Step 5: Save and Reuse Once refined, note the prompt as a template for future use. You can adapt it for similar tasks, making Grok Tasks a habitual part of your day. Finding Grok Tasks To discover existing Grok Tasks or inspiration for new ones, use X searches with tools like x_keyword_search or x_semantic_search (e.g., query: "Grok Tasks examples" with mode: Latest). Browse community-shared threads via x_thread_fetch, or web_search for tutorials on xAI features. Prompt Grok directly: "Show me popular Grok Tasks for productivity." 1 of 3show more

Brian Roemmele
152,242 views • 8 months ago
AI NEWS: OpenAI just launched 'Tasks', allowing users to... schedule actions and reminders within ChatGPT. Tasks can be one-time reminders or recurring actions (like a daily news rundown), with up to 10 active tasks able to be scheduled at a time. A new '4o with scheduled tasks' model will be available in the dropdown menu, and ChatGPT will also be able to suggest frequent tasks from a user's chat history. The beta feature is rolling out to Plus, Team, and Pro ChatGPT subscribers over the next few days. Is this the first (baby) step in ChatGPT's agentic era? 🤖show more

Rowan Cheung
355,674 views • 1 year ago
Jev + Kimi K3 for fraud detection! TLDR: Jev... classified 100 emails in 1.42 seconds, then I routed the uncertain cases to Kimi K3. The full pipeline got 96/100 correct for only ~$0.07. Video is not sped up, check out the live run! Here was my process: I gave Jev 100 emails to classify (a mix of 50 legit & 50 fraudelent emails). It classified all of them in 1.42 seconds. An underrated feature about Jev is it will give you the confidence score for a classification, so I routed any prediction under 95% confidence to Kimi K3 to be fully sure. 31 emails fell below that threshold. After routing those to Kimi K3, the combined pipeline reached 96% accuracy. The full run took 16 seconds & ~$0.07 in inference costs: - $0.068 from Kimi K3 on Together AI - $0.003 (1/3 of a cent) from Jev on TypeSafe AI. I think this is a really interesting pattern: use a fast specialized model like Jev for the narrow task, then route the uncertain cases to a larger LLM. I feel like this kind of approach could be a game changer for use cases like fraud or anything realtime. You can use the speed & low cost of Jev while having a larger LLM as a fallback to ensure high accuracy.show more

Hassan
59,800 views • 9 days ago
I was super intrigued when I noticed that JEV... returns not just probs for candidate choices, but also a CONFIDENCE SCORE of it's prediction... I was wondering if TypeSafe AI brought back Bayesian Networks or something to learn uncertainty. Turns out - NO. Acc to their docs, confidence is just a derived metric from the classification scores. It carries ZERO extra information. Their demo page shows: confidence = (3 × largest probability − 1) / 2 Confidence is tightly coupled with the classification scores it is measuring the uncertainty for. There is no magic here. The model can produce high probability on a wrong choice and just be confidently wrong.show more

AVB
66,354 views • 6 days ago
New open-source agent harness just landed! I got early... access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.show more

elvis
11,303 views • 1 month ago
Diffusion Transformers aren't just generative models, but also powerful... multi-modal encoders. ConceptAttention creates rich heatmaps of text concepts in images from DiT representations. This even works on real images, and can be applied to tasks like segmentation! Demo 👇show more

Alec Helbling
24,429 views • 1 year ago
GOOGLE 🔥: An upcoming Gemini Omni video model from... Google is expected to be much more advanced in video editing, capable of completing tasks like removing watermarks, replacing objects in the video, and more. It is also likely that Google will release 2 versions of this model, including a Pro variant. And I assume what we see isn't Pro? Anime sample 👀show more

🚨 AI News | TestingCatalog
179,938 views • 4 months ago
AudioLDM is a text-to-audio model by Haohe Liu et... al. It is available in the 🧨 Diffusers library from v0.15.0 onwards, and is capable of generating realistic music, sound effects and speech Try it in this Colab to generate music for yourself:show more

Sanchit Gandhi
27,143 views • 3 years ago
PhD Students - How to detect AI text in... your writing? We often use ChatGPT for writing. However, this leads to AI-plagiarized text. This can be problematic in many scenarios. For example, if you use AI text in your papers. Your research paper can get desk rejected. 🍁How to detect if there is AI text in your writing? 1. Go to and log in. 2. Click on 𝐴𝐼 𝑑𝑒𝑡𝑒𝑐𝑡𝑜𝑟 from the left menu 3. Insert your text and click on 𝐴𝑛𝑎𝑙𝑦𝑧𝑒. 4. will generate AI detection report This report shows the following. → Percentage of AI generated text → Options for converting AI text into non-AI text 🍁How good is this AI detector? SciSpace conducted a benchmarking study. In this study, the detection capability was compared with other AI-detectors. SciSpace AI detector was tested with 4000 samples. It showed an accuracy of 96%. This means it can detect AI-generated text with 96% accuracy. The study showed that SciSpace AI detector has outclassed AI detectors like GPTZero, ZeroGPT, and Grammarly. 🔴Anything you'd like to add?show more

Faheem Ullah
13,207 views • 11 months ago
Introducing Novo Launching today a new project I coded... for myself in 1 weekend in May and decided to finish this week. Novo is a dead simple to-do app that lets you "Speech-To-Tasks", or paste a huge text and organize for you. You can customize the AI and make it organize in any criteria: - Auto-tag by category - Schedule some types of tasks to certain days - Prioritize based on your own rules Try it:show more

Pedro
76,232 views • 1 year ago
"ChatGPT, find this person and follow them." GPT-6 Astra... can autonomously navigate a drone through our office to find and follow a specific person. It's the first AI model to beat a human on each of the 5 Drone-Bench tasks in at least one of its attempts.show more

Andon Labs
1,976,773 views • 16 days ago
They've added agentic commerce to the Contra ecosystem. That... means AI agents, the ones handling tasks on behalf of business owners, can now discover and buy from a creator's profile. Human clients still work the same way. This is just a new lane on top of it.show more

AI Frontliner
44,956 views • 7 months ago
subagents are just recursive agents where you can apply... different prompts + models depending on the task. since they’re just a primitive, Cursor cli can actually spawn subagents by calling cursor-agent in headless mode via shell commands. that’s what makes the cli so nice. you can extend it, experiment, and have a lot of fun exploring orchestration patterns. here’s one way to do it w. dynamic model selection: 1. create a subagents.mdc rule 2. drop in: ``` --- alwaysApply: true --- ALWAYS spawn subagents by running `cursor-agent -p [task] --output-format=text --force --model [model]` in the terminal. Each subagent should return a summary of the changes it made. Subagents should be used for ALL tasks You can adopt a fan-out pattern where you spawn subagents to perform parallel isolated tasks, and then fan-in the results. Use the following models: - `--model gpt-5` for reasoning, researching, and planning - `--model sonnet-4` for implementation ``` 3. start cursor cli and try it out you can also adjust the rule to be more explicit when it should use subagents, when not to, which models when etc.show more

eric zakariasson
57,554 views • 1 year ago
Woow Google has just rolled out the AI model... Flash Thinking 2.0 This is the first reasoning model capable of accessing YouTube and it changes everything: - Search for a video on your topic - Ask Gemini to think about the video - You'll have a tailor-made result in 10 sec. And it's even connected to Google Search and Maps!show more

Paul Couvert
79,489 views • 1 year ago
I used Jev to classify 1,018 AI research papers.... The result: $0.08 total cost and 256ms median end-to-end latency per paper. The pipeline was: 1. Summarize each paper with DeepSeek V4 Flash 2. Send the title + summary + 24 possible topics to Jev 3. Use Jev to classify each paper 4. Visualize everything on The summaries cost $3.99 on Together AI. The classifications cost $0.08 on TypeSafe AI. So for just over $4 of inference, I ended up with a pretty useful way to explore the top AI research papers from the past year. I think this is where things are heading: different models for different parts of the workflow, instead of using one model for everything. I’m running evals on the Jev classifications before replacing the current ones, but the site is already live:show more

Hassan
333,468 views • 9 days ago