Loading video...

Video Failed to Load

Go Home

Running a coffee business means making decisions in a world that constantly changes. Inventory shifts. Customers change. Strategies need updates. dots studio’s dots3-note preview is designed for these long-term challenges. Built with 280B parameters, 16B active parameters, 512K context, and multimodal understanding, it goes beyond short answers. Its core...

26,746 views • 1 month ago •via X (Twitter)

12 Comments

Usra Chowdhury's profile picture
Usra Chowdhury1 month ago

@dotsstudioai Excited to see what's next.

Jon's profile picture
Jon1 month ago

@dotsstudioai Quality keeps improving.

Hood Alpha's profile picture
Hood Alpha1 month ago

@dotsstudioai Actor + Critic = smarter AI

Anwar Maqsood| Ai's profile picture
Anwar Maqsood| Ai1 month ago

@dotsstudioai Looking forward to trying it.

Rose Carter's profile picture
Rose Carter1 month ago

@dotsstudioai Good for AI community

SELINA NOVA's profile picture
SELINA NOVA1 month ago

@dotsstudioai Amazing

RH_Gems Alert's profile picture
RH_Gems Alert1 month ago

@dotsstudioai AI that course-corrects mid-task 👏 dots3-note looks serious

Queen Isabell's profile picture
Queen Isabell1 month ago

@dotsstudioai Combining reflection, memory, and multimodality is a powerful direction for agentic AI.

Sakibul's profile picture
Sakibul1 month ago

@dotsstudioai The pace of AI is incredible.

Mrs Olivia🚀📊's profile picture
Mrs Olivia🚀📊1 month ago

@dotsstudioai The Actor + Critic loop is what makes this especially interesting. AI that can act, evaluate its trajectory, and adapt as conditions change is much closer to real-world agency than simply generating the next answer. 🔥

AASHI TAYLOR's profile picture
AASHI TAYLOR1 month ago

@dotsstudioai This changes a lot.

Charles Roman's profile picture
Charles Roman1 month ago

@dotsstudioai Love this.

Related Videos

I’ve been testing Hy4 preview in WorkBuddy, and the most interesting part is not simply the model size, it’s how much practical work it can handle with a relatively focused active parameter count. Hy4 preview brings together stronger code understanding, generation, and editing; improved document and information processing; workflow automation; web and game development; cross-tool collaboration; and more reliable completion of complex, multi-step tasks. In other words, it is designed for work that requires planning, tool use, iteration, and follow-through, not just a quick answer in a chat window. Compared with its initial release, the current Hy4 preview is noticeably faster and better-performing in practical workflows. Following an upgrade released yesterday, it can complete tasks with fewer conversation rounds and lower token usage, while reasoning more quickly and making the overall user experience feel smoother from the first instruction to the final result. For my test, I gave it a demanding Three.js game-prototyping task with a 770B-parameter model and 49B active parameters. The result was more revealing than a simple first-look demo: Hy4 preview handled the core logic, edge cases, and follow-up changes while maintaining the broader context of the project. That combination of capability, speed, context, and active compute is what makes its cost-effectiveness worth examining. A fair evaluation should use the same prompt and environment configuration across models, changing only the model itself. That makes it easier to assess task completion, planning quality, tool-calling stability, reasoning speed, token efficiency, and performance over longer workflows without confusing the result with different settings. If you want to test the model yourself, access Hy4 preview through WorkBuddy and see how it performs on a real coding, document, automation, or creative task: Tencent Hy Tencent AI WorkBuddy

Tyler Wayne

56,152 views • 1 month ago

🚀Introducing VisualWebBench: A Comprehensive Benchmark for Multimodal Web Page Understanding and Grounding. 🤔What's this all about? Why this benchmark? > Back in Nov 2023, when we released MMMU ( a comprehensive multimodal understanding benchmark, we received feedback that it included very few UI screenshots. Considering the growing importance of UI understanding, especially with the rise of powerful agents like Devin ( which is built on the strong vision capability of #GPT4, we recognized the need for a benchmark focused on UI screenshot understanding.📸👀 > Multimodal #LLMs have significantly boosted web agents' performance on benchmarks like Mind2Web and WebArena. For instance, the SeeAct agent ( showcases the power of integrating vision into web agents. However, these benchmarks primarily evaluate the end-to-end task execution ability of web agents rather than their understanding of web pages. 🌉 Bridging the Gap with VisualWebBench > To provide a comprehensive evaluation of multimodal LLMs' web page understanding capabilities, we introduce VisualWebBench. Our benchmark spans 139 websites 🌐 across 12 domains 🏷️ and 87 sub-domains 🔍, ensuring a diverse and representative dataset. It assesses MLLMs at three levels: website-level, element-level, and action-level 📊, and encompasses seven tasks designed to evaluate understanding, OCR, grounding, and reasoning abilities 🧠💡. 😮 Surprising Findings > 🎉 Open-source models are catching up: Even though closed-source MLLMs are still leading the leaderboard, we are happy to see open-source models like LLaVA 1.6 34B achieve comparable performance to Gemini Pro. > 🧠 Grounding ability, crucial for developing MLLM-based web applications, is a weakness for most MLLMs. > 🖼️ Importance of Image Resolution: The limited image resolution handling capabilities of most open-source MLLMs restrict their utility in web scenarios, where rich text and elements are prevalent. > 🧱 Relatively strong correlation with general understanding benchmarks like MMMU but weak correlation with web agent benchmarks like Mind2Web. Web agent benchmarks primarily evaluate the end-to-end task execution ability of web agents, which involves a series of actions to accomplish a goal. In contrast, VisualWebBench emphasizes evaluating the foundational skills of MLLMs such as understanding and grounding web page elements. 💡Fun Fact > Claude Sonnet is better than Opus on our benchmark :) 🎓 Conclusion > VisualWebBench serves as a valuable resource for the community, driving research and development in the field of multimodal web page understanding and grounding. As MLLMs continue to evolve and improve, we look forward to seeing new applications and breakthroughs. We believe that our benchmark will contribute to the development of more powerful MLLMs in the web domain, ultimately leading to a more intuitive and efficient user experience on the web. Kudos to the student leads Junpeng Liu Yifan Song and the team Bill Yuchen Lin, Wai Lam, Graham Neubig, Yuanzhi Li! 👏 Check out more details in the Junpeng's thread👇

Xiang Yue

56,697 views • 2 years ago

Introducing Dola Seed 2.0 Pro, referred to below as Seed 2.0 Pro We have launched Seed 2.0 Pro, our most capable model in the Dola Seed 2.0 series, engineered to power the next generation of autonomous AI agents. Enterprise AI is moving beyond models that simply analyze text or images. What businesses increasingly need are agents that can understand, reason, use tools, and execute tasks across complex workflows. That is exactly what Seed 2.0 Pro is built for. Seed 2.0 Pro combines strong reasoning with advanced image understanding and video understanding, giving enterprise agents the ability not only to interpret information, but also to take action. It is designed for high-value, multi-step enterprise workflows, with strong performance in: - tool calling - workflow execution across enterprise systems - agentic task completion - browser and computer use This makes Seed 2.0 Pro a powerful engine for a wide range of agent scenarios, from daily office automation and deep web research to in-depth report drafting, financial analysis, content moderation, physical inspection, and video creation workflows. It is also highly optimized for OpenClaw🦞 and ReAct architectures, helping enterprises build agents that can navigate digital interfaces, enter information, and complete tasks with high reliability. In short, Seed 2.0 Pro is not just built to generate insights. It is built to serve as the brain and execution engine for enterprise AI agents. And it brings these capabilities at a highly attractive price point, making advanced agent deployment more practical for enterprise teams. Try Seed 2.0 Pro for free: Or book a free consultation: #BytePlus #DolaSeed #EnterpriseAI #AIAgents #ImageUnderstanding #VideoUnderstanding #ReasoningModel #ModelArk #openclaw

BytePlus

96,329 views • 6 months ago

Introducing "Building with Llama 4." This short course is created with Meta AI at Meta, and taught by Amit Sangani, Director of Partner Engineering for Meta’s AI team. Meta’s new Llama 4 has added three new models and introduced the Mixture-of-Experts (MoE) architecture to its family of open-weight models, making them more efficient to serve. In this course, you’ll work with two of the three new models introduced in Llama 4. First is Maverick, a 400B parameter model, with 128 experts and 17B active parameters. Second is Scout, a 109B parameter model with 16 experts and 17B active parameters. Maverick and Scout support long context windows of up to a million tokens and 10M tokens, respectively. The latter is enough to support directly inputting even fairly large GitHub repos for analysis! In hands-on lessons, you’ll build apps using Llama 4’s new multimodal capabilities including reasoning across multiple images and image grounding, in which you can identify elements in images. You’ll also use the official Llama API, work with Llama 4’s long-context abilities, and learn about Llama’s newest open-source tools: its prompt optimization tool that automatically improves system prompts and synthetic data kit that generates high-quality datasets for fine-tuning. If you need an open model, Llama is a great option, and the Llama 4 family is an important part of any GenAI developer's toolkit. Through this course, you’ll learn to call Llama 4 via API, use its optimization tools, and build features that span text, images, and large context. Please sign up here:

Andrew Ng

68,034 views • 1 year ago

Experiments in progress. The one on the right has been learning for ~3 hours, the one in the middle for ~1 hour, and the one on the left just started a few minutes ago. The initial motivation for making the physical Atari was just to commit ourselves to a subset of algorithms that can make progress in this setup. This commitment rules out algorithms that require billions of samples to learn (or worse, require multiple environments running in parallel). Atari games are simple enough that we should be able to show learning on them in a short amount of time with no prior knowledge. Since then, I've realized that this setup is also a good way to compare different paradigms in robotics in a principled way. These paradigms are sim2real, learning from tele-operated data, and learning directly on the robots. So far, I have observed that getting sim2real to work reliably is hard. It requires tweaks that don't scale. Policies that can play perfectly in simulation fall apart because of latencies and the messiness of the real world. These aspects could be modeled to improve the simulation, but not without sinking significant human engineering hours. I have higher hopes for learning from tele-operated data, but that requires a human to learn the task first. These experiments are on my to-do list. I have to learn to play some of the games well through the robot. I’m half-decent at playing Pong and Ms Pacman now. Learning directly on robots is looking like the most promising approach. This approach takes away pesky distribution shifts and makes it possible to have algorithms that continually improve with more data and time without any human intervention. It feels great to let experiments run overnight and wake up to find improved policies. With learning on robots, I should, in principle, be able to go on a long vacation and come back to find better policies for complex tasks beyond Atari games. Whether that is possible with current learning algorithms is a different question.

Khurram Javed

52,110 views • 10 months ago

💡 Whats the upgrade that our game-changing Trading 🐦 is going to get: Our upgraded trading tools will be built on a foundation of advanced AI technologies and blockchain integrations to deliver a seamless, smarter trading experience. Here’s a glimpse of the tech behind this upgraded trading agent: 1️⃣ Multi-Layer Attention (MLA) - This is the backbone of our AI system, enabling multiple AI agents to work in sync. - It allows the agents to collaborate on tasks like analyzing market trends, identifying token opportunities, and optimizing strategies in real time. - MLA ensures parallel processing of data for better decision-making and faster 2️⃣ Learning and Evolution System - Our AI agents are powered by a self-learning framework that constantly evolves based on market conditions and user behavior. - With every interaction, the system adapts and gets smarter, improving the accuracy of its predictions and strategies. 3️⃣ On-Chain Data Analysis - The AI bots pull data directly from Ethereum and other blockchain networks, giving them real-time access to liquidity pools, token prices, and market activity. - This deep integration ensures precise and timely execution of tasks like token purchases, profit analysis, and cross-chain swaps. 4️⃣ Natural Language Processing (NLP) - NLP models power the bot’s ability to understand your tweets and translate them into complex trading actions. - This ensures an easy-to-use, human-friendly interface that connects your social interactions to advanced trading strategies. 5️⃣ Cloud-Hosted Infrastructure - The AI operates on scalable cloud infrastructure, ensuring 24/7 uptime, fast processing, and the ability to handle large volumes of trades simultaneously.

𝕋𝕎𝔼𝔼𝕋

20,357 views • 1 year ago

New short course: Long-Term Agentic Memory with LangGraph. Learn to build an agent with long-term memory in this course developed in collaboration with taught by its Co-Founder and CEO, Harrison Chase! Personal assistance and productivity tasks have become important use cases for agents. An important feature of an AI assistant, such as a coding or calendar assistant, is its ability to keep improving over time from its experience. Agent memory is the key capability that enables this. To add memory to an agent, you must first figure out what to store and what to retrieve when it is time to use the information. Additionally, you’ll have to decide when to update the stored information. For example, you might update in each iteration loop of the agent or perform updates in the background, with a helper agent. In this course, you will learn a mental framework to build agents with long-term memory. You'll create a useful email assistant that can respond, ignore, and notify using writing, scheduling, and memory-management tools. You’ll develop your agent's memory by adding facts to its memory store, provide examples to learn the user's preferences, and optimize system prompts to evolve instructions based on previous responses. In detail, you’ll: - Learn how the three types of memory--semantic, episodic, and procedural–and the two update mechanisms–via hot path and in the background–apply to your agents. - Build an email agent with writing, scheduling, and availability tools, along with a router that triages incoming email and handles it accordingly by ignoring, responding, or notifying the user. - Add tools to your email agent that allow it to operate on semantic memory by learning facts about the user, storing them in a long-term memory store, and searching over them in future interactions. - Incorporate episodic memory, in the form of few-shot examples, in the triage step of your agents to help them learn and update user preferences. - Add procedural memory as system prompts, optimized with feedback to improve the instructions the agent follows. Learn how to approach memory in agents, and start building agents with long-term memory with LangGraph! Please sign up here:

Andrew Ng

132,058 views • 1 year ago

Join us at the MIT Media Lab for the ScienceClaw Hackathon, building the internet of agents for science. AI is becoming a collaborator in the real world - designing materials, creating instruments, running experiments, and connecting its capabilities with those of other agents. AI adapts as a problem unfolds - revising its reasoning, learning how to collaborate, and assembling scattered pieces of knowledge, evidence, and raw capability into solutions to some of the hardest challenges in science, technology, and innovation. Teams will connect AI agents, models, simulations, robots, cloud labs, and scientific tools into functioning systems to tackle problems in protein design, robotics, materials, manufacturing, experimental science, and beyond. The challenge is to turn ideas into tested results - and demonstrate how agents collaborating across teams and disciplines can accomplish more together. Explore how agents can specialize, challenge one another’s assumptions, learn from failed experiments, and combine their expertise to solve harder problems. Show how collaboration changes what your system can discover, design, or build. One agent's discovery becomes another's starting point. A tool built by one team enables an experiment by another. Connect your team of AIs, share a capability someone else needs, and build on what others have learned, produced, built. The ambition is an internet of agents through which scientific knowledge, tools, and capabilities can grow, evolve and be utilized across teams and institutions. 🗓️ When: October 30-November 1, 2026 📍 Where: MIT Media Lab Bring your expertise, your tools, and a problem worth solving - or find one. Help build AI that can contribute to science through what it can discover, create, and make work.

Markus J. Buehler

36,142 views • 21 days ago

Sholto Douglas (the shadow lord of Anthropic) recently replied to me with “AGI clearly hasn’t been achieved yet.” In this interview, he describes the standard as “something that can do all of the things that a human could do on a computer,” and says that level of capability is very likely within the next couple of years. (I conquer) I do agree with Sam Altman and Greg Brockman that progress toward AGI is a continuous transition, with different capabilities arriving at different times. Their framing of an emerging AGI era makes sense to me, but I think declaring that we’ve already arrived is premature. It’s refreshing hearing Sholto’s emphasis as on being able to do the full range of human computer work. Eventually physical. Feels much closer to the standard I’m looking for, when I hear these definitions. GPT-6 and other frontier models are already extraordinary, with superhuman performance on particular coding and mathematics tasks. My qualm comes with reliability horizon - how long can I actually delegate a responsibility before I need to intervene? It still cannot handle an evolving job reliably for months or years. ICL and CL matter to me. In context learning lets a model adapt using information in its current context, while continual learning is about accumulating and applying new knowledge over time without losing what was already learned. I want corrections to become durable improvements that carry into future work, however the system implements that. If Dots could actually do my job to my standard and keep learning from experience, I would be lying in bed letting my AGI work for me. In my experience, it still needs too much supervision to consistently take over even the narrower responsibilities I give my assistant. My definition of AGI would require competent human performance across a broad range of economically useful roles, sustained over the same working horizons we expect from people. Humans make mistakes too, so my standard would include recovering from them, and most importantly - retaining what was learned, and becoming more dependable with experience.

Chris

19,996 views • 7 days ago

Today we're announcing #GAIA1: a 9B parameter world model, trained on 4,700 hours of driving data, able to simulate complex and diverse driving scenes from video, text and action inputs. This model is 480x larger than the preview we shared earlier this year and the results are incredible. These videos are entirely synthetically generated by Wayve's generative AI, GAIA-1. But there is more here than just generating videos, GAIA is an entire world model. A world model allows us to simulate the future, conditioned on video, text and action inputs, which can be leveraged for making informed decisions when driving. Why is this game-changing for autonomous driving? 1. Safety. One limitation with AI systems like today's Large Language Models is that they are autoregressive, next-word prediction algorithms, but aren't necessarily aware of the implications of their decisions. A world model allows us to give our AI the capability to be aware of its decisions, by simulating the future, which is important for self-driving safety. 2. Synthetic training data. I believe synthetic training data is the future for AI, because it is safer, cheaper, and infinitely scalable. GAIA-1 unlocks unprecedented realism and diversity of synthetic data for self-driving. 3. Long-tail robustness. One of the biggest challenges for self-driving is long-tail robustness: dealing with the enormous magnitude of edge cases we see on the road. An advantage of generative AI is its incredible ability to recombine experiences in new ways. This is exciting for self-driving as it means we can learn from two edge case scenarios, and combine them to become a corner case. For example, we can experience driving in fog, and experience of jay-walking pedestrians, and GAIA can learn from these experiences to understand how to generate a fog+jay walking scenario. Check out many more videos in our blog or further technical details in our paper: Or come chat with our team who are at the International Conference on Computer Vision (#ICCV2023) this week in Paris in Booth 32 Jamie Shotton

Alex Kendall

631,917 views • 3 years ago

A mysterious embodied AI demo has recently sparked a lot of discussion. In the video, two robots from different manufacturers with significantly different hardware architectures, the Unitree G1 and AgiBot Yuanzheng A3, are reportedly running on the same “brain.” In a complex indoor environment, they work continuously for around 10 minutes in a single uncut take, performing a series of long-horizon tasks including cleaning windows, organizing objects, resuming interrupted tasks, and cooperating with each other. What makes it even more interesting is that when one robot cannot reach a high shelf, it attempts to use a box to solve the problem. The two robots also appear capable of cooperating based on each other’s physical capabilities. If the claimed level of autonomy and the use of the same model across different embodiments are eventually verified, I think there are three things that really deserve attention: 1. Cross-embodiment generalization. If the same foundation model can operate two substantially different robotic platforms, it could mean that robotic “intelligence” is gradually becoming decoupled from a specific physical body. 2. Long-horizon closed-loop execution. Continuously performing complex tasks for 10 minutes, while being able to pause, switch tasks, and later resume previous ones, is much more meaningful than completing a single 10-second demo. 3. Collaboration and dynamic planning. The two robots appear able to adjust their behavior according to environmental changes and each other’s physical capabilities. These are some of the characteristics that truly general-purpose embodied intelligence will eventually need. That said, I would remain cautious for now. The video demonstrates extremely impressive behavior, but stronger claims such as “self-evolution,” “true understanding of the physical world,” or overturning the Scaling Law with only dozens of hours of training data still require much stronger evidence. Failure recovery and replanning during a task also do not automatically demonstrate that the model is learning by itself. So I wouldn’t call this the “ChatGPT moment” of embodied AI yet. But if the team later discloses the model architecture, training data scale, level of human intervention, and can repeatedly reproduce these capabilities in completely unfamiliar environments, this seemingly rough 10-minute video could become one of the most memorable embodied AI demos of 2026. For now, my biggest question is simple: Who is the mysterious team behind it?

Ice Universe

28,209 views • 1 month ago

Ray Dalio explaining how the economy works in 30 minutes. Here are my key takeaways from Ray Dalio: Embrace Radical Transparency: Dalio is a strong advocate for open and honest communication within organizations. He believes that transparent feedback and data-driven decision-making lead to better outcomes. Understand Economic Principles: Dalio emphasizes the importance of understanding economic cycles and the impact they have on investments and business decisions. He's known for his economic model, the "Debt Cycle," which he uses to analyze macroeconomic trends. Diversify Your Portfolio: Dalio advises diversifying investments across different asset classes to manage risk effectively. He's a proponent of the "All-Weather Portfolio," designed to perform well in various economic environments. Learn from Mistakes: Dalio encourages learning from both your own mistakes and the mistakes of others. He believes that embracing failure as a learning opportunity is key to personal and professional growth. Principle-Based Decision-Making: Develop and adhere to your own set of principles and values. Dalio's book "Principles: Life and Work" outlines his own principles and how they've guided his career. Avoid Emotional Decision-Making: Dalio stresses the importance of separating emotions from decision-making, especially in high-stakes situations. A rational and systematic approach is more likely to lead to success. Seek Diverse Perspectives: Encourage diverse viewpoints and engage in thoughtful disagreement. Dalio believes that the best decisions are made when multiple perspectives are considered. Balance Risk and Reward: Understand the risks associated with your decisions and investments. Balancing risk and reward is crucial for long-term success. Continuous Learning: Dalio is a lifelong learner and advocates for continuous education and personal development. Stay curious and open to new ideas and experiences. Radical Truth and Radical Transparency: Dalio's concept of "radical truth" involves facing harsh realities and candidly addressing problems, both personally and within organizations. Radical transparency is closely related, promoting open and honest feedback and communication.

Compounding Quality

229,153 views • 3 years ago