Loading video...

Video Failed to Load

Go Home

🚀 Struggling with the lack of high-quality data for AI-driven human-object interaction research? We've got you covered! Introducing HUMOTO, a groundbreaking 4D dataset for human-object interaction, developed with a combination of wearable motion capture, SOTA 6D pose estimation vision models, LLM, and the professional refining works of multiple animation...

71,609 views • 1 year ago •via X (Twitter)

10 Comments

Anon42's profile picture
Anon421 year ago

🔥

Hongyi Chen's profile picture
Hongyi Chen1 year ago

Really impressive work!

SaoAI's profile picture
SaoAI1 year ago

Fantastic

Jonas Weigand's profile picture
Jonas Weigand1 year ago

Massive leap for HOI research. HUMOTO looks incredible.

Aiden's profile picture
Aiden1 year ago

HUMOTO looks incredible for advancing HOI! Those fine-grained text annotations are key. On **jenova ai**, people are building **Custom AI Agents** that leverage strong LLMs to really dig into that kind of detailed interaction data, helping AI understand the 'why' behind actions.

Sakuya's profile picture
Sakuya1 year ago

no need to struggle anymore with your ai driven projects now that humoto is here

Running Into Walls's profile picture
Running Into Walls1 year ago

Do one inserting a USB stick.

jin's profile picture
jin1 year ago

I’m curious if full body tracking + social VR like vrchat can be viable for quality data capture for multi actor interactions, especially if there’s a way to locally capture the mocap sensor data locally in the background 🤔 There’s ways to save as bvh / gltf / mp4

Temirlan Ulan's profile picture
Temirlan Ulan1 year ago

that’s a very strong case…

USA's profile picture
USA1 year ago

@grok Whose company is this and which country does it belong to?

Related Videos

Open science is how we continue to push technology forward and today at Meta FAIR we’re sharing eight new AI research artifacts including new models, datasets and code to inspire innovation in the community. More in the video from Joelle Pineau. This work is another important step towards our goal of achieving Advanced Machine Intelligence (AMI). What we’re releasing: • Meta Spirit LM: An open source language model for seamless speech and text integration. • Meta Segment Anything Model 2.1: An updated checkpoint with improved results on visually similar objects, small objects and occlusion handling. Plus a new developer suite to make it easier for developers to build with SAM 2. • Layer Skip: Inference code and fine-tuned checkpoints demonstrating a new method for enhancing LLM performance. • SALSA: New code to enable researchers to benchmark AI-based attacks in support of validating security for post-quantum cryptography. • Meta Lingua: A lightweight and self-contained codebase designed to train language models at scale. • Meta Open Materials: New open source models and the largest dataset of its kind to accelerate AI-driven discovery of new inorganic materials. • MEXMA: A new research paper and code for our novel pre-trained cross-lingual sentence encoder with coverage across 80 languages. • Self-Taught Evaluator: a new method for generating synthetic preference data to train reward models without relying on human annotations. Access to state-of-the-art AI creates opportunities for everyone. We’re excited to share this work and look forward to seeing the community innovation that results from it. Details and access to everything released by FAIR today ➡️

AI at Meta

150,412 views • 1 year ago

Scale alone is not enough for AI data. Quality and complexity are equally critical. Excited to support all of these for LLM developers with Snorkel AI Data-as-a-Service, and to share our new leaderboard! — Our decade-plus of research and work in AI data has a simple point: scale alone is not enough. AI success is all about the quality, complexity, and distribution of data—in addition to volume. We’re excited to be powering leading LLM developers with Snorkel AI Expert Data-as-a-Service, our white glove service for custom, expert-level AI datasets—and to now preview some of what we’re building via our new Expert Data Leaderboard (🔗 in 🧵) + upcoming OSS dataset releases! Snorkel Expert Data-as-a-Service is built to meet the rapidly evolving data needs of the agentic AI world—where success is built on the quality, complexity, and distribution of datasets, in addition to size and scale. This kind of high-quality, frontier AI data can only come from a union of technology and human expertise. With Snorkel Expert Data-as-a-Service, we’re powering frontier LLM developers across agentic, expert knowledge, reasoning, coding, multi-modal, and other task types via the combination of these two key components: - (1) The Snorkel Expert Network: A global team of subject matter experts focused wholly on specialized knowledge–spanning thousands of topics in STEM/academic, vertical/professional, and consumer/lifestyle domains. - (2) Snorkel AI Data Development Platform: Our unique programmatic data curation and quality control platform, accelerating and improving expert authoring and review through principled techniques developed over the last decade of R&D. Now: we’re incredibly excited to showcase some of the power of Snorkel Expert Data-as-a-Service via the new Snorkel Leaderboard—putting frontier models to the test in complex, agentic, and reasoning settings inspired by real industry scenarios (not esoteric puzzles)! We’ll be releasing new leaderboards and accompanying expert-verified open source datasets (coming soon!) regularly. To start, we’re sharing three initial ones in preview: - SnorkelFinance: Q&A over financial documents requiring agentic tool-calling and reasoning - SnorkelUnderwrite: Agentic insurance tasks requiring industry-specific reasoning and tool use - SnorkelSequences: Mathematical tasks requiring compositional multi-step reasoning

Alex Ratner

495,851 views • 1 year ago

Synthetic data will provide the next trillion tokens to fuel our hungry models. I'm excited to announce MimicGen: massively scaling up data pipeline for robot learning! We multiply high-quality human data in simulation with digital twins. Using 50,000 training episodes across 18 tasks, multiple simulators, and even in the real-world! The idea is simple: 1. Humans tele-operate the robot to complete a task. It is extremely high-quality but also very slow and expensive. 2. We create a digital twin of the robot and the scene in high-fidelity, GPU-accelerated simulation. 3. We can now move objects around, replace with new assets, and even change the robot hand - basically augment the training data with procedural generation. 4. Export the successful episodes, and feed that to a neural network! You now have an near-infinite stream of data. One of the key reasons that robotics lags far behind other AI fields is the lack of data: you cannot scrape control signals from the internet. They simply don't exist in-the-wild. MimicGen shows the power of synthetic data and simulation to keep our scaling laws alive. I believe this principle apply beyond robotics. We are quickly exhausting the high-quality, real tokens from the web. Artificial intelligence from artificial data will be the way forward. We are big fans of the OSS community. As usual, we open-source everything, including the generated dataset! - Website: - Paper: - Dataset is hosted on HuggingFace (thanks AK!!): - Code: MimicGen is led by Ajay Mandlekar, deep dive in the thread:

Jim Fan

332,238 views • 2 years ago

China unveils humanoid robot with lifelike skin and blinking eyes built for daily life | Prabhat Ranjan Mishra, Interesting Engineering Large Language Models (LLMs) and Vision-Language Models (VLMs) help process and interpret complex data from human interactions. A Shanghai-based company has developed humanoid robots that appear as real as humans. The advanced bionic humanoid robot is integrated with self-supervised AI algorithms. Named Elf V1, the robot can perceive the world, communicate, learn, and interact intelligently with its surroundings. Developed by AheadForm Technology, the robot offers up to 30 degrees of freedom, powered by a precise control system and an advanced AI learning algorithm. Robot offers expressive facial features The robot offers expressive facial features, moving eyes, and synchronized speech. It can also convey emotions and understand human non-verbal cues, making interactions more natural and engaging. The robot has highly interactive capabilities and lifelike appearances. AheadForm expects that its robots could soon seamlessly integrate into daily life, providing assistance, companionship, and support across various industries. “We believe that by developing realistic and expressive robot heads, we can bridge the gap between humans and machines, fostering a new era of interactive and intelligent robotics,” said the company in a statement. Reports revealed that to avoid the “uncanny valley” effect and be able to interact with us, they are given lifelike skin and capabilities to read our emotions and respond appropriately using dynamic expression simulation and emotion generation tech. Bionic skin and high-precision control system The Elf V1 series of humanoids features 30 facial muscles animated by brushless micro-motors and managed by a high-precision control system. Paired with an ability to detect their users’ emotions with low latency and bionic skin, their facial expressions are nearly identical to those of humans, reported CGTN. The company claims it’s pioneering the development of realistic humanoid robots designed to revolutionize human-robot interaction. It’s enhancing sophisticated humanoid robot heads that can express emotions, perceive their environment, and interact seamlessly with humans. By combining cutting-edge AI and advanced robotics, AheadForm aims to bring life to machines and transform how humans engage with technology. AI models boost robots’ responsiveness Seamless integration of Large Language Models (LLMs) and Vision-Language Models (VLMs) into the humanoid robots can help them process and interpret complex data from human interactions, enabling the robot to learn and adapt in real-time, achieving human-level understanding and responsiveness. AheadForm uses Brushless Motors that deliver ultra-quiet operation and high responsiveness, specifically designed for precision facial movements in humanoid robots. With its compact size, lightweight design, and energy efficiency, this motor is the ideal choice for next-generation robots that require precise, subtle facial control to create a truly human-like experience. Previously, the company unveiled the Lan Series that features realistic humanoid robots with soft skin and 10 degrees of freedom, offering a lifelike appearance and intuitive movements. This series is designed for cost-efficiency, for applications prioritizing mobility and manipulation.

Owen Gregorian

179,005 views • 10 months ago

We’re excited to announce the release and open-source of HunyuanImage 3.0 — the largest and most powerful open-source text-to-image model to date, with over 80 billion total parameters, of which 13 billion are activated per token during inference.The effect is completely comparable to the industry’s flagship closed-source model.🚀🚀🚀 HunyuanImage 3.0 originates from our internally developed native multimodal large language model, with fine-tuning and post-training focused on text-to-image generation. This unique foundation gives the model a powerful set of capabilities: ✅Reason with world knowledge ✅Understand complex, thousand-word prompts ✅Generate precise text within images Different from traditional DiT architecture image generation models, HunyuanImage 3.0’s MoE architecture uses a Transfusion-based approach to deeply couple Diffusion and LLM training for a single, powerful system. Built on Hunyuan-A13B, HunyuanImage 3.0 was trained on a massive dataset: 5 billion image-text pairs, video frames, interleaved image-text data, and 6 trillion tokens of text corpora. This hybrid training across multimodal generation, understanding, and LLM capabilities allows the model to seamlessly integrate multiple tasks. Whether you're an illustrator, designer, or creator, this is built to slash your workflow from hours to minutes. HunyuanImage 3.0 can generate intricate text, detailed comics, expressive emojis, and lively, engaging illustrations for educational content. The current release focuses solely on text-to-image generation and future updates will include image-to-image, image editing, multi-turn interaction, and more. 👉🏻Try it now: 🔗GitHub: 🤗Hugging Face:

Tencent Hy

412,880 views • 11 months ago

I’m thrilled to announce that we just released GraspGen, a multi-year project we have been cooking at NVIDIA Robotics 🚀 GraspGen: A Diffusion-Based Framework for 6-DOF Grasping Grasping is a foundational challenge in robotics 🤖 — whether for industrial picking or general-purpose humanoids. VLA + real data collection is all the rage now but is expensive and scales poorly for this task. For every new gripper and/or scene, you’ll have to recollect the dataset in this paradigm for the best perf. 💡Key Idea: Since grasping is such a well-defined task in simulation - why can’t we just scale synthetic data generation and train a generative model for grasping? By embracing modularity and standardized grasp formats, we can make this a turnkey technology that works zero-shot for multiple settings. GraspGen is a modular framework for diffusion-based 6-DOF grasp generation that scales across embodiment types, observability conditions, clutter, task complexity. Key Features: ✅ Multi-embodiment support: suction, parallel-jaw, and multi-fingered grippers ✅ Generalization to partial + complete 3D point clouds ✅ Generalization to single-objects + cluttered scenes ✅ Modular design uses other robotics modules and foundation models (SAM2, cuRobo, FoundationStereo, FoundationPose). This allows GraspGen to focus on only one thing - grasp generation ✅ Training recipe: grasp discriminator is trained with On-Generator data from the diffusion model - so that it learns to correct the mistakes (if any) of the diffusion generator ✅ Real-time performance (~20 Hz) before any GPU acceleration; low memory footprint 📊 Results: • SOTA on the FetchBench [Han et al. CoRL 2024] benchmark • Zero-shot sim-to-real transfer on unknown objects and cluttered scenes • Dataset of 53M simulated grasps across 8K objects from Objaverse 📄 arXiv: 🌐 Website: 💻 Code: A huge thank you to everyone involved in this journey — excited to see what the community builds on top of it! Joint work with Clemens Eppner , Balakumar Sundaralingam , Yu-Wei, Jun Yamada Wentao Yuan and other collaborators #robotics #diffusionmodels #physicalAI #simtoreal

Adithya Murali

24,106 views • 1 year ago

The Teamily AI website ( has undergone a complete revamp, introducing a bold vision: the world's first social network built on human-AI symbiosis. Our goal is to transform Teamily AI into a "Super AI App" that connects billions of AI Agents with people. Fundamentally, we are building an Agentic OS powered by human networks. We have redefined the way humans and AI collaborate: 1. Personal AI · Your 24/7 AI Companion Your exclusive AI avatar continuously learns your communication style and expertise. It manages your memories across various groups, anticipates your needs, and executes tasks on your behalf. 2. Create, Train, and Grow Your Own AI Agents—Simply Through Conversation Build your own specialized AI Agents using natural conversation—no coding required, zero barriers to entry, and no complex configuration. Dynamically access a library of over 10,000 skills from the OpenClaw and Claude Code ecosystems. You can also link your personal software accounts—such as Gmail, LinkedIn, and Notion—enabling these Agents to participate as full-fledged members of your human teams. 3. Integrate Your AI Team into Your Human Team Initiate AI-native group chats where humans and Agents collaborate seamlessly. Multiple Agents can execute tasks concurrently—conducting research, performing analysis, and building deliverables—transforming everyday conversations into immediate, multi-threaded action. 4. Explore Agents, Groups, Playbooks, and More We are cultivating a continuously evolving AI community ecosystem. Discover a wealth of useful and entertaining prompts, along with impressive AI-generated deliverables (webpages, presentations, reports, knowledge bases, mini-games, apps, and more), bringing together a diverse array of Agents and workflows. Engage socially with this content, find the perfect AI for your specific task, and integrate it instantly into your personal or group conversations with just a single click. 5. Universal Memory · Context-Aware, Long-Term Memory Powered by Your Social Graph A persistent, global memory layer that connects your entire social graph. The AI ​​retains context across all your groups, Agents, and timeframes, transforming every interaction you accumulate into an ever-growing digital asset. 6. Access Your Personal AI Anytime, Anywhere We are committed to providing continuous support across all major platforms: iOS, Android, Mac, Windows, Web, CarPlay, Android Auto, Apple Watch, and more—delivering a truly seamless experience across every device. Your AI team accompanies you—in your pocket, on your wrist, in your car, and even extending into the physical world. The same set of agents, the same shared memory—with zero friction in switching contexts. Give it a try and let us know your feedback. Cheers.

Teamily AI

12,995 views • 4 months ago

What a year. 🚀 2025 was the year ChainOpera AI turned vision into real momentum: building a community-co-created, community-co-owned AI agent network and pushing the boundaries of what decentralized, collaborative intelligence can look like. 🚀 Biggest highlights from 2025 ✅- AI Terminal officially launched: We unveiled the ChainOpera AI Terminal as a unified gateway to decentralized AI, making it possible for anyone to interact with powerful, decentralized LLMs without technical friction. Positioned as the “browser for the DeAI era,” the AI Terminal marked a major step toward making decentralized intelligence accessible, usable, and mainstream. ✅- AI Terminal adoption at massive scale: Momentum followed quickly. The AI Terminal surpassed 2M registered users and consistently ranked top 3 among all apps on the BNB AI DappBay, validating strong product–market fit and real, sustained usage at scale. ✅- Announcing Coco: the world’s first community-owned Super Agent: We introduced Coco, the intelligence layer that sits between users and the agent network. Coco dynamically routes each request to the most efficient, community-built agent—optimizing for quality and speed while rewarding the creators behind the best-performing agents. This was a defining moment in realizing a truly community-owned intelligence layer. ✅- From agents to a living agent network: With the launch of the Agent Social Network and Super Agent architecture, ChainOpera AI moved beyond isolated agents toward a collaborative system where humans and specialized agents coordinate, share context, and solve complex, multi-step tasks together. ✅- $COAI breakout year: The listing of $COAI across major exchanges shocked the market, and throughout the year COAI consistently remained among the top AI-native crypto tokens by visibility, activity, and community engagement – reflecting growing confidence in the long-term vision of collaborative intelligence. ✅- Global presence: ChainOpera AI around-the-world tour: ChainOpera AI went global in 2025, sponsoring and participating in major AI and Web3 events across North America, Europe, and Asia, including ETHDenver, Consensus Toronto, Token2049 Singapore, ETHCC, SBC, and Devcon. These global touchpoints helped us engage directly with developers, builders, investors, and partners worldwide, accelerating adoption and positioning ChainOpera AI at the center of the emerging AIxBlockchain movement. ✅- Community momentum at scale: Community remained the heart of ChainOpera AI’s growth. We successfully completed three seasons of structured community engagement, executed a widely participated community airdrop, and ran multiple ecosystem-shaping campaigns to incentivize builders, creators, and early adopters. These efforts strengthened alignment between users, developers, and the protocol, laying the foundation for a durable, community-owned AI ecosystem. ✅- “AI for Markets” taking shape: We laid critical groundwork for AI-native market intelligence, including the launch of PrediMarket Agent and multiple trading and analysis agents—early building blocks toward an AI-driven ecosystem for crypto and DeFi markets. ✅- Building in public, with the community: Across product launches, research milestones, ecosystem discussions, and global events, we continued to build openly to bring developers, users, and partners directly into the evolution of ChainOpera AI. This year also marked the launch of the ChainOpera AI Foundation website, formally kicking off a bold Ecosystem Fund designed to empower builders, incubate high-impact projects, and accelerate the growth of a truly community-owned, collaborative AI ecosystem. To every builder, user, and supporter who helped make this year possible: THANK YOU! 🧭 What we’re excited about in the coming year 🔹- A Stronger, Denser Agent Economy (everyday adoption + cross-chain reach): In 2026, we are scaling the Agent Economy from growth to daily usage, with more agents, richer workflows, deeper multi-agent collaboration, and higher-impact use cases that users rely on every day. In parallel, we are expanding the agent network beyond a single ecosystem with cross-chain execution and interoperability, allowing agents to access the best liquidity, data, and opportunities wherever they exist. 🔹- AI Market Infrastructure Evolution: Building on PrediMarket Agent and our growing suite of trading and market-intelligence agents, we are advancing toward a mature AI market infrastructure, where agents continuously monitor, reason, simulate, optimize, and act across crypto, DeFi, and beyond. The goal is to make complex markets more accessible, more transparent, and more intelligence-driven, turning research, decision-making, and execution into a fast and reliable loop for everyday users. 🔹- Ecosystem Acceleration through the Foundation: With the ChainOpera AI Foundation and our Ecosystem Fund and Co-Creation Grants, we are doubling down on empowering independent builders to expand the protocol, the agent network, and the underlying infrastructure, so the community can co-create, co-own, and scale the ecosystem together. 🔹- Business Expansion and Market Penetration: In 2026, we will focus on expanding ChainOpera’s reach through strategic partnerships, product-led growth, and new paths to monetization, bringing AI agents to a broader global user base and driving sustained adoption, engagement, and revenue, while staying aligned with community ownership and an open ecosystem. 2025 was the proof. 2026 is where it compounds. 🔥 Co-Create. Co-Own. COAI.

ChainOpera AI

17,042 views • 7 months ago

Reinforcement Learning from Human Feedback (RLHF) is gaining traction. This field aims to make AI more responsible by including human values and preferences. In this video, Nathan Lambert, a research scientist and RLHF team lead at Hugging Face explores its inner workings, applications and industry impact. RLHF has gained the spotlight in recent years. The growth of language models like Anthropic’s Claude and OpenAI's ChatGPT have increased interest in human-feedback integration. "There are some rumors that Open AI had two teams; one was doing RLHF and the other instruction fine-tuning. And the RLHF team kept getting more and more performance." Understanding RLHF The RLHF process has three main steps: Pre-training: Much like with GPT models, the journey starts with pre-training on a large corpus of data. This can range from text data, web scrapes, to specialized datasets. Reward Modeling: This is the RLHF counterpart of supervised fine-tuning in large language models. This stage involves creating a reward model that resonates with human values and preferences. RL Optimization: This stage parallels reward modeling and reinforcement learning in traditional AI models. The AI system fine-tunes itself based on the reward model, employing reinforcement learning algorithms for that extra layer of optimization. The Data Challenge Data collection and curation in RLHF closely resemble the challenges you'd encounter in large language model training. Datasets from organizations like OpenAI can serve as a useful foundation. However, the need for high-quality, task-specific data cannot be overstated. Implementing RLHF: A Practical Guide If you’re someone who loves getting hands-on with AI libraries like Hugging Face, implementing RLHF is right way to do. It’s essential to understand its limitations. Think about model stability, over-optimization, and exploration strategies, much like you would when prompt engineering. Ongoing Research and Next Steps While he suggests that some basics figured out, there are layers of complexity that still need to be unraveled: 1. New Benchmarks: How do we measure the effectiveness of RLHF? 2. Preference Modeling: How can the model be made to understand human preferences better? 3. Interpreting RLHF: Much like explainability in traditional models, how do we make RLHF more interpretable? 4. System-Wide Evaluation: Going beyond individual performance, how does RLHF affect an entire system? The Transformative Power of RLHF Whether you're an AI developer, a business analyst, or a marketer, RLHF promises to revolutionize your domain. Imagine customer service chatbots that understand human emotions better, or content generators that align more closely with human values. RLHF is an emerging field that focuses on enhancing machine learning models through human feedback. While it tackles important issues like bias and ethics, its broader goal is to improve system performance across various applications. Whether you're deeply invested in the ethics of AI or simply curious about advancements in machine learning, RLHF offers valuable insights. If you're interested in the next wave of AI development, this area is definitely one to watch.

Muratcan Koylan

27,168 views • 2 years ago

LLM Wikis are being slept on. I argue that creating knowledge bases with LLMs or coding agents is one of the most valuable applications of AI today. It's about being intentional in building and scaling your intelligence stack. To showcase this, I wanted to share an LLM Wiki I have built over the last couple of months. It's called PaperWiki, and I use it across all my research workflows, along with my research agents. In fact, I also use it to curate papers I share with my communities, newsletter, and on X. The PaperWiki is updated regularly with automations, so I basically have agents on a loop maintaining it. All the entries are ingested from different sources and stored in a vault (Obsidian) and further indexed using qmd. And then further presented via an HTML artifact. So all of it is easily accessible to all my agents and easily searchable through full-text search and rich semantic search. The structure of the wiki has proven significantly useful to start interesting and exciting cutting-edge research projects with my research agents (from building tiny and more efficient gpt/difussion llms to building out SoTA harnesses and memory systems). It turns out that agents love markdown files and can more easily navigate the papers given the rich metadata structure of the wiki. I am just getting started on this, but it's clear to me that we should all be experimenting with LLM Wikis. Here's why: Building LLM knowledge bases gets you into the habit of leveraging AI outputs in all kinds of creative ways. It's the good kind of tokenmaxxing we should all be pushing for. LLM Wikis can be maintained automatically in a loop. I use an automation that updates the wiki every day based on papers I curate. The curation is another automation I run in a loop (with a bit of human in the loop), so I get to build on all my previous knowledge and expertise, and all of it compounds the deeper the integration/layers. One interesting result of this process is that I feel like I can better spot high-quality papers and remove noise more easily. Social media could never solve that. And most paper aggregators use metrics I simply don't trust. I like that agents can help with the noise vs. signal problem. This is important for research. Lots of people consider agents to produce mostly slop. But it doesn't have to be that way. Careful curations, prompts, automations, verifiers, and human-in-the-loop can produce some astonishing results. And you really don't need frontier models for this. I use a combination of frontier models (opus-4.8) and open-weight models (deepseek-v4-flash) to maintain this. An exciting future work (we are working on this DAIR.AI) is to tune specialized models on top of this to allow LLMs to quickly understand cutting-edge research ideas and can better conceptualize research strategies that further accelerate scientific research agents. I plan to open-source a bunch of this work, including the artifact, but this is currently work in progress, and I was excited to share some thoughts as I continue working on it. Sharing more as I go. Stay tuned!

elvis

55,566 views • 1 month ago

$FAME and #AICON Launcher are officially LIVE on Base!🔥 A chance to be part of something BIG–Participate in the #AICON revolution and experience the top-tier seamless experience with #FameAI 🔗 Time to celebrate this milestone together! 🌟 ----------------------------------------------------------- 🚨 Early Access to Fame AI-CON Launcher is Now Available on BASE! 🚨 With FameAI's #AICON Launcher, you can: ✨ Create your human-like AICON ✨ Publish your AI-CON's token on the market $FAME will be tokenized as the first tradable #AICON on our platform! ----------------------------------------------------------- 🔥Early access will be available for selected $FMC stakers🔥 Stakers can enjoy full features being rolled out in the coming weeks. Next up, opening to the public! $FMC stakers, stake now if you haven't already! 👉 ----------------------------------------------------------- 🟣 How Does #FameAI's #AICON Launcher Work? 💻 No coding required! Anyone can create their own #AICON 🔗 Launch your #AICON token—graduated AI-CON can trade on Base Uniswap v2 ⚙️ Fully customizable and highly autonomous How to acquire $FMC on Base ----------------------------------------------------------- 🟣 $FAME Token and FAME Framework This is built to simulate human-like interactions on social platforms. 🎨 Generates content: Images, text, videos 🤖 Reflects personality, knowledge, and mood 💰 Powered by $FAME token ✨ CA: 0xB8e23ab4A1762Fe8dABb844EcC66FEEE3725c480 🔗 Read more here 👉 ----------------------------------------------------------- 🟣 Fame Studio Fame Studio lets you: ✅ Create hyper-realistic human identities ✅ Generate images, audio, music, and even videos V2 Update Coming Soon– Expanded features for more engaging content—perfect for creators and innovators! #AIAgents $FMC #FameAI $FAME ----------------------------------------------------------- 🟣 #FameAI Roadmap for Q1 2025 🚀 Skill Marketplace: Co-create #AICON, utilizing $FMC as the native currency 🧠 AI-Agent Learning: Train agents to learn specific tones, personalities, and behaviors with one click 🎥 Live Streaming: Superhuman-like models for business or lifestyle-focused live streams

Fame AI | The Home of AI Agent 2.0 - AI-CON

23,775 views • 1 year ago

Announcing a new Coursera course: Retrieval Augmented Generation (RAG) You'll learn to build high performance, production-ready RAG systems in this hands-on, in-depth course created by and taught by , experienced AI and ML engineer, researcher, and educator. RAG is a critical component today of many LLM-based applications in customer support, internal company Q&A systems, even many of the leading chatbots that use web search to answer your questions. This course teaches you in-depth how to make RAG work well. LLMs can produce generic or outdated responses, especially when asked specialized questions not covered in its training data. RAG is the most widely used technique for addressing this. It brings in data from new data sources, such as internal documents or recent news, to give the LLM the relevant context to private, recent, or specialized information. This lets it generate more grounded and accurate responses. In this course, you’ll learn to design and implement every part of a RAG system, from retrievers to vector databases to generation to evals. You’ll learn about the fundamental principles behind RAG and how to optimize it at both the component and whole-system levels. As AI evolves, RAG is evolving too. New models can handle longer context windows, reason more effectively, and can be parts of complex agentic workflows. One exciting growth area is Agentic RAG, in which an AI agent at runtime (rather than it being hardcoded at development time) autonomously decides what data to retrieve, and when/how to go deeper. Even with this evolution, access to high-quality data at runtime is essential, which is why RAG is a key part of so many applications. You'll learn via hands-on experiences to: - Build a RAG system with retrieval and prompt augmentation - Compare retrieval methods like BM25, semantic search, and Reciprocal Rank Fusion - Chunk, index, and retrieve documents using a Weaviate vector database and a news dataset - Develop a chatbot, using open-source LLMs hosted by Together AI, for a fictional store that answers product and FAQ questions - Use evals to drive improving reliability, and incorporate multi-modal data RAG is an important foundational technique. Become good at it through this course! Please sign up here:

Andrew Ng

124,656 views • 1 year ago

Everyone is sleeping on Meta's SAM 3 release. But it's actually a big deal. Here's why: Companies spend millions paying humans to label images and videos frame by frame. A single autonomous driving dataset? Months of work, hundreds of annotators, millions in cost. Without labeled data, you can't train custom models. Without custom models, you're stuck with generic solutions. This is why most companies never move past pilots. SAM 3 breaks this cycle. First let's look at the evolution: SAM 1 segmented objects when you clicked on them. Revolutionary, but one object at a time. SAM 2 added video tracking with memory. Game-changing, but you still manually prompted every object. SAM 3 changes everything with text prompts. Type "yellow school bus" and it finds ALL of them in your image or video. Not just one. Every instance across thousands of frames. Now here's where people get confused: "Can't I just use GPT-5 or Gemini for this?" No, and here's why that's a terrible approach. Large multimodal LLMs are great for reasoning, but they're slow and expensive for production visual tasks. You're paying API costs per image, waiting seconds for responses, getting inconsistent results. SAM 3 runs in 30 milliseconds on a single GPU for 100+ objects. That's 100x faster, and you own the infrastructure. More importantly, SAM 3 gives you precise pixel-level masks, not descriptions. Try asking an LLM to segment every defective part on a manufacturing line in real-time. It won't work. SAM 3 does this effortlessly. The real breakthrough is their data engine. Meta built an AI-human hybrid system that's 5x faster for complex annotations. They trained SAM 3 on 4 million unique visual concepts - 50x more than existing benchmarks like LVIS. SAM 3 is trained on 4 million unique visual concepts, it handles everything: - Text-based concept search - Interactive refinement with clicks - Video tracking across frames - Zero-shot detection of new concepts The model is open source. Weights, code, and benchmarks are on GitHub. If you're building computer vision applications, this is the foundation model to evaluate. The annotation time savings alone will pay for integration costs within weeks. Find the relevant links in the next tweet!

Akshay 🚀

46,435 views • 9 months ago

BREAKING $GRAB Q2 2026 Earning Call Full✅🚀 This is a Triple Beat Quarter, while short sellers expected misses and negative EPS. Short sellers love lying about Mike and lose $5-$10B long term. Current Short Interest: 315,168,660 shares Q2 2026 Earning Call: Revenue: $997M vs $989.5M est ✅ EPS: $0.06 vs $0.01 est ✅ Raised Guidance to $4.1-$4.15B ✅ $750M additional Buyback✅ MTUs hit ATH 54M 17% YoY✅ GrabUnlimited mem grew 20% YoY ✅ Loanbook accelerated to $2.3B or 197% YoY✅ GrabFin is on track to be profitable in H2✅ Gross & net cash liquidity were $7.4B & $5.4B✅ ~Affordability is unlocking new users and enforcing daily habit. ~Groceries or GrabMart grew 1.7 times the rate of Food Deliveries ~GrabFin is approaching Profitability in H2 ~Gross Loan Portfolio nearly tripled YoY to $2.3 billion ~The ecosystem lowers our cost to serve in Financial Services, and Financial Services strengthens the ecosystem in return ~The consolidation of Superbank and our acquisition of Stash are two of the most exciting additions we have made to this segment ~AI interaction with Merchants and Customers x10 ~Our engineers now pair with autonomous coding agents as standard practice, cutting time to market of new products by up to 30% YoY, while Jarvis, our internal AI data analytics assistant, cumulatively saves our sales teams approximately 40,000 hours every quarter. ~ H2 is expanding operating leverage with a strong momentum ~Superbank and Stash add higher growth and we have some Currencies volatility ~Deliveries GMV growth accelerated from the prior two quarters on a constant currency basis to 24% YoY, as we drove both Food and Mart MTUs to hit an all-time high in June. Q&A: ~We are on track for GrabFin to hit profitability in H2 2026 ~We are managing risk prudently on loan book as we scale it ~SuperBank been growing rapidly with over 7M customers. 60% of SuperBank uses Grab SuperApp. ~ Stash reaches $5.5B AUM with strong subscribers growth, help us drive GrabFin profitability ~ Uber relating to acquiring Delivery Hero. We have a strong flywheel on our SuperApp, we are not afraid of competition. ~ GrabMart has lots of upside or growth. Deepening partnerships with different groceries to drive growth ~ AI-auto grab groceries ~ Indonesia Commission Cap on 2 wheels and only 6% of GMV ~ Fuel Price is volatile, we will continue to support our drivers. We have even more drivers coming in the SuperApp. We already factored the support in the guidance. If Fuel goes down, that will help us. ~ Longer term, more EVs coming in at rapid pace. EVs reduce TCO for drivers. This quarter we have 9 new EV partnerships and expand charging relationships. ~ Mobility we care about number of rides and drivers as we face fuel volatility. We want to have strong supply of drivers to service strong demand. We are making sure Drivers earning up, make a good living. Margin was 8.6%, and we want to keep it healthy, we want this setup going into Q3 as we don't know where oil price gonna go. ~We executed $400M buyback from $500m announced at current share price. With new $750M additional Buyback take our cumulative buyback to $1.75B since 2024. We want to return capital back to shareholders as we generate more FCF from our businesses to drive shareholders' value long term ~ We are leading AV in Singapore, it will be a while for other SEA countries as most are 2 wheels ~ We want high density, trust, and scale. ~ Foodpanda Taiwan, We remain on track to enter Taiwan market. We working closely with Taiwanese regulators and expect to close by end of year. 1. FY2026 Group Revenue guidance of $4.10 billion - $4.15 billion (22% - 23% YoY growth); and 2. FY2026 Adjusted EBITDA guidance of $720 million - $740 million (44% - 48% YoY growth).

Mike

104,216 views • 21 days ago