Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I usually skip benchmark comparisons. They're almost always picked to flatter. But EgoSchema is different. It measures first-person, egocentric video understanding. The AI equivalent of “look through the robot's eyes and tell me what happened.” I spent some time testing Perceptron Mk1 (“Mark One”) on the robotics workflows Perceptron...

116,425 Aufrufe • vor 2 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

I am stocked to announce that I won the OpenAI Developers Codex x Mollie Hacka Worldwide Hackathon in Paris. 60+ builders, every one of us working solo, one day to ship. I built mine around a single question: who gets to own intelligence? The default answer is scary. You hand your data to a handful of labs, they train the model, they own it, and you rent back a thin slice of what your own data made possible. That is the bargain on the table today. I do not accept it. So I built Lensemble: a Tapestry like distributed training platform for JEPA based World Models. What does it enable: World Models that a community improves together, keeps sovereign, and co-owns. Two bets sit underneath it. First, the paradigm. Language models predict the next token. Powerful for text, a dead end for the physical world. A robot does not need to autocomplete sentences, it needs to predict what happens next in the world. That is what JEPA does: it learns by predicting representations instead of pixels or tokens. I am convinced world models are the most underrated paradigm in AI right now, and the closest thing we have to a ChatGPT moment for robotics. Second, the politics. Your raw trajectories never leave your machine. Each participant trains locally against a shared protocol and ships only an update, never the data. A federated round folds those updates into one shared world model, a LeWorldModel based model, and the gain is measured, not claimed: a 12k-parameter adapter on a frozen backbone, held-out prediction error down about 12 percent, the model measurably less surprised by the world. Then the upside is split by contribution weight, so the people who improved the model own a share of what it earns. This is the thesis behind Project Tapestry, the AI Alliance and Yann LeCun's push for federated, sovereign frontier AI, carried into world models and robotics. Call it Tapestry for the physical world. All of it built solo, in a single day, with Codex as my pair the whole way. Thank you to OpenAI Codex and Mollie for backing builders who ship real things, and to Boris and the organizing crew for the room and the standard you set. Intelligence the world improves, and the world owns. That is the future I want for my kids, and the one I will keep building.

abdel

17,370 Aufrufe • vor 1 Monat

🚀 Three Next-Gen AI & Web3 Projects Are Launching on Mindo AI A new chapter for community-powered intelligence, prediction markets, and open AI infrastructure The AI + Web3 landscape is entering a decisive phase — one where real usage, real revenue, and real ownership matter more than hype. Today, MindoAI is proud to welcome three groundbreaking projects that represent this shift clearly and powerfully: Perceptron Network Space DeepNode AI Each project tackles a different bottleneck in the AI economy — data, forecasting, and infrastructure — but they all share the same vision: decentralization, community ownership, and sustainable value creation. Let’s take a deeper look 👇 🧠 Perceptron Network The world’s first community-powered AI data engine Perceptron Network is redefining how AI data is sourced, validated, and delivered. Instead of relying on expensive, closed, and slow legacy data providers, Perceptron unlocks community-powered data pipelines that are: Faster Cheaper Revenue-generating from day one This isn’t experimental AI infrastructure — Perceptron already serves real clients with real revenue, proving that decentralized data engines can outperform traditional incumbents. Why Perceptron matters: AI models are only as good as their data Centralized data monopolies slow innovation Communities can produce higher-quality data at scale By aligning contributors, validators, and clients through incentives, Perceptron turns unused human and network potential into a living data engine for AI. Launching on Mindo AI gives Perceptron access to a broader AI-native community — accelerating adoption, partnerships, and ecosystem growth. 🌌 intodotspace The first 10× leveraged prediction market on Solana intodotspace is pushing the boundaries of on-chain prediction markets. Built by the $1.5B UFO team, this platform introduces: 10× leveraged predictions Ultra-fast execution on Solana Deep liquidity and composable market design The market’s confidence is already clear — the project completed a record-breaking raise that was oversubscribed by 1,360%. What makes intodotspace different: Leverage amplifies conviction, not noise On-chain transparency replaces opaque odds Markets become real-time intelligence engines Prediction markets are often called “truth machines.” intodotspace upgrades them into high-signal, high-efficiency forecasting layers — useful for traders, protocols, DAOs, and even AI systems that need probabilistic insights. Launching on positions intodotspace at the intersection of AI-driven decision-making and on-chain market intelligence. 🌐 DeepNode AI Infrastructure for open intelligence DeepNode AI is tackling one of the biggest problems in modern AI: centralized ownership. Today, AI is dominated by a handful of corporations. DeepNode flips that model by building open intelligence infrastructure where: Anyone can deploy AI models Builders earn directly from usage Intelligence is co-owned, not extracted Backed by leading validators, miners, and ecosystem builders, DeepNode transforms AI from a closed monopoly into a shared utility. DeepNode’s core philosophy: “Own what you build — or someone else will.” This is more than infrastructure. It’s an economic redesign of AI itself: Builders keep ownership Contributors share upside Networks replace platforms Launching on connects DeepNode to creators, researchers, and communities who believe intelligence should belong to everyone — not just Big Tech. 🤝 Why This Matters for With the launch of Perceptron Network, intodotspace, and DeepNode AI, #MindoAI is rapidly becoming: A hub for AI-native Web3 innovation A launchpad for real, revenue-backed projects A meeting point for data, markets, and intelligence infrastructure These three projects don’t compete — they complement each other: Perceptron supplies data intodotspace produces market intelligence DeepNode powers open AI execution Together, they form the backbone of a decentralized intelligence economy. 🔥 The future of AI is open, composable, and community-owned — and it’s launching now on Which of these projects are you most excited about? And how do you see decentralized intelligence reshaping the next AI cycle? 👇 Share your thoughts and join the conversation.

Hồng Ngọc | Ruby💎

12,837 Aufrufe • vor 5 Monaten

Dr. Fei-Fei Li just called out the biggest blind spot in the entire AI industry. We have been building half of human intelligence. And calling it the finish line. Li: “If you look at human intelligence, it pretty much boils down to two buckets.” The first bucket is language. Symbolic reasoning. Communication. The ability to think in words and abstractions. That’s what every major AI lab has spent the last decade building. The second bucket is the one the industry has almost entirely ignored. Li: “We call that in AI spatial intelligence.” How humans and animals perceive, navigate, and interact with the three-dimensional physical world. How we reach for objects. How we move through space. How we build and manipulate physical reality. From painting masterpieces to constructing the pyramids, non-verbal spatial intelligence is what actually shapes the world. Language describes reality. Spatial intelligence acts on it. And the gap between those two things is the gap between a chatbot and a robot. Li: “When this technology is ready, the robotic revolution is gonna start. We’re already seeing that trend.” Every robot is a moving agent. Every moving agent requires spatial intelligence to function in the real world. The humanoid robots being deployed in factories right now are hitting the ceiling of what language models alone can power. Spatial intelligence is the unlock. But Li didn’t stop at robotics. Li: “From a geopolitics point of view, this is part of the technology that goes straight into weapons.” Autonomous drone swarms. Battlefield navigation. Physical target acquisition without human oversight. Every military application of AI that operates in the real world runs on spatial intelligence. The nation that masters the transition from static text to dynamic three-dimensional perception doesn’t just win the software race. It commands the physical battlefield. The AI arms race just broke out of the data center. It’s operating in three dimensions now.

Dustin

122,680 Aufrufe • vor 5 Monaten

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕

1,078,282 Aufrufe • vor 2 Monaten

AI has transformed how video is created. We think the next wave is about understanding it. Over the past few years, we've seen remarkable advances in video generation, editing, avatars, and creative tooling. An increasingly important problem is teaching machines to search, analyze, reason over, and extract insight from video - across massive libraries and live streams alike. We're calling this video intelligence, and we're actively looking to back founders building here. We're most excited about companies pushing on the core capabilities: - Video-native models - multimodal embeddings, temporal reasoning, and retrieval built specifically for video rather than adapted from image or text - Real-time and large-scale pipelines - infrastructure for processing, indexing, and querying video at the speed and scale enterprises actually need - Agentic and reasoning layers - systems that don't just retrieve clips but answer questions, surface anomalies, and take action on what they see The models and infrastructure to make this real are appearing to be crossing a capability threshold right now. Multimodal foundation models are maturing, storage costs have collapsed, and enterprises are sitting on years of unstructured video with no way to use it. That infrastructure unlocks a wide range of applications including media and sports workflows, security and physical operations, enterprise knowledge management, advertising analytics, robotics, and consumer products, where video has historically been dark data. If you're building in video intelligence at the model layer, the platform layer, or in a vertical application, we'd love to talk!

Jason Cui

36,205 Aufrufe • vor 3 Monaten

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,044 Aufrufe • vor 14 Tagen