Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Today we're announcing microSLAM, a monocular SLAM system built by our computer vision team in Zürich with ETH Zurich's Computer Vision and Geometry Lab. It ranks first on the LaMaria benchmark for monocular SLAM, and holds up against systems that carry many more cameras and IMUs. Ours runs on...

67,975 Aufrufe • vor 13 Tagen •via X (Twitter)

20 Kommentare

Profilbild von Pablo Vela
Pablo Velavor 13 Tagen

Cool! What’s the runtime compared to orbslam 3?

Profilbild von toolshed
toolshedvor 13 Tagen

The extra cameras and IMUs are not only about accuracy. One moving camera recovers geometry up to an unknown scale factor, and a stereo baseline or an accelerometer is what gets the metres back. A robot judging a 40cm gap needs metres. Where does microSLAM get its scale?

Profilbild von /
/vor 13 Tagen

closed source?

Profilbild von Ricardo de Azambuja
Ricardo de Azambujavor 13 Tagen

Link to the repo or it doesn't exist 🙃

Profilbild von mansin
mansinvor 13 Tagen

Awesome! Would love to test it

Profilbild von Azeem
Azeemvor 13 Tagen

microSLAM uses a single RGB stream and still outperforms systems with multiple cameras and IMUs. The LaMaria benchmark is one of the most rigorous for monocular SLAM. The fact that ETH Zurich's lab is involved adds significant credibility.

Profilbild von Vasco 🇳🇱🇹🇭
Vasco 🇳🇱🇹🇭vor 13 Tagen

👀👀👀

Profilbild von The Canaanite
The Canaanitevor 13 Tagen

Wonderful, wen robot for housework and how much?

Profilbild von Armeen
Armeenvor 13 Tagen

The hard part is not monocular geometry; it is keeping it stable under motion and layout change. We can help capture the rare failure slices, then keep #heldout evaluation on unseen plant configurations rather than another pass through the same walk.

Profilbild von AI Mastery Guide
AI Mastery Guidevor 13 Tagen

Ranking first with one camera, impressive

Profilbild von indoor positioning
indoor positioningvor 13 Tagen

开源没

Profilbild von Senthilnathan K
Senthilnathan Kvor 13 Tagen

Would love to test this and integrate into our platform, monocular depth estimation if done accurately is super useful

Profilbild von Alvaro L
Alvaro Lvor 13 Tagen

How does it compare to Lingbot Map

Profilbild von DAY☀️
DAY☀️vor 13 Tagen

love a low calorie solution

Profilbild von Daniel Mika
Daniel Mikavor 13 Tagen

Cool work guys

Profilbild von Aryan Dhawan
Aryan Dhawanvor 13 Tagen

Getting a first-place result from one camera is honestly exciting. I’d look next at motion blur and fast body rotation, since those are exactly where monocular SLAM gets uncomfortable on a walking robot.

Profilbild von remco_bos
remco_bosvor 12 Tagen

Release source and open source MicroRover

Profilbild von Julien Blanchon 🇺🇦
Julien Blanchon 🇺🇦vor 13 Tagen

Link ?

Profilbild von Roko’s Basilica
Roko’s Basilicavor 13 Tagen

Cool, but do you really need to use AI to write a slop tweet?

Profilbild von Dredge
Dredgevor 12 Tagen

Monocular SLAM ranks first on LaMaria - benchmark conditions. Real factory: uncontrolled light, fast movements, occlusions. Does geometry recovery hold? If microSLAM degrades 10-20% on messy data, how much does that hurt downstream model training?

Ähnliche Videos

When we started Score, the standard computer vision tools already existed. About a million people use them every day. Most of those people are still waiting on labels, running training jobs by hand, and watching models fail once they leave the test set. Most of those people are still waiting on labels, running training jobs by hand, and watching models fail once they leave the test set. Most of those people are also still waiting on verified computer vision models, evaluated against real life conditions and ready to be deployed for them to deliver value for their teams, clients or users. Score Studio is the full computer vision path in one place. A team describes the problem. The system can generate the missing scenes, label them, train the candidates, evaluate which ones actually hold, and deploy the winner. Data, labels, training, eval, ship. One loop. If no model exists for that job yet, they can put a bounty on the subnet. Anything from a small vision brick to a full VLM. Miners compete on the task. Only the winning work comes back. Same path for software agents. Any agent can call it. Built to be fully agent-accessible. Built for the people who already do this work: computer vision engineers and the small teams around them in plants, warehouses, farms, robotics, sport, and security. And for the agents those teams will run. That is the part that changes the job. Not another training screen. The stretch that used to take a lab and a calendar, footage, boxes, versions, failed runs, a separate deploy project, sits behind one starting point. And if the network needs a new model, that request is part of the same path. We spent more than a year building it. Then we had a choice. Keep it for us, or commoditize the whole subnet and make it available 24/7, in permissionless and open-source way. And we knew we couldn't keep it for us. It had to live on Bittensor. Open source software already showed how this should work. Infrastructure should not sit inside one company. Same idea as open AI before the phrase changed meaning: inspect it, fork it, keep building. That is what SN44 is for. Open vision intelligence, powered by Bittensor. Miners do the work. Studio is how that gets monetized. Profit does not stay in a company account. It goes back into the subnet through buyback and burn. We built the tool we wanted on day one. It will live on the network now, and for ever. Waitlist is open.

Score

11,292 Aufrufe • vor 27 Tagen

Can United States manufacture robots? Matic Robots says "yes." It makes the best floor cleaning robot, that has won many perfect scores from Wired to many others. We love ours. But my trip there to get a tour from AI pioneer Navneet Dalal Navneet Dalal provided some real insights into how hard it is for a hardware company to make hardware in the United States. And how deeply AI is changing consumer electronics products that are going to be in many more homes soon. In this first part (Part II coming tomorrow) we get a look at how long it took for this company to go through prototypes to a shipping product. In the second part, you'll see the scaling hell that it takes to even ship a few thousand robots and the kinds of problems that scaling up a factory brings. Matic is one of my favorite small Silicon Valley companies. It has found what we call "product market fit." I just came back from CES where I saw many of its competitors, and the Matic wins because of not just the product thinking of Mehul and Navneet Dalal but because of their AI leadership. In a way their robot took many lessons from Tesla, from where to put the batteries to its bet on computer vision, which Navneet has been a pioneer in for years, working quietly behind the scenes. It is about to move into a new location that will allow it to grow to meet the demand that now is showing up (the boxes in its lobby show that it's outgrowing its current facilities). In terms of AI, it has aspirations of making a humanoid too, but it is taking a far more measured approach to getting there. By starting on the floor it can not just build world models based on real world data (customers are given a choice whether to allow its data to be used that way. Most customers choose to keep their data on the robot only, for privacy reasons, but if you opt in you can help them improve their models). They are using that data to understand homes. Navneet told me they hit very unusual situations in people's homes already that they couldn't really predict in simulators, like full-wall mirrors that confuse computer vision systems, or pools and water features in people's homes. Having real customers brings a ton of customer feedback about how to further improve the robot, and, as Navneet demonstrates in the second video, forces them to build a manufacturing muscle memory. Getting teams to work together, figuring out how to solve supply chain problems, from Trump's tarriffs, to a new one that showed up over the past couple of weeks. A supplier for its bags (one of the cheaper parts that goes into the robot) changed the glue it used, which caused robots to fail quality tests and the manufacturing line to stop. Reminds me a lot of the hell Elon Musk faced in its Fremont factory when Tesla was first starting to manufacture its Model 3, which almost bankrupted the company. Off the record Mehul and Navneet 🇮🇳 showed me some of the prototypes and plans for its next products that will show up over the next few years. Certainly not as sexy as Tesla, Figure, 1x_tech, and all the Chinese manufacturers are showing off already, but far better thought out for the typical Western home and AI plays a huge role in its future. It is the product that speaks for itself. It's amazing, and is about to get better this year due to AI. It's the first real vision-only robot to be in my home and I bet it won't be the last from this company. Real honor that they invited me over with my Insta360 camera (another company launched in my home, just like Matic was last year). In Part II we go into the factory.

Robert Scoble

69,229 Aufrufe • vor 8 Monaten

Aravind Srinivas just described a future most founders are pretending they are ready for. One person. One machine. A company that runs itself. Srinivas: “Buy a Mac mini, set up a Perplexity personal computer, and run their business on that.” Not a side project. Not a pitch deck. A real business with real revenue while the founder is not in the building. AI runs the ads. Handles SEO. Integrates Stripe. Ships features. Answers customers. All of it executing without a single employee. Srinivas: “Have this all working while you can be sipping wine in Napa.” But before he sold the dream he killed the one most people are already chasing. Srinivas: “Everybody talks about this one-person one-billion-dollar company. It’s not truly moving the GDP by one billion. It’s not truly creating new value.” One researcher collecting a billion in equity does not grow an economy. It rearranges numbers between balance sheets. Nothing gets built. No customer gets served. That is not value creation. That is valuation creation. Srinivas wants no part of it. What he described is the opposite. The person driving Uber between shifts who has the idea but not the payroll. Not the engineering. Not the marketing. Not the support staff. That person gets a machine that replaces all of it. Hundreds of thousands in revenue. Millions. Generated by autonomous systems doing the work that used to require ten employees and a burn rate. Not paper wealth. Not valuation theater. Output that moves through an economy and touches real customers. That is what moves GDP. Not one person worth a billion dollars. A million people each building something worth a million. That math rewrites a country. Then Srinivas said the part that separates him from every hype merchant in the room. Srinivas: “Everybody thinks AI is already there. It’s not there yet. Someone has to do that hard work.” The vision is real. The infrastructure is not. The agents are not autonomous. The integrations are not seamless. The plumbing is not finished. Someone has to wire the APIs. Connect the billing. Build the bridge between what a founder wants and what a machine can deliver. That work is not a keynote. It is not a tweet thread. It is engineering that nobody wants to do and everybody will depend on. Whoever finishes it first does not just build a product. They hand every ambitious person on Earth a company they can run alone. The corporations that need five hundred people to do what one founder with the right infrastructure could do are not efficient. They are exposed. And the person building the thing that exposes them just told you exactly what it looks like. He also told you it is not going to build itself.

Dustin

64,593 Aufrufe • vor 6 Monaten

Want to create an avatar from a single image? FlexAvatar is a transformer model that creates full 360°, high-quality, and expressive 3D head avatar from just a single portrait image in minutes. Real-time Demo: FlexAvatar's lightweight architecture allows both animation and rendering in real-time, enabling interactive user experiences. To create a new 3D head avatar, only one image is required, e.g., from a webcam. The final avatar is ready after 2 minutes. Architecture: Under the hood, FlexAvatar adopts a transformer-based encoder-decoder design. The encoder maps the input image onto a latent avatar space, while the decoder produces 3D Gaussian attribute maps by incorporating the animation signal via cross-attention. The model learns all facial animations directly from the data without relying on pre-built 3D face models. This equips the avatars with realistic facial expressions. The internal avatar latent space can be conveniently used to integrate additional observations of a person via fitting. This enables use-cases where more than one image of a person is available, e.g., from a phone scan of the person. We train jointly on 2D monocular videos and multi-view data. However, in monocular videos, the animation signal leaks the target viewpoint, causing the model to produce incomplete 3D heads. We call this phenomenon entanglement of driving signal and target viewpoint. To prevent entanglement, we introduce bias sinks. These are learnable tokens that indicate whether a training sample stems from a monocular or a multi-view dataset. During training, the model learns to produce incomplete 3D heads only when the monocular token is present. During inference, FlexAvatar then always uses the multi-view token for which the model has learned to produce complete 3D heads. This simple design allows to combine the generalizability from monocular data with the quality of multi-view data. FlexAvatar summary: - Input: Single-image, phone scan, or monocular video - Output: Full 360° head avatar - Expressive animations - Real-time rendering and animation - Generalization to any portrait - Create a new avatar in 2 minutes - Use bias sinks to combine 2D and 3D data 🏠 🌍 🎥 Great work by Tobias Kirschstein and Simon Giebenhain!

Matthias Niessner

96,371 Aufrufe • vor 9 Monaten

WOW. 😳 Apple just quietly won the 3D maps war at WWDC. Gaussian Splatting is coming to Apple Maps Flyover this fall. Apple Maps Flyover covers 300+ cities. Until yesterday, every single one was built on standard drone photogrammetry. The technology captures photos from the air and reconstructs 3D geometry from them. Gaussian Splatting does not reconstruct geometry. It represents the scene as millions of tiny 3D ellipsoids, each one carrying its own color and opacity information based on how light actually behaves in that location. The output is not a mesh model. It is a field of light. When you move through it, it does not crumble at the edges. The detail holds because it was never geometry to begin with. Apple has been hiring for this for years. Their SHARP model, published in research last year, generates photorealistic 3D scenes from a single image in under a second. Google has more sensor data than anyone. More Street View cars, more satellites, more capture history. On navigation accuracy and geodata depth, Google Maps is still ahead by most measures. But fidelity in 3D city rendering is a different competition, and Apple just set a bar in that. Most people will experience this in the fall without knowing the name of the technology. They will open Flyover, look at a city they know, and notice it looks different. Real, not rendered. That is the moment Gaussian Splatting stops being a research term and becomes something a billion people use. Bookmark this. It will look prescient by October.

Shruti

19,832 Aufrufe • vor 3 Monaten

Mark Zuckerberg just described the obsolescence of every institution on Earth and delivered it like a product update. Zuckerberg: “I just think in the future almost everyone is gonna have the power of a 10,000-person organization.” He did not say better tools. He did not say smarter apps. He said the full cognitive output of ten thousand human beings. Packaged into a product. Handed to one person. That is not an upgrade. That is the end of the reason human beings organize at all. Companies exist because one brain is not enough. Governments exist because coordination requires hierarchy. Universities exist because knowledge demands infrastructure. Every institution ever built was a workaround for the same limitation. No single person could do it alone. Zuckerberg is telling you that limitation is about to disappear. The 500-person startup becomes one founder and an AI stack. The law firm becomes one attorney and a system that never sleeps. The hospital becomes one doctor carrying every specialist in their pocket. That is not speculation. That is a deployment schedule. And the man writing it runs a 70,000-person company. He employs 70,000 people and just told the world one person will soon need none of them. That is not a prediction. That is a confession from the man who will be first to act on it. But the part nobody is discussing matters more. This technology does not land on everyone equally. It lands first on the people who already command 10,000-person organizations. Zuckerberg does not get the power of 10,000 people. He already has that. He gets the power of 10,000 organizations. Every revolution in history was sold as liberation. The printing press was supposed to democratize knowledge. It built media empires. The internet was supposed to democratize commerce. It built trillion-dollar platforms. The tool always arrives as liberation. It always settles as leverage. And leverage always consolidates upward. Zuckerberg is not wrong about the capability. One person will do what ten thousand once did. But the question nobody is asking is the only one that matters. If everyone wields the output of 10,000 people, what is a single person actually worth? And then Zuckerberg answered his own question without realizing it. Zuckerberg: “If the intelligence of a 10,000-person company is not greater than the intelligence of a single person, then what are we doing here?” He meant it as a case for AI. That is the most brutal thing a CEO has ever said about the people who work for him.

Dustin

54,061 Aufrufe • vor 4 Monaten