Загрузка видео...

Не удалось загрузить видео

На главную

Here it is! I rebuilt a visual iniertial multicamera pipeline in rust + python, alll with support for Rerun OSS data catalog + visualization. code here - < The core is rust to make sure it runs fast, but you can also easily call it from python. I've been...

17,561 просмотров • 10 дней назад •via X (Twitter)

Комментарии: 12

Фото профиля Pablo Vela
Pablo Vela10 дней назад

1. Multi Camera - I'm convinced 4 or more cameras are optimal here, sure you can do purely monocular or stereo, but IMO most robots and actual deployed systems out there that need to run at realtime on low powered hardware should AND need to be robust against occlusion/motion blur/ect work best with a multicam setup. Theres a reason aria gen2 switched to 4 global shutter cameras. The cuvslam paper argues the same point for general robotics.

Фото профиля Pablo Vela
Pablo Vela10 дней назад

2. IMU support - bascially the same point as above, realtime robotics needs to be robust, and imu + cameras are perfect for each other. They each cover the others weakness

Фото профиля Pablo Vela
Pablo Vela10 дней назад

3. Diverse hardware - I didn't want this to just run on pytorch/nvidia devices. I want this to work on my linux machine, or my mac mini, or my Raspberry Pi! 4. Hardware acceleration - In the same way I want diversity of hardware, I also want to make sure we can take advantage of hardware acceleration. If I have a 5090, I should get the most out of it =] and if I have a rpi5 I should also get the most out of it! It has a GPU on it after all

Фото профиля Pablo Vela
Pablo Vela10 дней назад

5. Opensource - I'm sure theres plenty of great vio/slam pipelines out there that do the above, but none that are open that I know of =/

Фото профиля Pablo Vela
Pablo Vela10 дней назад

So thats exactly what I did, taking inspiration from basalt/cuvslam/glide this repo has imu+multicamera support, works on diverse hardware (tested on my mac mini + 5090 linux machine + rpi5 + rockchip 3588) AND support GPU thanks to the awesome CubeCL library that lets you write kernels on rust and support wgpu/metal/vulkan I also leaned heavily on @kornia_foss and the great work done by the folks there. I used many of the components in kornia-rs and got lots of inspiration from kornia-slam I plan to rip out parts that make sense and contribute them back to kornia. The biggest problem now is that this is VIO, so it has lots of drift. I tried walking for a mile or two and returning to the same spot and the drift is pretty bad. That'll a problem for later me, but shouldn't be too hard to add loop closure. I'm looking at cuSFM + colmap for this.

Фото профиля clankr
clankr10 дней назад

@rerundotio Looks great. Could models like World Labs’ Atlas be used to test SLAM algorithms? What do you think?

Фото профиля Pablo Vela
Pablo Vela10 дней назад

@rerundotio I’m not sure, if I get access I’d be up to try

Фото профиля Mateo de Mayo
Mateo de Mayo10 дней назад

@rerundotio This has many similarities to what we submitted to ICRA two days ago 😅

Фото профиля Pablo Vela
Pablo Vela10 дней назад

@rerundotio Awesome, ya'll do great work! I'd be curious to see how we differ in implementation. I still think theres a bunch that could be done on the GPU, I really only got a good speedup on my 5090 machine.

Фото профиля Mateo de Mayo
Mateo de Mayo10 дней назад

@rerundotio I'll do a post when we release the code :)

Фото профиля Carlos Pinheiro
Carlos Pinheiro9 дней назад

@rerundotio Nice! And btw thanks for the reference, GLidE-SLAM will be presented at iros26.

Фото профиля Damir Wallener
Damir Wallener10 дней назад

@rerundotio Have you ever watched a pigeon walk?

Похожие видео

Well that was an adventure. Toad's fuzzy file search boosts first characters, smallest number of groups, and it highlights the best scoring match. So you can locate a file with the bare minimum of keypresses. It's an expensive algorithm as it scores every possible combination of the query characters. For any given path there could be many indices which match, and each of them is scored separately. Still, it was fast enough to search as you type for my largest repository. Alas, it was slow for very large repos. Like Microsoft's Typescript, which has more than 84,000 files. I figured I'd reach the limit of what I could do with Python. So my options were to use something like ripgrep or re-implement the fuzzy searching in a faster language. I didn't want to add the ripgrep dependency, so I built a Rust version (with the help of AI) which eventually yielded a ~14X speedup. Nice. Better. But it could still become painful to type under certain pathological conditions. An inefficient algorithm can always be slow if you add more items. So I scrapped the Rust solution and built an index in Python—so the fuzzy search could throw away most paths that wouldn't match or have low scores. And its faster than the Rust solution was. The only downside is that there is a little additional work upfront (done in the background). I think this will do for a while. Until I get an issue that it is slow with 10million files. Rust is still an option, as is doing the work in parallel. This will be in the next version of Toad. Here I am searching the Typescript repo... #Python

Will McGugan

12,261 просмотров • 7 месяцев назад

There's so much focus on "how can AI do my work for me?" I think the more important question is "what work can I now do with AI that I would have never attempted before?" Earlier this year I wrote freestiler, a vector tiling engine for R and Python, with the help of Claude and Codex. I knew what the ideal engine looked like and how it would work at a high level. I didn't know how to put it together, and I don't know Rust, the language I wanted under the hood. Previously I would never have attempted this project as the ROI wasn't there. It would have taken me a year or more to learn the internals of a vector tiling engine and enough Rust to implement one. With Opus-level models, I could take it on. freestiler now powers all my vector tiling pipelines, including the map below rendering 143 million jobs from LODES, and it has 114 GitHub stars. Building this way has required a different set of skills. I don't review the code line by line. I set up adversarial agents to do that and write the test suites. What I review is the architecture, the behavior, and the results. Agent teams surface findings and explain their reasoning; I evaluate and critique. My job isn't to stress over code formatting, but instead to focus on questions like whether the engine is designed right, whether the output is correct, and if the UX makes sense. This means that I haven't "replaced my work." I've taken on entirely new work, with the help of agents, that I would have never done otherwise. It has taken some getting used to shipping code I haven't personally typed. In the old way of working, I built understanding through writing that code. Now I build understanding through managing the project - writing a spec, reviewing structure, evaluating UX. And that's helped me think a whole lot bigger in terms of what I can now do.

Kyle Walker

13,867 просмотров • 2 месяцев назад