正在加载视频...
视频加载失败
I have released gem-x.cpp on GitHub! This produces a 3D skeleton from video, including live video, which is what interests me. NVIDIA showed the original GEM-X teleoperating humanoids, so the next thing I want to do is get that working in simulation all in GGML/C++.
22,353 次观看 • 18 天前 •via X (Twitter)
26 条评论

Source is here:

GEM-X is related to SAM3D-BODY, in that they do similar things and in offline mode GEM-X uses SAM3D. So checkout sam3d.cpp too:

I wonder how nice could be to have this in games. In C++ and if working on low end devices.. that opens up a lots of new ways for entertainment

It will very soon be in a game ;-)

looking forward to try it 👀

nice! strong follow :)

Thanks!

what are the 2 colored sets of lines? Different models for keypoint detection?

Yes, one is ViTPose which locates the joints in 2D, then GEM-X builds on that to create 3D motions which is the other.

Is the cyan-colored line the 3D version?

Yes

Looks awesome! from your experience, how does it compare to Mediapipe or similar?

No idea, I'm interested to hear from people how it compares to other solutions too

From the looks alone , yours seem to detect finger bones much better

VitPose does in 2D, I'm not sure the final 3D joint positions are tracking fingers well

Fun! I remember researching pose estimation almost a decade ago for augmented reality games and sports analysis software and was so sad we weren't close to being able to do it well (and in real time). It's come a long way. Very cool.

This looks pretty cool. Does it require Nvidia GPU?

No, runs on Vulkan by default and other backends (e.g Metal) can be added

Amazing work as always! :)

Thanks!

I’d love to know difference vs mediapipe

海报帧 GEM-X:左栏 Live webcam / Start live,9.2 fps;画面褐衫 FOSDEM 22 logo,身上叠青蓝 Projected 3D skeleton 和橙 ViTPose 点,手指关节细线一根根。开源句写在 FOSDEM 22 圆 logo 和青蓝手指骨架上。

Not bad, I made a similar version with a lot less latency.

Feel free to share the details of that

How's performance for this vs other skeleton tracking older models? Natively 3d? How's it handle occlusion, best fit overall? Dropout?

It produces 3d poses. I find that it tries to infer the location of body parts when they are occluded and it doesn't make the best guesses if large parts of your body are not visible, but self-occlusion it seems to handle well. Otherwise I don't know how it compares
