Загрузка видео...

Не удалось загрузить видео

На главную

Meta releases VGGSfM Visual Geometry Grounded Deep Structure From Motion Structure-from-motion (SfM) is a long-standing problem in the computer vision community, which aims to reconstruct the camera poses and 3D structure of a scene from a set of unconstrained 2D images. Classical frameworks solve this problem in an incremental...

96,527 просмотров • 2 лет назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Wonderland: Navigating 3D Scenes from a Single Image Contributions: • First, we introduce a representation for controllable 3D generation by leveraging the generative priors from camera-guided video diffusion models. Unlike image models, video diffusion models are trained on extensive video datasets. This enables them to capture comprehensive spatial relationships within scenes across multiple views and embed a form of "3D awareness" in their latent space, which allows us to maintain 3D consistency in novel view synthesis. • Second, to achieve controllable novel view generation, we empower video models with precise control over specified camera motions. We introduce a novel dual-branch conditioning mechanism that effectively incorporates desired diverse camera trajectories into the video diffusion model. This enables expansion of a single image into a multi-view consistent capture of a 3D scene with precise pose control. • Third, to achieve efficient 3D reconstruction, we directly transform video latents into 3DGS. We propose a novel latent-based large reconstruction model (LaLRM) that lifts video latents to 3D in a feed-forward manner. With this design, during inference, our model directly predicts 3DGS from a single input image, effectively aligning the generation and reconstruction tasks—and bridging image space and 3D space—through the video latent space. Compared with reconstructing scenes from images, the video latent space offers a 256× spatial-temporal reduction while retaining essential and consistent 3D structural details. Such a high degree of compression is crucial, as it allows the LaLRM to handle a wider range of 3D scenes within the reconstruction framework, with the same memory constraints.

MrNeRF

52,849 просмотров • 1 год назад

Colmap 4.0 was very recently released, so it inspired me to do some work to better understand it and its new capabilities with Rerun. I want to really understand how Colmap, and in particular, pycolmap, works outside of just calling it via the CLI. So my goal is to use the low-level pycolmap API to log every part of the pipeline. The explicit goal is to have an alternative to the SQLite database that I can utilize. Instead of SQLite, I want to try logging everything directly to rerun and use RRD. This means I can have deep inspectability and still save the features/matches/2D view geometry, but be able to view it directly in rerun. I think this is one of the superpowers that rerun provides; data and visualizations are deeply integrated. As I'm often working with sequential data (videos), I'm going to specifically focus on four things: 1. Monocular Video Simple: Calls high-level APIs such as pycolmap.extract_features, pycolmap.match_sequential, pycolmap.incremental_mapping. These are basically identical to the CLI options and provide a good baseline. 2. Monocular Video Streamed: Take the above high-level APIs and break them down to their iterator version, logging each component in a streamed manner. This way, I can stream the intermediate features to rerun while the extraction/matching/mapping is happening. 3. Rig with unknown calibration: <- WHAT THE VIDEO SHOWS This is probably the most interesting version and the first one I've been working on. It allows one to set a rig between known sensors, such as in VR/AR devices, leading to much better reconstructions with multiple cameras. This is the case where we don't know the calibration a priori, so we have to run a reconstruction twice: once as a normal Colmap reconstruction with no rig constraints, use this to generate the constraints, and then do it again with the newly found rig. 4. Rig with known calibration: This is the RoboCap example, where we have a pre-calibrated set of sensors, so we don't need to run the two reconstructions and also gain better matching between cameras, both spatially and temporally. Again, this leads to a much better reconstruction! Along with all this, GLOMAP has become a first-class global mapper, making it super easy to use directly within pycolmap! I'm excited to do more with this and compare it to things like pycuvslam, vipe, and other alternatives.

Pablo Vela

30,070 просмотров • 4 месяцев назад

The U.S. Special Presidential Envoy for Ukraine, General Keith Kellogg Keith Kellogg, is in Kyiv today. I am grateful for a constructive meeting. We discussed various vectors of cooperation – how to achieve real peace and guarantee Ukraine’s security. These include projects within the PURL initiative for financing production and procurement of Patriot systems, strong bilateral agreements on co-production of drones and weapons that we have proposed to America. We count on a positive response from the United States. We had a substantive discussion on stepping up pressure on the Russians and what we can do together with partners in tariff and sanctions policy to enable a meeting at the leaders’ level at the earliest and bring this war to an end. A trilateral leaders’ format is undoubtedly the most effective. We also discussed the return of abducted Ukrainian children, international cooperation on this track, and the conditions in which our children are being held. I expressed condolences to the American people over the horrific murder of Charlie Kirk and thanked President Trump for his condolences and response to the brutal murder of Ukrainian citizen Iryna Zarutska in North Carolina. It is important that justice prevail every time violence seeks to take hold. We are also preparing for the 80th session of the UN General Assembly in New York. We discussed planned events, coordination between Ukraine and the U.S., and work within the Coalition of the Willing. We are working on potential meetings and various formats.

Volodymyr Zelenskyy / Володимир Зеленський

228,172 просмотров • 10 месяцев назад

In the summer of 2023, I cold emailed Jensen Huang and asked to capture a NeRF of him at SIGGRAPH. He responded in about an hour and said yes. A radiance field is, in the simplest terms, akin to a 3D photograph. A moment in time, so completely reconstructed that you can move through it and see it from angles the original cameras never occupied. NeRFs were the original method. Gaussian splatting, which debuted at that same SIGGRAPH, has since become the dominant form of radiance field. I called my late friend James, who told me we needed to begin practicing immediately. We ran capture after capture for weeks until we consistently got the capture time down to ~30 seconds with one camera. Later, in a hallway at the LA Convention Center during SIGGRAPH, I captured the portrait you're seeing now, a full 360° gaussian splat of Jensen, rendered here as a 2D flythrough. Afterward, I continued the conversation with him and members of his team to make the case for radiance fields as a foundational representation for imaging. To my surprise, they listened. Three years later, NVIDIA has several works, including NuRec, fVDB, 3DGRUT, and gsplat all utilizing radiance fields. The landscape has evolved enough that the reasoning is obvious. Gaussian splatting has begun to ship across some of the world’s largest industries, including autonomous vehicles, AEC, geospatial, media and entertainment, robotics, e-commerce, hospitality. It’s become clear that lifelike 3D is here to stay. And yet I think we will look back and be disappointed by how late we started taking 3D portraits of the people around us, just like how we have sparse 2D photos of our grandparents and great grandparents. We have billions of photographs of the people we know and love, but almost no radiance fields of them. I'll be returning to SIGGRAPH in LA where this was initially captured three years ago, with the landscape looking significantly different. Radiance fields are more under deployed than ever relative to what they can do. I'm excited for the future of imaging, and for 2D to transition into 3D. I have a few things up my sleeve that I think will make that case plainly.

Radiance Fields

17,663 просмотров • 1 месяц назад

3D Gaussian Splatting for Real-Time Radiance Field Rendering paper page: Radiance Field methods have recently revolutionized novel-view synthesis of scenes captured with multiple photos or videos. However, achieving high visual quality still requires neural networks that are costly to train and render, while recent faster methods inevitably trade off speed for quality. For unbounded and complete scenes (rather than isolated objects) and 1080p resolution rendering, no current method can achieve real-time display rates. We introduce three key elements that allow us to achieve state-of-the-art visual quality while maintaining competitive training times and importantly allow high-quality real-time (>= 30 fps) novel-view synthesis at 1080p resolution. First, starting from sparse points produced during camera calibration, we represent the scene with 3D Gaussians that preserve desirable properties of continuous volumetric radiance fields for scene optimization while avoiding unnecessary computation in empty space; Second, we perform interleaved optimization/density control of the 3D Gaussians, notably optimizing anisotropic covariance to achieve an accurate representation of the scene; Third, we develop a fast visibility-aware rendering algorithm that supports anisotropic splatting and both accelerates training and allows realtime rendering. We demonstrate state-of-the-art visual quality and real-time rendering on several established datasets.

AK

633,532 просмотров • 3 лет назад

Dear François Nzanga Mobutu, Good evening. I bet your father is turning in his grave, writhing in pain and shame after reading your tweet. He, at least, knew that there are authentic and fully-fledged Congolese Tutsis. He proved it until people like you and others came along and misled him. Do I need to inform you, in case you're unaware, that the people who demonstrated yesterday in Washington DC, roughly 4,000 of them, others the same day in Nairobi, and still others in England shortly before, are Banyamulenge, these Congolese Tutsis from the highlands, who are protesting against what Tshisekedi and the FARDC, the Wazalendo, Évariste Ndayishimiye and the FDNB, as well as the FDLR and mercenaries, are doing to their relatives in Minembwe and throughout the highlands? Their Sukhoi fighter jets and drones bomb daily, killing children, women, the elderly, and men. They destroy homes, churches, schools, hospitals, and community radio stations. They even kill cows and sheep and destroy fields and crops. They have imposed a blockade on Minembwe, with no entry or exit. They have done and are doing the same thing to the Tutsi in Masisi and Rutshuru in North Kivu. Instead of listening to their cries and examining their demands, the excuse is quickly found; the preferred shortcut is Rwanda. Always and only Rwanda. Too easy, isn't it? Your thinking, which is also the regime's narrative, can be summarized in these two sentences: - all the problems facing the DRC come from elsewhere, particularly from Rwanda; we will end up being told that even the migrants are here because of Rwanda. “And all the solutions must come from elsewhere, especially from the United States of America, from Papa Trump.” Under these conditions, what is the point of the regime you serve? Two small truths to remember, dear François: “Such an attitude, this kind of ideology, this denial of nationality to Congolese Tutsis of origin, from North and South Kivu, as well as to all those who are victims of the same persecution, like the Hema and others, this easy rejection based on appearance (racial profiling), are among the root causes of the crisis the country is going through;” “As long as we haven’t decided to be sufficiently responsible, to sit down as a nation and rigorously assess our share (of responsibility) in what is happening to us, we will continue to wait for solutions from others, solutions that may never come.” Furthermore, there are satanic verses that must be banished immediately, in the interest of everyone and the country, such as: "There are no Congolese Tutsis, every Tutsi is Rwandan, therefore a foreigner," etc. Either we will be able to put an end to exclusion, discrimination, hate speech, and ethnic hatred, and live together according to the law and history, or this deep-seated problem risks haunting us for a long time. But to achieve this, we need leadership capable of understanding and transcending differences and turning them into assets for living together. This is possible.

Me Moise Nyarugabo

22,470 просмотров • 3 месяцев назад