Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

stitching artifact as an #echofirst tool These demarcations appear at interface of subvolumes that are incorrectly juxtaposed but can id exact location of 2D cross-sectional cuts in 3D image since they always occur parallel to TEE 2D multiplane rotation

17,461 Aufrufe • vor 1 Jahr •via X (Twitter)

2 Kommentare

Profilbild von Prasannasimha🇮🇳
Prasannasimha🇮🇳vor 1 Jahr

@ASE360 @JournalASEcho @rajdoc2005 @NMerke @echo_stepbystep @bwoody58 @CASivaram1 @DrRajeshG1 @purviparwani @NadeenFaza Nice - using an"error" or an "artifact" to ones advantage

Profilbild von Brian Wood
Brian Woodvor 1 Jahr

@ASE360 @JournalASEcho @rajdoc2005 @NMerke @echo_stepbystep @CASivaram1 @DrRajeshG1 @purviparwani @NadeenFaza 🙏 cool makes sense but had never thought about that

Ähnliche Videos

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,708 Aufrufe • vor 3 Jahren

Sora animates with this wonderful wonky, sketchy dreamlike quality that perfectly captures the nostalgic atmosphere of a hazy 90's suburban summer afternoon. While Seedance is just as sophisticated and in some aspects superior, I will miss the quality of Sora. Sora feels like 35mm film, with the nuanced way it captures lighting, color, and perspective, while Seedance feels digital - a bit too clean. I am confident that I can replicate this sketchy quality in Seedance by tweaking my prompts, but Sora has a unique way of animating that really should be preserved. There are many more bugs and mistakes with Sora, but when it gets it right, it REALLY hits a level of artful magic that sets it above all other models. I am shocked that I still appear to be the only person who is creating AI-Generated 2D animation with original characters while developing a unique Western style. Practically no one is doing 2D AI-animation at all, and when they do, it is usually an attempt to mimic an Anime style. 2D really should be attempted more by AI Filmmakers!! Especially with a Western model like Sora, before it is terminated in the fall. Really, everyone should be taking advantage of the way this model so skillfully animates, with the heart and soul of a seasoned professional. I personally feel AI-generated animation is superior to 3D/realistic AI filmmaking. While AI-generated realism is an attempt to mimic real life through a lens, 2D animation IS what it is - not a replica of anything, just a cartoon. I love all my AI filmmaker bros and all of the cutting edge work going on, but I urge all of you to give 2D animation a chance. : ) I'm still not giving up hope that we can save Sora somehow. Now that I see that Seedance is able to animate 2D brilliantly and beautifully, I know that it's not some hidden secret, and we can replicate the weights and maths of Sora. If we can't preserve Sora, we can make a comparable model. It's not just in the interest of AI-Animators like me to preserve this model, but in the interest of the entire American AI community if they want Western AI to be superior. If not, Chinese video models will dominate. And while I am perfectly happy to use a Chinese model and I am just eternally grateful that this technology exists at all, I would love to see American audiovisual models continue to be developed! #AIAnimation #AIFilmmaking #AIArt #SaveSora #SummerofSora #WillStancilShow Elon Musk Marc Andreessen 🇺🇸 Sora Bill Peebles NVIDIA Sam Altman

Emily Youcis

18,940 Aufrufe • vor 2 Monaten

Alibaba presents MIMO Controllable Character Video Synthesis with Spatial Decomposed Modeling Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D works typically require multi-view captures for per-case training, which severely limits their applicability of modeling arbitrary characters in a short time. Recent 2D methods break this limitation via pre-trained diffusion models, but they struggle for pose generality and scene interaction. To this end, we propose MIMO, a novel framework which can not only synthesize character videos with controllable attributes (i.e., character, motion and scene) provided by simple user inputs, but also simultaneously achieve advanced scalability to arbitrary characters, generality to novel 3D motions, and applicability to interactive real-world scenes in a unified framework. The core idea is to encode the 2D video to compact spatial codes, considering the inherent 3D nature of video occurrence. Concretely, we lift the 2D frame pixels into 3D using monocular depth estimators, and decompose the video clip to three spatial components (i.e., main human, underlying scene, and floating occlusion) in hierarchical layers based on the 3D depth. These components are further encoded to canonical identity code, structured motion code and full scene code, which are utilized as control signals of synthesis process. The design of spatial decomposed modeling enables flexible user control, complex motion expression, as well as 3D-aware synthesis for scene interactions. Experimental results demonstrate effectiveness and robustness of the proposed method.

AK

148,998 Aufrufe • vor 1 Jahr