Cinematic Mindscapes: High-quality Video Reconstruction from Brain Activity propose... Mind-Video that learns spatiotemporal information from continuous fMRI data of the cerebral cortex progressively through masked brain modeling, multimodal contrastive learning with spatiotemporal attention, and co-training with an augmented Stable Diffusion model that incorporates network temporal inflation paper page:show more

AK
255,257 次观看 • 3 年前
🏆 We're thrilled to announce that Meta FAIR’s Brain... & AI team won 1st place at the prestigious Algonauts 2025 brain modeling competition. Their 1B parameter model, TRIBE (Trimodal Brain Encoder), is the first deep neural network trained to predict brain responses to stimuli across multiple modalities, cortical areas, and individuals. The approach combines pretrained representations of several foundational models from Meta – text (Llama 3.2), audio (Wav2Vec2-BERT from Seamless) and video (V-JEPA 2) – to predict a very large amount (80 hours per subject) of spatio-temporal fMRI brain responses to movies acquired by the Courtois NeuroMod project Download the code: Read the paper: Learn about the challenge: Download the data:show more

AI at Meta
1,094,017 次观看 • 1 年前
DimensionX: Create Any 3D and 4D Scenes from a... Single Image with Controllable Video Diffusion TL;DR: Create 3/4DGS from Video Diffusion Note: Some first inference code released (not all yet). Contributions (cited): • We present DimensionX, a novel framework for generating photorealistic 3D and 4D scenes from only a single image using controllable video diffusion. • We propose ST-Director, which decouples the spatial and temporal priors in video diffusion models by learning (spatial and temporal) dimension-aware modules with our curated datasets. We further enhance the hybriddimension control with a training-free composition approach according to the essence of video diffusion denoising process. • To bridge the gap between video diffusion and real-world scenes, we design a trajectory-aware mechanism for 3D generation and an identity-preserving denoising approach for 4D generation, enabling more realistic and controllable scene synthesis. • Extensive experiments manifest that our DimensionX delivers superior performance in video, 3D, and 4D generation compared with baseline methods.show more

MrNeRF
17,062 次观看 • 1 年前
Wonderland: Navigating 3D Scenes from a Single Image Contributions:... • First, we introduce a representation for controllable 3D generation by leveraging the generative priors from camera-guided video diffusion models. Unlike image models, video diffusion models are trained on extensive video datasets. This enables them to capture comprehensive spatial relationships within scenes across multiple views and embed a form of "3D awareness" in their latent space, which allows us to maintain 3D consistency in novel view synthesis. • Second, to achieve controllable novel view generation, we empower video models with precise control over specified camera motions. We introduce a novel dual-branch conditioning mechanism that effectively incorporates desired diverse camera trajectories into the video diffusion model. This enables expansion of a single image into a multi-view consistent capture of a 3D scene with precise pose control. • Third, to achieve efficient 3D reconstruction, we directly transform video latents into 3DGS. We propose a novel latent-based large reconstruction model (LaLRM) that lifts video latents to 3D in a feed-forward manner. With this design, during inference, our model directly predicts 3DGS from a single input image, effectively aligning the generation and reconstruction tasks—and bridging image space and 3D space—through the video latent space. Compared with reconstructing scenes from images, the video latent space offers a 256× spatial-temporal reduction while retaining essential and consistent 3D structural details. Such a high degree of compression is crucial, as it allows the LaLRM to handle a wider range of 3D scenes within the reconstruction framework, with the same memory constraints.show more

MrNeRF
52,849 次观看 • 1 年前
Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation paper page:... Large text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring temporal consistency across video frames remains a formidable challenge. This paper proposes a novel zero-shot text-guided video-to-video translation framework to adapt image models to videos. The framework includes two parts: key frame translation and full video translation. The first part uses an adapted diffusion model to generate key frames, with hierarchical cross-frame constraints applied to enforce coherence in shapes, textures and colors. The second part propagates the key frames to other frames with temporal-aware patch matching and frame blending. Our framework achieves global style and local texture temporal consistency at a low cost (without re-training or optimization). The adaptation is compatible with existing image diffusion techniques, allowing our framework to take advantage of them, such as customizing a specific subject with LoRA, and introducing extra spatial guidance with ControlNet. Extensive experimental results demonstrate the effectiveness of our proposed framework over existing methods in rendering high-quality and temporally-coherent videos.show more

AK
375,160 次观看 • 3 年前
Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos... with Spatio-Temporal Diffusion Models Contributions: • We introduce Diffuman4D, a novel diffusion model that generates spatio-temporally consistent and high-resolution (1024p) human videos from sparse-view video inputs. • We propose a sliding iterative denoising mechanism that enhances both the spatial and temporal consistency of generated long-term videos while maintaining efficient inference. • We design a human pose conditioning scheme to enhance the appearance quality and motion accuracy of generated human videos. • We plan to release our processed version of the DNA-Rendering dataset, which we believe will benefit future research in this area.show more

MrNeRF
24,729 次观看 • 1 年前
Depth Any Video with Scalable Synthetic Data AI physicists... and chemists continue to make strides in depth estimation from video. Check out this new paper featuring some impressive examples. See the thread for more details (unfortunately no code yet). Abstract: Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge through two key innovations. First, we develop a scalable synthetic data pipeline, capturing real-time video depth data from diverse game environments, yielding 40,000 video clips of 5-second duration, each with precise depth annotations. Second, we leverage the powerful priors of generative video diffusion models to handle real-world videos effectively, integrating advanced techniques such as rotary position encoding and flow matching to further enhance flexibility and efficiency. Unlike previous models, which are limited to fixed-length video sequences, our approach introduces a novel mixed-duration training strategy that handles videos of varying lengths and performs robustly across different frame rates 0 - even on single frames. At inference, we propose a depth interpolation method that enables our model to infer high-resolution video depth across sequences of up to 150 frames. Our model outperforms all previous generative depth models in terms of spatial accuracy and temporal consistency.show more

MrNeRF
27,428 次观看 • 1 年前
Chop the gradients ✂️! We found that truncating decoder... gradients in latent video diffusion to a fixed window allows us to finetune on videos with pixel-wise perceptual losses without running out of memory. Pixel losses have been essential for image generation and reconstruction, but until now, they haven't scaled to long-duration, high-resolution video diffusion due to recursive activation accumulation in causal decoders, leading to OOM during training 💥📉. Project: Video diffusion models can do a lot more 🚀 when you can backprop the decoder! Post-process neural rendered scenes, super-resolve videos, harmonize lighting in controlled synthetic driving scenes, and inpaint videos — all in a single step ⚡ with a quick finetune from a standard diffusion model.show more

Felix Heide
28,399 次观看 • 4 个月前
3. The Return Phase: Was Earth’s Recovery Smooth or... Episodic? The third paper examines the final stage of the sequence: how Earth returns toward equilibrium after large-scale disturbance, and whether that return is best described as smooth or episodic. Classical geophysical models typically assume continuous, gradual relaxation governed by viscoelastic timescales. However, geological and paleoclimate records frequently show bursts of rapid change separated by quieter intervals. The study compares a smooth continuous return model with an event-guided model in which most adjustment occurs during discrete episodes derived independently from the geological record. Rather than relying solely on geophysical proxies, the paper uses archaeological site distributions as an external diagnostic. If surface stability matters, occupied sites should preferentially cluster in regions that remain stable across modeled return phases, and their timing should not align randomly with adjustment episodes. The results show that both early Homo and early civilization sites are non-random with respect to modeled stability fields. Spatial and temporal null tests indicate statistically significant alignment with the event-guided return model. Crucially, model parameters are fixed a priori and not tuned to archaeological data. This paper addresses the return phase of the ECDO sequence by testing whether recovery dynamics inferred from geophysics are consistent with independent records of long-term human occupation. Full Paper :show more

Craig Stone
21,369 次观看 • 7 个月前
The Hidden Language of Diffusion Models paper page: tackle... the challenge of understanding concept representations in text-to-image models by decomposing an input text prompt into a small set of interpretable elements. This is achieved by learning a pseudo-token that is a sparse weighted combination of tokens from the model's vocabulary, with the objective of reconstructing the images generated for the given concept. Applied over the state-of-the-art Stable Diffusion model, this decomposition reveals non-trivial and surprising structures in the representations of concepts. For example, we find that some concepts such as "a president" or "a composer" are dominated by specific instances (e.g., "Obama", "Biden") and their interpolations. Other concepts, such as "happiness" combine associated terms that can be concrete ("family", "laughter") or abstract ("friendship", "emotion"). In addition to peering into the inner workings of Stable Diffusion, our method also enables applications such as single-image decomposition to tokens, bias detection and mitigation, and semantic image manipulationshow more

AK
41,830 次观看 • 3 年前
Meta open-sourced a brain-to-text system that reaches 78% word... accuracy without surgery. Brain2Qwerty v2 converts non-invasive brain recordings into text with 61% average word accuracy and 78% for its strongest participant. The system reads MEG signals from a helmet, not electrodes placed inside brain tissue. 9 volunteers typed about 22,000 sentences while researchers recorded 10 hours of neural activity each. Brain2Qwerty v1 mostly mapped brain signals to single typed characters. It tries to recover characters, words, and full sentence meaning together. The system studies those brain signals and tries to turn them into the words you wanted to type. - 61% average word accuracy across all participants - 78% word accuracy for the top participant - 50%+ of sentences decoded with no more than 1 word error Performance improves as the data pile grows Raw brain signals are messy because many mental and physical processes fire at once. Deep learning handles that mess by learning patterns directly from the original recordings. A fine-tuned LLM then uses language context to repair likely word and sentence errors. This explains why the system beats earlier non-invasive methods reporting 8% word accuracy. More than half of sentences from the strongest participant had one word error or less. Accuracy also improved as training data grew, suggesting more recordings may close more of the gap.show more

Rohan Paul
17,127 次观看 • 2 个月前
Scientists suspect that at the threshold of death, the... human brain may display activity patterns similar to memory replay. One published study reported intriguing findings from EEG recordings of an 87-year-old patient who died of a heart attack. The study observed elevated gamma brain waves, which are often associated with memory, dreaming, and conscious thought. This phenomenon is sometimes referred to as a “life recall experience,” in which a person subjectively feels as if they are replaying moments from their life immediately before or after the heart stops beating. This may occur because as the brain begins to lose its oxygen supply, it can release bursts of intense neural activity in parts of the cerebral cortex, potentially triggering a surge in memory-related processing. Although more research is needed, these findings offer a new perspective on how the brain may still show organized activity even after the heart has stopped. This could provide a biological basis for various reports of near-death experiences, which often include life review or vivid flashbacks.show more

Kekius Maximus
1,079,882 次观看 • 7 个月前
I am blown away 🤯. Check this out! CameraCtrl... II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models TL;DR: "To enable broader exploration of dynamic scenes, our model can generate new video clips of the same scene based on previously generated content and user-provided camera trajectories. This approach maintains dynamic capabilities, accurate camera control, and scene consistency throughout the extended exploration." "Our model enables precise camera control across diverse scenarios while preserving dynamic scene elements, e.g." "Our method can generate videos with strong 3D consistency, which enables high-quality 3D reconstruction using the camera-controlled videos." Contributions: 1) A systematic data curation pipeline for constructing a dynamic video dataset with camera trajectory annotations; 2) A lightweight camera control injection module and corresponding training strategy that preserves dynamic video generation capabilities while adding camera control effect; 3) A clip-wise autoregressive generation recipe that enables extended range exploration of generated scenes.show more

MrNeRF
12,633 次观看 • 1 年前
💫Sneak Peek: Meet the Future OptimAI Core Node! We're... thrilled to give you an exclusive sneak peek of the upcoming OptimAI Core Node for Desktop PCs—our next-level node packed with powerful capabilities. 🔎Inside the OptimAI Core Node: Autonomous Data Mining Agent 🔸 Intelligent AI-driven mining that autonomously gathers, structures, and processes high-quality data directly from any platforms, dramatically improving AI training and model accuracy. Enhanced Contribution & Rewards 🔸Provide more computational resources to the network, unlocking greater rewards and maximizing your earning potential. Strengthened Ecosystem & Community 🔸Your contributions will fuel advanced AI development, creating an even stronger, more decentralized, and more intelligent OptimAI Network. This is just the beginning—stay tuned for more exciting updates, innovations, and ways you can grow with OptimAI Network. #DePIN #Miningshow more

OptimAI Network
49,527 次观看 • 1 年前
3D Gaussian Splatting for Real-Time Radiance Field Rendering paper... page: Radiance Field methods have recently revolutionized novel-view synthesis of scenes captured with multiple photos or videos. However, achieving high visual quality still requires neural networks that are costly to train and render, while recent faster methods inevitably trade off speed for quality. For unbounded and complete scenes (rather than isolated objects) and 1080p resolution rendering, no current method can achieve real-time display rates. We introduce three key elements that allow us to achieve state-of-the-art visual quality while maintaining competitive training times and importantly allow high-quality real-time (>= 30 fps) novel-view synthesis at 1080p resolution. First, starting from sparse points produced during camera calibration, we represent the scene with 3D Gaussians that preserve desirable properties of continuous volumetric radiance fields for scene optimization while avoiding unnecessary computation in empty space; Second, we perform interleaved optimization/density control of the 3D Gaussians, notably optimizing anisotropic covariance to achieve an accurate representation of the scene; Third, we develop a fast visibility-aware rendering algorithm that supports anisotropic splatting and both accelerates training and allows realtime rendering. We demonstrate state-of-the-art visual quality and real-time rendering on several established datasets.show more

AK
633,674 次观看 • 3 年前
Introducing Kaleido💮 from AI at Meta — a universal... generative neural rendering engine for photorealistic, unified object and scene view synthesis. Kaleido is built on a simple but powerful design philosophy: 3D perception is a form of visual common sense. Following this idea, we formulate rendering purely as a sequence-to-sequence generation problem, successfully unifying neural rendering with the architecture principles behind modern language and video models. Unlike traditional neural rendering methods, Kaleido learns 3D purely in a data-driven way, without explicit 3D representations or structures. It acquires spatial understanding directly through large-scale video pretraining, then multi-view 3D data finetuning, inspired by how LLMs acquire textual common sense from large corpora before specialising in domains like coding. Through extensive ablations, we progressively modernised the architecture design and training strategies and tackled key scaling challenges in sequence-to-sequence generative rendering, arriving at a design that’s simple, versatile, and scalable. Kaleido significantly outperforms prior generative models in few-view settings, and remarkably is the first zero-shot generative method matches InstantNGP-level rendering quality in multi-view settings. We view Kaleido also as an alternative step towards world modeling that flexibly spans a spectrum of “realities": with many views, it faithfully reconstructs grounded reality; with fewer views, it imagines plausible unseen details. 🔗 Explore more results and paper:show more

Shikun Liu
22,442 次观看 • 11 个月前
Building on the previous paper, in this study we... compare a continuous “smooth return” S2>S1 model with an event-driven one, where long periods of relative calm are punctuated by short, intense episodes of global reorganisation. Both models cover the same time window. Neither uses archaeological data in its construction. When compared against where early humans and early civilizations actually appear and persist, the difference is statistically robust. The smooth model behaves like background noise. The event-driven model lines up in time and space far better than chance allows, even after aggressive temporal and spatial randomization tests. Statistically, the event-driven model lines up with where and when early civilizations appear far better than a smooth, continuous model, even after we randomize both timing and location to test what could arise by chance. The event timeline itself was built independently from well-known late-glacial disruptions - such as Heinrich events, meltwater pulses, and abrupt deglacial transitions - rather than from any archaeological data. Nothing here claims that specific events caused specific cultures. It does suggest that history may not unfold on a smooth clock. Human societies seem to flourish during recovery phases between disruptions, not during the disruptions themselves. The animation contrasts the two return models. Draft paper : Source & Results : (coming soon)show more

Craig Stone
10,899 次观看 • 7 个月前