Today we released the code for our CVPR 2026... paper, Flowception. Flowception bridges fully bidirectional sequence modeling and autoregressive generation by inserting frames via learned order, then denoising them with continuous flow. Website: Code:show more

John Nguyen @ ICML
18,980 次观看 • 2 个月前
We have released the code and weights for our... #CVPR2023 paper "Avatars Grow Legs: Generating Smooth Human Motion from Sparse Tracking Inputs with Diffusion Model"! code: abs: project: The demo is below:show more

Artsiom Sanakoyeu
35,711 次观看 • 3 年前
(1/n) 🚀 With FastVideo, you can now generate a... 5-second video in 5 seconds on a single H200 GPU! Introducing FastWan series, a family of fast video generation models trained via a new recipe we term as “sparse distillation”, to speed up video denoising time by 70X! 🖥️ Live demo: (Thanks to @gmicloud for the support!) 🔗 Blog: 🔓 We fully open-source our models, code, and data with Apache-2.0 licensesshow more

Hao AI Lab
78,660 次观看 • 11 个月前
DimensionX: Create Any 3D and 4D Scenes from a... Single Image with Controllable Video Diffusion TL;DR: Create 3/4DGS from Video Diffusion Note: Some first inference code released (not all yet). Contributions (cited): • We present DimensionX, a novel framework for generating photorealistic 3D and 4D scenes from only a single image using controllable video diffusion. • We propose ST-Director, which decouples the spatial and temporal priors in video diffusion models by learning (spatial and temporal) dimension-aware modules with our curated datasets. We further enhance the hybriddimension control with a training-free composition approach according to the essence of video diffusion denoising process. • To bridge the gap between video diffusion and real-world scenes, we design a trajectory-aware mechanism for 3D generation and an identity-preserving denoising approach for 4D generation, enabling more realistic and controllable scene synthesis. • Extensive experiments manifest that our DimensionX delivers superior performance in video, 3D, and 4D generation compared with baseline methods.show more

MrNeRF
17,047 次观看 • 1 年前
Today, we are releasing Stable Video Diffusion, our first... foundation model for generative AI video based on the image model, Stable Diffusion. As part of this research preview, the code, weights, and research paper are now available. Additionally, today you can sign up for our waitlist to access a new upcoming web experience featuring a Text-To-Video interface. To access the model & sign up for our waitlist, visit our website here:show more

Stability AI
1,024,498 次观看 • 2 年前
Bio-inspired #TrueAI continues the journey, today we Jose Sánchez... & David Vivancos - e/acc are very glad to introduce for Qubic #OpenScience #MultiNeuraxon 2.0 hibridized with #Aigarth Come-from-Beyond Code as allways at GitHub Demo and #BrainBuilder at Hugging Face Demo Video Explainer later today. Paper will be presented in the following months, stay tuned for updates. 🧠💻Why does it matter? It bridges the gap between artificial and biological intelligence by replacing rigid, layer-by-layer AI pipelines with interconnected neural modules (spheres) that function like distinct regions of the brain. By utilizing continuous-time processing and trinary logic (excitatory, neutral, and inhibitory states), it paves the way for energy-efficient AI capable of real-time adaptation and lifelong learning without suffering from catastrophic forgetting. Stay tuned evolution towards #AGI just started...show more

David Vivancos - e/acc
19,258 次观看 • 4 个月前
Gaussian Head Avatar: Ultra High-fidelity Head Avatar via Dynamic... Gaussians paper page: Creating high-fidelity 3D head avatars has always been a research hotspot, but there remains a great challenge under lightweight sparse view setups. In this paper, we propose Gaussian Head Avatar represented by controllable 3D Gaussians for high-fidelity head avatar modeling. We optimize the neutral 3D Gaussians and a fully learned MLP-based deformation field to capture complex expressions. The two parts benefit each other, thereby our method can model fine-grained dynamic details while ensuring expression accuracy. Furthermore, we devise a well-designed geometry-guided initialization strategy based on implicit SDF and Deep Marching Tetrahedra for the stability and convergence of the training procedure. Experiments show our approach outperforms other state-of-the-art sparse-view methods, achieving ultra high-fidelity rendering quality at 2K resolution even under exaggerated expressions.show more

AK
65,847 次观看 • 2 年前
Bug fixes shipping to Grok Build 0.2.13 (release notes... will be available in the TUI and on change-log website) We are leveraging the alt-screen to better handle your background tasks, subagents, monitors with smart grouping allowing you to navigate between them quickly • Group Subagents → Tasks → Watchers, within subagents, order by agent type (Explore, General, Plan) • Show timestamps on pinned user messages at top of scrollback • Fix command highlighting when prompt has paste chips • Update the context-usage indicator to show tokens by default • ANSI16 fallback for themes • Better context usage breakdown/rendering • Update highlighting for /loop, monitor, tag colors • Tab now cycles Prompt → Scrollback → Tasks → Prompt • Order gateway turn-completion after streamed content • Formatting: fix language-tagged fenced code blocks and fix code under a list itemshow more

skcd
15,075 次观看 • 1 个月前
MaskINT: Video Editing via Interpolative Non-autoregressive Masked Transformers paper... page: Recent advances in generative AI have significantly enhanced image and video editing, particularly in the context of text prompt control. State-of-the-art approaches predominantly rely on diffusion models to accomplish these tasks. However, the computational demands of diffusion-based methods are substantial, often necessitating large-scale paired datasets for training, and therefore challenging the deployment in practical applications. This study addresses this challenge by breaking down the text-based video editing process into two separate stages. In the first stage, we leverage an existing text-to-image diffusion model to simultaneously edit a few keyframes without additional fine-tuning. In the second stage, we introduce an efficient model called MaskINT, which is built on non-autoregressive masked generative transformers and specializes in frame interpolation between the keyframes, benefiting from structural guidance provided by intermediate frames. Our comprehensive set of experiments illustrates the efficacy and efficiency of MaskINT when compared to other diffusion-based methodologies. This research offers a practical solution for text-based video editing and showcases the potential of non-autoregressive masked generative transformers in this domain.show more

AK
25,449 次观看 • 2 年前
🧬 We have many foundation models or language models... for DNAs, but can we control them? We introduce Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RL — a reinforcement learning framework for controllable cis-regulatory sequence generation. Paper: Code: 🔬What’s the challenge? Designing regulatory DNA that is both highly expressive in target cell types and inactive in others is essential for synthetic biology, gene therapy, and precision medicine. Yet, controlling these trade-offs is challenging due to sparse, sequence-level rewards and biological constraints. 🔥Why Ctrl-DNA? Ctrl-DNA fine-tunes pre-trained DNA language models using a value model free, Lagrangian-guided RL framework, enabling flexible and customizable constraint optimization. Users can define application-specific thresholds across cell types, balancing expression strength with specificity. ✅ Maximize target-cell expression ✅ Constrain off-target activity under user-defined thresholds ✅ Preserve cell-type-specific TF motif structure Benchmarked on human enhancer and promoter datasets, Ctrl-DNA consistently outperforms prior methods, achieving stronger specificity, higher fitness, and more biologically grounded sequence generation — all with direct control over regulatory trade-offs. Shoutout to the PhD students Xingyu Chen (Xingyu Chen ) and Rex Ma (Rex Ma) for their amazing work leading this project!show more

Bo Wang
30,719 次观看 • 1 年前
🚀 Self-speculation brings 6.75x real speedup for LLM generation... with SGLang inference! Same model drafts future tokens in Diffusion mode → then verifies them in AR (causal) mode. One model and one KV cache. Just different attention masks. Thanks to perfect alignment, we get 2× longer acceptance lengths than MTP techniques (Eagle-3, MTP, dFlash). We run 2 forward passes… but the 2× higher acceptance means we break even - and with zero overhead from extra drafter, KV cache, or LM head that comes with MTP - those are not free. Last week we released Nemotron-Labs-Diffusion + Tri-mode LLMs! We did continued pre-training on Ministral-3 models by switching attention patterns (block causal bidirectional). Result: one model that runs AR mode, Diffusion mode, and Self-Speculation. Diffusion mode already shows high benchmark accuracy - excited to see what happens when someone beats left-to-right acceptance! 🔥 Github: Paper: SGLang inference: Try the models on HF:show more

Pavlo Molchanov
66,554 次观看 • 1 个月前
Hey #NeuraxonMini is literally out! , we manage to... "transplant" a Neuraxon 2 bioinspired #AI brain to a physical robot the #SpheroMini moving from our last Scientific Paper (link bellow) by David Vivancos - e/acc & Jose Sánchez for Qubic #OpenScience hybridized with #Aigarth to the real World. First you need a Sphero Education Mini robot about 50$ Then you can try the first cool demos at Hugging Face: 1.- Neuraxon2MiniControl to drive the sphero robot 2.- Neuraxon2MiniWrite to write letters or words with physical moves of the sphero robot using Neuraxon Video Tutorials on youtube later today. Why this matters? Remember we are not building "dead" LLMs we are building #AliveAIs and for that we need to explore how it behaves in reality, from how it learns to how it fails, and what better way that in the emerging field of #robotics , time will tell if your next #HumanoidRobot have a #Neuraxon brain... Read the Paper: Explore the Neuraxon code here: Are you ready for #TrueAI ?show more

David Vivancos - e/acc
29,293 次观看 • 5 个月前
Depth Any Video with Scalable Synthetic Data AI physicists... and chemists continue to make strides in depth estimation from video. Check out this new paper featuring some impressive examples. See the thread for more details (unfortunately no code yet). Abstract: Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge through two key innovations. First, we develop a scalable synthetic data pipeline, capturing real-time video depth data from diverse game environments, yielding 40,000 video clips of 5-second duration, each with precise depth annotations. Second, we leverage the powerful priors of generative video diffusion models to handle real-world videos effectively, integrating advanced techniques such as rotary position encoding and flow matching to further enhance flexibility and efficiency. Unlike previous models, which are limited to fixed-length video sequences, our approach introduces a novel mixed-duration training strategy that handles videos of varying lengths and performs robustly across different frame rates 0 - even on single frames. At inference, we propose a depth interpolation method that enables our model to infer high-resolution video depth across sequences of up to 150 frames. Our model outperforms all previous generative depth models in terms of spatial accuracy and temporal consistency.show more

MrNeRF
27,428 次观看 • 1 年前
For years Yann LeCun has argued that generative video... models can't truly learn physics. DeepMind's Physics-IQ benchmark proved him right, reporting "a striking lack of physical understanding in current generative video models". The best one scored just 29.5%. This CVPR 2026 paper finds the fix in LeCun's own playbook. "Inference-time Physics Alignment" doesn't retrain the generator. It steers a video model's denoising at inference using a reward from VJEPA-2, LeCun's Joint Embedding Predictive Architecture. - Won first place in the ICCV 2025 PhysicsIQ Challenge with a 62.64% score, beating the previous state of the art by 7.42%. - Repurpose VJEPA-2's "surprise" score as a reward, then search and rank multiple candidate denoising trajectories at test time. - Why it matters: It's a neat vindication of LeCun's thesis. The pixel-prediction generator alone doesn't get physics, but the JEPA world model approach he champions supplies the physics the video model lacks.show more

Vai Viswanathan
44,186 次观看 • 1 个月前
World modeling and imitation learning have largely been considered... two disparate worlds. In our recent work, Unified World Models, just accepted to #RSS2025, Chuning Zhu provides a dead-simple unifying solution: just train a joint diffusion model over actions and future states, but with *decoupled* diffusion time steps across these modalities. Manipulating these decoupled time steps then allows for marginalization or conditioning on actions or states; a single model can serve as a policy, forward dynamics model, video prediction model, or inverse dynamics model by simply setting diffusion timesteps carefully. The resulting model can leverage video datasets along with robot training data much more effectively, and shows improved robustness, generalization, and flexibility. This is exciting because it is frustratingly simple, scalable, and shows strong improvement on real-world robotics problems. Please refer to Chuning Zhu 's excellent thread for more details! More details/code can be found on our website and in the paper -show more

Abhishek Gupta
11,430 次观看 • 1 年前
We release Diamond Maps💎 unlocking accurate and efficient guidance... for diffusion models. Our experiments show that our methods scale incredibly well. Excited to see what people will build with this! Accurate guidance has been a notoriously hard problem, but in this work, we’re bringing TWO (!) solutions to the table. The recipe for success: 1️⃣ Speed: Use distilled models (flow maps, mean flows, consistency models). 2️⃣ Exploration: Inject stochasticity to properly explore your search space. Because this fundamentally improves anything using flow matching and diffusion, we see a lot of potential for applications across audio, robotics, molecules, and beyond. Paper: Code: Huge thanks to an amazing team: Douglas Chen, Luca Eyring @ ICML26, Ishin Shah, Giri Anantharaman, Yutong (Kelly) He, Zeynep Akata, Tommi Jaakkola, Nicholas Boffi, and Max Simchowitz. It was awesome bringing this to life together!show more

Peter Holderrieth
60,179 次观看 • 3 个月前
🔥 Phoenix is officially live on Solaris AI Flow.... You can now trade Phoenix perps inside a Solaris AI workflow. No code. Drop a node onto the canvas, pick an operation, and wire it to anything: AI signals, price feeds, schedules, alerts. The full order-book DEX from Ellipsis Labs, now programmable. Built so you can trade with confidence: ✦ Paper mode is on by default. Every order is checked and simulated against the live order book, real depth and real slippage, but nothing is signed or broadcast. ✦ Paper behaves exactly like Live. If an order would be rejected on-chain, it is rejected in simulation too. No false fills, no surprises. ✦ Going Live is one switch, and it asks for confirmation before any real funds move. Build it. Test it. Trade it. A full perps strategy, proven on paper before a single dollar moves. 30 operations in one node: ✦ Read live markets, order book depth, candles, and funding rates ✦ Track your positions, collateral, and PnL, realized and unrealized ✦ Place limit, market, and stop-loss orders, plus conditional triggers ✦ Cancel orders, manage margin, and move collateral, all from the workflow No scripts. No backend. No terminal to babysit. Just a workflow that trades. You can try it for free. Demo + Link down below 👇show more

Solaris AI
10,075 次观看 • 2 个月前
After the good development and involvement shown, we are... hosting Rebirth in an AMA today at 6PM UTC with a $500 giveaway in Rebirth token for a handful of winners live. Rebirth comes with a cool concept to not only revive dead tokens but also give back to the crypto community as their old dusted tokens are reborn through their dapp receiving Rebirth tokens in exchange. They are doing things the right way and their Dapp has been released and is being fully audited by Assure. To know more about them and their utilities, check out and join their community at Do not miss our AMA at today at 6PM UTC, see you there! Disclaimer:show more

Travladd 𐤊
15,155 次观看 • 2 年前
Up, Up, Down, Down, Left, Right, Left Right, B,... A…❤️ As a kid I called the Konami Code the “Contra Code.” I played with friends and instinctively recall “Select, Start” at the end, which gives 30 lives for 2 players. Yet “Select” scrolls to the 2 player option and “Start” is necessary to start the game. Thus, it’s technically not part of the code. Pressing only “Start” after goes to 1 player mode with 30 lives. Notably, the 1st Nintendo Power issue’s “Classified Information” section explicitly included “Start” when disclosing the sequence. The Konami Code was initially created by Kazuhisa Hashimoto while testing the arcade to console port of Gradius, which he found too difficult. (It also “includes” Start to start the game.) There, it gave access to all power ups ups. Why it stayed in Gradius is unclear, and it appears in future Konami games, most famously Contra. It confers 30 lives in Contra, Life Force, and others, and is sometimes called “the 30 Lives Code.” The effect can vary in different games. As kids, knowing the code was like magic. We imparted it to our friends like a fabled wisdom. Like riding a bike, I’ll never forget it, and I smile inputting it now just as much as I did as a child. No shame in it—Nintendo codes were and still are rad. Homages to the Konami Code outside of gaming are ubiquitous, including in Wreck it Ralph, unlocking Amazon’s “Super Alexa,” and even unlocking a chiptune version of the National Anthem on a Bank of Canada website! Also, Contra’s soundtrack is so good…I hope I’m not the only dork who takes a moment to jam out to it 🤭show more

Christina Rose
14,993 次观看 • 1 年前
Introducing ASAL: Automating the Search for Artificial Life with... Foundation Models Artificial Life (ALife) research holds key insights that can transform and accelerate progress in AI. By speeding up ALife discovery with AI, we accelerate our understanding of emergence, evolution, and intelligence–core principles that can inspire the next generation of AI systems! We proudly collaborated with MIT, OpenAI, Swiss AI Lab IDSIA, and Ken Stanley on this exciting project. Full Paper (Website): Full Paper (arxiv): Code: In this work, we propose a new algorithm called Automated Search for Artificial Life (“ASAL”) to automate the discovery of artificial life using vision-language foundation models. Instead of tediously hand-designing every tiny rule of an Alife simulation, simply describe the space of simulations to search over, and ASAL will automatically discover the most interesting and open-ended artificial lifeforms! Because of the generality of foundation models, ASAL can discover new lifeforms across a diverse range of seminal ALife simulations, including Boids, Particle Life, Game of Life, Lenia, and Neural Cellular Automata. ASAL even discovered novel cellular automata rules that are more open-ended and expressive than the original Conway’s Game of Life. We believe this new paradigm may reignite ALife research by overcoming the bottleneck of manually designed simulations, thus advancing beyond the limits of human ingenuity.show more

Sakana AI
750,711 次观看 • 1 年前
Introducing Kaleido💮 from AI at Meta — a universal... generative neural rendering engine for photorealistic, unified object and scene view synthesis. Kaleido is built on a simple but powerful design philosophy: 3D perception is a form of visual common sense. Following this idea, we formulate rendering purely as a sequence-to-sequence generation problem, successfully unifying neural rendering with the architecture principles behind modern language and video models. Unlike traditional neural rendering methods, Kaleido learns 3D purely in a data-driven way, without explicit 3D representations or structures. It acquires spatial understanding directly through large-scale video pretraining, then multi-view 3D data finetuning, inspired by how LLMs acquire textual common sense from large corpora before specialising in domains like coding. Through extensive ablations, we progressively modernised the architecture design and training strategies and tackled key scaling challenges in sequence-to-sequence generative rendering, arriving at a design that’s simple, versatile, and scalable. Kaleido significantly outperforms prior generative models in few-view settings, and remarkably is the first zero-shot generative method matches InstantNGP-level rendering quality in multi-view settings. We view Kaleido also as an alternative step towards world modeling that flexibly spans a spectrum of “realities": with many views, it faithfully reconstructs grounded reality; with fewer views, it imagines plausible unseen details. 🔗 Explore more results and paper:show more

Shikun Liu
22,332 次观看 • 9 个月前