Decentralized Diffusion Models power stronger models trained on more... accessible infrastructure. DDMs mitigate the networking bottleneck that locks training into expensive and power-hungry centralized clusters. They scale gracefully to billions of parameters and generate photorealistic images with just a week of training on eight independent GPU nodes. They’re easy to implement, adopt DiT hyperparameters directly and outperform standard models FLOP-for-FLOP.show more

David McAllister
46,480 Aufrufe • vor 1 Jahr
Announcing our $320M Series A at a $2.3B valuation,... led by Khosla Ventures, with General Catalyst, Eric Schmidt and Jeff Bezos. General Intuition is the frontier lab for acting in space and time. We build large action foundation models trained on billions of ground truth action-labeled gameplay clips from 17M monthly active users on Medal, and push the frontier of world models to generate infinite training environments.show more

General Intuition
530,267 Aufrufe • vor 2 Monaten
Photorealistic Object Insertion with Diffusion-Guided Inverse Rendering discuss: The... correct insertion of virtual objects in images of real-world scenes requires a deep understanding of the scene's lighting, geometry and materials, as well as the image formation process. While recent large-scale diffusion models have shown strong generative and inpainting capabilities, we find that current models do not sufficiently "understand" the scene shown in a single picture to generate consistent lighting effects (shadows, bright reflections, etc.) while preserving the identity and details of the composited object. We propose using a personalized large diffusion model as guidance to a physically based inverse rendering process. Our method recovers scene lighting and tone-mapping parameters, allowing the photorealistic composition of arbitrary virtual objects in single frames or videos of indoor or outdoor scenes. Our physically based pipeline further enables automatic materials and tone-mapping refinement.show more

AK
19,101 Aufrufe • vor 2 Jahren
LayerAI Compute: Decentralized GPU Power for AI 🧬 LayerAI... is stepping into the physical AI infrastructure game with our latest: DePin for AI Computing. This move tackles a key AI/ML market gap: instant, on-demand compute power, available worldwide with just a few clicks. We're not just decentralizing data with AI2Earn; we're now also decentralizing the compute, making the AI ecosystem stronger and more resilient. The first iteration of this product will go live in the 2nd week of February.show more

LayerAI | AI2Earn
23,656 Aufrufe • vor 1 Jahr
🔥Nexera & Aethir: Unleashing AI’s Next Frontier Through Tokenized... GPU Power 🤝 Nexera is proud to join forces with Aethir in a strategic partnership to make cutting-edge AI infrastructure globally accessible. By tokenizing fractional GPU ownership, we’re enabling developers, enterprises, and investors everywhere to harness the explosive growth of deep learning and generative AI without being limited by geography, scale, or cost. With transparent tokenization, innovators can access powerful GPUs for faster model training and more advanced applications. GPU providers gain streamlined funding for expansion and upgrades, and investors tap into a high-growth market with secure, compliant opportunities that can provide higher yields than other RWA products. It’s an entirely new ecosystem where everyone can thrive, fueling AI’s evolution at an unprecedented pace. By 2030, the global GPU market is projected to exceed hundreds of billions of dollars, driven by the explosive demand for AI-powered applications, deep learning, and increasingly sophisticated generative models, ensuring that tokenizing these invaluable resources is poised to tap into a massive, rapidly expanding opportunity. $NXRAshow more

Nexera
27,776 Aufrufe • vor 1 Jahr
First look at SPARTA, a distributed AI training algorithm... that avoids synchronization by randomly exchanging sparse sets of parameters ( 1,000x reduction in inter-GPU communication, enabling training of large models over slow bandwidths without specialized infrastructure. SPARTA works on its own but can also be combined with sync-based low communication training algorithms like DiLoCo for even better performance.show more

EXO Labs
99,414 Aufrufe • vor 1 Jahr
Generative models can’t discover what they can’t reach. We’re... excited to introduce ActFlow: a continued pre-training scheme that actively expands the valid design space reachable by flow and diffusion models. We call this generable set expansion — a new learning principle for out-of-distribution generative modeling, and a step toward evolvable search spaces for scientific discovery. (1/5)show more

Riccardo De Santi
40,016 Aufrufe • vor 12 Tagen
POWER IS PLANE-SPECIFIC! In other words, you’re going to... develop much more sport-specific POWER by training in planes closely related to your sport. Baseball is a sport that relies heavily on being powerful, and efficient in the frontal and transverse plane. With that being said, a lot of our training focuses on developing power in these planes.show more

Alex Simone
28,220 Aufrufe • vor 10 Monaten
This is... not a remotely accurate description of what... the 2023 Al executive order did? Undersecretary Emil Michael: "If you remember the Biden executive order on Al, which was this crazy executive order that limited the amount of compute any model company could do and was essentially grandfathering in a small number of ai companies that they were gonna designate as the winners, and everyone else was out" Its not true that the EO limited the compute that AI companies could do. What it did do was require companies who were training models above a certain very high compute threshold (10^26 FLOP or 10^23 FLOP for models trained primarily on biological sequence data) to notify the government and share what testing and red teaming they were doing for certain national security risks. People are free to dislike the Biden AI EO! But it seems good to factually describe what the policy said.show more

Nathan Calvin
58,745 Aufrufe • vor 6 Monaten
Something NVIDIA & Google do better than anyone else... is software-hardware-system co-design, and not just optimizing hardware for current model architectures, but predicting future ones. Back in early 2022, when NVIDIA started the design process for NVL72, MoE (Mixture of Experts) models were not yet the standard, and dense models were still dominant for frontier models. However, NVIDIA's strong software-hardware co-design culture enabled them to make a calculated bet that MoEs were the future, and they built NVL72 specifically for best MoE performance per TCO (Total Cost of Ownership). Furthermore, back in 2022, disaggregated prefill and wide expert parallelism (wideEP) MoE inference optimizations hadn't been invented yet, but it turns out that these MoE inference optimizations work best on large-scale systems like NVL72. While most other AI chip companies' in-house AI labs focus on training small 5B models that mainly use data parallelism, NVIDIA and Google's in-house AI labs continuously push the boundaries of model architecture and training recipes, such as NVFP4 training. Just like Super Idol & IShowSpeed, there must be a strong partnership between software engineers and hardware engineers to deliver the best systems that maximize performance per TCO.show more

SemiAnalysis
51,021 Aufrufe • vor 9 Monaten
Holy sh!t ! OpenAI will have their custom inference... chips ready in just a few months and deployed at scale by the end of the year! 🤯 Training chip = The heavy lifters that require massive amounts of data and power to build and teach the AI models from scratch. Inference chip = The specialized, highly efficient chips that actually run the AI and generate the answers in real-time when you use it. This is going to help OpenAI drastically cut down their massive compute costs, speed up model reasoning times, and finally break free from relying entirely on Nvidia to scale their operations.show more

Chris
60,278 Aufrufe • vor 5 Monaten
Traditional physics-based models struggle with extreme rainfall, often overestimating... light events while underestimating extreme precipitation. Our new hybrid climate model, NeuralGCM, more accurately simulates extreme rainfall, mean precipitation (with a 40% bias reduction compared to CMIP6 models), and the diurnal cycle of precipitation. These improvements are a direct result of training the model on NASA satellite observations, not just reanalysis data.show more

Google Research
16,267 Aufrufe • vor 6 Monaten
(1/n) 🚀 With FastVideo, you can now generate a... 5-second video in 5 seconds on a single H200 GPU! Introducing FastWan series, a family of fast video generation models trained via a new recipe we term as “sparse distillation”, to speed up video denoising time by 70X! 🖥️ Live demo: (Thanks to @gmicloud for the support!) 🔗 Blog: 🔓 We fully open-source our models, code, and data with Apache-2.0 licensesshow more

Hao AI Lab
78,660 Aufrufe • vor 1 Jahr
Today, we are excited to announce an integration with... XION (XION), the L1 built for mainstream adoption through chain abstraction. Nesa will be powering XION’s rich ecosystem of decentralized applications with fast and secure AI inference. The first project, XionSpace (Xion Ecosystem), has integrated with Nesa to use AI image generation models such as Nesa’s SDXL on-chain to power visuals through its entertainment dApp. We now look forward to supporting XION's full network as they execute on their vision for a truly accessible blockchain for the masses. Stay tuned for more updates on this groundbreaking integration and upcoming releases.show more

Nesa
122,062 Aufrufe • vor 2 Jahren
A major question in multimodal modeling is how to... leverage strong pre-trained models, such as industry-scale LLMs and VLMs, during training to avoid starting from scratch. This is important as the demand for rich and domain-specific multimodal models continues to increase, while training them from scratch is obviously impractical due to limited data and compute. The so-called “any-to-any” multimodal models can model a large diverse dictionary of modalities. That’s good. Their downside is that their architectures are often not decoder-only, which has limited their performance in practice and prevents them from leveraging strong pre-trained decoder-only models as priors. We are releasing MODUS (ICML Conference '26) a decoder-only any-to-any multimodal model to address some of these questions. A single transformer decoder predicts any modality from any others with no modality-specific heads, losses, or task pipelines. We show efficient adaptation of established models, e.g., BAGEL, to rich any-to-any multimodal modeling. We are releasing 14B to 77B parameter models. All materials are open-source. The download links and demos here 🧵show more

Amir Zamir
23,927 Aufrufe • vor 2 Monaten
JUST IN: Meta AI introduces Voicebox, an all-in-one generative... speech model. Voicebox is an impressive breakthrough! It could do for speech what other models like GPT-3 and Stable Diffusion have done for text and images. Some key details: - Voicebox can synthesize speech across 6 languages - It's a general-purpose model that can perform tasks it wasn't trained on. It can perform noise removal, content editing, style conversion, and more - Supports in-context text-to-speech synthesis and cross-lingual style transfer - It's 20x faster than current models and outperforms single-purpose models through in-context learning paper: blog:show more

elvis
88,518 Aufrufe • vor 3 Jahren
Most recent diffusion language model research (that I’ve seen)... seems to be using masking as the noising process. It looks like, however, most closed-source models (Google Gemini Diffusion and possibly Inception Labs’ Mercury) use a different noising process, where instead of masking tokens, they replace them with different tokens (either with a random token or a semantically similar token). I wondered how they were getting such high throughput with the latter noising process, since I believed that optimizing inference with KVCache approximation would be more difficult (for various reasons). I visualized this noising process with tiny-diffusion and compared it to normal unmasking, and was very surprised to see how fast the generation “settles” into a reasonable output, and then only slightly refines afterwards, requiring much fewer steps in total. Unmasking (where tokens are never remasked, the typical implementation) is inherently limited in generation speed by the fact that an increase in tokens decoded per step leads to more errors due to the mismatch between individual and marginal token probability distributions we sample from. The token replacement noising process seems to have a much different set of characteristics. Because we sample each token per step, every token makes “progress” towards the final output each iteration (in addition to *potentially* giving other tokens more information in future steps). Generally, masking has outperformed other noising processes, which is probably why most research focused on it (using smaller models). But the paper referred to in the retweet shows that random replacement as a noising process may scale better as model size increases. Big labs might have noticed these results much earlier (due to having drastically more training resources and being able to test larger models), which may explain the discrepancy in the choice of noising process. I’m gonna test this with larger models, since tiny-diffusion only has 10M parameters.show more

nathan (in sf)
40,440 Aufrufe • vor 7 Monaten
Meet Stable Audio 3.0, the open-weight model family built... for artistic experimentation. This is our open invitation to experiment with generative audio. We believe the best innovations are still waiting to be built. The 4-1-1 on 3.0: 📣 You own your outputs, and can distribute and commercialize them under the Stability AI Community License (up to $1 million in revenue). 🎵 New and improved capabilities include variable-length generation up to six minutes, and full song composition on portable devices, no GPU required. ✅ Trained on a fully licensed dataset. 🎨 You can customize the models on your own library with support for LoRa training, which we’ve documented for the first time. More on the models 👇show more

Stability AI
166,625 Aufrufe • vor 3 Monaten
Earn Bitcoin rewards! 🤑 We’ve just launched #Bitcoin rewards... at 30% APY in the Wallet app and on decentralized exchange $Verse DEX. This self-custody Bitcoin rewards program enables anyone in the world to earn Bitcoin — without asking for permission, and without relying on a centralized exchange. Bitcoin rewards are in $tBTC, ThreshHold Network’s breakthrough Bitcoin bridge that brings the power of decentralized finance to Bitcoin, enabling faster transactions, flexible trading, lending, and more.show more

Bitcoin.com
20,264 Aufrufe • vor 1 Jahr