Loading video...

Video Failed to Load

Go Home

Nvidia's PID (Pixel diffusion Decoder) It decodes latent representations into high-resolution images, replacing the decode–then–super-resolve cascade while achieving lower latency and higher visual quality page: HF repo:

37,247 views • 1 month ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 views • 10 months ago

Hey Joe, I've been away from the space for some time, but it seems you still haven't fixed what are clearly development errors on your platform. Earlier this year (or possibly late 2025), I returned to Duelbits and was quite active on the platform. One day, while playing HiLo, I noticed something unusual. At first, I thought I was imagining it because I was deep into a gambling session, but considering I've wagered well over $5 billion on HiLo alone across various sites, I decided to pay closer attention. What was the issue? I discovered what appeared to be a critical UI error. The "Lower" option was displayed as having a 92% hit chance when, based on the card shown, it should have been approximately 15% (or whatever the correct probability was). Instead, the platform displayed "Lower" as the higher-probability outcome and "Higher" as the lower-probability outcome. Any experienced HiLo player makes decisions based on the displayed percentages rather than waiting to calculate the probabilities manually. If the percentages shown are incorrect, that has the potential to mislead players and affect their decisions. You might not believe me, but I started recording my screen to see if it would happen again. It did. I captured it on video, sent the evidence to you, and what happened next? Nothing. So I'd like to ask the public for their opinion. Below, I'll attach screenshots and recordings of the gameplay. I asked you to investigate the issue, fix the error, and compensate me because I believe this mistake cost me approximately $20,000. What's even more concerning is that, to this day, another issue still exists: if you have an active HiLo session and refresh the page, the round can disappear entirely. How is that considered normal behavior? And while we're discussing this, let's also address the RTP claims. You promote your house games as having a 99% RTP, yet in the video, a HiLo sequence starting on a 5 (bet High, then Low on a J, then High on a 2, followed by a cashout) resulted in a net payout of 1.79x. Performing the same sequence and achieving the same outcome on Stake pays 1.83x. Can you explain that discrepancy?

Scarface

17,310 views • 1 month ago

PHOTON COUNTING CT is NOT a better CT It is a NEW imaging modality Photon Counting CT (PCCT) represents a transformative leap in medical imaging, not only as a molecular imaging modality but also as a technology offering ultra-high resolution and functional imaging capabilities. It is fundamentally more than just an enhanced version of traditional CT—PCCT introduces new ways of seeing and understanding the human body, providing critical insights at the molecular, structural, and functional levels. This positions PCCT as a unique imaging modality that requires a fresh approach to technical implementation, operational workflows, and financial planning. Despite the larger upfront investment, PCCT’s ability to drastically reduce downstream healthcare costs makes it a highly valuable investment in the long run. 1. Technical Innovations • Molecular Imaging and Energy Discrimination: Unlike traditional CT, which simply measures the total absorbed energy, PCCT counts individual X-ray photons and differentiates their energy levels. This allows for precise molecular imaging, revealing the composition of tissues and materials at a biochemical level. By distinguishing between different tissue types and contrast agents, PCCT opens up new diagnostic possibilities, such as identifying molecular biomarkers in tumors or distinguishing between stable and unstable plaque in coronary arteries. This capability shifts the focus of imaging from purely anatomical to both anatomical and molecular, offering more comprehensive diagnostic information. • Ultra-High Spatial Resolution: PCCT features significantly smaller detector elements compared to conventional CT scanners, allowing for ultra-high resolution imaging. This means clinicians can visualize fine structures such as microcalcifications in arteries, small lesions in soft tissues, or the intricate architecture of bones. This level of detail was previously unattainable with traditional CT. When combined with molecular imaging, this ultra-high resolution allows for the precise localization and characterization of disease at very early stages, which is essential for early diagnosis and intervention. • Functional Imaging Capabilities: PCCT also excels as a functional imaging modality. By capturing energy-resolved information, PCCT can provide insights into tissue functionality and dynamic physiological processes. For instance, it can detect changes in blood flow, tissue perfusion, and oxygenation without the need for additional contrast agents or scans. This functionality allows for real-time assessment of physiological processes, making it particularly valuable in cardiology, oncology, and neurology for evaluating organ function and monitoring disease progression. • Reduced Noise and Artifact Reduction: Photon-counting technology dramatically reduces electronic noise and imaging artifacts, such as beam hardening, resulting in clearer and more accurate images. The ability to deliver ultra-high resolution images with minimal artifacts improves diagnostic accuracy, reducing the need for repeat scans and ensuring that even subtle abnormalities are detected. 2. Operational Considerations • New Workflow for Molecular, High-Resolution, and Functional Imaging: The integration of molecular, ultra-high resolution, and functional imaging into routine clinical workflows introduces complexity that requires adaptation. Radiologists and technicians need specialized training to interpret and analyze multi-energy datasets that include molecular and functional information. PCCT produces a vast amount of detailed data, requiring clinicians to adopt new imaging protocols and refine their diagnostic approaches to fully leverage its capabilities. • Post-Processing and Data Management: PCCT generates richer, more complex datasets, which necessitates advanced post-processing tools and data management systems. Existing PACS and imaging software may not be equipped to handle such large volumes of data or to process functional and molecular information effectively. This means healthcare institutions must invest in robust IT infrastructure, including upgraded software and storage solutions, as well as provide additional training for staff on new imaging analysis techniques. • Revised Clinical Protocols: The molecular, functional, and ultra-high resolution imaging capabilities of PCCT will likely prompt changes in clinical protocols. For instance, the need for contrast agents may be reduced, simplifying patient preparation and decreasing the risk of adverse reactions. Additionally, the ability to monitor physiological functions in real-time through functional imaging could lead to more dynamic diagnostic procedures, such as assessing the effectiveness of interventions or treatments in real-time. 3. Financial Impact • Higher Initial Investment: PCCT systems are more expensive than traditional CT scanners due to their advanced technology, which includes photon-counting detectors and the computational power required for high-resolution, molecular, and functional imaging. While this upfront cost is significant, it is crucial to view it in the broader context of the downstream benefits and cost reductions that PCCT offers. • Downstream Cost Reductions: Although the initial capital investment is higher, PCCT’s ability to combine molecular, functional, and ultra-high resolution imaging leads to substantial reductions in downstream healthcare costs. Its superior diagnostic accuracy minimizes the need for follow-up tests, repeat scans, or invasive diagnostic procedures, such as diagnostic coronary angiographies. For example, in cardiology, PCCT can precisely differentiate between types of coronary plaque, reducing the need for invasive procedures to assess risk. • Lower Overall Healthcare Expenditures: By enabling earlier, more accurate diagnoses, PCCT can reduce the overall cost of patient care. Early detection of disease, particularly through its molecular and functional imaging capabilities, allows for more targeted treatments, potentially preventing the need for more aggressive and expensive interventions down the line. For instance, early-stage tumor detection via molecular imaging could lead to less invasive treatments, reducing hospital stays and improving patient outcomes, ultimately driving down healthcare costs. • Increased ROI Through Enhanced Patient Outcomes: Over time, the combination of molecular, functional, and ultra-high resolution imaging enhances diagnostic precision, which translates into better patient outcomes. Improved diagnostic accuracy reduces the incidence of unnecessary procedures, minimizes treatment delays, and results in more personalized and effective care. This leads to increased patient satisfaction, better healthcare outcomes, and greater patient throughput—all factors that improve the institution’s return on investment (ROI). • Competitive Advantage and New Revenue Streams: By adopting PCCT, healthcare institutions position themselves at the forefront of advanced imaging technologies. The ability to offer molecular, functional, and ultra-high resolution imaging creates a competitive advantage, attracting more complex and high-value cases. This can boost the institution’s reputation for excellence in diagnostics, leading to increased referrals, new patient populations, and expanded revenue opportunities. Summary Photon Counting CT (PCCT) is not just an evolution of existing CT technology—it is a molecular, ultra-high resolution, and functional imaging modality that fundamentally transforms the diagnostic landscape. Its ability to capture detailed molecular data, visualize minute anatomical structures with ultra-high resolution, and provide real-time functional imaging opens new possibilities for earlier and more precise diagnoses. While the financial investment in PCCT is larger, the reduction in downstream healthcare costs through improved diagnostic accuracy, fewer unnecessary interventions, and earlier disease detection far outweighs the initial expense. For institutions committed to advancing patient care and improving long-term financial outcomes, PCCT is an essential investment in the future of medical imaging. The video attached shows a patient accessing the Hospital for ACS. PCCT can provide ALL the imaging information of the concurrent imaging modalities (CXR, CAG, Echo, CMR) that you see around it... that's a lot! #PhotonCountingCT #MolecularImaging #UltraHighResolution #FunctionalImaging #FutureOfImaging #AdvancedMedicalImaging #EarlyDiseaseDetection #InnovativeCT #CuttingEdgeHealthcare #PrecisionDiagnostics #HealthcareInnovation #MedicalTechnology #CostEffectiveImaging #NextGenCT #PatientCareRevolution

Dr. Filippo Cademartiri

11,820 views • 1 year ago

This BlenderFusion paper basically says "screw trying to describe 3D edits through text" and just... use Blender :-) The idea is pretty straightforward -- instead of trying to cram 3D understanding into a diffusion model, use depth estimation & segmentation to project 2D images into 2.5D meshes, edit them in actual 3D software, then use a fine-tuned diffusion model to make the results photorealistic again. The clever bit is their "dual-stream architecture" -- the model sees both the original scene AND the edited Blender render in parallel, learning to preserve what matters while fixing the inevitable artifacts from transforming imperfect 2.5D/3D reconstructions. They train it with smart masking strategies so it learns when to ignore the original scene (for removals/replacements) and can manipulate objects independently of camera motion. What you get is pretty impressive control -- not just moving objects around, but changing materials, deforming shapes, swapping backgrounds, all while maintaining visual coherence. Neural Assets (one of my favorite papers last year) tried to crack this with learned object tokens, but it struggled with overlapping objects and loses fine details (due to low res DINO encodings). BlenderFusion just sidesteps the whole problem -- want to rotate something 173.5 degrees? Just rotate it in Blender. Want to duplicate an object 8 times? Copy paste away. The diffusion model's only job is making it look photorealistic, not figuring out the 3D underpinnings. The catch? Lacks temporal consistency for animation. Each viewpoint is generated independently, so while a single edit looks great, smoothly animating a car or camera down the street won't work -- you'd get flickering and inconsistencies between frames. That said, this approach is so much more intuitive for finer grain image editing than trying to describe your changes in text prompts. It's the kind of thing that makes you wonder why we're trying to do everything inside neural networks when perfectly good 3D tools already exist -- giving you the best of both worlds.

Bilawal Sidhu

34,440 views • 1 year ago

Most video-action robot models are a content-creation video generator with an action module attached. LingBot-VA 2.0 from Robbyant, a video-action foundation model, throws that starting point out and trains the whole stack natively for control. And it runs closed-loop at a peak 225 Hz. It's so important because A robot cannot move responsively when its controller pauses to imagine the next few frames. LingBot-VA 2.0 predicts during execution, then corrects using each real observation. And it carries only about 13B video parameters while activating roughly 1.9B per token. Bigger robot models usually mean slower reactions, creating a direct conflict between intelligence and control. LingBot-VA 2.0 is trained from scratch for robot control rather than adapted from a video generator built for content creation. Robbyant, an embodied AI company under Ant Group, built it to learn how scenes change under actions, predict what should happen next, and turn those predictions into real-time robot movements. Most video-action systems inherit a tokenizer and video backbone trained mainly to reproduce visual appearance. LingBot-VA 2.0 rebuilds both parts around physical control. Its semantic visual-action tokenizer maps observations toward features from a frozen vision foundation model and learns compact latent actions from frame-to-frame changes using self-supervised inverse and forward dynamics. Unlabeled web video can therefore carry action-relevant training signals without robot action labels. The policy is causal from the start, so every prediction can use only past observations. Its sparse Mixture-of-Experts video backbone has about 13B total parameters, while about 1.9B are active per token, keeping the compute lower during each step. A high-level vision-language planner breaks long tasks into smaller instructions, while the low-level video-action policy handles continuous movement. Foresight Reasoning predicts future visual states while the robot is already acting, then replaces imagined states with every new real observation. Combined with few-step distillation and systems acceleration, the paper reports a peak asynchronous execution frequency of 225 Hz. The model adapts from 10–15 demonstrations, transfers across robot embodiments, and handles some new tasks zero-shot. In the paper’s own evaluations, it reaches 93.6 average on RoboTwin 2.0 and reports stronger real-world results than LingBot-VA and π0.5 across the tested tasks. 🧵 1.

Rohan Paul

10,996 views • 8 days ago

🔴 Finally! NVIDIA has finally made the code for Neuralangelo public! It has the ability to transform any video into a highly detailed 3D environment, and it's a technology related to but DIFFERENT from NeRF. 💡 Here's how it works: It takes a 2D video as input, showing an object, monument, building, landscape, etc., from various perspectives and analyzes details such as depth, size, and the shapes of objects. From this, the AI sketches an initial 3D model, similar to how an artist molds a figure. This representation is then refined to highlight more details, just as an artist would make the final touches when sculpting. The result is a 3D environment/model, perfect for use in any environment. Imagine the applications it will have for video games, cinema, virtual environments, VR, and more! 📽️🎮 💡 More details: A year ago, an article was presented on a groundbreaking technique called NVIDIA's Instant NeRF. This technique turns images into stunning 3D scenes in a short time, ideal for creating realistic models for video games and other applications. Although Instant NeRF had a lot of potential, the generated models were not perfect and often lacked detailed structures, appearing somewhat cartoonish. A year on, NVIDIA releases a new technique based on Instant NeRF, named Neuralangelo. This enhances the fidelity of surface structures. While NeRF reconstructs real objects in virtual environments from images or videos, Instant NeRF speeds up this process, and Neuralangelo further improves the quality, making the generated objects appear even more realistic when examined up close. Neuralangelo improves Instant NeRF's approach in two key ways related to the hash grid encoding technique: 1⃣ Numerical gradients have been used to compute higher-order derivatives as a smoothing operation. This optimizes the "hash grid" encoding using numerical rather than analytical gradients, providing a smoother input to the network that produces the 3D model. 2⃣ A "coarse-to-fine" optimization has been implemented in the hash grids to control different levels of detail. That is, they first focus on a smoothed version of the scene, and then refine it with more detailed updates. Well, as Arthur C. Clarke said, "Any sufficiently advanced technology is indistinguishable from magic."

Javi Lopez ⛩️

689,128 views • 2 years ago