NVIDIA AI's banner
NVIDIA AI's profile picture

NVIDIA AI

@NVIDIAAI336,394 subscribers

Teaching your AI new tricks.

Shorts

We took a 30B model and split it in two to write tokens in parallel instead of one at a time. Introducing Nemotron-Labs-TwoTower: a diffusion language model from NVIDIA Research adapted from Nemotron-3-Nano-30B-A3B. Here’s how it works: one half holds the context, the other writes the tokens, with both reusing the pretrained model instead of training a new one from scratch. We found it kept 98.7% of the original model’s quality at 2.42× faster generation.

We took a 30B model and split it in two to write tokens in parallel instead of one at a time. Introducing Nemotron-Labs-TwoTower: a diffusion language model from NVIDIA Research adapted from Nemotron-3-Nano-30B-A3B. Here’s how it works: one half holds the context, the other writes the tokens, with both reusing the pretrained model instead of training a new one from scratch. We found it kept 98.7% of the original model’s quality at 2.42× faster generation.

761,701 görüntüleme

This #CVPR2026 paper from our research team is trending #1 on Hugging Face 🤗 Meet LocateAnything: a vision-language detection model that rethinks bounding box prediction. For AI agents and robots, “seeing” is only useful if a model can pinpoint where something is fast enough to act. Trained on 138M high-quality samples, LocateAnything decodes bounding boxes in parallel instead of one coordinate at a time, improving localization accuracy while dramatically increasing throughput for visual grounding and detection. Project page:

This #CVPR2026 paper from our research team is trending #1 on Hugging Face 🤗 Meet LocateAnything: a vision-language detection model that rethinks bounding box prediction. For AI agents and robots, “seeing” is only useful if a model can pinpoint where something is fast enough to act. Trained on 138M high-quality samples, LocateAnything decodes bounding boxes in parallel instead of one coordinate at a time, improving localization accuracy while dramatically increasing throughput for visual grounding and detection. Project page:

340,312 görüntüleme

Happy Friday! We just put DeepSeek-V4-Pro up on It’s the world’s largest open source model at 1.6T parameters, and you can run it for free running on NVIDIA Blackwell GPUs. Try the NVIDIA NIM API →

Happy Friday! We just put DeepSeek-V4-Pro up on It’s the world’s largest open source model at 1.6T parameters, and you can run it for free running on NVIDIA Blackwell GPUs. Try the NVIDIA NIM API →

206,258 görüntüleme

What if your model handled its own tuning? ⚙️ With NVIDIA TAO 7, just prompt your coding agent what you want in plain language and let it do the rest. • Agent skills plug into your coding agent and improve accuracy. • AutoML removes the guesswork from hyperparameters. LLM-guided tuning finds strong configs up to 2x faster. • Fine-tune any Hugging Face CV or VLM model on local NVIDIA GPUs. • Data enhanced fine-tuning helps your agent identify why the model fails, then fix it. Try TAO skills ➡️

What if your model handled its own tuning? ⚙️ With NVIDIA TAO 7, just prompt your coding agent what you want in plain language and let it do the rest. • Agent skills plug into your coding agent and improve accuracy. • AutoML removes the guesswork from hyperparameters. LLM-guided tuning finds strong configs up to 2x faster. • Fine-tune any Hugging Face CV or VLM model on local NVIDIA GPUs. • Data enhanced fine-tuning helps your agent identify why the model fails, then fix it. Try TAO skills ➡️

42,238 görüntüleme

Artificial Analysis Image-to-video generation is just as impressive. Input image: "Generate a 16:9 image from a dashcam view of a formula 1 racing event" Video prompt: "A high-speed racing event where a car navigates multiple winding turns" 🔊 Sound on - generated by Cosmos 3.

Artificial Analysis Image-to-video generation is just as impressive. Input image: "Generate a 16:9 image from a dashcam view of a formula 1 racing event" Video prompt: "A high-speed racing event where a car navigates multiple winding turns" 🔊 Sound on - generated by Cosmos 3.

49,260 görüntüleme

“Every enterprise needs a claw strategy.” How did LangChain go from a weekend project to 1B+ downloads in 3 years? We sat down with CEO and co-founder Harrison Chase (Harrison Chase) to talk deep agents, evolving agent architectures, and what’s coming next. 🎧 Full episode:

“Every enterprise needs a claw strategy.” How did LangChain go from a weekend project to 1B+ downloads in 3 years? We sat down with CEO and co-founder Harrison Chase (Harrison Chase) to talk deep agents, evolving agent architectures, and what’s coming next. 🎧 Full episode:

47,817 görüntüleme

Selected as a best paper finalist at #CVPR2026: PixelDiT from NVIDIA Research In most image generation models, a pretrained autoencoder compresses the image before any diffusion happens, causing quality loss that accumulates across the entire pipeline. PixelDiT, or Pixel Diffusion Transformers, removes this step entirely. It's a single-stage model that learns the diffusion process directly in pixel space, end-to-end.

Selected as a best paper finalist at #CVPR2026: PixelDiT from NVIDIA Research In most image generation models, a pretrained autoencoder compresses the image before any diffusion happens, causing quality loss that accumulates across the entire pipeline. PixelDiT, or Pixel Diffusion Transformers, removes this step entirely. It's a single-stage model that learns the diffusion process directly in pixel space, end-to-end.

28,364 görüntüleme

Factories are getting a new AI brain 🧠 Introducing NVIDIA Factory Operations Blueprint (FOX), a reference design for building factory manager agents that monitor operations, reason across real-time data, and coordinate specialized AI agents to help resolve issues at scale. Early adopters including Hon Hai Technology Group (Foxconn), Pegatron, Advantech USA, and Wistron AiEdge Corporation are already seeing gains in productivity, quality, and efficiency. Read more: #NVIDIAGTC

Factories are getting a new AI brain 🧠 Introducing NVIDIA Factory Operations Blueprint (FOX), a reference design for building factory manager agents that monitor operations, reason across real-time data, and coordinate specialized AI agents to help resolve issues at scale. Early adopters including Hon Hai Technology Group (Foxconn), Pegatron, Advantech USA, and Wistron AiEdge Corporation are already seeing gains in productivity, quality, and efficiency. Read more: #NVIDIAGTC

25,688 görüntüleme

With 3M+ downloads and counting, NVIDIA Cosmos is redefining physical AI. Announced at #CORL25, new Cosmos updates are allowing developers to generate diverse data for accelerating training robot models at scale. 👏 Cosmos Predict 2.5 will combine three models into one powerful model—reducing complexity, powering up to 30s video generation, and enabling multi-view simulations. 👏 Cosmos Transfer 2.5 will be 3.5x smaller yet faster and sharper—generating photorealistic synthetic data from 3D scenes or spatial inputs. 🔗

With 3M+ downloads and counting, NVIDIA Cosmos is redefining physical AI. Announced at #CORL25, new Cosmos updates are allowing developers to generate diverse data for accelerating training robot models at scale. 👏 Cosmos Predict 2.5 will combine three models into one powerful model—reducing complexity, powering up to 30s video generation, and enabling multi-view simulations. 👏 Cosmos Transfer 2.5 will be 3.5x smaller yet faster and sharper—generating photorealistic synthetic data from 3D scenes or spatial inputs. 🔗

40,359 görüntüleme

The next wave of AI agents won’t just read text—they’ll reason over video and extract actionable insights.💡 In this #NVIDIAGTC session, be the first to hear about the new features powering visual AI agents with NVIDIA Blueprint for Video Search and Summarization (VSS), open VLMs including the latest Cosmos Reason, for development across edge, on‑prem, and cloud workloads. Add to schedule 👉 📅 Tues, March 17 | 10:00 a.m PT Speakers🎤: Roopa Prabhu and Adam Ryason

The next wave of AI agents won’t just read text—they’ll reason over video and extract actionable insights.💡 In this #NVIDIAGTC session, be the first to hear about the new features powering visual AI agents with NVIDIA Blueprint for Video Search and Summarization (VSS), open VLMs including the latest Cosmos Reason, for development across edge, on‑prem, and cloud workloads. Add to schedule 👉 📅 Tues, March 17 | 10:00 a.m PT Speakers🎤: Roopa Prabhu and Adam Ryason

12,850 görüntüleme

Videos