Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Every AI audio provider ships a different SDK. TanStack AI just unified them. Gemini Lyria, fal MiniMax, Stable Audio. Same streaming call, any model. Typed. Tree-shakeable. Hooks for React, Solid, Vue, Svelte. Try it → Blog 👇

88,933 Aufrufe • vor 3 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕

1,078,418 Aufrufe • vor 2 Monaten

NVIDIA JUST DROPPED A FREE AI MODEL THAT READS PDFS, WATCHES VIDEOS, LISTENS TO AUDIO, AND UNDERSTANDS YOUR SCREEN SIMULTANEOUSLY. Not one at a time. ALL AT ONCE. In a single pass. It is called Nemotron 3 Nano Omni and it runs 9 times faster than every other multimodal model currently available. Think about what that actually means for how you work. Right now you are switching between tools constantly. One tool for transcribing your call recordings. A different tool for analyzing your client PDFs. Another tool for processing your training videos. A separate workflow for understanding what is happening on your screen. Four tools. Four contexts. Four different outputs you have to manually synthesize into one decision. Nemotron 3 Nano Omni does all of it in one model. One pass. One output. The use cases that just got dramatically simpler: Meeting recordings where you need the transcript, the visual context, and the document references all analyzed together. Training videos where the audio, the slides, and the on-screen demonstrations all feed into one coherent summary. Client PDFs where you need the document content cross-referenced against your screen data and your call notes simultaneously. Sales call transcripts analyzed alongside the proposals and the CRM data in one unified pass. This is not a marginal improvement on existing multimodal models. It is a 9x speed increase on a capability that was already changing how people work. Free. From NVIDIA. Available right now. Bookmark this before everyone catches on. Follow CyrilXBT for every AI capability shift the moment it drops.

CyrilXBT

37,847 Aufrufe • vor 3 Monaten

Holy sh*t! I just swapped a UGC creator with an AI character in 5 minutes 🤯 Same video. Same expressions. Same movements. Completely different person. Perfect for DTC brands & agencies who want to test multiple UGC variations without filming 10 different people. The use case: You've got a winning UGC video. The script works. The pacing works. But you want to test it with different creators—different ages, genders, ethnicities—to see what resonates best with your audience. Traditionally, you'd need to: Hire 5-10 different creators → Brief them all → Hope they match the original energy → Pay $500+ per person → Wait weeks This AI workflow does it differently: → Start with your original UGC video → Create AI character image with Nano Banana → Upload both to FAL AI's Wan Animate model → AI swaps the creator while maintaining all facial expressions, movements, sync → Get variation in ~10 minutes The motion tracking is insane. Every head tilt, smile, gesture from the original is replicated perfectly on the AI character. What you can test: → Same winning script, 5 different creator demographics → A/B test which "face" drives highest CTR → Localize content for different markets → Refresh creative without reshooting I recorded a full Loom walkthrough showing the exact process step-by-step. Want the complete tutorial? > Comment "SWAP" > Like this post And I'll send the Loom over (must be following so I can DM) (Obviously, only use this with videos you own 100% rights to + get creator permission upfront if you plan to make AI variations.)

Mike Futia

185,285 Aufrufe • vor 8 Monaten