正在加载视频...

视频加载失败

Udio v1.5 is here with advanced tools for creators! Make extraordinary music with AI & tag #UdioMusic on your creations.🤘 🔊 Improved Audio Quality 🎛️ Stem Downloads 🔑 Key Guidance 🎨 New Creation Workflows 🌍 Global Language Support 🎥 Shareable Lyric Videos 🎶 Audio to Audio

102,739 次观看 • 2 年前 •via X (Twitter)

10 条评论

udio 的头像
udio2 年前

Audio to Audio: Remix, edit, and extend your own audio. Here is an amazing example from Maxbarzel.

udio 的头像
udio2 年前

Superior Audio Quality: Headphones on for the incredible bass in "Reach the Stars" from the talented @Grotestique.

udio 的头像
udio2 年前

Key Guidance: Using the same prompt & indicating keys, you can now create songs in major and minor.

udio 的头像
udio2 年前

Global Language Support: Enjoy this catchy track in Mandarin showcasing improved global language support.

udio 的头像
udio2 年前

Improved Creation Workflows: Watch this quick overview of the new streamlined creation experience.

udio 的头像
udio2 年前

Shareable Lyric Videos: Generate engaging shareable lyric videos. Here's how.

udio 的头像
udio2 年前

Stem Downloads: Learn how to download stems of your creations.

udio 的头像
udio2 年前

Read our v1.5 blog post:

Moreno ☕ Ⓜ️ 的头像
Moreno ☕ Ⓜ️2 年前

As a producer I am a bit intimidated by ai but at the same time this also creates new opportunities claiming your own samples. I always had a hard time finding samples for Hiphop or Electro Swing music now I can orchestrate and sample them myself.

Cody Savage-Dragon/acc 的头像
Cody Savage-Dragon/acc2 年前

Love your program! Keep up the excellent work!

相关视频

Type a sentence, get any sound - from talking cats to singing saxophones. Brilliant release by NVIDIA ✨ NVIDIA just unveiled Fugatto, a groundbreaking 2.5B parameter audio AI model that can generate and transform any combination of music, voices, and sounds using text prompts and audio inputs Fugatto could ultimately allow developers and creators to bring sounds to life simply by inputting text prompts, → The model demonstrates unique capabilities like creating hybrid sounds (trumpet barking), changing accents/emotions in voices, and allowing fine-grained control over sound transitions - trained on millions of audio samples using 32 NVIDIA H100 GPUs 👨‍🔧 Architecture Built as a foundational generative transformer model leveraging NVIDIA's previous work in speech modeling and audio understanding. The training process involved creating a specialized blended dataset containing millions of audio samples → ComposableART's Innovation in Audio Control Introduces a novel technique allowing combination of instructions that were only seen separately during training. Users can blend different audio attributes and control their intensity → Temporal Interpolation Capabilities Enables generation of evolving soundscapes with precise control over transitions. Can create dynamic audio sequences like rainstorms fading into birdsong at dawn → Processes both text and audio inputs flexibly, enabling tasks like removing instruments from songs or modifying specific audio characteristics while preserving others → Shows capabilities beyond its training data, creating entirely new sound combinations through interaction between different trained abilities 🔍 Real-world Applications → Allows rapid prototyping of musical ideas, style experimentation, and real-time sound creation during studio sessions → Enables dynamic audio asset generation matching gameplay situations, reducing pre-recorded audio requirements → Can modify voice characteristics for language learning applications, allowing content delivery in familiar voices NVIDIA AI Developer

Rohan Paul

96,354 次观看 • 1 年前

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,265,502 次观看 • 1 年前

VoxCPM 2 just dropped by OpenBMB Only 2B-param open-source TTS (Text-to-Speech) model built for production-grade multilingual voice work. Apache-2.0 license, Can run on only 8GB VRAM. • Eliminates the "robotic" feel of traditional TTS, delivering prosody and emotional depth suitable for high-stakes professional environments like filmmaking, gaming, animation, and audiobooks. • 30-language multilingual: no language tag needed, just type in a supported language and generate directly. • Voice design: create a brand-new voice from a text description alone, like age, tone, pace, or emotion. No reference audio required. Describe the desired voice characteristics (gender, age, tone, emotion, pace …) in Control Instruction, and VoxCPM2 will craft a unique voice from your description alone. • Controllable cloning: clone from a short clip, then steer delivery style without losing the speaker’s core voice. • Ultimate cloning: use reference audio + transcript for continuation-style cloning that keeps the tiny vocal details. • 48kHz output: takes 16kHz reference audio and produces studio-quality speech without an external upsampler. • Real-time ready: around 0.3 RTF on RTX 4090, even lower with Nano-VLLM. • Commercial use: Apache-2.0 licensed. Developer-Friendly Infrastructure: - Native Torch Inference: Direct support for PyTorch-based workflows. - Training Flexibility: Supports both full-parameter and LoRA fine-tuning for specific domain adaptation. - Production Readiness: Compatible with voxcpm-nanovllm for large-scale, high-concurrency deployment.

Rohan Paul

13,541 次观看 • 4 个月前

Seedance 2.0 is allowing us to enter a new era of music video creation. Here is how I created HONEY. It was a quick test to see how well this workflow holds up. 🐝 1 - Write your song and generate the music with Suno 5.5. 2 - Use an image generator of your choice. For HONEY I combined both Grok Imagine for aesthetics and Nano Banana Pro for refined editing. 3 - In Capcut I import my audio and just save out a blank video video containing the audio. This step is important because this video file containing audio will now be used with Seedance 2.0 as a video reference with Omni. This allows the AI to apply automatic and realistic lipsync and movement to the music, it's extremely powerful! 4 - Once I have a both my image and video with audio as reference, I use Seedance 2.0 Omni and upload my starting image and then the video reference with the audio. 5 - From here I'm simply prompting like normal, specifying what's happening in my scene with detailed instructions, mentioning multi shots and camera angle changes and then specifying that the person is singing along to the song. I type out the lyrics that are present to have better lipsync accuracy. 6 - Once I have generated a video and like the result, I do video to video, so i upload that video that just got generated and type "The scene continues" and prompt new actions to take place. This allows you to expand on a narrative. These new shots can be used as B-ROLL and since I uploaded my video as reference I have full consistency of everything it saw in the video. This is also extremely powerful. 7 - This is actually the most difficult part. Edit in Capcut. This is where you need to understand pacing and shot selection from all the scenes you generated to bring it all together. You must be strategic with the editing. Goodluck! I'll probably record a video tutorial at some point as it's easier to see what is being done.

Travis Davids

19,236 次观看 • 4 个月前