正在加载视频...

视频加载失败

Next gen AI audio separation is here 🤯 AudioSep is a model that can separate audio events, musical instruments, and even enhance speech with natural language queries which makes this a versatile tool for different audio tasks.

126,038 次观看 • 2 年前 •via X (Twitter)

10 条评论

Dreaming Tulpa 🥓👑 的头像
Dreaming Tulpa 🥓👑2 年前

If you like Tweets like this, you might enjoy my weekly newsletter, #aiartweekly. A free, once–weekly e-mail round-up of the latest AI art news, interviews with artists and useful tools & resources. Join 2500+ subscribers here:

Sylvain Filoni 的头像
Sylvain Filoni2 年前

And you can test AudioSep on @huggingface 🤗

Boise 🍄 Midtrip Art 的头像
Boise 🍄 Midtrip Art2 年前

Ooof this is crazy Seems AI is mostly talked about when it comes to visual creation, but soon it will be even more prominent in music creation

Dreaming Tulpa 🥓👑 的头像
Dreaming Tulpa 🥓👑2 年前

Because it’s something we can grasp instantly when we look at it. Listening takes time. But like with images, AI enhanced or generated music will be everywhere in the next 2-5 years.

Coopdville.eth (Coop D'Ville) 的头像
Coopdville.eth (Coop D'Ville)2 年前

This is a game changer for my upcoming music releases that use samples... 🔥🔥🔥🎶🎶🎶🎹🎹🎹

Kickiniteasy 的头像
Kickiniteasy2 年前

so rad this tech in plugins for music production is very useful for isolating tracks historically meh... but with the new AI isolation in plugins 🤯 made a song last week using AI track isolation & our minds were blown in the studio opens up whole new worlds for sampling

Dreaming Tulpa 🥓👑 的头像
Dreaming Tulpa 🥓👑2 年前

Got an example? Would love to give it a listen!

Brennan Erbz 的头像
Brennan Erbz2 年前

I wonder why Bytedance is working on audio separation 🤔

CMO333 的头像
CMO3332 年前

Wow

M.A. Buth 的头像
M.A. Buth2 年前

Sound cool!

相关视频

Type a sentence, get any sound - from talking cats to singing saxophones. Brilliant release by NVIDIA ✨ NVIDIA just unveiled Fugatto, a groundbreaking 2.5B parameter audio AI model that can generate and transform any combination of music, voices, and sounds using text prompts and audio inputs Fugatto could ultimately allow developers and creators to bring sounds to life simply by inputting text prompts, → The model demonstrates unique capabilities like creating hybrid sounds (trumpet barking), changing accents/emotions in voices, and allowing fine-grained control over sound transitions - trained on millions of audio samples using 32 NVIDIA H100 GPUs 👨‍🔧 Architecture Built as a foundational generative transformer model leveraging NVIDIA's previous work in speech modeling and audio understanding. The training process involved creating a specialized blended dataset containing millions of audio samples → ComposableART's Innovation in Audio Control Introduces a novel technique allowing combination of instructions that were only seen separately during training. Users can blend different audio attributes and control their intensity → Temporal Interpolation Capabilities Enables generation of evolving soundscapes with precise control over transitions. Can create dynamic audio sequences like rainstorms fading into birdsong at dawn → Processes both text and audio inputs flexibly, enabling tasks like removing instruments from songs or modifying specific audio characteristics while preserving others → Shows capabilities beyond its training data, creating entirely new sound combinations through interaction between different trained abilities 🔍 Real-world Applications → Allows rapid prototyping of musical ideas, style experimentation, and real-time sound creation during studio sessions → Enables dynamic audio asset generation matching gameplay situations, reducing pre-recorded audio requirements → Can modify voice characteristics for language learning applications, allowing content delivery in familiar voices NVIDIA AI Developer

Rohan Paul

96,354 次观看 • 1 年前