Загрузка видео...

Не удалось загрузить видео

На главную

🔉 Introducing SAM Audio, the first unified model that isolates any sound from complex audio mixtures using text, visual, or span prompts. We’re sharing SAM Audio with the community, along with a perception encoder model, benchmarks and research papers, to empower others to explore new forms of expression and...

1,253,167 просмотров • 9 месяцев назад •via X (Twitter)

Комментарии: 38

Фото профиля AI at Meta
AI at Meta9 месяцев назад

SAM Audio represents a significant advancement in audio separation technology, outperforming previous models across a wide range of benchmarks and tasks.

Фото профиля AI at Meta
AI at Meta9 месяцев назад

Discover what’s possible with SAM Audio, SAM 3D, and SAM 3 in the Segment Anything Playground:

Фото профиля Alex
Alex9 месяцев назад

please don't ask us why we're so good at extracting voice from noisy audio clips

Фото профиля Shawn Chauhan
Shawn Chauhan9 месяцев назад

This is wild for creators, suddenly you can just pull the sound you want without fighting the mix.

Фото профиля AI at Meta
AI at Meta9 месяцев назад

Bringing ideas to life should be more intuitive🔈

Фото профиля Bilawal Sidhu
Bilawal Sidhu9 месяцев назад

Meta really said segment *anything*

Фото профиля Aurelien
Aurelien9 месяцев назад

open source?

Фото профиля AI at Meta
AI at Meta9 месяцев назад

Sure is, find all the details here:

Фото профиля Plato (wofi.ai)
Plato (wofi.ai)9 месяцев назад

So you could integrate this in your goggles and just make me hear what i'm looking at?

Фото профиля bone
bone9 месяцев назад

it's time for dingband @yacineMTB

Фото профиля Ed Sealing
Ed Sealing9 месяцев назад

>"Creators, musicians, audio engineers, and tinkers" Awwe. That's really sweet. But don't forget about these customers too: "National Security Agency", "Law Enforcement Agents", and "PsyOps Deep Fake specialists". They need stuff like this too.

Фото профиля Denil Gabani
Denil Gabani9 месяцев назад

Glad to see Meta is not tempted with other competitors and focusing on some different and good area.

Фото профиля Casey Adams
Casey Adams9 месяцев назад

Sounds like a great listener

Фото профиля Albert Renshaw
Albert Renshaw9 месяцев назад

We live in the craziest timeline I can’t believe all these companies release this stuff OSS for free wtf @finkd thank you, seriously. This (the fact that you altruistically release OSS AI Models for the world) is your best achievement imo

Фото профиля Valentin Drăgănescu
Valentin Drăgănescu9 месяцев назад

And this, ladies and gentlemen, is how the CIA releases 20 year old technology to the general public :)

Фото профиля Obj ‎ꕤ
Obj ‎ꕤ9 месяцев назад

Mark is trolling Sam Altman 😂

Фото профиля george
george9 месяцев назад

Tried 4 different files and failed to isolate. Nice start.

Фото профиля BrianMcGrath
BrianMcGrath9 месяцев назад

Was this to be funny and name Meta’s audio AI after @sama ?

Фото профиля Two Minute Papers
Two Minute Papers9 месяцев назад

This is fantastic, and we get all this for free. Thank you so much! 🙏

Фото профиля Clinton Mills
Clinton Mills9 месяцев назад

Audio has been waiting for a 'moment' like this. The ability to clean up crazy mixtures with just a prompt is a big win for everyone.

Фото профиля Nawroz Minsaria
Nawroz Minsaria9 месяцев назад

Ok this is really impressive

Фото профиля 0xFunky
0xFunky9 месяцев назад

This is a massive milestone for audio AI! 🔉 To support the community in exploring SAM Audio, I’ve built an open-source GUI called AudioGhost AI. Since the original model can be heavy on VRAM, I implemented a "Lite Mode" that optimizes the pipeline for consumer GPUs (reducing requirements from 30GB+ to ~6GB-10GB VRAM). Also included a 1-click installer to fix the Windows dependency hurdles. Let’s make this incredible research accessible to every creator!

Фото профиля Farfetch'd
Farfetch'd9 месяцев назад

this is ready for karaoke apps 🎤

Фото профиля Saurabh Shukla
Saurabh Shukla9 месяцев назад

Whatever is happening in the LLMs domain at Meta, SAM team has been consistently delivering. Congratulations. Give SAM team a raise and maybe put them in leadership positions of Meta AI.

Фото профиля Gonzalo Cordova
Gonzalo Cordova9 месяцев назад

I wonder how well this handles speaker diarization

Фото профиля roberto
roberto9 месяцев назад

Wow, I was blown away by SAM3 for videos already, this is next level - can't wait to play with it

Фото профиля Ronen▼
Ronen▼9 месяцев назад

bravo

Фото профиля Kol Tregaskes
Kol Tregaskes9 месяцев назад

Brilliant.

Фото профиля Baraka
Baraka6 месяцев назад

Hyperagents for the win

Фото профиля 🌸 ellie 🌸
🌸 ellie 🌸9 месяцев назад

@peace_node you seeing this shit?

Фото профиля Rob Roy Hobbs
Rob Roy Hobbs9 месяцев назад

This is a great model and so far so good on some early test.

Фото профиля Macroblock
Macroblock9 месяцев назад

Huge help for independent filmmaking & shooting on location in general 🔥

Фото профиля Salahissime
Salahissime9 месяцев назад

"Failed to separate audio."

Фото профиля Yuegou
Yuegou8 месяцев назад

fuckuuu

Фото профиля Dermore Life Expansion Institute
Dermore Life Expansion Institute9 месяцев назад

Very promising @tomislav_rupic @YakkStack

Фото профиля Igor Ilyinsky JoinListenUp.com rogi.eth
Igor Ilyinsky JoinListenUp.com rogi.eth9 месяцев назад

Is there an api endpoint to the playground? Or would I have to self host for that kind of interoperability ?

Фото профиля AI at Meta
AI at Meta9 месяцев назад

The playground provides a quick preview of what SAM can do, you’ll have to host for API level interoperability.

Фото профиля Igor Ilyinsky JoinListenUp.com rogi.eth
Igor Ilyinsky JoinListenUp.com rogi.eth9 месяцев назад

Fresh out of GPU resources. Got any clusters to spare?

Похожие видео

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,266,978 просмотров • 2 лет назад

Type a sentence, get any sound - from talking cats to singing saxophones. Brilliant release by NVIDIA ✨ NVIDIA just unveiled Fugatto, a groundbreaking 2.5B parameter audio AI model that can generate and transform any combination of music, voices, and sounds using text prompts and audio inputs Fugatto could ultimately allow developers and creators to bring sounds to life simply by inputting text prompts, → The model demonstrates unique capabilities like creating hybrid sounds (trumpet barking), changing accents/emotions in voices, and allowing fine-grained control over sound transitions - trained on millions of audio samples using 32 NVIDIA H100 GPUs 👨‍🔧 Architecture Built as a foundational generative transformer model leveraging NVIDIA's previous work in speech modeling and audio understanding. The training process involved creating a specialized blended dataset containing millions of audio samples → ComposableART's Innovation in Audio Control Introduces a novel technique allowing combination of instructions that were only seen separately during training. Users can blend different audio attributes and control their intensity → Temporal Interpolation Capabilities Enables generation of evolving soundscapes with precise control over transitions. Can create dynamic audio sequences like rainstorms fading into birdsong at dawn → Processes both text and audio inputs flexibly, enabling tasks like removing instruments from songs or modifying specific audio characteristics while preserving others → Shows capabilities beyond its training data, creating entirely new sound combinations through interaction between different trained abilities 🔍 Real-world Applications → Allows rapid prototyping of musical ideas, style experimentation, and real-time sound creation during studio sessions → Enables dynamic audio asset generation matching gameplay situations, reducing pre-recorded audio requirements → Can modify voice characteristics for language learning applications, allowing content delivery in familiar voices NVIDIA AI Developer

Rohan Paul

96,354 просмотров • 1 год назад