Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

🔉 Introducing SAM Audio, the first unified model that isolates any sound from complex audio mixtures using text, visual, or span prompts. We’re sharing SAM Audio with the community, along with a perception encoder model, benchmarks and research papers, to empower others to explore new forms of expression and...

1,253,167 görüntüleme • 9 ay önce •via X (Twitter)

38 Yorum

AI at Meta profil fotoğrafı
AI at Meta9 ay önce

SAM Audio represents a significant advancement in audio separation technology, outperforming previous models across a wide range of benchmarks and tasks.

AI at Meta profil fotoğrafı
AI at Meta9 ay önce

Discover what’s possible with SAM Audio, SAM 3D, and SAM 3 in the Segment Anything Playground:

Alex profil fotoğrafı
Alex9 ay önce

please don't ask us why we're so good at extracting voice from noisy audio clips

Shawn Chauhan profil fotoğrafı
Shawn Chauhan9 ay önce

This is wild for creators, suddenly you can just pull the sound you want without fighting the mix.

AI at Meta profil fotoğrafı
AI at Meta9 ay önce

Bringing ideas to life should be more intuitive🔈

Bilawal Sidhu profil fotoğrafı
Bilawal Sidhu9 ay önce

Meta really said segment *anything*

Aurelien profil fotoğrafı
Aurelien9 ay önce

open source?

AI at Meta profil fotoğrafı
AI at Meta9 ay önce

Sure is, find all the details here:

Plato (wofi.ai) profil fotoğrafı
Plato (wofi.ai)9 ay önce

So you could integrate this in your goggles and just make me hear what i'm looking at?

bone profil fotoğrafı
bone9 ay önce

it's time for dingband @yacineMTB

Ed Sealing profil fotoğrafı
Ed Sealing9 ay önce

>"Creators, musicians, audio engineers, and tinkers" Awwe. That's really sweet. But don't forget about these customers too: "National Security Agency", "Law Enforcement Agents", and "PsyOps Deep Fake specialists". They need stuff like this too.

Denil Gabani profil fotoğrafı
Denil Gabani9 ay önce

Glad to see Meta is not tempted with other competitors and focusing on some different and good area.

Casey Adams profil fotoğrafı
Casey Adams9 ay önce

Sounds like a great listener

Albert Renshaw profil fotoğrafı
Albert Renshaw9 ay önce

We live in the craziest timeline I can’t believe all these companies release this stuff OSS for free wtf @finkd thank you, seriously. This (the fact that you altruistically release OSS AI Models for the world) is your best achievement imo

Valentin Drăgănescu profil fotoğrafı
Valentin Drăgănescu9 ay önce

And this, ladies and gentlemen, is how the CIA releases 20 year old technology to the general public :)

Obj ‎ꕤ profil fotoğrafı
Obj ‎ꕤ9 ay önce

Mark is trolling Sam Altman 😂

george profil fotoğrafı
george9 ay önce

Tried 4 different files and failed to isolate. Nice start.

BrianMcGrath profil fotoğrafı
BrianMcGrath9 ay önce

Was this to be funny and name Meta’s audio AI after @sama ?

Two Minute Papers profil fotoğrafı
Two Minute Papers9 ay önce

This is fantastic, and we get all this for free. Thank you so much! 🙏

Clinton Mills profil fotoğrafı
Clinton Mills9 ay önce

Audio has been waiting for a 'moment' like this. The ability to clean up crazy mixtures with just a prompt is a big win for everyone.

Nawroz Minsaria profil fotoğrafı
Nawroz Minsaria9 ay önce

Ok this is really impressive

0xFunky profil fotoğrafı
0xFunky9 ay önce

This is a massive milestone for audio AI! 🔉 To support the community in exploring SAM Audio, I’ve built an open-source GUI called AudioGhost AI. Since the original model can be heavy on VRAM, I implemented a "Lite Mode" that optimizes the pipeline for consumer GPUs (reducing requirements from 30GB+ to ~6GB-10GB VRAM). Also included a 1-click installer to fix the Windows dependency hurdles. Let’s make this incredible research accessible to every creator!

Farfetch'd profil fotoğrafı
Farfetch'd9 ay önce

this is ready for karaoke apps 🎤

Saurabh Shukla profil fotoğrafı
Saurabh Shukla9 ay önce

Whatever is happening in the LLMs domain at Meta, SAM team has been consistently delivering. Congratulations. Give SAM team a raise and maybe put them in leadership positions of Meta AI.

Gonzalo Cordova profil fotoğrafı
Gonzalo Cordova9 ay önce

I wonder how well this handles speaker diarization

roberto profil fotoğrafı
roberto9 ay önce

Wow, I was blown away by SAM3 for videos already, this is next level - can't wait to play with it

Ronen▼ profil fotoğrafı
Ronen▼9 ay önce

bravo

Kol Tregaskes profil fotoğrafı
Kol Tregaskes9 ay önce

Brilliant.

Baraka profil fotoğrafı
Baraka6 ay önce

Hyperagents for the win

🌸 ellie 🌸 profil fotoğrafı
🌸 ellie 🌸9 ay önce

@peace_node you seeing this shit?

Rob Roy Hobbs profil fotoğrafı
Rob Roy Hobbs9 ay önce

This is a great model and so far so good on some early test.

Macroblock profil fotoğrafı
Macroblock9 ay önce

Huge help for independent filmmaking & shooting on location in general 🔥

Salahissime profil fotoğrafı
Salahissime9 ay önce

"Failed to separate audio."

Yuegou profil fotoğrafı
Yuegou8 ay önce

fuckuuu

Dermore Life Expansion Institute profil fotoğrafı
Dermore Life Expansion Institute9 ay önce

Very promising @tomislav_rupic @YakkStack

Igor Ilyinsky JoinListenUp.com rogi.eth profil fotoğrafı
Igor Ilyinsky JoinListenUp.com rogi.eth9 ay önce

Is there an api endpoint to the playground? Or would I have to self host for that kind of interoperability ?

AI at Meta profil fotoğrafı
AI at Meta9 ay önce

The playground provides a quick preview of what SAM can do, you’ll have to host for API level interoperability.

Igor Ilyinsky JoinListenUp.com rogi.eth profil fotoğrafı
Igor Ilyinsky JoinListenUp.com rogi.eth9 ay önce

Fresh out of GPU resources. Got any clusters to spare?

Benzer Videolar

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,266,978 görüntüleme • 2 yıl önce

Type a sentence, get any sound - from talking cats to singing saxophones. Brilliant release by NVIDIA ✨ NVIDIA just unveiled Fugatto, a groundbreaking 2.5B parameter audio AI model that can generate and transform any combination of music, voices, and sounds using text prompts and audio inputs Fugatto could ultimately allow developers and creators to bring sounds to life simply by inputting text prompts, → The model demonstrates unique capabilities like creating hybrid sounds (trumpet barking), changing accents/emotions in voices, and allowing fine-grained control over sound transitions - trained on millions of audio samples using 32 NVIDIA H100 GPUs 👨‍🔧 Architecture Built as a foundational generative transformer model leveraging NVIDIA's previous work in speech modeling and audio understanding. The training process involved creating a specialized blended dataset containing millions of audio samples → ComposableART's Innovation in Audio Control Introduces a novel technique allowing combination of instructions that were only seen separately during training. Users can blend different audio attributes and control their intensity → Temporal Interpolation Capabilities Enables generation of evolving soundscapes with precise control over transitions. Can create dynamic audio sequences like rainstorms fading into birdsong at dawn → Processes both text and audio inputs flexibly, enabling tasks like removing instruments from songs or modifying specific audio characteristics while preserving others → Shows capabilities beyond its training data, creating entirely new sound combinations through interaction between different trained abilities 🔍 Real-world Applications → Allows rapid prototyping of musical ideas, style experimentation, and real-time sound creation during studio sessions → Enables dynamic audio asset generation matching gameplay situations, reducing pre-recorded audio requirements → Can modify voice characteristics for language learning applications, allowing content delivery in familiar voices NVIDIA AI Developer

Rohan Paul

96,354 görüntüleme • 1 yıl önce