Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

🔉 Introducing SAM Audio, the first unified model that isolates any sound from complex audio mixtures using text, visual, or span prompts. We’re sharing SAM Audio with the community, along with a perception encoder model, benchmarks and research papers, to empower others to explore new forms of expression and...

1,253,167 Aufrufe • vor 9 Monaten •via X (Twitter)

38 Kommentare

Profilbild von AI at Meta
AI at Metavor 9 Monaten

SAM Audio represents a significant advancement in audio separation technology, outperforming previous models across a wide range of benchmarks and tasks.

Profilbild von AI at Meta
AI at Metavor 9 Monaten

Discover what’s possible with SAM Audio, SAM 3D, and SAM 3 in the Segment Anything Playground:

Profilbild von Alex
Alexvor 9 Monaten

please don't ask us why we're so good at extracting voice from noisy audio clips

Profilbild von Shawn Chauhan
Shawn Chauhanvor 9 Monaten

This is wild for creators, suddenly you can just pull the sound you want without fighting the mix.

Profilbild von AI at Meta
AI at Metavor 9 Monaten

Bringing ideas to life should be more intuitive🔈

Profilbild von Bilawal Sidhu
Bilawal Sidhuvor 9 Monaten

Meta really said segment *anything*

Profilbild von Aurelien
Aurelienvor 9 Monaten

open source?

Profilbild von AI at Meta
AI at Metavor 9 Monaten

Sure is, find all the details here:

Profilbild von Plato (wofi.ai)
Plato (wofi.ai)vor 9 Monaten

So you could integrate this in your goggles and just make me hear what i'm looking at?

Profilbild von bone
bonevor 9 Monaten

it's time for dingband @yacineMTB

Profilbild von Ed Sealing
Ed Sealingvor 9 Monaten

>"Creators, musicians, audio engineers, and tinkers" Awwe. That's really sweet. But don't forget about these customers too: "National Security Agency", "Law Enforcement Agents", and "PsyOps Deep Fake specialists". They need stuff like this too.

Profilbild von Denil Gabani
Denil Gabanivor 9 Monaten

Glad to see Meta is not tempted with other competitors and focusing on some different and good area.

Profilbild von Casey Adams
Casey Adamsvor 9 Monaten

Sounds like a great listener

Profilbild von Albert Renshaw
Albert Renshawvor 9 Monaten

We live in the craziest timeline I can’t believe all these companies release this stuff OSS for free wtf @finkd thank you, seriously. This (the fact that you altruistically release OSS AI Models for the world) is your best achievement imo

Profilbild von Valentin Drăgănescu
Valentin Drăgănescuvor 9 Monaten

And this, ladies and gentlemen, is how the CIA releases 20 year old technology to the general public :)

Profilbild von Obj ‎ꕤ
Obj ‎ꕤvor 9 Monaten

Mark is trolling Sam Altman 😂

Profilbild von george
georgevor 9 Monaten

Tried 4 different files and failed to isolate. Nice start.

Profilbild von BrianMcGrath
BrianMcGrathvor 9 Monaten

Was this to be funny and name Meta’s audio AI after @sama ?

Profilbild von Two Minute Papers
Two Minute Papersvor 9 Monaten

This is fantastic, and we get all this for free. Thank you so much! 🙏

Profilbild von Clinton Mills
Clinton Millsvor 9 Monaten

Audio has been waiting for a 'moment' like this. The ability to clean up crazy mixtures with just a prompt is a big win for everyone.

Profilbild von Nawroz Minsaria
Nawroz Minsariavor 9 Monaten

Ok this is really impressive

Profilbild von 0xFunky
0xFunkyvor 9 Monaten

This is a massive milestone for audio AI! 🔉 To support the community in exploring SAM Audio, I’ve built an open-source GUI called AudioGhost AI. Since the original model can be heavy on VRAM, I implemented a "Lite Mode" that optimizes the pipeline for consumer GPUs (reducing requirements from 30GB+ to ~6GB-10GB VRAM). Also included a 1-click installer to fix the Windows dependency hurdles. Let’s make this incredible research accessible to every creator!

Profilbild von Farfetch'd
Farfetch'dvor 9 Monaten

this is ready for karaoke apps 🎤

Profilbild von Saurabh Shukla
Saurabh Shuklavor 9 Monaten

Whatever is happening in the LLMs domain at Meta, SAM team has been consistently delivering. Congratulations. Give SAM team a raise and maybe put them in leadership positions of Meta AI.

Profilbild von Gonzalo Cordova
Gonzalo Cordovavor 9 Monaten

I wonder how well this handles speaker diarization

Profilbild von roberto
robertovor 9 Monaten

Wow, I was blown away by SAM3 for videos already, this is next level - can't wait to play with it

Profilbild von Ronen▼
Ronen▼vor 9 Monaten

bravo

Profilbild von Kol Tregaskes
Kol Tregaskesvor 9 Monaten

Brilliant.

Profilbild von Baraka
Barakavor 6 Monaten

Hyperagents for the win

Profilbild von 🌸 ellie 🌸
🌸 ellie 🌸vor 9 Monaten

@peace_node you seeing this shit?

Profilbild von Rob Roy Hobbs
Rob Roy Hobbsvor 9 Monaten

This is a great model and so far so good on some early test.

Profilbild von Macroblock
Macroblockvor 9 Monaten

Huge help for independent filmmaking & shooting on location in general 🔥

Profilbild von Salahissime
Salahissimevor 9 Monaten

"Failed to separate audio."

Profilbild von Yuegou
Yuegouvor 8 Monaten

fuckuuu

Profilbild von Dermore Life Expansion Institute
Dermore Life Expansion Institutevor 9 Monaten

Very promising @tomislav_rupic @YakkStack

Profilbild von Igor Ilyinsky JoinListenUp.com rogi.eth
Igor Ilyinsky JoinListenUp.com rogi.ethvor 9 Monaten

Is there an api endpoint to the playground? Or would I have to self host for that kind of interoperability ?

Profilbild von AI at Meta
AI at Metavor 9 Monaten

The playground provides a quick preview of what SAM can do, you’ll have to host for API level interoperability.

Profilbild von Igor Ilyinsky JoinListenUp.com rogi.eth
Igor Ilyinsky JoinListenUp.com rogi.ethvor 9 Monaten

Fresh out of GPU resources. Got any clusters to spare?

Ähnliche Videos

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,266,978 Aufrufe • vor 2 Jahren

Type a sentence, get any sound - from talking cats to singing saxophones. Brilliant release by NVIDIA ✨ NVIDIA just unveiled Fugatto, a groundbreaking 2.5B parameter audio AI model that can generate and transform any combination of music, voices, and sounds using text prompts and audio inputs Fugatto could ultimately allow developers and creators to bring sounds to life simply by inputting text prompts, → The model demonstrates unique capabilities like creating hybrid sounds (trumpet barking), changing accents/emotions in voices, and allowing fine-grained control over sound transitions - trained on millions of audio samples using 32 NVIDIA H100 GPUs 👨‍🔧 Architecture Built as a foundational generative transformer model leveraging NVIDIA's previous work in speech modeling and audio understanding. The training process involved creating a specialized blended dataset containing millions of audio samples → ComposableART's Innovation in Audio Control Introduces a novel technique allowing combination of instructions that were only seen separately during training. Users can blend different audio attributes and control their intensity → Temporal Interpolation Capabilities Enables generation of evolving soundscapes with precise control over transitions. Can create dynamic audio sequences like rainstorms fading into birdsong at dawn → Processes both text and audio inputs flexibly, enabling tasks like removing instruments from songs or modifying specific audio characteristics while preserving others → Shows capabilities beyond its training data, creating entirely new sound combinations through interaction between different trained abilities 🔍 Real-world Applications → Allows rapid prototyping of musical ideas, style experimentation, and real-time sound creation during studio sessions → Enables dynamic audio asset generation matching gameplay situations, reducing pre-recorded audio requirements → Can modify voice characteristics for language learning applications, allowing content delivery in familiar voices NVIDIA AI Developer

Rohan Paul

96,354 Aufrufe • vor 1 Jahr