Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Bark Text-to-Audio Model Full Text Input: "Why was six afraid of seven?" Ignore Bark's "I'm done with this input" token and tell Bark to just keep generating more audio anyway.

461,950 görüntüleme • 3 yıl önce •via X (Twitter)

7 Yorum

J.J. Larrea profil fotoğrafı
J.J. Larrea3 yıl önce

Like Max Headroom for real: A stuttering glitchy simulation of human vocal performance, rather than a human enactment of a stuttering glitchy simulation of a human vocal performance. The future is now! And this sits right in the uncanny valley — in this context, deliberately.

Utah teapot 🫖 profil fotoğrafı
Utah teapot 🫖3 yıl önce

Get this into physical robots immediately. We need this kind of output coming from creepy uncanny valley machines at grocery stories yesterday.

Bepis™ 🔀 profil fotoğrafı
Bepis™ 🔀3 yıl önce

What are the consultants of seven tho

Neverduft profil fotoğrafı
Neverduft3 yıl önce

"..that is.. abusive.." LMAO

Simon Willison profil fotoğrafı
Simon Willison3 yıl önce

What did you use to generate the video?

Adrian Blake profil fotoğrafı
Adrian Blake3 yıl önce

Oh cool thanks now I can have some new nightmares instead of the old ones.

Alex Balfanz profil fotoğrafı
Alex Balfanz3 yıl önce

What model is generating the video!?

Benzer Videolar

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,265,399 görüntüleme • 1 yıl önce