正在加载视频...

视频加载失败

Bark Text-to-Audio Model Full Text Input: "Why was six afraid of seven?" Ignore Bark's "I'm done with this input" token and tell Bark to just keep generating more audio anyway.

461,950 次观看 • 3 年前 •via X (Twitter)

7 条评论

J.J. Larrea 的头像
J.J. Larrea3 年前

Like Max Headroom for real: A stuttering glitchy simulation of human vocal performance, rather than a human enactment of a stuttering glitchy simulation of a human vocal performance. The future is now! And this sits right in the uncanny valley — in this context, deliberately.

Utah teapot 🫖 的头像
Utah teapot 🫖3 年前

Get this into physical robots immediately. We need this kind of output coming from creepy uncanny valley machines at grocery stories yesterday.

Bepis™ 🔀 的头像
Bepis™ 🔀3 年前

What are the consultants of seven tho

Neverduft 的头像
Neverduft3 年前

"..that is.. abusive.." LMAO

Simon Willison 的头像
Simon Willison3 年前

What did you use to generate the video?

Adrian Blake 的头像
Adrian Blake3 年前

Oh cool thanks now I can have some new nightmares instead of the old ones.

Alex Balfanz 的头像
Alex Balfanz3 年前

What model is generating the video!?

相关视频

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,265,314 次观看 • 1 年前