Загрузка видео...
Не удалось загрузить видео
I'm still convinced that you can get a better acting performance with AI by generating dialogue first and generating the video to sync with the audio. As audio generation is so much cheaper you're also able to validate the performance for a few cents before using credits for the... show more
38,110 просмотров • 2 месяцев назад •via X (Twitter)
Комментарии: 49

In discussions I'm having about a feature film, we're hands down getting actors in a recording studio before a single shot is generated.

Excellent - I think a decent actor will be the gold standard for a while yet. Having said that I think we're getting to the point where a decent AI performance can work better than a poor human performance.

totally agree

the male voice sounds a bit like a mix of Liam Neeson and Liam Cunningham

Hah - I think the priest must be from a similar part of Ireland. Dublin or further North.

This is well done. I'm almost 40 min into an AI original drama based on the "Epic of Gilgamesh" I've done all the dialogue scenes with Grok Imagine. I was a struggle but it came out ok. This clip is pretty long but the Uruk general near the end is pretty good.

Soooo fucking boring. You don't even know the meaning of performances. You are trying to mimic, trying to surface level this shit. They are pretty mannequins man. Nothing dynamic in here at all.

Need to introduce a little motion noise in the body movements. Humans sway a bit imperceptibly - nobody is completely still. Perhaps some metrics on that can be introduced in the prompts - if things are too “still” you have simulated the uncanny valley of muscle control

I completely agree with the audio-first workflow.

Yes - I'm looking forward to seeing how SeeDance 2.0 handles references. Would be great if longer audio refs were possible combined with video refs.

You can suck my ass

Totally agree. Here's one I did starting from voice (Clip has strong language). The AI models produce generic performances.

Even better than that is recording a video with someone, even yourself, acting the scene. You'll get all the life and micro expressions missing from AI acting. In my opinions it's much much better, the uncanny valley is almost non existent

Yeah, this is fantastic. Also a strong argument for filmmakers using generative AI to work with real actors, dial in the performance like an animated film in the recording booth, and then iterate the visuals. Best of both worlds.

Acting and performance are two words that should not be allowed in conjunction with AI.

Agreed. When i built for AI anime production, one of the craziest quality gains i got was allowing users to upload/pre-generate their voicelines before sending it to seedance 2.0. You get 1000% better facial animations and way better acting authenticity.

I can only speak for grok imagine; just about anything would be an improvement from that. Everyone sounds like they are saying something for the fifth time to their near-deaf grandparen

Lol - yep. I wish @SpaceXAI would train Imagine to accept an audio reference it could sync. Visually Grok Imagine can be so strong with a detailed prompt.

I am going to try this! The acting and the entire scene is so natural. Its so good!

Thanks Jennifer and yes, even without the video generation it's pretty cool to have a cast read your scripts and take your direction.

The audio also gives way better context to the video gen

That’s how I’ve been using it too. Made Spielberg react rather amusingly to some ludicrous accusations in my last short.

And you can record your own voice and change it for perfect emotion and intonation.

That isn’t acting. It’s 1’s and 0’s

It may also allow two-camera setups -- Seedancing it twice from a different perspective -- which would let you more freely edit the results together.

Yeah - good point. With this one I prompted for a multishot video and prompted cuts so it did the edit for me, but totally possible to manually edit with multiple cameras.

I’d probably not use Liam Neeson’s voice unless you want him to find you and demonstrate his particular set of skills.

Where are you using seedream? I might try it out

One shot is for slop, applying proper workflows is what makes the difference! (gosh this looks like written by AI, what's happening to me?)

When I spend time writing something it's too precise and sounds as though it's been written by AI. I sometimes find myself wanting to add some first person uncertainty to sound less AI. (Need to change my system prompt - lol)

This looks pretty cool. What I want to see is someone take the raw video and then run it through post, using something like Resolve.

I did actually make some subtle changes using a custom tool. Added film grain, tiny bit of grading. This is a rejected clip from a series I'm building up so I didn't spend too much time on it but I could have worked on sound design and look a bit more.

I just don't like AI acting yet. This clip is really good... for AI. I'm sticking with music videos and action until a real advance is made with an under-the-skin personality suite for AI actors.

I agree.. Thoughts on best audio models? I've been using elevenlabs. Also don't sleep on LTX, it gets surprisingly really good outputs

@fableshowrunner I’ve yet to try that order out myself. I hear syncing to music works as well.

@fableshowrunner Yeah - music lip sync was one of the first things I tried with SeeDance 2. It works very well.

I've been thinking about this a lot lately for something I've been working on. I've taken to calling it semantic compositing. Fast is sometimes faster than than fastest because you treat everything like a compositing problem. Audios a great example.

Nice work! I gotta figure out how to do this with Higgsfield and eleven labs

I agree with you, tried multiple combinations, this is the best by far for me

Do you feed the audio dialogue to Seedance during video generation?

Yep. Just generated at FAL Downloaded audio and used that along with prompt and imagery to generate videos.

does the emotion in the audio hold once the video syncs to it, or does the model flatten the performance back out?

Usually the performance holds. I used the same timestamped dialogue section in audio and video prompt so that helped with consistent performance. 0:00 – PRIEST (low, measured; a pause after her name) “Jo… look at me. Not the corner. Me.” 0:05 – JO (quietly, with controlled irritation) “I am looking at you.” 0:08 – PRIEST (softening, almost a whisper) “Then tell me what it said.” 0:11 – JO (a small breath; faint amusement) “It said you’d ask that.” 0:15 – PRIEST (a beat; trying not to react) “What else did it tell you?” 0:18 – JO (softly, almost kindly) “That you’re frightened.” 0:22 – PRIEST (firmer now) “I’m not frightened of you.” 0:25 – JO (looks toward the corner; a whisper) “No. You’re frightened it knows your name.”

This is the best I have seen 😮

Damn I want to see the rest.

I’ve been building a story in the same world involving exorcism and the unseen. Still unsure what I will do with it.

Publish episodes?

you know, i think seedaudio 1.0 is what we'll get on sd 2.5, so we probably won't need this that much. but you're right about it being a lot cheaper to get right.

Good point - a while ago I was wondering about the routing of the SeeDance models and how it works internally.
