正在加载视频...

视频加载失败

many friends keep asking me which AI model is the best for lip sync or talking characters, so here’s my honest personal breakdown based on my experience using these tools. this is purely from my perspective, so feel free to agree, disagree, or add your own insight. seedance 1.5...

31,627 次观看 • 6 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

🇨🇳 Another great Chinese Model, OmniHuman-1.5 from ByteDance Turns 1 image plus a voice track into expressive avatar video by pairing a System 1 and System 2 inspired planner with a Diffusion Transformer, Produces coherent motion for over 1 minute with moving camera and multi character scenes. Most avatar models move to the beat of the audio but miss meaning, so gestures feel generic and emotions feel shallow. The fix here is a Multimodal LLM planner that listens to the speech and drafts a structured plan describing intent, emotions, beats, and high level actions, which gives the motion engine clear semantic targets instead of only rhythm. The motion engine is a Multimodal Diffusion Transformer that fuses the plan with audio, the single reference image, and optional text prompts, then synthesizes continuous body, face, and head motion that matches both words and tone. A key trick is a Pseudo Last Frame, a synthetic target that summarizes the next expected state, which stabilizes fusion across modalities and keeps motion consistent over long spans. From just 1 image and speech, the system outputs speaking avatars with synchronized lips, context aware gestures, and continuous camera movement, and it also supports multi character interactions without manual choreography. Reported results show strong lip sync accuracy, high video quality, natural motion, and close match to text prompts, and the same setup works on nonhuman characters too.

Rohan Paul

63,859 次观看 • 10 个月前

You have to really give it to OpenAI because Sora 2 is very impressive on a lot of fronts: - high quality video model with great physics - high quality audio in each video - high character consistency - multiple characters in one scene - accurate characters voice - social platform attached to it Before today the best AI video models were dominated by Chinese companies like ByteDance and Kuaishou and Google with Veo3. ByteDance makes TikTok, Kuaishou makes Kwai (similar app) and Google has YouTube to train on But none of these models had great character consistency, if it was a feature at all, let alone multiple characters in one scene. Generally you'd make a video and the face would slowly change into someone else, just not good On top of that Google was struggling with allowing people to upload characters scared it'd get abused for deep fakes, and just generally nerfing their model so you can't really use it for anything OpenAI solved that by re-thinking ownership over your characters smartly with Cameo, which is essentially "train yourself as a AI model" which we've all been doing in our apps for years, but in a more smart way, where you can control if only you make content with your appearance, or others too They've also added voice training to it immediately, which people would have to do separate on for ex ElevenLabs before On top of that the social platform aspect: Google's Veo 3 didn't have ANY community at all, while the Chinese video models did, but it was all more like weekly themed contests to win free credits, they never really managed to make it more than that, and it kinda stayed in this nerdy AI hacking vibe This vibe fits how hard it was/is to simply make a video featuring you or your friends with proper voice and audio and everything that Sora 2 does for you. You'd have to go to ElevenLabs to train your voices, then go to for ex Photo AI to train yourself as a person, then make videos, then add audio and voices, then edit them together, a lot of work! We don't know if Sora 2's social platform features will actually be used or take off, but it's a real cool experiment in trying to find a way to build a community around AI in a more Instagram-like way Being able to tag your friends and then add them as multiple characters is innovative in both the social and technical aspect So TL;DR OpenAI essentially took a lot of stuff that was already technically possible, then added new things that weren't possible yet, and then put it all together in a very friendly interface that even my mom can use, with generation times of just a few minutes which is extremely fast if you think of the pipeline behind it (multiple video generation + voice + audio etc.) And also importantly, it doesn't look like they nerfed it much for safety which is also very cool considering the legal risks So yes very very very impressive

@levelsio

178,100 次观看 • 9 个月前

250709 | #ATEEZ #Hongjoong on how creative expression beyond music inspires his growth as an artist , TOKTOQ pop (voice) live (rough translation): I’m also studying design and slowly creating things on my own, step by step. I’ve said something similar before, but honestly - who knows what might happen in the distant future, right? For now, though, I’m still in the process of learning more about myself - my tastes, my design style, and how I work. And I know that if I ever do create something, our ATINYs would definitely take interest and support it. But as I continue getting to know myself, I just want to say - and I’ll say this clearly - I have absolutely no intention of starting a brand or selling anything at this point. Not even a little bit. Right now, I just see this - working and designing - as another way of expressing myself. That’s all it is. At least for now, I don’t have any plans beyond that. So I know there are people who hope I might do something more with this, and on the other hand, there may also be some fans who start to wonder, “Is he planning something?” - and maybe feel a bit uneasy about it. Because it could seem like I’m taking on too much or not focusing on my main work. But I’m very aware of that myself, and honestly, I don’t want that to happen. I really don’t. So to be clear - I’ll say it firmly - I don’t have any such plans right now. It all started simply because I wanted to try wearing clothes from different brands, and eventually, I thought, “I want to wear what I want,” or “I want to create something I’d like to wear.” That’s the situation I’m in. I just want to keep expressing myself. As long as it doesn’t become a burden for me or interfere with my schedule, I’d love to keep doing fun and creative things and share them with our ATINYs. So… it’s really just that. Since I’ve been using something like a stylized “HJ” - kind of like a personal mark - some people might start thinking, “Oh, is he launching a brand?” But absolutely not. That’s not the case at all. I’ve just been adding that mark to the clothes I make because I think it looks nice, and it kind of makes it feel like it’s mine. That’s really all there is to it. To be honest, I do want to make a tag eventually, but the design isn’t fully clear in my head yet - I haven’t figured it out. So for now, I’m just using the logo that’s in my mind. And honestly, it’s not like I’m trying to hide anything or doing something secretly behind my members’ backs. I just wanted to talk about it openly and put it out there. Because that way, I can really have fun with it. And if our ATINYs say, “Oh, that looks nice,” then I can just feel happy about it as it is. And even if I end up making something that doesn’t turn out so great sometimes, if ATINYs say, “You made that?” - even that, I can just laugh and enjoy it for what it is. So that’s what it is. That’s really the reason. Continuously creating - not just in music, but in other areas too - gives me so much energy. And I truly believe that this kind of creativity brings new inspiration to my performances as well. I think that’s what it is - the process of constantly making something new gives me another kind of drive, another kind of motivation. That’s what it feels like to me. So… that’s why I enjoy it. And honestly, that’s also why - even more so - I feel more motivated when it comes to things like choreography practice, or even just the basics of rapping. It makes me want to put in even more effort.

Irene | AhgaTiny🍋

27,502 次观看 • 1 年前

the lyrics from the song woonhak made as jaehyun’s birthday present 🧸🎵 “when the burden sitting on your shoulders feels heavy it's okay to put it down for a moment congrats your birthday just for today, it's okay to let go of what's been holding you tight and sleep peacefully brother has it been 4 years since i met hyung those strange jeans and heavy scented perfume the young me fell for that crooked charm asking you this and that about everything i had my first drink (in life) with hyung too and cried for a long time together too now we're sharing life together brother it might sound cheesy but can you get my sincerity? even if there comes a day when you feel like you’re on your own at least i'll always be by your side i know on those days when you feel like crying tell me, and let's share a drink you know you are my best friend it's strange, i can tell just by looking at hyung's face you shut yourself in the room, not saying a word again i know that anxiety won't fade away so easily that's why i'll hold on with you even if it breaks us your spring is coming soon so lift your shoulders high because hyung is my pride even if we died and were born again a thousand times i know some we'd still found each other my friend there's so much i'm sorry for and so much i'm grateful for but i'm sorry i don't express it well myung jaehyun i love you! even if there comes a day when you feel like you’re on your own at least i'll always be by your side i know on those days when you feel like crying tell me, and let's share a drink you know you are my best friend thank you for being the hyung i've always wanted to get since i was young i guess we can just grow old like this and share a coffin together”

노이

71,473 次观看 • 7 个月前

I am watching the new Netflix series Roger D. Parish and I was honestly surprised, in a good way at first, to hear Zimbabwe being mentioned right from the very first episode the Zimbabwean Flag features too. Even more interesting was the inclusion of Shona lines. As a Zimbabwean, moments like that usually bring a sense of pride because it feels like a small part of our culture is being recognised on a global platform. I couldn’t help but feel a bit conflicted about how the Shona language was portrayed. The way the lines were delivered sounded unnatural, almost mechanical. It did not feel like real, lived in Shona the way native speakers talk every day. Instead, it came across as forced and slightly exaggerated, which took away from what could have been a really powerful and authentic moment. What confused me most is that there are actual Zimbabwean actors and actresses based in the United States who speak fluent, natural Shona and understand the cultural depth behind the language. It would have made so much sense to involve them or at least consult native speakers to make those scenes feel real. When language is done properly, it adds texture and authenticity to a story. When it is not it can feel like a missed opportunity. That said, I still think Parish is a good series overall. The storyline is engaging and the production quality is strong. But I genuinely believe that Africa’s languages and cultures deserve better representation. If you are going to include Shona, it should sound like Shona as it is spoken by real people, not like robots reading words off a script. I hope that as more global platforms continue to tell diverse stories, they will also learn to do it with deeper respect and cultural accuracy.

Setfree Nherera Mafukidze 🇿🇼

33,293 次观看 • 7 个月前

🐺: As Nu said, I also read the feedback about me. I feel a bit shy talking about it. So, regarding my hairstyle, I’ve actually been thinking about it for a while and discussing with my stylist whether I should change it or try something new. Because depending on the work… what do you call it? Confidence in yourself. Sometimes, if the style is too much, I may feel less confident, or if it’s too much, it’s not suitable for the event. But now I’m trying to be more diverse and trying to change more. I’m trying more with some events because some styles are really about my confidence. Because sometimes, when I have long hair, I really want to get a haircut. I feel like I have to guess my hair. And I feel confident about my hair like this. For anyone who really knows me, they’ll understand that I take my hair seriously. I touch it so much that my stylist even complains, and Nu complains too. Because I’m confident in that style. But sometimes I don’t stick to that style all the time. I understand the feedback people give me, and I’m open to it 😽: Nu isn’t complaining when Hia touched it 🐺: Actually, I do want to do a style that shows my forehead. Huh? “Nu isn’t complaining?” Nu is complaining ka 😽: Nu is just teasing, not complaining 🐺: Oh, complaining Hia means teasing 😽: No 🐺: There are some hairstyles that everyone wants me to show my forehead, and honestly, I really want to wear that style. But it only works for still photos. It’s not handsome from every angle, or from certain angles, it doesn’t look good. Can you imagine? Because I don’t have a face that is heaven-given, handsome, that much, but it’s about right. Yes… #ZeePruk Z

Zee Pruk Vietnam | บ้าน ซี พฤกษ์ พานิช เวียดนาม

50,889 次观看 • 1 年前

JOONG LEGEND100 NEW FACE #100ListSocialClub2026 #จุงอาเชน #JoongArchen 🌞as for this year i recently wrapped filming for “Loveless Heroine” . it took around 40 to almost 50 shooting days. i feel that the whole team put so much effort into it. it’s a very big project so everyone can definitely look forward to it. it’s going to be fun 🌞: in this series i play Plew Kham. it’s a role many people have been waiting for because the character already has fans from the novel and is quite well-known. that made me feel quite pressured, wondering whether i could portray the character well enough. but i want everyone to believe that i truly put my heart into everything. please look forward to and support Plew Kham as well. everyone might see me shirtless throughout most of the series but i believe the number of my mom fans definitely won’t decrease—it’ll only increase … 🌞i eat 3 boiled eggs in the morning chicken at noon and eggs again at night : just because you have to take off your shirt? 🌞no i mean i have to take care of myself too. after wrapping filming i feel like i can finally fly. i’m really becoming jub chicken now. anyway please look forward to it and support us …. 🌞i feel that the hardest thing about taking on a period drama like this is actually doing my homework on the script and dialogue. cause as an actor, there are some things we can’t just improvise based on our own feelings. since it uses an ancient form of language i have to really study and absorb the script a lot. that part is quite difficult. it’s the dialogue. i feel like it really is old language. words and phrases like feel quite distant from my everyday life

🇻🇳Jaidee’s aunt Bamnie🐣

16,714 次观看 • 8 天前

New! ✨ Lex 🤝 Kit ✨ We built a new technique to train AI to write in your voice—using your Kit newsletters—that's the closest I've ever gotten AI to sound like me. 👉 > "Damn. This is really solid. Immediately obvious that it is trained on my newsletters." — Nathan Barry > "damn finally just read this and those subject lines and that example newsletter, it feels 90% nailed, super intrigued how we can build this tone of voice into the app more, feels like a total gear shift for any AI suggestions." — fred rivett 🇬🇧📈 (usually a skeptic of the "write a draft for me" approach) The way we did it is cool and (I think?) new! We all know if you go to ChatGPT or Claude and ask it to write for you, it's gonna sound like AI. Maybe you've tried uploading some examples or a style guide and still get disappointing results. Lots of AI writing apps purport to "write in your style," but they typically just generate a short summary of your style using a prompt like "Analyze the style and tone of these writing samples" and then stick it into the prompt. Maaaaybe they'll throw in a few examples. We've tried this and didn't think it was great, so we killed it. But when Nathan Barry pushed us to think about this problem again, we came up with a subtly new technique that had a big impact on the results. (The reason I'm sharing it is because our goal is to build the world's best interface for collaborating on text, and the world's best platform for saving, sharing, and running prompts. Proprietary prompting techniques are not our thing.) Instead of just asking what the "style" is (a very fuzzy question) we ask AI what patterns it can find. Specifically we ask it to look for patterns in structure and tone. Then—and this is crucial—we ask the AI to generate a detailed set of instructions that a new writer could use to consistently reproduce those patterns. We include those instructions and a bunch of examples in a prompt (often quite a large one, it's kinda expensive for us tbh). I think it works so much better than just giving examples or giving a broad overview of "style" because LLMs are trained to pay close attention to instructions and thrive on specificity. Here we ask the LLM to look for very specific patterns and generate equally specific instructions to reproduce those patterns. The other cool thing is unlike a fine-tuned model this is powered by a big-ass prompt that you can inspect and modify to your liking. Does it sound a little too enthusiastic? A bit cheesy? Just edit the prompt. Of course it's not perfect, it's still gonna need careful editing, but to us this feels like an obvious leap. I'd be really curious to hear if it feels the same to you. To start this is only available via our Kit integration but we'll start rolling it out more broadly soon. You can sign up at 👉 Would love your honest feedback!

Nathan Baschez

16,279 次观看 • 1 年前

Geoffrey Hinton says a big language model runs on about 1% of your brain's connections and still ends up knowing more than you: "So in your brain, you have a hundred trillion connections, roughly speaking. Okay. That's a lot. And you only live for about two billion seconds. That's not much." "If you compare how many seconds you live for, with how many connections you've got, you have a whole lot more connections than experiences." "Now with these neural nets, it's sort of the other way round. They only have of the order of a trillion connections. So like 1% of your connections, even in a big language model, many of them fewer, but they get thousands of times more experience than you." "So the big language models are solving the problem with not many connections, only a trillion. How do I make use of a huge amount of experience?" "And back propagation is really, really good at packing huge amounts of knowledge into not many connections." "But that's not the problem we're solving. We've got huge numbers of connections, not much experience. We need to sort of extract the most we can from each experience." Two to three billion seconds is the whole budget. Everything you know, you learned inside it. So evolution built you to squeeze a lot out of very little. Hinton's point is that a language model has the opposite problem and the opposite fix, and backprop turned out to be extremely good at that fix. Worth noticing what this predicts about failure. A system running on 1% of your wiring and thousands of times your experience is not going to fail the way you do. You fail from having seen too few examples. It fails from compressing too many into too little, and the compression is where the errors get made. That is a strange thing to be deploying into hospitals and courts with no way to inspect it. We test these systems by asking them questions, which tells you what came out. Nobody can yet look at a trillion connections and say what got packed in. - Geoffrey Hinton, Nobel laureate and Turing Award winner, on StarTalk (StarTalk) with Neil deGrasse Tyson.

Karl Mehta

510,252 次观看 • 22 小时前

#BaiLu on choosing to dub Li Peiyi with her own voice “I’d like to talk a bit about the dubbing. It was mainly because of the filming environment in Hengdian there were a lot of noise issues. When we first started shooting this drama, the plan was to use the original on-set audio entirely. But as filming continued, the environmental interference became too much. We had a huge amount of dialogue every day, and eventually the director told us, ‘Forget it either you dub it yourselves in post-production, or we’ll bring in professional voice actors.’ By the end, voice actress Qiao Shiyu had already recorded my lines three times. She worked incredibly hard! Later on, though, the director messaged me on WeChat and said, ‘Lulu, come to the editing room. Don’t just listen on your phone listen in surround sound and then make your choice.’ Since this character is somewhat similar to my role in The Legends (Zhao Yao), Teacher Qiao Shiyu had dubbed it very diligently three times. In the end, I spent an entire afternoon in the editing room, and the whole team decided together that my natural voice might suit Pei Yi better. My voice isn’t perfect, and my delivery isn’t the absolute best, but perhaps for Pei Yi who isn’t meant to sound overly polished or perfectly ‘pretty’ my vocal tone fit the character more. So we ultimately chose to use my own voice. As actors, we know that using original audio is a real test of our abilities. We’ll keep working hard, and for any areas where I didn’t do well enough, I hope everyone can be understanding.”

21,805 次观看 • 5 个月前