Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

🚀 Create stunning, character-consistent videos with Hailuo's Subject Reference (S2V-01) Model 🚀: ✨ Single-Image Simplicity: Unlock precise facial details from just one reference image. ⚡ Ultra-Low Cost: Achieve results at less than 1% of traditional computational demands. 🌟 High Flexibility: Enjoy creative freedom with simple text prompts. ⏳ Faster...

630,934 Aufrufe • vor 1 Jahr •via X (Twitter)

11 Kommentare

Profilbild von cyberyogi
cyberyogivor 1 Jahr

It’s incredible how easy this feature is to use compared to other AI models that require tons of tokens and multiple inputs for similar functionality. Amazing work! 🙌

Profilbild von Freepik
Freepikvor 1 Jahr

🚀 Introducing Freepik AI video generator. Everything you need to create high-quality, physically accurate videos in one place. 🤩

Profilbild von イッチ@AI術士
イッチ@AI術士vor 1 Jahr

This feature is very wonderful 👍

Profilbild von Shine by Nous ✨
Shine by Nous ✨vor 1 Jahr

Great update 🔥

Profilbild von Zerocarbon
Zerocarbonvor 1 Jahr

tried it results are amazing, but if the prompt contains any text generation in the video, it does not generate correctly..

Profilbild von OscarAI
OscarAIvor 1 Jahr

This new feature is awesome 👌

Profilbild von Maskai
Maskaivor 1 Jahr

This is amazing

Profilbild von AI ArtProdigy
AI ArtProdigyvor 1 Jahr

Being able to keep characters consistent with just one image is a game changer. I'm really excited to try this out and see what kind of videos I can whip up. Thanks for this awesome feature! 🎥✨

Profilbild von B.M. Cifer
B.M. Cifervor 1 Jahr

Awesome update. Enticing feature.

Profilbild von Özge Döner
Özge Dönervor 1 Jahr

❤️❤️

Profilbild von ysf.ai
ysf.aivor 1 Jahr

👍👍👍👍👍

Ähnliche Videos

🎥 Introducing Hailuo's Subject Reference: Revolutionizing Character Consistency in Video Creation 🔥 We’re excited to present Hailuo's S2V-01 model, a groundbreaking innovation in AI video generation that tackles one of the industry’s biggest challenges: maintaining consistent, realistic facial features and identity across dynamic video content, regardless of camera angles or movements. 💡 Why It’s a Game Changer: - Pioneering Technology: The first-of-its-kind to ensure character consistency in dynamic video generation, surpassing even fine-tuned models in performance. - Minimal Input, Maximum Impact: Generate character-consistent videos from just one reference image. Every frame remains true to the original identity with unmatched accuracy and reliability. - Enhanced Flexibility: Adjust more than just facial features—modify posture, expressions, lighting, and more, all with simple text-based prompts. 🌟While the new model enhances subject consistency, it may occasionally follow prompts less precisely than T2V or I2V, with some environmental morphing. Despite these early-stage challenges, Hailuo Subject Reference marks a significant leap in AI video generation. We’re committed to continual improvements including multi-subject references, objects references, and complex, multi-layered scenes. Explore the future of creative, consistent video production with Hailuo S2V-01 today. 🔥We believe the possibilities are endless.

Hailuo AI-MiniMax Hub

692,515 Aufrufe • vor 1 Jahr

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 Aufrufe • vor 10 Monaten

🇨🇳 Another great Chinese Model, OmniHuman-1.5 from ByteDance Turns 1 image plus a voice track into expressive avatar video by pairing a System 1 and System 2 inspired planner with a Diffusion Transformer, Produces coherent motion for over 1 minute with moving camera and multi character scenes. Most avatar models move to the beat of the audio but miss meaning, so gestures feel generic and emotions feel shallow. The fix here is a Multimodal LLM planner that listens to the speech and drafts a structured plan describing intent, emotions, beats, and high level actions, which gives the motion engine clear semantic targets instead of only rhythm. The motion engine is a Multimodal Diffusion Transformer that fuses the plan with audio, the single reference image, and optional text prompts, then synthesizes continuous body, face, and head motion that matches both words and tone. A key trick is a Pseudo Last Frame, a synthetic target that summarizes the next expected state, which stabilizes fusion across modalities and keeps motion consistent over long spans. From just 1 image and speech, the system outputs speaking avatars with synchronized lips, context aware gestures, and continuous camera movement, and it also supports multi character interactions without manual choreography. Reported results show strong lip sync accuracy, high video quality, natural motion, and close match to text prompts, and the same setup works on nonhuman characters too.

Rohan Paul

63,859 Aufrufe • vor 10 Monaten

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,264,973 Aufrufe • vor 1 Jahr

🔥HOLY SMOKES! $TAO holders! 🚀 SUBNET 19 (VISION) ON BITTENSOR IS ABSOLUTELY CRUSHING IT! In my 5+ years covering crypto and AI, this is one of the most impressive implementations I've seen. The combination of scale, performance, and decentralization is absolutely next level! 🚀 @namoray_dev @Corcel_X 💨 INSANE Speed Performance: - Llama 3.1 8B: 196.18 tokens/s with +107.23% advantage - Llama 3.1 70B: 124.96 tokens/s with +154.96% advantage - Llama 3.2 3B: 166.69 tokens/s with +21.66% advantage 🔥 Top Tier Model Integration: - Meta-Llama-3-70B & 8B Instruct - FLUX.1-schnell for Text-to-Image - ProteusV0.4-Lightning (Text & Image) - Multiple model variations for redundancy 🔥 What Makes This INSANE: - Complete decentralization - No single point of failure - Multiple model choices for redundancy - Real-time performance tracking - Transparent incentive structure The incentive distribution curve shows a healthy network with: - Strong rewards for top performers - Fair distribution across all participants - Clear path for growth and improvement - Sustainable economic model What's truly MIND-BLOWING is how they've managed to: 1. Scale to millions of operations 2. Maintain high quality across multiple tasks 3. Create a fair, competitive marketplace 4. Build in redundancy and reliability 5. Achieve true decentralization This isn't just another subnet - this is the future of decentralized AI inference happening RIGHT NOW! 🔥 1. MASSIVE Scale & Adoption: - We're seeing 7M+ tokens being processed - 14K+ processing steps being executed - Multiple AI models running simultaneously - Incredible miner participation across the network 2. Revolutionary Task Distribution: - Llama 3.1 70B leading with 20% weighting - Avatar Generation at 15% - Perfectly balanced task distribution for optimal network performance - Multiple specialized tasks including Text-to-Image and Image-to-Image processing 3. Elite Performance Metrics: - Top miners hitting 0.00775 incentive rates - Consistent performance across the network - Impressive scaling from top to bottom performers - Strong incentive curve maintaining network quality 📈 Network Performance: - Consistent upward trend in tokens/s - Quality scores maintaining high levels (>0.9) - Steady improvement in miner performance - Rock-solid network reliability ⚡ Platform Highlights: - Permissionless, serverless architecture - Global network of Always-On GPUs - Instant API access - Full decentralization - Multi-model support with seamless switching What makes this TRULY SPECIAL is the consistent upward trajectory in both speed and quality, while maintaining a decentralized architecture. The performance advantages over industry standards (+154.96% for 70B!) are absolutely mind-blowing! 🚀 This isn't just another AI subnet - it's a glimpse into the future of decentralized AI inference! The combination of speed, reliability, and model variety makes this one of the most impressive implementations in the space! 🔥 📽 Watch Now on YouTube and TikTok: Source 🔗

Andy ττ

11,616 Aufrufe • vor 1 Jahr

After years of being absolutely tortured by expensive ad creative pipelines, I think I finally found the ultimate savior for brand growth. Most AI video tools are great at generating flashy but random clips. The real challenge starts when you need to produce high-converting ecommerce content at scale without burning your budget. I’ve been testing Wizstar_official, and it honestly feels less like a simple AI generator and more like serious production infrastructure for scaling digital businesses. What stood out to me is how their ecosystem completely automates the two biggest bottlenecks in growth marketing: Bulk Testing and Creative Adaptation. First, their Agent setup paired with Fast Mode is a cheat code for volume. Instead of spending days scripting and storyboarding, the system intelligently extracts your product selling points, writes algorithm-friendly influencer scripts, and batch-produces massive ad variations in one day. It’s ultra-low-cost, built for rapid listing, and keeps your brand logos and product textures 100% consistent and lossless across the board. Second, the Video Reference workflow is an absolute game-changer. Instead of rebuilding every ad from scratch or guessing what works, you can reference any existing successful e-commerce video. The AI reverse-engineers its exact pacing, structure, camera movement, and storytelling style, and applies that winning DNA to a completely different product. That completely changes the production workflow from: prompt → random output into something closer to: reference → structured production → scalable content system Under the hood, Wizstar doesn't just rely on one platform; it supports flexible multi-model orchestration. Driven by their newly integrated Seedance 2.0 engine, it allows direct face input, meaning your character consistency and scene continuity stay rock-solid with absolutely none of that creepy AI face warping across complex cuts. If you are running global campaigns, you can also utilize their Video Translation tool to flip master clips into 12 languages with flawless, natural lip sync in minutes. ✨ New users get free credits upon registration 💸 First month subscription is only $19 (includes a complimentary 30-second E-commerce Agent experience to test features like Product to Video) Stop letting slow pipelines bottleneck your global growth. Try it here: #Wizstar #GrowthMarketing #AIVideo

FELIX

97,682 Aufrufe • vor 1 Monat

Let this chocolate brown Countach heal your algorithm! This commercial started as nothing more than an idea. I wanted to create something that felt timeless a chocolate brown Lamborghini Countach, an old-money atmosphere, a powerful woman, a luxury mansion party, natural daylight, authentic human interactions, and a commercial so realistic it wouldn't feel AI-generated at all. So I sat down with ChatGPT and explained the vision in detail. First, we built the story. I described the world, the mood, the characters, the setting, and the feeling I wanted viewers to experience. From there, ChatGPT helped transform that vision into a complete narrative, turning scattered ideas into a cinematic concept. Next came the planning stage. The story was broken down into a professional commercial timeline, scene by scene, shot by shot, exactly how a luxury automotive campaign would be storyboarded by a creative director and videographer. Then we went even deeper. Every scene was translated into detailed prompts, camera directions, movement instructions, lighting setups, slow-motion sequences, environmental details, character positioning, and luxury automotive close-ups. Once the sequence felt perfect, those ideas were converted into structured JSON prompts to maintain consistency throughout the project. After that, I generated individual storyboard images for every scene. Each image became a visual reference frame representing a specific moment of the commercial the mansion, the audience, the old-money styling, the Countach hero shots, the interior reveals, the scissor doors, and the final cinematic departure. Those reference images were then brought into Yapper where I used Seedance 2.0 to generate the actual video sequences. Multiple generations. Multiple revisions. Multiple refinements. Frame by frame, the vision started becoming reality. Finally, I edited both video sequences together, refined the pacing, synchronized the transitions, selected the music, and shaped everything into one seamless luxury automotive commercial. What you're watching isn't just an AI video. It's the result of storytelling, creative direction, prompt engineering, shot planning, storyboard creation, image generation, video generation, editing, and countless creative decisions. A chocolate brown Lamborghini Countach. An old-money lifestyle. A cinematic luxury fantasy brought to life from a single idea. Built through imagination. Designed through collaboration. Created frame by frame.

Julia Clark

26,529 Aufrufe • vor 1 Monat

As I promised yesterday, I'll briefly explain LoRA training and share a workflow I made so you can do it quickly. First, let me answer a very common question: 'Why train LoRAs when we have such advanced models?' Even though we have incredibly advanced models now (like NBP), we still can't always get them to do specific things we want. Simplest example: the spritesheet LoRA I made the other day. I generated 1000 images with Nano Banana and only 100 were what I wanted. The LoRA I trained using those 100 images gives me nearly 100% consistent results. Second point is cost and speed. With LoRA, we can cut costs by 4-5x. And while doing that, we're generating 4-5x faster. How many images do you need for a good LoRA? This depends on your LoRA's complexity. For example, when I training the spritesheet LoRA, even though I used 100 images, I didn't include buildings in the training data, so this LoRA doesn't work for buildings. So think about your LoRA's use cases and add examples for as many use cases as possible to improve quality. What are paired images and how to train LoRAs for image-editing? When training LoRAs for image editing on fal, we call each edit example paired images - one with _start suffix, one with _end suffix. For example, if you're training a background remove LoRA, the unedited original photo will be your '_start' image. The image with background removed will be the '_end' image. Simply put: images we want to edit or use as reference get _start, target images we want to achieve get '_end'. Important: save both images with the same name. Like image332_start.jpg and image332_end.jpg. This way the system knows which images pair together. What about training LoRAs for models with multiple image inputs? Same logic. We still use _start and _end suffixes, but with one difference. Since there are multiple input images, we can number them: _start, _start1, _start2. Example: start images, 1st image = Woman portrait (image35_start.jpg) 2nd image = Glasses photo (image35_start1.jpg) 3rd image = Hat photo (image35_start2.jpg) Output image = portrait of woman wearing glasses and hat (image35_end.jpg) Can we do more detailed captioning? Yes. Similarly, you can improve training quality by creating a txt file for each set with the caption inside. Example: create image35.txt and write: 'Recreate the image by putting the glasses from the second image and the hat from the third image on the woman in the first image.' What are Steps? How many should I use? What's Learning Rate? Steps determines how many times the model sees and processes your training data (your images). Each step, the model learns a bit more. But as steps increase, so does the risk of overfitting. So there's no real default. But for a simpler LoRA with 20 paired images, 1000 steps is ideal. Here's a metaphor for the Steps and Learning Rate relationship: Imagine you have a balloon. Our goal is to inflate it to the optimal size. Steps = How many times we blow into the balloon Learning rate = How hard we blow each time If we blow too softly, we need to blow many more times. If we blow too hard, we risk popping it quickly and can't reach optimal size. Of course training won't explode, but it won't work as intended because it wasn't trained optimally. Training's done, now what? Once training's complete, you'll have a safetensors file. Every model you train on fal has a LoRA inference endpoint. In that inference, add your safetensors file link to the LoRA url input, and you can use your LoRA. Thanks for the read! The workflow in the video: If I forgot anything, let me know in the replies.

ilker

15,133 Aufrufe • vor 6 Monaten

Everyone is sleeping on Meta's SAM 3 release. But it's actually a big deal. Here's why: Companies spend millions paying humans to label images and videos frame by frame. A single autonomous driving dataset? Months of work, hundreds of annotators, millions in cost. Without labeled data, you can't train custom models. Without custom models, you're stuck with generic solutions. This is why most companies never move past pilots. SAM 3 breaks this cycle. First let's look at the evolution: SAM 1 segmented objects when you clicked on them. Revolutionary, but one object at a time. SAM 2 added video tracking with memory. Game-changing, but you still manually prompted every object. SAM 3 changes everything with text prompts. Type "yellow school bus" and it finds ALL of them in your image or video. Not just one. Every instance across thousands of frames. Now here's where people get confused: "Can't I just use GPT-5 or Gemini for this?" No, and here's why that's a terrible approach. Large multimodal LLMs are great for reasoning, but they're slow and expensive for production visual tasks. You're paying API costs per image, waiting seconds for responses, getting inconsistent results. SAM 3 runs in 30 milliseconds on a single GPU for 100+ objects. That's 100x faster, and you own the infrastructure. More importantly, SAM 3 gives you precise pixel-level masks, not descriptions. Try asking an LLM to segment every defective part on a manufacturing line in real-time. It won't work. SAM 3 does this effortlessly. The real breakthrough is their data engine. Meta built an AI-human hybrid system that's 5x faster for complex annotations. They trained SAM 3 on 4 million unique visual concepts - 50x more than existing benchmarks like LVIS. SAM 3 is trained on 4 million unique visual concepts, it handles everything: - Text-based concept search - Interactive refinement with clicks - Video tracking across frames - Zero-shot detection of new concepts The model is open source. Weights, code, and benchmarks are on GitHub. If you're building computer vision applications, this is the foundation model to evaluate. The annotation time savings alone will pay for integration costs within weeks. Find the relevant links in the next tweet!

Akshay 🚀

46,404 Aufrufe • vor 8 Monaten

🚨 The Next Evolution of AI Music is Here 🚨 We haven’t been standing still. We’ve been building at an incredible pace, with laser-sharp focus, pushing the boundaries of AI-powered music creation like never before. Our latest upgrade isn’t just more powerful—it’s more versatile, precise, and deeply creative than anything before. 🎶 Proof is in the sound: This song was generated from a simple prompt—“Blues with slight Arabian influence about a man lost in the desert searching for his bride.” Listen to the end and hear how $SUEDE AI captures emotion, style, and storytelling like never before. But this is just the beginning. Our new features are built for artists who want total creative control. Get extremely granular with how you craft and shape your sound: 🎛️ Full Production Control – Download an entire pack of every isolated instrument. 🎤 Use Your Own Voice – Or someone else’s. 📝 Exact Lyrics, Your Way – Have your words set to music seamlessly. 🎶 Reference Songs – Upload one for style analysis, extraction, and modeling—or simply note a publicly available track. 🔊 Text-to-Speech & AI Vocalists – Shape voices like never before. 🎼 Melody Collaboration – Upload a melody idea and let others build around it—or vice versa. However, due to cost considerations, we’ve capped it at 3 free songs per trial until subscription payments roll out in the next day or two. We’ll be launching a new payment gateway soon, so stay tuned for more details. And remember, all of this is powered by the $SUEDE token. It fuels the entire ecosystem, allowing artists to generate, own, and monetize their work like never before. We’re still working out a few kinks—like image generation—but prepare to be impressed. A major post is coming soon, breaking down these game-changing features and the revenue model behind them. Thread dropping soon. Turn notifications on. $SUEDE powers the future of culture. #SuedeAI #Web3Music

Suede Labs

17,421 Aufrufe • vor 1 Jahr

🚀 Announcing Echo — our new frontier model for 3D world generation. Echo turns a simple text prompt or image into a fully explorable, 3D-consistent world. Instead of disconnected views, the result is a single, coherent spatial representation you can move through freely. This is part of a bigger shift in AI: from generating pixels and tokens to generating spaces. Echo predicts a geometry-grounded 3D scene at metric scale, meaning every novel view, depth map, and interaction comes from the same underlying world — not independent hallucinations. Once generated, the world is interactive in real time. You control the camera, explore from any angle, and render instantly — even on low-end hardware, directly in the browser. High-quality 3D world exploration is no longer gated by expensive equipment. Under the hood, Echo infers a physically grounded 3D representation and converts it into a renderable format. For our web demo, we use 3D Gaussian Splatting (3DGS) for fast, GPU-friendly rendering — but the representation itself is flexible and can be easily adapted. Why this matters: consistent 3D worlds unlock real workflows — digital twins, 3D design, game environments, robotics simulation, and more. From a single photo or a line of text, Echo builds worlds that are reliable, editable, and spatially faithful. Echo also enables scene editing and restyling. Change materials, remove or add objects, explore design variations — all while preserving global 3D consistency. Editing no longer breaks the world. This is only the beginning. Echo is the foundation for future world models with dynamics, physical reasoning, and richer interaction — environments that don’t just look right, but behave right. Explore the generated worlds on our website and sign up for the closed beta. The era of spatial intelligence starts here. 🌍 #Echo #WorldModels #SpatialAI #3DFoundationModels Check it out:

SpAItial AI

176,105 Aufrufe • vor 7 Monaten

This BlenderFusion paper basically says "screw trying to describe 3D edits through text" and just... use Blender :-) The idea is pretty straightforward -- instead of trying to cram 3D understanding into a diffusion model, use depth estimation & segmentation to project 2D images into 2.5D meshes, edit them in actual 3D software, then use a fine-tuned diffusion model to make the results photorealistic again. The clever bit is their "dual-stream architecture" -- the model sees both the original scene AND the edited Blender render in parallel, learning to preserve what matters while fixing the inevitable artifacts from transforming imperfect 2.5D/3D reconstructions. They train it with smart masking strategies so it learns when to ignore the original scene (for removals/replacements) and can manipulate objects independently of camera motion. What you get is pretty impressive control -- not just moving objects around, but changing materials, deforming shapes, swapping backgrounds, all while maintaining visual coherence. Neural Assets (one of my favorite papers last year) tried to crack this with learned object tokens, but it struggled with overlapping objects and loses fine details (due to low res DINO encodings). BlenderFusion just sidesteps the whole problem -- want to rotate something 173.5 degrees? Just rotate it in Blender. Want to duplicate an object 8 times? Copy paste away. The diffusion model's only job is making it look photorealistic, not figuring out the 3D underpinnings. The catch? Lacks temporal consistency for animation. Each viewpoint is generated independently, so while a single edit looks great, smoothly animating a car or camera down the street won't work -- you'd get flickering and inconsistencies between frames. That said, this approach is so much more intuitive for finer grain image editing than trying to describe your changes in text prompts. It's the kind of thing that makes you wonder why we're trying to do everything inside neural networks when perfectly good 3D tools already exist -- giving you the best of both worlds.

Bilawal Sidhu

34,440 Aufrufe • vor 1 Jahr

一番最後の[Prompt for original image]の部分に画像生成に使用したPromptを入れると一貫性が増します。不要な場合は3行削ってしまっても大丈夫です。 --- Extreme wide-angle perspective and dynamic pose remix edit. This is an EDIT of the original image, not a new character. Use the original image as a strict reference for: – the person’s identity, hairstyle, and overall fashion style, – the general type of background and location (same street, same room, same beach, same kind of architecture, etc.). You are allowed to completely change the camera position, angle, and pose, but you must keep the scene in the SAME location and keep the SAME person and outfit design. Camera and perspective: – Use an ultra wide-angle or fisheye feeling lens (around 12–18mm full-frame look). – The camera angle MUST change significantly from the original: use dramatic angles such as • worm’s-eye view from directly below looking up, • bird’s-eye view from directly above looking down, • very low angle from the ground, • high angle from above, • tilted Dutch angles. – Always create strong foreshortening: body parts close to the lens look huge, while the rest of the body falls away in perspective. – The final result must look like a bold fashion or street photo, fully photorealistic, not illustration or anime. Background consistency: – Keep the same location as the original image: same street, same bridge, same room, same studio, same beach, same general structures and materials. – Do NOT replace the background with a completely different place. – Because the camera angle changes, it is allowed and expected that different parts of the environment become visible. – When new areas appear, extend the original environment logically (same buildings, fences, road markings, walls, colors, materials, lighting style), as if the camera moved within the same place. Body parts near the lens (1–2 parts, sometimes 3): – In each edit, choose ONE or TWO main body parts to be extremely close to the lens (sometimes even THREE in more complex poses). – Vary them from image to image, do NOT always use the same body part. – Allowed near-the-lens parts include: • one or both hands / fingers reaching toward the camera, • one or both feet / shoes / boots near the lens, • knees or thighs, • face very close to the lens, • shoulders or chest close to the lens in a leaning pose. – The chosen body parts should come extremely close to the lens, almost touching it, with visible skin texture, fabric texture, and realistic wide-angle distortion. Pose and overall body (complex and varied): – Create strong, cool, dynamic poses that match the extreme perspective. – Randomly use different pose types, including: • standing with one leg or one arm reaching toward the camera, • crouching or squatting low to the ground, • sitting on the floor or on objects, • lying on the ground with legs or feet toward the lens, • leaning forward aggressively toward the camera, • twisting the body, crossing legs, or arching the back for more dynamic lines. – Allow complex poses where: • both hands are near the lens forming shapes (peace signs, triangles, frames, pointing toward the viewer), • both feet are toward the lens, • one hand and one foot are both large in the foreground, • the face is close to the lens while hands or feet are also visible in perspective. – Maintain believable anatomy even with extreme foreshortening. Angle and attitude (randomized): – Randomize camera angle and orientation (up, down, side, Dutch tilt) while keeping the composition visually balanced and powerful. – Keep the vibe cool, confident, and fashion/editorial or street style, depending on the original outfit. – Facial expressions can vary (serious, playful, confident, mysterious), but must still look like the same person. Lighting and rendering: – Keep the general time of day and lighting mood similar to the original (night vs day, indoor vs outdoor, soft vs hard light), but you may enhance contrast and color to make the image punchy and dramatic. – Maintain realistic shadows and contact points with the ground or floor. – High-resolution, sharp details with clear skin texture, fabric weave, and material highlights. Variation and randomness: – Each edit should look noticeably different from the original image and from other edits, with different: • camera angles, • pose types, • which body parts are closest to the lens, • orientation (straight, tilted, from above, from below). – Avoid repeating the exact same single-foot-close-up composition; produce a wide variety of dynamic poses and angles. Strict rules: – Do NOT change the person into someone else. – Do NOT change the outfit type; only restyle it through pose, perspective, and small natural movement of clothing. – Do NOT move the scene to a completely different location; always stay in a plausible extension of the original place. – Do NOT add text, logos, watermarks, or graphic design elements. – Do NOT switch to painting, illustration, or anime style; keep it photorealistic. Overall: Transform the original photo into a dramatic, photorealistic, ultra wide-angle shot with an extreme camera angle (including views from directly below or above), where one or more body parts are right next to the lens and look huge, the rest of the body recedes in perspective, and the same person strikes a stylish, complex, powerful pose in a consistent, expanded version of the original environment. Also, below is the prompt for generating the original image. Please use it as a reference. [Prompt for original image] #nanobanana2

AI Girl's Photo Studio

20,684 Aufrufe • vor 7 Monaten

🎉 new skill unlocked: 20s uninterrupted, unstitched, single render from our new ai video engine: Nami. This is my birb (#7531) from the Moonbirds collection, idling in the library. patent: "Intra-Latent Semantic Injection via Cross-Spatial Encoding and Decoding during Multi-Pass Inference for Generative AI Video Creation" At Scrypted we've been quietly working on an agentic generative AI stack for two years: • integrating and testing w/ partners across the games & entertainment sectors • stealthily building a community of early believers through AVB • showcasing some of what we're doing with amazing projects like H011yw00d Agent. -- about Nami -- Nami is an agentic orchestration layer for AI video models: it unlocks their inner superpowers without making them rely on custom LoRAs or fine-tunings. Instead of throwing raw training power and tens of millions of dollars at training yet another ai video model: we figured out new ways to use what we have. Nami harnesses a multi-agent system to perform the work needed in taking a simple prompt or image and turning it into something bigger - much bigger. The agentic steps are allowed to manipulate latent space, digging into tensors, yet doing so in semantically aware chunks - meaning that Nami inherently supports video generation of arbitrary length, though it's bound to O(n) rendering time. (We do have some cool sharding tech that allows us to cut the generative time in half for a reference pose idle-animation like this demo). It's also fairly agnostic, picking and choosing the right tools for the job, and plays really well with emerging tech like FLUX Kontext, FramePack, or <- without being limited by any of them. -- use cases -- Even just a year or two ago the 20 second render below would cost a company, paying an agency, around $10k start-to-finish. This one cost me $6.25 on our dev hardware in an unoptimized environment. There's something mind-blowing about the state-of-the-art when we reduce costs to 0.0625% - less than 1% - of what we used to pay. It's also empowering. For creators. Game developers. Content influencers: you name it. -- superpowers -- 1. it does the things you ask for, in the order you asked for it 2. consistency is king 3. single-shot text or image-to-video 4. future videos can reference previous ones to seamlessly maintain style 5. semantic stitching: can't wait to showcase this -- gtm -- We think Generative AI Video, like image generation, like text, like games, should be a publicly accessible common good. We believe democratizing access to Nami in web3, via x402 payments proposed by Drew Coffman, or in World's mini-apps, is a bold step forward for digital freedom. Permissionless, decentralized, generative ai video. Naturally, we'll also soon release a web platform for using Nami in a traditionally SaaSy way: bring your own images, videos, or prompts and we'll take care of the rest. In the mid-term, Scrypted is building a stack of agentic skills (we call it AVB) and making them available to projects like H011yw00d Agent on Virtuals Protocol and other platforms. -- long-term vision -- Scrypted's mission is to decentralize the things that can't be decentralized. We participated in a16z crypto's CSX (London 2024) during our pre-seed specifically to research a new consensus protocol for hard things like AI video and AI agents: where there's no "one right answer". When Zero-Knowledge Proofs (ZKP) can't secure it, and Trusted Execution Environments (TEEs) are too small, we've got you covered with our upcoming Inori Network. -- how you can help -- 1. Are you a GPU farm? We're gonna need more flops. 2. Do you represent an L1 or L2? We want to build bridges. 3. Do you represent a Wallet or App creator? Let's get an endpoint exposed. 4. Are you an investor? Let's chat. 5. Like, repost, share! -- team background -- We come from a background of AI in the Video Game industry with each founder having over 20 years of experience at companies like Electronic Arts & Square Enix. -- contact -- DMs are open, reach out if you want to be an early tester for your site, game, collection, or project! -- try it out -- Go anywhere on X and tag H011yw00d Agent with a prompt and she'll give you a free 2 second render. Have fun making cinematic shorts or meme videos! -- thanks -- AWS Startups has been an incredible help scaling our prototypes. Also, shout out to all loyal beans 🫘 in the Autonomous Virtuals Beings (AVB) community. Nami has a very important role in the upcoming XP agent platform, can't wait to show you all. AVbeings

Tim Cotten

12,617 Aufrufe • vor 1 Jahr