Playing with Nvidia Cosmos3 Super models for image and... video generation. Here's obligatory Will Smith eating spaghetti. First few renders were pretty boring, so I went all out on an "energetic stuffing face with spaghetti prompt" here. That's some bottomless spaghetti.show more

Harrison Kinsley
27,729 views • 3 months ago
Excited to announce that my new single #YesMan, which... I've been playing in all my sets over the past few months, is out TODAY! I went back for a more old school rave vibe and am super happy with all the response on it so far. Go check it out now!show more

Ferry Corsten
11,015 views • 2 years ago
🚨 one person can now do the work of... an entire creative team. i just tested it on a real one. a friend needed an ad for his brand, so I opened the new Runway Agent 2.0 to try it out. here's how it went: → it generated the music and the key image first, so I could approve the direction → once I gave the ok, it built the full video around it → and when something was off, i changed just that one piece, without redoing the rest one prompt, and I had the ad we needed, work that used to take weeks. this is what it made 👇 if you want to try it → · 30% off 3 months with code RUNWAYAGENT — made with Runway · #MadeWithRunway · #adshow more

brenz.
28,698 views • 2 months ago
As always everyone is blind staring at the progress... of LLMs for coding and chat But meanwhile the new SOTA video model Seedance 2.5 has been slowly rolling out and it's really quite exceptional It's made by ByteDance (TikTok) who of course have lots of training data With just a few reference pics, it can get quite close to how you look IRL and you can do quite professional video shots with just a prompt I'd say it's the first video model that's now at the level of image models with the level of character likeness, cracking that in image models also took about 3 years (2022-2025) Generating 15 seconds takes about 4 minutes I put it live now on Photo AI, you can use it under [ Make video ] from the sidebar with just a prompt and your model selected! So you don't need to take an AI photo first and then turn that into a video! Saves lots of time :D It's more expensive than but I kept the credits the same (30 for 1 video) It also works inside the new video editor and you can make changes in your video with [ Magic edit ] in both the main app and the video editor Also a message for my server guy Daniel Lockyer (it can do voice too and you can even submit a voice sample of yourself, but I didn't here)show more

@levelsio
710,766 views • 1 month ago
I topped up $5 on an API aggregator ToAPIs... Then I found out GPT Image 2 costs only around $0.015 per image. If you do a lot of testing or batch-generate commercial AI images, that difference adds up fast. I think I just found the secret to generating more, testing more, and spending less. And it’s not just one model. With the same key, you can access 50+ models for image, video, and text, including GPT Image 2, Gemini Omni, Seedance 2.0, Kling AI 3.0, grok-video-1.5-preview, and more. Some models are priced up to 80% lower than official platforms. Just top up and test what you need: Made on ToAPIs with GPT Image 2 + Seedance 2.0show more

Shami
22,991 views • 3 months ago
DimensionX: Create Any 3D and 4D Scenes from a... Single Image with Controllable Video Diffusion TL;DR: Create 3/4DGS from Video Diffusion Note: Some first inference code released (not all yet). Contributions (cited): • We present DimensionX, a novel framework for generating photorealistic 3D and 4D scenes from only a single image using controllable video diffusion. • We propose ST-Director, which decouples the spatial and temporal priors in video diffusion models by learning (spatial and temporal) dimension-aware modules with our curated datasets. We further enhance the hybriddimension control with a training-free composition approach according to the essence of video diffusion denoising process. • To bridge the gap between video diffusion and real-world scenes, we design a trajectory-aware mechanism for 3D generation and an identity-preserving denoising approach for 4D generation, enabling more realistic and controllable scene synthesis. • Extensive experiments manifest that our DimensionX delivers superior performance in video, 3D, and 4D generation compared with baseline methods.show more

MrNeRF
17,062 views • 1 year ago
MiniMax H3 Instead of sharing the prompts for each... of these videos, I thought it would be more useful to share how I created that prompts. All of the videos were generated with text-to-video. First, find an image with the kind of scene, composition and mood you want to recreate. I used a few YouTube playlist thumbnails as references but Pinterest is also a great place to find inspiration. You can even use your own old or nostalgic photographs. Then upload the image to ChatGPT and ask it to describe the scene. The description it gives you can essentially become your text-to-video prompt. From there, you can generate completely new scenes with a similar composition, atmosphere and cinematic language. You can of course use the reference image directly with image-to-video or as a first frame. But if the original image isn't yours, I prefer using it only as visual inspiration and recreating the scene through text-to-video. This is the prompt I use with ChatGPT: "Describe the scene in this image in English, focusing primarily on what is happening, the characters, their actions and body language, the setting and the overall atmosphere. Also briefly describe the composition, framing, camera angle, approximate lens choice, lighting, color palette and cinematic aesthetic. Keep it concise and scene-focused rather than overly technical."show more

Kōda
54,352 views • 25 days ago
So sans I said i would give you all... an update on what was on yesterdays deportation flight out of Dublin to South Africa. First one doesn't look san face blurred out amd giving some sob story, and second one is a video taken by one of the deported heading through Dublin Airport for the free flight. Note in the video Muslim traditional dress so pakistani or Bangladeshis where on the flight.show more

Dave Jones ( Mdeva) 😄
128,718 views • 2 months ago
Yup, a football video. The World Cup made us... do it Luma rebuilt image generation from scratch — reasoning first, pixels second. And it beats Google's Nano Banana 2 and GPT Image 1.5 on reasoning benchmarks All 3 new models are now live on AI/ML API luma/uni-1 plans before it draws. The model generates autoregressively: it works out layout, composition and text placement first, then renders the pixels. $0.052/image luma/uni-1-max — same prompts, same params, max fidelity. 2K output + editing with up to 9 reference images. Built for hero shots and ad creative. $0.13/image luma/ray-3-2 — up to 16 keyframes per clip, 20s, 1080p, native HDR + 16-bit EXR export. The video in this post came straight out of it model ids "luma/uni-1" "luma/uni-1-max" "luma/ray-3-2" Luma cooked. We serveshow more

AI/ML API
19,375 views • 1 month ago
✨ Grok's new Imagine video model also comes with... an Edit model We know edit models for images, you submit an image, write a prompt what to change, but this is the first time I've seen a proper edit model for video And it kinda works, not great yet though but it does something Here I had to remove the old name "Nomad List" in the video for my site First it said "Go nomad -> Nomad List", so I prompted it "remove the text Nomad List. do not change anything else", it didn't remove it but it replaced it with just "Go nomad" again, okay good enough Useful because otherwise I'd have to scour my backups for the original video in Final Cut Pro, and this is faster One thing you see is it changes the pattern on the door also, but that's okay for now if I fade it inshow more

@levelsio
73,183 views • 7 months ago
I’ve used all the recent GenAI video models extensively... & here’s my 2¢: 🎬 Runway Gen3 Alpha - best image quality & motion for text-to-video & embedded words. Great at prompt travel changes over the course of 10 sec. And I’m super bullish on how gen3 will evolve, hopefully adopting the features listed below. Kling - best quality for image-to-video with prompt control, like eating food. Great clip extension that accounts for character (ie walking stride) & camera movement (speed & angle), rather than just using final frame. But it’s limited availability & Chinese native language is limiting. Used for Spider-Man video below (via Midjourney). LumaLabs - best for keyframe start & end control (it can not be overstated how important this is. other services should add it ASAP!) and their high dynamic action movements are really fun. Luma was used in my viral Multiverse of Memes video. PikaLabs - they haven’t gotten as much attention as others lately. But they did update their video model a few weeks ago and it looks great. Also, they are notable for their unique & AWESOME features, like video in-painting & out-painting. My perfect AI video platform would have the following features: 1) Gen3’s quality, prompt control & text embedding. 2) KLing’s image-to-video quality, prompt control & clip extension quality. 3) Luma’s multi-keyframe control & dynamic movement ability. 4) Pika’s inpainting & outpainting ability. And a video-to-video (aka next-gen Runway gen1) could be a game changer, too. It’s an exciting time to be alive 🫶 Who will get there first? 🔉🔉show more

Blaine Brown
26,535 views • 2 years ago
“The Last Meatball” Animation: jboogx.creative & enigmatic_e Movement: jboogx.creative... & enigmatic_e Reference Imagery: FLUX on Civitai Jboogx and I have known each other for the better part of the last 12 months. I think we've both been inspired by each other's work and dedication to exploring all the ways we can push these tools to their limits for VFX and animation. We didn't want to force a collaboration, but recently, Jboogx was inspired by the work I've been doing and how I incorporate myself into my animations. This idea is a playful take on 'Atlas carrying the boulder,' but with a twist—let's make Atlas out of spaghetti and the boulder a meatball, to fit more with what Jboogx has been doing on the food side of things lately (shoutout to our good friend James Gerde -@gerdegotit @gerdedoesit @gerdemadeit, who did the first spaghetti animation). Since I was going to record myself, it wouldn't have been right if Jboogx didn't as well! This is just a taste of what’s possible. I'm located in Germany, and Jboogx is in Hawaii. A world apart, and we pulled this together in 48 hours. BTS coming soon 😉show more

enigmatic_e
11,422 views • 2 years ago
This type of video animation for a gut supplement... ad would normally cost $500 from your ad agency. I built it in 20 minutes for $10 with this AI system. Here's exactly how I did it: → Ad scripting in Claude → Image generation of digestion track with Fabric → Image to video animation with Fabric → Talking head UGC avatar with Fabric → Edit all clips together & ad captions with CapCut With this tool + system, I can generate 12 variations of this ad concept in an afternoon and be testing all of them by tomorrow morning. By the end of the week, I can see which is driving highest ROAS, and then iterate on the winners again. In a month, I will have tested hundreds of high quality video concepts across dozens of personas. This is how you scale eComm in 2025 and beyond. Those who don't learn how to do this will fall behind and be outcompeted by competitors who do who can outbid you on facebook. If you want access to all the prompts and the tool I used to build this ad: Like, RT, and comment "FABRIC", and I'll dm it to you. (must be following so I can dm)show more

David Roberts
65,056 views • 9 months ago
most of you don't know how hard hermes agent... is optimized for local AI at the system level. watch the full setup flow on screen. you paste an openai-compatible v1 endpoint, hermes auto-detects every model running behind it. doesn't matter if it's llama.cpp or vllm or any compatible server, all your models surface and become selectable in seconds. no config gymnastics, no manual model list. then it goes deeper. hermes ships with per model parsers, prompt template auto-handling, tool call format detection per model architecture, thinking mode awareness, all the small friction points other harnesses leak on. these were not built for cloud apis with one canonical model. they were built for builders running 10 different local models across 10 different stacks. cloud first harnesses bolt local support on top. hermes agent is local first from the architecture out. that's the system level gap. if you're getting started on local AI, this is the harness you start with. try for yourself and find out. anyone serious about local AI lands here eventually.show more

Sudo su
14,407 views • 3 months ago
We made full detail, life-sized statues of some of... our hero characters for the marketing of #Transformers ROTB. To create these, we sent 16K turntable renders of the chosen characters in a studio environment to the manufacturer along with the 3D models so that they could replicate them as closely as possible. I worked on the Optimus Primal 3D delivery and can confidently say that those last few days of preparing the character for renders and send-off were some of the most stressful of the entire production. I remember working until 4am on the Friday delivery day to get the character sent down the pipe, ultimately having to skip the MPC Summer Party which was held on the same day. Shout out to everyone involved in the creation of these statues, they turned out amazing!show more

Rassoul Edji
31,167 views • 1 year ago
🚨 JUST IN: THIS FREE TOOL JUST REPLACED FOUR... AI IMAGE AND VIDEO SUBSCRIPTIONS AT ONCE. Midjourney. Krea. Higgsfield. Openart. One repo. 200+ models. Zero dollars a month. Here is what it actually does. It is a full image and video studio that runs in your browser or as a desktop app. Text to image, image to image, text to video, image to video, lip sync, cinema mode with real camera controls. All of it. 4,500 people already starred this. What you get for free: → 50+ image models including Flux, Midjourney v7, Ideogram, GPT-4o, Seedream → 60+ video models including Kling, Sora, Veo, Runway, Wan, Hailuo → lip sync studio with 9 dedicated models. upload a portrait and audio and it talks → cinema studio with real camera controls. lens, focal length, aperture, film stock → feed up to 14 reference images into one generation → self-hosted. your data never leaves your machine The crazy part is there is also a hosted version that needs zero setup. Just open the link and start generating. Now the math. Midjourney Standard: $30/month Krea AI Pro: $30/month Higgsfield Plus: $49/month Openart AI: $15/month That is $124 a month. $1,488 a year. This repo does everything all four do. With more models than any of them. For free. Forever. No subscription. No vendor lock-in. MIT licensed. Download it in one click on Mac or Windows. Someone should have told me about this sooner. I feel like an idiot. ( save this )show more

Kanika
14,769 views • 4 months ago
this effect is all over tiktok right now and... nobody's explaining how to actually do it properly... the 3d balloon character thing. where someone turns into a shiny inflatable version of themselves that still moves and talks. looks pretty smooth in feeds. the workflow is stupid simple once you see it. step 1: take any photo. drop it into an image gen tool (nano banana pro). prompt it with something like "make the person in the photo a plastic blow up balloon character with a shiny surface. keep the face details as 3d balloon details including the person in the background. don't change background" that's it for the image. don't overcomplicate the prompt. shorter = more consistent results. (learned this after wasting like 2 hours trying to get "perfect" prompts that kept giving me garbage) step 2: take that balloon image + your original video and drop both into kling motion control. prompt: "turn the motion and detailed mouth movement of the video to the setting of the image" that's literally it. kling maps the motion from the real video onto the balloon character. mouth moves. head turns. expressions transfer. the whole thing renders in a few minutes. the result looks like a $500 custom animation and costs you maybe $0.30 in kling credits. people are getting 500k+ views with these because the scroll-stop factor is insane. nobody expects to see a shiny inflatable version of someone giving a real speech or doing a product review. the play here is obvious btw. run this for client content (mix with the hook and real body, check the results yourself) or use it on your own faceless channels as a hook pattern before the algo catches up...show more

KNOX
25,773 views • 7 months ago
I wasn't expecting this to blow up 😭 We're... soo excited to be finally working on this after delaying it for over a yr — thank you so much for the love and interest! 💜🥹 Here's a little video of all the mascots we have so far ✨ and some info that might answer some of your questions: ⟡ For Twitch, Youtube, Ko-fi & Throne ⟡ All colors and fonts are fully customizable ⟡ Mascots will have personality features and react to other things, not just events ⟡ Not all pet suggestions fit since they need to work well in this very minimal frontal flat style with eyes only, but we're doing our best to come up with a good variety! ⟡ All shapes, expressions and animations are 100% code and adjusted to fit each body -- there's no video/image files involved so this widget won't support custom files.show more

Nani ⋆ ˙⟡ 💭
63,930 views • 4 months ago
On AEW Grand Slam Australia: Yes, the merch stand... was poorly planned out like last year. Yes, there were empty seats and some fans apparently were moved. Yes, a perhaps unwise choice was made with the booking of the first AEW match on the card IMO. But I'll say this, and you can't feel this through a phone screen, but it was such a fun atmosphere and a fun show for the most part. And my cousin (15), whose first ever live show this was, keeps talking about how cool certain spots and moments were (particularly the ladder match and crowd fighting). They made a fan out of him last night, and I just think that's pretty cool. Sure, it won't move the needle, but we were all that kid falling in love with pro-wrestling once upon a time. It's nice to see the tradition continue. So yeah, it was well done show on an emotional level, and on a wrestling level, if not entirely on a logistical level.show more

Alyssa
82,816 views • 6 months ago
What does all this mean for culture? If you... are watching my AI Influencer list here on X you'll see a ton of people posting various videos generated by AI, like this one that Justine posted last week, done with Kling AI Tonight I had some fun. My wife shot a video of me crawling on the floor being a goofball. Kling turned me into an evil 1X robot in a few minutes. Ahh, the kids and I are gonna have some fun with this one! Done with kling 2.6 motion control.show more

Robert Scoble
21,331 views • 7 months ago