Introducing Muse Image and Muse Video, the first media... generation models developed by Meta Superintelligence Labs. Muse Image is our most advanced image generation model yet. It follows instructions faithfully, edits with precision, composes from multiple references, and draws on Instagram for social context. It also brings agentic tool use capabilities to image generation and integrates with Muse Spark. You can try Muse Image in the Meta AI app and web, as well as in Instagram Stories and WhatsApp – starting in limited countries with more locations on the way. Today we’re also previewing Muse Video, which is built upon the same pretraining base as Muse Image to deliver exceptional visual fidelity with native audio support. Learn more about both models:show more

AI at Meta
848,991 views • 1 month ago
Meta just launched its 1st image model after Mark... Zuckerberg’s AI shake-up. Muse Image is Meta Superintelligence Labs' first image generator after Muse Spark. They said the model will power new editing features in Meta’s Instagram photo app and be added to its marketer tools for creating platform ads. Meta had relied on Midjourney and Black Forest Labs for generation inside Meta AI. Muse Image brings that layer in-house, so Meta controls quality, cost, and product timing. Consumers get free access through Meta AI, WhatsApp chats, and Instagram Stories. Power users need Meta One or another monthly plan when free limits run out. Users can start from prompts, add photos, annotate edits, or sketch changes directly. Meta says it can erase photobombers, create QR codes, and keep visual text readable. Advertisers get variants through Advantage+ creative, with edits, style swaps, and brand-matched versions. Meta says internal tests trail GPT Image 2 but beat Nano Banana 2 on editing. This is another move in Meta's effort in trying to convert AI infrastructure spending into revenue beyond social ads.show more

Rohan Paul
20,377 views • 1 month ago
Alongside the release of Muse Image, we’re sharing an... early preview of Muse Video. It offers competitive performance in prompt adherence, visual fidelity, and temporal consistency. We’re investing in areas with current performance gaps, such as audio-video synchronization and physically accurate fast motion.show more

AI at Meta
265,152 views • 1 month ago
1/ today, with the launch of muse image and... the preview of muse video, we wanted to share more samples from the model alongside research details and eval results see some more samples from muse video below (and see thread for research blog!)show more

Alexandr Wang
378,031 views • 1 month ago
It's safe to say Meta is back in the... game. We ran a blind "style-transfer" tournament with AI at Meta's new Muse Image vs >OpenAI's GPT Image 2 >Google Gemini's Nano Banana Pro >BlackForestLabsAI - Unofficial's FLUX.2. on 10 real world briefs, 55 tournaments, with every output ranked blind by professional working creatives. GPT won the most, but Muse placed top-two in 59%, more than any other model, already beating Nano Banana Pro 👀 (Image reference created with Muse Image btw)show more

ben
36,620 views • 1 month ago
Muse Image is now available on Meta Model API... and priced for production volumes at $0.01/image. It’s an agentic image model that reasons before it renders. Each call searches the web to refine the output through iterative passes and evaluates the prompt for precise elements like charts and QR codes. Text-to-image, single-image and multi-image editing, and multi-reference composition all live in one model, so there's no multi-step pipeline to stitch together. Start building →show more

Meta for Developers
134,067 views • 6 days ago
From product image to video with just one tool... - Dzine As you may have noticed, this is one of my favorite tools. It is also very underrated, as probably 50% of my tutorials include some workflow. I was testing the new image-to-video option today, and I love it. Step - by step guide in comments 🔽 I can do 95% of a workflow without switching between apps. Image generation, Image to image with style reference, background removal, background generation, and 2 frames image to video. The only other app I have been using for this video is CapCut so that I can stitch it together. Step by step 🔽show more

Teodora P L
28,523 views • 1 year ago
Today we’re introducing Meta AI Voice Conversations powered by... Muse Spark that let you talk naturally to Meta AI (interrupt, switch topics, or swap languages), and as you talk, Meta AI can generate images and pull up recommendations from Reels, maps, and more. We’re also bringing live AI to the app, so you can point your camera at the world and ask about what you’re seeing in real time.show more

Meta Newsroom
277,124 views • 3 months ago
FLUX.2 is live on Runware! D0 drop! 🔥 Built... by Black Forest Labs, FLUX.2 is the new SOTA model that brings better control over structure, text, and references in image generation. This is a huge step forward for image gen & image editing. And we’re here to offer you the best prices on D0. This model comes in three versions. Details below.show more

Runware
293,432 views • 9 months ago
GPT Image 2.0 is now available on Akool 🚀... This isn’t just an upgrade—it’s a major leap in AI image generation. With GPT Image 2.0, you get: • Sharper, high-resolution outputs with improved detail and realism • Much stronger text rendering (finally, text in images that actually looks right) • Better prompt understanding for more accurate and controllable results • Improved consistency across variations and multi-image generations • More precise editing & refinement for iterative creative workflows Whether you're creating marketing visuals, product images, or storytelling assets, GPT Image 2.0 delivers cleaner, more reliable, and production-ready results. Now live on Akool, try the latest in AI image generation today.show more

Akool Inc
1,541,129 views • 4 months ago
MiniMax H3 Instead of sharing the prompts for each... of these videos, I thought it would be more useful to share how I created that prompts. All of the videos were generated with text-to-video. First, find an image with the kind of scene, composition and mood you want to recreate. I used a few YouTube playlist thumbnails as references but Pinterest is also a great place to find inspiration. You can even use your own old or nostalgic photographs. Then upload the image to ChatGPT and ask it to describe the scene. The description it gives you can essentially become your text-to-video prompt. From there, you can generate completely new scenes with a similar composition, atmosphere and cinematic language. You can of course use the reference image directly with image-to-video or as a first frame. But if the original image isn't yours, I prefer using it only as visual inspiration and recreating the scene through text-to-video. This is the prompt I use with ChatGPT: "Describe the scene in this image in English, focusing primarily on what is happening, the characters, their actions and body language, the setting and the overall atmosphere. Also briefly describe the composition, framing, camera angle, approximate lens choice, lighting, color palette and cinematic aesthetic. Keep it concise and scene-focused rather than overly technical."show more

Kōda
51,449 views • 17 days ago
Player stats UI animation with GPT Image 2 and... MiniMax H3. This model is amazing for UI animations. Give GPT Image 2 your character image and ask it to create a stats UI. Then use that UI as a reference for MiniMax H3. You can check the prompt in the replies.show more

Kōda
41,330 views • 24 days ago
You direct the look. The motion. The sound. MiniMax... H3 is now available in Luma Agents. Generate up to 15 seconds of 2K video with native stereo sound, guided by text, image, video, and audio references. More creative range. One continuous workflow with Luma. Try it today →show more

Luma
19,213 views • 27 days ago
What makes Dola Seedream 5.0 Pro (hereafter referred to... as Seedream 5.0 Pro) different? It starts with control. Let’s start with one of the biggest shifts in AI image generation: Interactive Precise Editing. Instead of repeatedly rewriting prompts, edit exactly where changes should happen while preserving the rest of the image. From product edits and sketch-guided creation to color, material, and multi-image editing, Seedream 5.0 Pro gives creative teams greater precision throughout the design workflow.show more

BytePlus
3,974,972 views • 1 month ago
Yup, a football video. The World Cup made us... do it Luma rebuilt image generation from scratch — reasoning first, pixels second. And it beats Google's Nano Banana 2 and GPT Image 1.5 on reasoning benchmarks All 3 new models are now live on AI/ML API luma/uni-1 plans before it draws. The model generates autoregressively: it works out layout, composition and text placement first, then renders the pixels. $0.052/image luma/uni-1-max — same prompts, same params, max fidelity. 2K output + editing with up to 9 reference images. Built for hero shots and ad creative. $0.13/image luma/ray-3-2 — up to 16 keyframes per clip, 20s, 1080p, native HDR + 16-bit EXR export. The video in this post came straight out of it model ids "luma/uni-1" "luma/uni-1-max" "luma/ray-3-2" Luma cooked. We serveshow more

AI/ML API
19,375 views • 1 month ago
Uncensored image generation is live in OpenGradient. (∇, ∇)... Two new Seedream models, 5.0 Lite and 4.5, with no filter on legitimate creative work. We blurred the demo. The model didn't have to. And like everything here, it stays private. Your prompts and images are private.show more

OpenGradient (∇, ∇)
11,338 views • 2 months ago
🚨 JUST IN: THIS FREE TOOL JUST REPLACED FOUR... AI IMAGE AND VIDEO SUBSCRIPTIONS AT ONCE. Midjourney. Krea. Higgsfield. Openart. One repo. 200+ models. Zero dollars a month. Here is what it actually does. It is a full image and video studio that runs in your browser or as a desktop app. Text to image, image to image, text to video, image to video, lip sync, cinema mode with real camera controls. All of it. 4,500 people already starred this. What you get for free: → 50+ image models including Flux, Midjourney v7, Ideogram, GPT-4o, Seedream → 60+ video models including Kling, Sora, Veo, Runway, Wan, Hailuo → lip sync studio with 9 dedicated models. upload a portrait and audio and it talks → cinema studio with real camera controls. lens, focal length, aperture, film stock → feed up to 14 reference images into one generation → self-hosted. your data never leaves your machine The crazy part is there is also a hosted version that needs zero setup. Just open the link and start generating. Now the math. Midjourney Standard: $30/month Krea AI Pro: $30/month Higgsfield Plus: $49/month Openart AI: $15/month That is $124 a month. $1,488 a year. This repo does everything all four do. With more models than any of them. For free. Forever. No subscription. No vendor lock-in. MIT licensed. Download it in one click on Mac or Windows. Someone should have told me about this sooner. I feel like an idiot. ( save this )show more

Kanika
14,769 views • 4 months ago
As announced in partnership with NVIDIA at CES, we’re... excited to introduce Stable Point Aware 3D (SPAR3D), setting a new standard in 3D generation. Ideal for running on NVIDIA RTX AI PCs, SPAR3D enables real-time editing and complete structure generation of 3D objects from a single image in under a second. You can download the weights on Hugging Face and code on GitHub, or access the model through the Stability AI API. Learn more here: (1/3)show more

Stability AI
181,554 views • 1 year ago
Excited to launch a new way to upskill with... AI agents. This is how we are making it possible for anyone to learn to build with coding agents. To start, we are launching 4 new hands-on labs on the following topics: - Agent Skills - Agentic Image Generation - 30 Days of Hermes Agents - Prompt Engineering with Agents I am confident that with our new DAIR.AI platform, anyone can learn to become a top AI builder by building and acquiring highly-demanded AI skills. And there is a lot more landing in the coming weeks.show more

elvis
19,058 views • 2 months ago
Wonderland: Navigating 3D Scenes from a Single Image Contributions:... • First, we introduce a representation for controllable 3D generation by leveraging the generative priors from camera-guided video diffusion models. Unlike image models, video diffusion models are trained on extensive video datasets. This enables them to capture comprehensive spatial relationships within scenes across multiple views and embed a form of "3D awareness" in their latent space, which allows us to maintain 3D consistency in novel view synthesis. • Second, to achieve controllable novel view generation, we empower video models with precise control over specified camera motions. We introduce a novel dual-branch conditioning mechanism that effectively incorporates desired diverse camera trajectories into the video diffusion model. This enables expansion of a single image into a multi-view consistent capture of a 3D scene with precise pose control. • Third, to achieve efficient 3D reconstruction, we directly transform video latents into 3DGS. We propose a novel latent-based large reconstruction model (LaLRM) that lifts video latents to 3D in a feed-forward manner. With this design, during inference, our model directly predicts 3DGS from a single input image, effectively aligning the generation and reconstruction tasks—and bridging image space and 3D space—through the video latent space. Compared with reconstructing scenes from images, the video latent space offers a 256× spatial-temporal reduction while retaining essential and consistent 3D structural details. Such a high degree of compression is crucial, as it allows the LaLRM to handle a wider range of 3D scenes within the reconstruction framework, with the same memory constraints.show more

MrNeRF
52,849 views • 1 year ago