Qwen just released Qwen-Image-2.1, and this is exactly why... open models matter. I loaded it onto my NVIDIA DGX Spark and built Spark Image Lab, a local image generation and editing workspace designed for the DGX Spark ecosystem. QWEN-IMAGE-2.1 • Text-to-image generation • Image editing • Native transparent/RGBA workflows • Up to 10 reference images • Strong identity and product preservation • Improved typography, lighting, textures and detail SPARK IMAGE LAB • Clean Gradio interface • Width and height controls • Steps, seed and batch controls • Persistent generation history • Prompts and settings saved with every result • Reference images saved and restored • Docker setup for DGX Spark • Measured DGX Spark performance benchmarks Everything runs locally. Spark Image Lab is now open source under the MIT License. This is the first public alpha, so clone it, test it on your DGX Spark, open an issue and show me what you create. MODEL REPOshow more

Joey
23,210 просмотров • 4 дней назад
Introducing Muse Image and Muse Video, the first media... generation models developed by Meta Superintelligence Labs. Muse Image is our most advanced image generation model yet. It follows instructions faithfully, edits with precision, composes from multiple references, and draws on Instagram for social context. It also brings agentic tool use capabilities to image generation and integrates with Muse Spark. You can try Muse Image in the Meta AI app and web, as well as in Instagram Stories and WhatsApp – starting in limited countries with more locations on the way. Today we’re also previewing Muse Video, which is built upon the same pretraining base as Muse Image to deliver exceptional visual fidelity with native audio support. Learn more about both models:show more

AI at Meta
856,903 просмотров • 2 месяцев назад
GPT Image 2.0 is now available on Akool 🚀... This isn’t just an upgrade—it’s a major leap in AI image generation. With GPT Image 2.0, you get: • Sharper, high-resolution outputs with improved detail and realism • Much stronger text rendering (finally, text in images that actually looks right) • Better prompt understanding for more accurate and controllable results • Improved consistency across variations and multi-image generations • More precise editing & refinement for iterative creative workflows Whether you're creating marketing visuals, product images, or storytelling assets, GPT Image 2.0 delivers cleaner, more reliable, and production-ready results. Now live on Akool, try the latest in AI image generation today.show more

Akool Inc
1,541,129 просмотров • 5 месяцев назад
Poolside is hosting a 2-day model research hackathon in... London. Join us to push an open-weight agent model as far as you can. RL and fine-tune Laguna XS.2, our latest-generation model, on Prime Intellect Lab. Dates: May 29–30 Partners: NVIDIA + Prime Intellect + Hugging Face Prize: NVIDIA DGX Spark Agents need better models. Better models need cracked researchers. Link below.show more

Poolside
94,714 просмотров • 4 месяцев назад
From product image to video with just one tool... - Dzine As you may have noticed, this is one of my favorite tools. It is also very underrated, as probably 50% of my tutorials include some workflow. I was testing the new image-to-video option today, and I love it. Step - by step guide in comments 🔽 I can do 95% of a workflow without switching between apps. Image generation, Image to image with style reference, background removal, background generation, and 2 frames image to video. The only other app I have been using for this video is CapCut so that I can stitch it together. Step by step 🔽show more

Teodora P L
28,523 просмотров • 2 лет назад
Meta just launched its 1st image model after Mark... Zuckerberg’s AI shake-up. Muse Image is Meta Superintelligence Labs' first image generator after Muse Spark. They said the model will power new editing features in Meta’s Instagram photo app and be added to its marketer tools for creating platform ads. Meta had relied on Midjourney and Black Forest Labs for generation inside Meta AI. Muse Image brings that layer in-house, so Meta controls quality, cost, and product timing. Consumers get free access through Meta AI, WhatsApp chats, and Instagram Stories. Power users need Meta One or another monthly plan when free limits run out. Users can start from prompts, add photos, annotate edits, or sketch changes directly. Meta says it can erase photobombers, create QR codes, and keep visual text readable. Advertisers get variants through Advantage+ creative, with edits, style swaps, and brand-matched versions. Meta says internal tests trail GPT Image 2 but beat Nano Banana 2 on editing. This is another move in Meta's effort in trying to convert AI infrastructure spending into revenue beyond social ads.show more

Rohan Paul
20,377 просмотров • 2 месяцев назад
I tested MTPLX v2 with QWEN 3.6 27B and... compared it with oMLX without cache on M5 Max and DGX Spark on vllm using nvfp4 model version. More details in 🧵 I've reached 82.8 tps of max decoding speed! 🔥 Custom Metal Kernel design specifically for this model and for Apple Silicon is just perfect! This is the way forward! Great job Youssof Al Toukhi Look at the website here! 👇 Here a website with recap, built with GLM 5.2 running locally 💪 First chart and preview from the website.show more

Ivan Fioravanti ᯅ
15,885 просмотров • 2 месяцев назад
🚨 JUST IN: THIS FREE TOOL JUST REPLACED FOUR... AI IMAGE AND VIDEO SUBSCRIPTIONS AT ONCE. Midjourney. Krea. Higgsfield. Openart. One repo. 200+ models. Zero dollars a month. Here is what it actually does. It is a full image and video studio that runs in your browser or as a desktop app. Text to image, image to image, text to video, image to video, lip sync, cinema mode with real camera controls. All of it. 4,500 people already starred this. What you get for free: → 50+ image models including Flux, Midjourney v7, Ideogram, GPT-4o, Seedream → 60+ video models including Kling, Sora, Veo, Runway, Wan, Hailuo → lip sync studio with 9 dedicated models. upload a portrait and audio and it talks → cinema studio with real camera controls. lens, focal length, aperture, film stock → feed up to 14 reference images into one generation → self-hosted. your data never leaves your machine The crazy part is there is also a hosted version that needs zero setup. Just open the link and start generating. Now the math. Midjourney Standard: $30/month Krea AI Pro: $30/month Higgsfield Plus: $49/month Openart AI: $15/month That is $124 a month. $1,488 a year. This repo does everything all four do. With more models than any of them. For free. Forever. No subscription. No vendor lock-in. MIT licensed. Download it in one click on Mac or Windows. Someone should have told me about this sooner. I feel like an idiot. ( save this )show more

Kanika
14,769 просмотров • 5 месяцев назад
Yup, a football video. The World Cup made us... do it Luma rebuilt image generation from scratch — reasoning first, pixels second. And it beats Google's Nano Banana 2 and GPT Image 1.5 on reasoning benchmarks All 3 new models are now live on AI/ML API luma/uni-1 plans before it draws. The model generates autoregressively: it works out layout, composition and text placement first, then renders the pixels. $0.052/image luma/uni-1-max — same prompts, same params, max fidelity. 2K output + editing with up to 9 reference images. Built for hero shots and ad creative. $0.13/image luma/ray-3-2 — up to 16 keyframes per clip, 20s, 1080p, native HDR + 16-bit EXR export. The video in this post came straight out of it model ids "luma/uni-1" "luma/uni-1-max" "luma/ray-3-2" Luma cooked. We serveshow more

AI/ML API
19,375 просмотров • 2 месяцев назад
I built a plugin to generate 3D assets directly... from Codex, using the latest update that introduces image generation. It generates an image, calls the Tripo3D API, then starts an MCP server and a local viewer directly inside Codex. It uses the paid Tripo3D API (12 free generations, 600 credits included). Everything is open source on my GitHub. Link below. Feel free to contribute. Room for improvement: skeleton/animation support and multi-provider 3D support.show more

Defend Intelligence (Anis Ayari)
12,151 просмотров • 5 месяцев назад
Step 1 : input your char reference, poster and... your product into GPT Image 2 and then put the prompt : Use the woman on image 1 as the main subject. create a vertical poster ad inspired by the reference poster style. show the image 1 woman holding the kimchi jar with her left hand, while the right hand eating kimchi to her mouth. with framing wide lens, low angle from the reference poster. This step is where you lock the visual direction. Step 2 : take the generated image into Seedance 2.0 and convert it into motion. Set it to 1080p if your design includes a lot of typography, this helps preserve text clarity. You can also strengthen your prompt by adding keywords like: “dynamic motion design commercial advertisement” to push the result closer to a polished ad style. Important note : If your design contains heavy typography, expect some inconsistencies in text rendering. The best approach is to generate multiple variations and select the cleanest result. For this sample, I generated it 4 times before landing on the final version.show more

DStudioproject
11,413 просмотров • 5 месяцев назад
MiniMax H3 Instead of sharing the prompts for each... of these videos, I thought it would be more useful to share how I created that prompts. All of the videos were generated with text-to-video. First, find an image with the kind of scene, composition and mood you want to recreate. I used a few YouTube playlist thumbnails as references but Pinterest is also a great place to find inspiration. You can even use your own old or nostalgic photographs. Then upload the image to ChatGPT and ask it to describe the scene. The description it gives you can essentially become your text-to-video prompt. From there, you can generate completely new scenes with a similar composition, atmosphere and cinematic language. You can of course use the reference image directly with image-to-video or as a first frame. But if the original image isn't yours, I prefer using it only as visual inspiration and recreating the scene through text-to-video. This is the prompt I use with ChatGPT: "Describe the scene in this image in English, focusing primarily on what is happening, the characters, their actions and body language, the setting and the overall atmosphere. Also briefly describe the composition, framing, camera angle, approximate lens choice, lighting, color palette and cinematic aesthetic. Keep it concise and scene-focused rather than overly technical."show more

Kōda
55,398 просмотров • 1 месяц назад
SOMEONE GOT TIRED OF PAYING HIGGSFIELD AI'S SUBSCRIPTION SO... HE REBUILT THE WHOLE THING AND OPEN-SOURCED IT 200+ models. text-to-image, image-to-image, text-to-video, image-to-video all in one interface you configure a virtual camera in the Cinema Studio. pick the body, the lens, the focal length, the aperture and it writes the optimized cinematic prompt for you. completely in the background you never touch the camera keywords. you just set up the shot like a real cinematographer would Kling v3, Sora 2, Veo 3, Flux Dev, Midjourney v7, GPT-4o, Seedream 5.0, Runway Gen-3 all in there self-hosted. MIT licensed. runs on your machine. your data stays local the only thing you pay for is the model API calls themselves someone built this so you never have to pay Higgsfield AI againshow more

Rimsha Bhardwaj
101,547 просмотров • 5 месяцев назад
Kling AI 3.0 is the Nano Banana Pro moment... for video models. Highlight: Multi cut with up to 15s per run and enhanced lip sync. The performance of characters is the best I’ve seen so far! And you can literally use it like a reference model. This is the image I used:show more

Halim Alrasihi
74,989 просмотров • 7 месяцев назад
THE BLACK BOX OF AI VIDEO PRODUCTION IS FINALLY... WIDE OPEN Higgsfield just open-sourced their 'Originals' productions, making every single cinematic shot an open book 🤯 Instead of hiding their best workflows, they are giving creators the exact production blueprints. Here's what's unlocked: → raw prompts and exact generation settings → all image and audio reference files → workflows for 4K and native lip-sync You don't have to guess how to build cinematic, character-consistent AI video anymore. You can literally download the clips, copy the exact setups, and remix them instantly for your own projects 👊 Prompt and links in 🧵↓show more

Charly Wargnier ♨️
21,839 просмотров • 2 месяцев назад
Google Nano Banana 🍌 is crazy good at static... ads... But it only generates one image at a time. This n8n AI Agent helps you generate 1000s of winning ad variations in minutes, fully automated. → Built with the latest Nano Banana image model → Creates static ad images in bulk → Upload product reference image via n8n form → OpenAI Vision analyzes your product automatically → AI Agent generates custom image prompts (you choose how many) → Nano Banana creates static ad images on demand → Images auto-stored in Box. com for instant access You can request 50, 100, or even 1000 ad variations with one upload. Just specify the number in the form → AI does everything else. Built 100% in n8n. Zero manual work after setup. Want access to the template? → Like this post → Comment "ADS" And I'll send it right over.show more

Mike Futia
209,969 просмотров • 1 год назад
With Hunyuan3D World Model 1.0 now released and open-sourced,... we're excited to showcase the technical highlights behind this impressive innovation: ✅360° Panoramic Generation: Creates complete, immersive “world scenes”, far beyond localized views. ✅Explorable 3D Scene Generation: Generates diverse, spatially consistent 3D worlds from text/image for truly immersive exploration. ✅Interactive/Editable: Achieves separation of foreground objects, background terrain, ground, and sky, for seamless secondary editing. ✅Exportable Mesh: Generated scenes can be exported as 3D meshes for direct import into mainstream game engines and modeling software. ✅Industry-Leading SOTA Evaluation: Surpasses state-of-the-art open-source models in generation quality. As the industry's first open-source model for physical simulation and explorable world generation, Hunyuan3D World Model 1.0 aims to foster a collaborative community ecosystem with developers and enthusiasts. ✨ Try it now: 🤗 Hugging Face:show more

Tencent Hy
23,240 просмотров • 1 год назад
Midjourney sref + Sora 2 Pro is the sauce.... With one Midjourney style image, you can give a specific style for your entire project. I created two different 12-second clips and edited them together. Some details aren’t fully consistent, like the iPod or AirPods because the clips were made separately from a single image (Character in a specific style). It could be fixed in post-production, but that would take more time, and this was more of an experimental test. It would be great to add the actual product image with the current one to maintain product consistency. I feel like if there were a way to add 2–4 images into this workflow, it could open up a lot more possibilities and consistency. With an API, it could be possible. Or let’s see what Veo 3.1 has to offer.show more

Allar Haltsonen
10,141 просмотров • 11 месяцев назад
DeepSeek-V4-Flash-Vision, Qwen 3.8 Max, and $200 in AI credits... all FREEEEE right now😳 access: multiple platforms (see below) bonus: the gap between paid and open keeps shrinking you can now use DeepSeek's first 305B multimodal model, Qwen 3.8 Max with a massive 1M context window, and claim up to $200 in credits for Claude and computer-use agents at zero cost what you get: -deepSeek-V4-Flash-Vision → (305B multimodal, MIT open weights) -Qwen 3.8 Max → (1M context, image input, zero card needed) -KkToken → (signup + daily check-in for up to $100 Claude credits) -novita AI → ($100 Agent Sandbox credits for browser automation) every week the free tier grows, bookmark this list and claim your free access while it lastsshow more

Nahid
15,438 просмотров • 24 дней назад
Most AI video tools generate clips. The interesting ones... help you build a complete creative workflow. MiniMax H3 stood out because it can use text, images, audio, and video together with Omni Reference. That makes it easier to keep the style, branding, and creative direction consistent from the first prompt to the final video. What caught my attention: • Text + image + audio + video inputs • Native 2K video generation • Better control over edits and refinements • Built for ads, e-commerce, gaming, and branded content AI video is moving beyond single prompts. It is becoming a creative workspace.MiniMax Design (H3) Try here 👇 #hailuoai #hailuo #MiniMaxH3 #H3 #AIVideo #GenerativeAIshow more

Alina Davy
91,979 просмотров • 1 месяц назад
What did you do to my friend !! 😡... This is the trending “Fight Prompt” going viral Prompt : Use the first uploaded image as the main reference for the school uniform, body proportions, pose, posture, background, camera angle, framing, and overall composition. Use the second uploaded image as the identity reference for the face and hairstyle. Create a realistic Korean influencer-style school uniform portrait where the person from the second image naturally appears wearing the school uniform from the first image, photographed in the same studio setting. Important: Keep the school uniform, blazer, shirt, tie, skirt or pants, and overall outfit design from the first image. Keep the body proportions, standing pose, hand placement, posture, camera angle, framing, and studio background from the first image. Replace the face with the person from the second image. Also preserve the hairstyle from the second image, including bangs, hairline, hair part, hair length, hair framing around the face, and overall hair silhouette. Do not use the hairstyle from the first image if it differs from the second image. Identity: The face from the second image must remain clearly recognizable. Preserve the second person’s face shape, eyes, nose, lips, skin tone, jawline, and overall facial impression. Do not turn the face into a generic attractive face. Do not beautify too heavily. Preserve the person’s recognizable identity, but do not copy the face too rigidly. Reinterpret it naturally so it looks like a realistic photo of the same person in this new school-uniform scene. Keep the same overall facial impression and identity while allowing natural refinement and seamless adaptation to the lighting, angle, and mood of the target image. Hair: Follow the hairstyle from the second image. Preserve the second person’s bangs, hairline, hair part, hair texture, hair length, and overall hairstyle impression. Only adapt the hair naturally so it fits the pose, lighting, and composition of the first image. Korean influencer mood: clean modern Korean influencer portrait polished but natural beauty soft photogenic expression subtle editorial mood stylish, slightly chic, youthful, and confident atmosphere refined but believable skin texture clear eyes with soft catchlights naturally pretty, not over-retouched avoid stiff ID-photo mood Lighting: soft Korean beauty lighting gentle facial brightness clean skin tone soft natural highlights on the face natural shadow transition subtle glow, but realistic skin texture avoid harsh flash avoid flat passport-photo lighting avoid dramatic studio glamour lighting Style: realistic photography clean studio portrait quality Korean influencer-style school portrait mood natural skin texture high detail seamless face and hair integration polished but believable Negative prompt: no identity loss no generic attractive face no over-beautified face no first-image hairstyle if different no awkward face blending no mismatched skin tone no mismatched hairline no distorted facial features no blurry eyes no deformed hands no extra fingers no change to the school uniform no change to the body pose no change to the background no cartoon style no anime style no text no watermarkshow more

Ai Arainz
105,275 просмотров • 4 месяцев назад