Can a VLM see without a vision encoder? We... trained one for $100, inspired by Gemma 4 12B. Latency on an M3 Pro MacBook: 112 ms -> 1.1 ms for the image path 30% lower end-to-end image+LLM The architecture is just: patchify the image -> linear projection with pos embeddings -> LLM Writeup:show more

Andi Marafioti
60,236 Aufrufe • vor 1 Monat
i just ran Google's brand new Unsloth Gemma4 12B... dense GGUF on my RTX 4060 using llama.cpp + CUDA 13.2 21 tokens per second. on a budget consumer GPU. locally. no API. no cloud. no subscription. and the benchmarks are absolutely cooked # first let's talk architecture because this is genuinely different every multimodal model you've used has a frozen vision encoder + frozen audio encoder + LLM backbone glued together Gemma 4 12B is different it's a single decoder only transformer. that's it. vision? raw 48×48 pixel patches → one matmul → projected directly into the LLM audio? raw 16kHz signal sliced into 40ms frames → linear projection → same LLM input space no encoder tax. no latency penalty. no fragmented memory to put the encoder savings in perspective: old Gemma 4 26B approach: - 550M param vision encoder (frozen) - 300M param audio encoder (frozen) - LLM backbone Gemma 4 12B: - 35M param vision embedder (a single matmul) - no audio encoder at all - LLM backbone handles EVERYTHING 550M → 35M for vision alone. that's a 15x reduction this is why the gemma-4-12b-it-Q4_K_M.gguf is just 6.6 GBs!!! and it has 256K native context context # Benchmarks: AIME 2026 (math olympiad): 77.5% GPQA Diamond (expert science): 78.8% LiveCodeBench v6 (real code): 72% Codeforces ELO: 1659 MMLU Pro: 77.2% MATH-Vision: 79.7% BigBench Extra Hard: 53% inference → llama.cpp, LM Studio, vLLM, SGLang llamacpp flags: -m "gemma-4-12b-it-Q4_K_M.gguf" -ngl 99 -c 8000 -v --port 8080 Available on huggingface now! Link belowshow more

Alok
279,768 Aufrufe • vor 1 Monat
Snap presents MoA Mixture-of-Attention for Subject-Context Disentanglement in Personalized... Image Generation We introduce a new architecture for personalization of text-to-image diffusion models, coined Mixture-of-Attention (MoA). Inspired by the Mixture-of-Expertsshow more

AK
47,488 Aufrufe • vor 2 Jahren
P-Image-Upscale is the fastest and cheapest image upscaler in... the world: supporting outputs up to 128 MP under 1 seconds. A couple of weeks ago we released p-image-upscale, and it now works better than ever. It’s the fastest image upscaler in the world, supporting outputs up to 128 MP while keeping pricing simple and predictable: - $0.005/image for 1–4 MP - $0.01/image for 4–8 MP - $0.02/image for 8–16 MP - $0.04/image for 16–32 MP - $0.06/image for 32–64 MP - $0.12/image for 64–128 MP That means you can go from low-res input to production-ready output with extreme speed, while preserving detail and keeping costs easy to understand. Available on: - Pruna AI | - each::labs | - inference.sh | - Replicate | - Runware | - Segmind | - WaveSpeedAI | - Wiro | If you’re building workflows where image quality, price, latency, and scale all matter, p-image-upscale is built for you.show more

Pruna AI
23,723 Aufrufe • vor 1 Monat
Testing the new Gemma 4 12B (QAT) vision and... OCR capabilities locally with LM Studio. # The setup: - GPU: NVIDIA RTX 4060 (8GB VRAM) - CPU: Intel i7 - Runner: LM Studio - Config: 32k context, 38 layers offloaded, Flash Attention enabled - Speed: ~14 tokens/sec decode throughput # The test: I gave it a screenshot of Google AI Studio. Prompt: "clone this. give me a single html file" # The result: A solid one shot replication. It successfully mapped out the layout, recognized the UI text, and structured the divs correctly, with only minor differences from the original. Results available at the end of the video. Quite capable for a 12B model running on budget consumer hardware. A gpu that costs only $300. # Why the architecture under the hood is notable: Unlike traditional models that rely on heavy, separate vision and audio encoders, Gemma 4 12B uses a unified, encoder free architecture. It bypasses separate multi stage encoders. Uses a 35M parameter vision embedder to project raw 48x48 pixel patches directly to the LLM hidden dimension. Local multimodal development is becoming highly accessible on standard hardware. If you've spun up Gemma 4 12B locally, what setup are you using and what kind of throughput are you seeing?show more

Alok
25,717 Aufrufe • vor 1 Monat
3. Modify image Use this prompt to modify the... image you have chosen: "Modify image [1] with seed [1470033597]: add a parrot on her shoulder" Dall-E 3 identifies the image and makes the changes for you! Tips and tricks: - You can generate as many variations as you like. - It's also possible to remove an element from the image using the same method. - Sometimes the image isn't 100% identical, but still looks very similar.show more

Paul Couvert
44,073 Aufrufe • vor 2 Jahren
Create a 3D model from a single image, set... of images or a text prompt in < 1 minute 😮💨 This new AI paper called CAT3D shows us that it’ll keep getting easier to produce 3D models from 2D images — whether it’s a sparser real world 3D scan (a few photos instead of hundreds) or your favorite 2D image generator like Midjourney (just an image). How does this magic work? “This architecture is similar to video diffusion models, but with camera pose embeddings for each image instead of time embeddings. The generated views are passed into a robust 3D reconstruction pipeline to create the 3D representation (Zip-NeRF or 3DGS)”show more

Bilawal Sidhu
92,792 Aufrufe • vor 2 Jahren
From product image to video with just one tool... - Dzine As you may have noticed, this is one of my favorite tools. It is also very underrated, as probably 50% of my tutorials include some workflow. I was testing the new image-to-video option today, and I love it. Step - by step guide in comments 🔽 I can do 95% of a workflow without switching between apps. Image generation, Image to image with style reference, background removal, background generation, and 2 frames image to video. The only other app I have been using for this video is CapCut so that I can stitch it together. Step by step 🔽show more

Teodora P L
28,523 Aufrufe • vor 1 Jahr
let’s create the most dank image library for tap.fun.... i quickly vibed together a frontend so everyone can submit images. i’ll drop some cash for the help: - $50 for the most unique meme - $50 for the craziest pic - top 20 will get added to the taplab contributor tg here’s what i’m looking for: images that can transform any image into a unique new one. think memes, crazy visuals, unique outfits, weird energy, funny shit. no text. single image only (not a grid). multiple characters in one pic is totally fine. how to submit: 1. go to 2. submit your image or meme with a name 3. download the image and post it as a comment here so we can see itshow more

Will Mexi
11,376 Aufrufe • vor 6 Monaten
Create a short film like this in just 1... minute with GPT Image 2.0 + Seedance 2.0. GPT Image 2.0 can naturally combine multiple photos into one single image, while Seedance 2.0 can use that image as a reference to automatically separate the scenes, generate a coherent video sequence, and add suitable background music. This workflow greatly improves the overall creative efficiency. When using this method, simply provide the merged image as a reference for Seedance 2.0 and briefly describe each scene with a simple prompt. This can significantly increase the success rate of the final video. All of the above was created on GPT Image Prompt: Seedance Prompt:show more

Midjourney Sref and prompt Library
40,572 Aufrufe • vor 2 Monaten
GPT Image 1.5 by OpenAI is live on invideo.... Free. Unlimited. 365 days. Image generation just levelled up. RT + Comment 'GPT' for a chance to win the highest generative plan + $1200 in credits on invideo.show more

Invideo
497,006 Aufrufe • vor 7 Monaten
Gen-3 Alpha Image to Video now supports using an... image as either the first or last frame of your video generation. This feature can be used on its own or combined with a text prompt for additional guidance. All examples below demonstrate using an image as the last frame. (1/5)show more

Runway
143,494 Aufrufe • vor 1 Jahr
With the launch of Nano Banana Pro, we're also... rolling out the ability for Gemini users to check whether an image was generated with or edited by Google AI using SynthID, our digital watermarking technology. Now you can upload any image into the Gemini app and ask "Is this AI-generated?" Gemini then scans the image for the imperceptible SynthID watermarks that are embedded into all Google AI-generated images - including those by Nano Banana Pro! Learn more about Google’s efforts to increase AI transparency:show more

Google Gemini
190,528 Aufrufe • vor 8 Monaten
Real-time image editing is insanely good for architecture. You... can take a sketch, render a photorealistic building, and then change the materials, weather, or environment by simply adjusting the prompt. Done with the new KREA AI model in seconds 🤯show more

Justine Moore
181,851 Aufrufe • vor 5 Monaten
Unpopular opinion: Most agent evals are theatre. You run... them once before the deployment. It'll take 800ms+ as another LLM would be judging your LLM. Most annoying part - no one tells where in the chain things went wrong. I wasted a lot of time in this loop. And then I came across Future AGI bringing 5 different tools under one umbrella, best part - the platform is completely open source. They open sourced their entire platform and the eval layer is noticeably different. It is multimodal - works on everything text, image, audio, pdf. Not an LLM-as-judge adding latency but an agent with memory and tools. The biggest win are learned classifiers trained on actual production failure patterns to run evals at low cost. It also runs across the full reasoning chain, not just the final response. Check out → Try it here →show more

Swapna Kumar Panda
49,557 Aufrufe • vor 2 Monaten
Everyone's sleeping on image-to-3D AI models. They can make... your app look incredibly unique, with just a little effort. Here's how. This is my calorie tracker, built in a week with nothing but prompting. Just Claude Code + a couple APIs. The visuals are all AI-generated. I'll be sharing the full workflow + all the crazy technical stuff Claude and I did to make this work, so nobody has to struggle through it like me. Deep dive coming soon! Till then, this is the high-level idea: 1. Get a clean image of the food (or whatever your asset is) - In my app, the user describes foods via text, or attaches images (or both) - If text, an LLM extracts the food description and formats it into a specific prompt I tuned for this design, and we generate an image using Z-Image Turbo through fal - If image, we do the same thing but with FLUX.2 [dev] to edit the user image into our reference design - Originally, both used Google Nano Banana, but switching to open models cut costs and latency a ton 2. Gaussian splatting (2D image → 3D model) - I tried various 2D-to-3D options on fal and ended up with TripoSplat as my preferred balance of speed, cost, latency; this turns an image into a 3D model that looks super high quality (link below) - The app displays the 2D image while our backend generates the 3D splat - We "groom" the splat to reduce size and load time by culling low-opacity/scale points 3. Render efficiently on device Originally, it looked great but ran at 10 FPS. Getting to 120 FPS was a crazy journey. TL;DR: - SwiftUI had to go; it forced us to render each asset in independent MTKViews, which wasn't workable - Instead, we composite every dish into one full-bleed CAMetalLayer using MetalSplatter (link below) - We had to make some optimizations within MetalSplatter's code too, to reduce the overhead of sorting points per render Then I added some finishing touches like the subtle rotation and parallax as they move around. I think it turned out pretty cool :) Overall, this took some effort, but we still got it done in less than a day. Hopefully your agent can follow in the footsteps of mine and do it much faster. Keep an eye out for the bigger writeup, which'll give your agent everything it needs. If you have any questions, drop em below!show more

Anshu
19,931 Aufrufe • vor 1 Monat
🎥 Comparing AI video models: Image to video •... Gen-3 • Kling AI 1.5 • Hailuo MiniMax • Luma Dream Machine I used a Midjourney image in each model 4 times with no text prompt. This type of image is difficult for the AI to separate the subject in the front from the people behind her - they tend to move the group as if they are a single unit. But the results were interesting, as you can see! I chose my favorite results, below.show more

Heather Cooper
41,354 Aufrufe • vor 1 Jahr
I topped up $5 on an API aggregator ToAPIs... Then I found out GPT Image 2 costs only around $0.015 per image. If you do a lot of testing or batch-generate commercial AI images, that difference adds up fast. I think I just found the secret to generating more, testing more, and spending less. And it’s not just one model. With the same key, you can access 50+ models for image, video, and text, including GPT Image 2, Gemini Omni, Seedance 2.0, Kling AI 3.0, grok-video-1.5-preview, and more. Some models are priced up to 80% lower than official platforms. Just top up and test what you need: Made on ToAPIs with GPT Image 2 + Seedance 2.0show more

Shami
22,991 Aufrufe • vor 1 Monat
It's safe to say Meta is back in the... game. We ran a blind "style-transfer" tournament with AI at Meta's new Muse Image vs >OpenAI's GPT Image 2 >Google Gemini's Nano Banana Pro >BlackForestLabsAI - Unofficial's FLUX.2. on 10 real world briefs, 55 tournaments, with every output ranked blind by professional working creatives. GPT won the most, but Muse placed top-two in 59%, more than any other model, already beating Nano Banana Pro 👀 (Image reference created with Muse Image btw)show more

ben
36,282 Aufrufe • vor 15 Tagen
Quick tip for those doing FPV shots on Seedance... 2 🔥 A lot of people draw a trajectory on an image to guide the camera movement… But if you really want to unlock the full power of the model without limiting it, create your trajectory directly with a strong, detailed prompt. The difference is massive. Are you team “draw trajectory on image” or team “detailed prompt” for your FPV shots? Tell me in the comments 👇show more

Pierrick Chevallier | IA
24,072 Aufrufe • vor 1 Monat