First test of LatentSync; an open source lip sync... tool 🔥 Link in thread 🧵 It only requires a source video and a spoken audio clip. 🔊 Google Veo2: video clips E2-F5-TTS: Text to Speech (o/s) LatentSync: Lip sync (o/s) MMAudio: Background SFX (o/s)show more

Blaine Brown
21,262 görüntüleme • 1 yıl önce
🔴 SOME CHINESE DEVELOPERS JUST HUMILIATED THE ENTIRE PAID... AI VIDEO INDUSTRY WITH A FREE TOOL they released LongCat-Avatar, an open-source AI that turns a photo and audio file into a realistic talking video with synchronized lip movements. you can generate videos that run for minutes, completely free. no camera or studio needed. upload the image, add the audio, and let the model do the rest. it’s open source, FREE to use, and the repo is public. I’ll leave the repo in the comments.show more

MIKE
73,957 görüntüleme • 12 gün önce
I recently saw a thread of an AI tool... (D-ID) that can turn a single image into a video when given text or an audio file✨ I decided to see if I could build something similar using only #python and #opensource models🐍💕 I managed to do it in a few hours!!! Here's the result👩🏾💻show more

Marlene Mhangami
211,557 görüntüleme • 3 yıl önce
Facial expressions are really well handled in Kling AI... 2.0’s text-to-video mode 🎭🎬 It brings a whole new level of emotion and realism to the scenes 👀🔥show more

Pierrick Chevallier | IA
35,032 görüntüleme • 1 yıl önce
LTX-2.3 is now live on OpenArt. 🎬 The most... capable open video model just got a major upgrade and you can use it right now. What's new in 2.3: → Sharper fine detail: Hair, textures, text, edges. All of it. → Tighter prompt adherence: Complex multi-subject prompts? Handle it. → Stronger image-to-video: less freezing, less Ken Burns drift, more actual motion. → Cleaner audio: fewer artifacts, tighter sync across text-to-video and audio workflows. → Native portrait: up to 1080×1920, trained on vertical data.show more

OpenArt
2,151,122 görüntüleme • 4 ay önce
This AI just turned me into a film director…... No editing skills. No timeline headaches. Just one prompt. This is Seedance 2.0 🎬 You can literally combine: → Text → Images → Videos → Audio And it understands everything. Even crazier? You can control it like this: Image → character Video → camera movement audio1 → music/voice It doesn’t just generate clips… It builds full cinematic scenes with: → Consistent characters → Smooth transitions → Realistic motion → Built-in lip sync Basically… From a single prompt → you get a multi-shot story. Not AI video. AI filmmaking. Go try it before everyone catches on 👇show more

Kshitij Mishra | AI & Tech
60,393 görüntüleme • 4 ay önce
🚨 JUST IN: THIS FREE TOOL JUST REPLACED FOUR... AI IMAGE AND VIDEO SUBSCRIPTIONS AT ONCE. Midjourney. Krea. Higgsfield. Openart. One repo. 200+ models. Zero dollars a month. Here is what it actually does. It is a full image and video studio that runs in your browser or as a desktop app. Text to image, image to image, text to video, image to video, lip sync, cinema mode with real camera controls. All of it. 4,500 people already starred this. What you get for free: → 50+ image models including Flux, Midjourney v7, Ideogram, GPT-4o, Seedream → 60+ video models including Kling, Sora, Veo, Runway, Wan, Hailuo → lip sync studio with 9 dedicated models. upload a portrait and audio and it talks → cinema studio with real camera controls. lens, focal length, aperture, film stock → feed up to 14 reference images into one generation → self-hosted. your data never leaves your machine The crazy part is there is also a hosted version that needs zero setup. Just open the link and start generating. Now the math. Midjourney Standard: $30/month Krea AI Pro: $30/month Higgsfield Plus: $49/month Openart AI: $15/month That is $124 a month. $1,488 a year. This repo does everything all four do. With more models than any of them. For free. Forever. No subscription. No vendor lock-in. MIT licensed. Download it in one click on Mac or Windows. Someone should have told me about this sooner. I feel like an idiot. ( save this )show more

Kanika
14,769 görüntüleme • 4 ay önce
I just tried out the newly launched Pika Audio... Models and it is genuinely impressive. Pika launched four audio models in one drop- SFX, speech, music and soundtrack. They are the least expensive in the market and faster than everything else out there. I tested Pika Soundtrack on a silent AI video I generated- a woman walking through a rainy city intersection at night. I gave it a prompt describing the sounds I wanted and it delivered- rain on wet asphalt, soft footsteps, city ambience and a full cinematic score that matched the mood of the video perfectly. I didn't have to source separate sound effects or manually sync anything to the visuals. Just a prompt and it handled the rest. If you create AI videos this removes so much stress from the process. It is worth trying.-show more

Oluwatimileyin✨🦋
218,884 görüntüleme • 19 gün önce
Before the week ends, let's acknowledge one of the... most INSANE week ever for open AI, with 25+ notable open-weight drops across every modality: 🧠 LLMs → NVIDIA Nemotron 3 Ultra: 550B hybrid Mamba-MoE, only 55B active, 1M context, MMLU 89.1. NVFP4 variant claims ~5x throughput on Blackwell. First openly-weighted 550B hybrid Mamba-Transformer, closing the gap with frontier closed models. → Google Gemma 4 12B: fully open dense any-to-any (text/image/audio/video), 256k context, encoder-free, 140+ languages, AIME 2026 at 77.5. Shipped with a 23-checkpoint QAT wave (mobile ONNX + MLX). Most deployable model of the week. → StepFun Step-3.7-Flash: 198B sparse MoE VLM, ~11B active, SWE-Bench PRO 56.3. Apache 2.0. → Liquid AI LFM2.5-8B-A1B: edge MoE, just 1.5B active, 128k ctx, MATH500 88.8, MLX-ready. Best on-device option this week. → JetBrains Mellum2-12B-A2.5B-Thinking: their first open MoE, near-Qwen3-14B coding at 2.5B active. Apache 2.0. 🎨 Image gen (the surprise of the week) → Ideogram 4: their FIRST-EVER open weights. 9.3B flow-matching DiT trained from scratch. #2 overall behind GPT Image 2, top open-weight model on Design Arena + LMArena. Strongest open checkpoint for text-rich images, full stop. It has taste. Still can't believe this is open weights. 🔊 Audio & Speech (a breakout week for open TTS, 4 labs shipped) → Boson Higgs Audio v3 4B: 102 languages, 21 emotions, singing/whispering/shouting, sub-second TTFA. → RedNote dots.tts: the only fully continuous (no codec) open TTS pipeline, Apache 2.0. → Google Magenta RealTime 2: real-time music gen, <200ms latency, text+audio+MIDI. multimodalart ported it to PyTorch within hours with live ZeroGPU demos. → NVIDIA Nemotron-3.5 ASR: 600M streaming, 17x more concurrent streams vs Parakeet RNNT 1.1B. 👁️ Vision & VLMs → PaddleOCR-VL-1.6: SOTA document parsing at 1B params, Apache 2.0. → Baidu NAVA: 6.3B joint audio-video gen, best-in-class A/V sync, Apache 2.0. 🎬 Video, 3D & World Models → NVIDIA Cosmos3-Super: 64B omnimodal world model coupling action trajectories with video+audio gen, for Physical AI. → JD JoyAI-Echo: up to 5-min multi-shot text-to-video on LTX-2.3. → ByteDance Bernini-R + VAST TripoSplat (single-image-to-3D Gaussian splats, MIT).show more

Victor M
541,566 görüntüleme • 2 ay önce
Building RAG is easy. Parsing real, unstructured data is... the hard part. Most tools fail when documents get complicated. RAGFlow by InfiniFlow makes the entire process visual and flawless 🔥 It is an (open-source!) engine built specifically to find the exact needle in a data haystack, even across literally unlimited tokens. The platform comes packed with: → "Quality in, quality out" parsing for highly complex formats → Multiple recall paired with fused re-ranking → A built-in Python and JavaScript code executor for agents → An orchestrable ingestion pipeline Here's why it stands out: 1️⃣ Structural Understanding Instead of just scraping text, it handles tables across pages, scanned copies, slides, and Excel sheets natively using deep document understanding. 2️⃣ Grounded Citations Every answer is verifiable. The UI highlights the exact chunks used, allowing you to trace any response directly back to the source material. 3️⃣ Enterprise Synchronization Keep your context constantly updated with native data sync from Google Drive, Notion, Discord, and Confluence. Stop letting bad document parsing ruin your RAG systems. Best part? It's 100% Free and open-source. Link to the repo in 🧵↓show more

Charly Wargnier
19,220 görüntüleme • 5 ay önce
I just turned one UGC ad into 5 different... languages in 20 minutes 🤯 This workflow is insane: Take any winning UGC ad... Clone the creator's voice... Translate it to any language... And lip sync so it looks like they actually speak it. Perfect for e-comm brands & agencies scaling ads into new markets without reshooting anything. Say you've got a winning ad in English. You want to run it in Germany, France, Spain, Japan—wherever. But that means hiring new creators, re-filming the same script, hoping the energy matches the original. Or running the English ad and hoping people don't care. (They do.) ElevenLabs handles all of it: → Upload your original video → Clone the creator's voice instantly → Translate to any target language → AI generates the dubbed audio in their voice → Lip sync warps their mouth to match the new language perfectly Same creator. Same energy. Completely different language. What this unlocks: - One winning UGC ad → every market you want to test - No reshooting for localization - No hiring native-speaking creators per country - Lip sync that actually looks real I recorded a 12-minute Loom walking through exactly how to do this step-by-step inside ElevenLabs. > Comment "LABS" > Like this post And I'll send it over (must be following so I can DM)show more

Mike Futia
11,134 görüntüleme • 8 ay önce
We are in an insane run of open-weight drops.... Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: 🧠 LLMs & Reasoning → DeepSeek-V4-Flash-0731 (my king 👑): 304B MoE refresh, Terminal-Bench 2.1 jumps 61.8→82.7 over the preview, DeepSWE 7.3→54.4. Closes in on Opus-4.8 on Agents' Last Exam (25.2 vs 25.7). MIT. → Muse-Glimmer-30B, from Meta (they are back!!): their first open agentic model. ~29.6B dense + perception encoder, 131k+ context, built to run fully local, no cloud. Apache 2.0. → Liquid AI LFM2.5-2.6B: 2.69B params, 131k context, 220 tok/s on an M5 Max in under 2.5GB RAM. Competitive with models 4x larger on agentic tasks. → inclusionAI Ling-3.0-flash: 124B total, only 5.1B active, ~12% the size of their old 1T flagship Ring-2.6, matches it on key benchmarks. MIT. → inclusionAI Ling-3.0-tiny: 7.9B total, 1.3B active, 86-90 tok/s on an M4 Pro MacBook at ~8GB peak memory. MIT. → NVIDIA Nemotron-3.5-Lightning-30B-A3B: hybrid Mamba-2+MoE+Attention, up to 1M context, runs on a single H100 or DGX Spark, SWE-bench Verified 52.8. → deepgrove maple-preview: 20B-A1B ternary-weight reasoner, 218 tok/s on a Mac mini M4, 5.3GB checkpoint. MIT. → BigBang-v1 (endless-frontier): fine-tuned from Qwen3.6-35B-A3B via a self-evolving generator/critic synthetic-data loop. Lands aggregate performance between DeepSeek V4 Flash (284B) and V4 Pro (1.6T), at 35B. Apache 2.0. 🎬 Video → MiniMax-H3: 33B dense omni model, native stereo audio, up to 2K/15s. 3.6k+ likes already. → Minimax-H3-Turbo (lightx2v): Apache-2.0 turbo distillation of H3 for fast inference. → Lightricks LTX-2.5: image-to-video update, custom Gemma-4-12B text encoder, a markedly stronger distilled model. 🔊 Voice → NVIDIA NemotronLabs VoiceChat-11B: full-duplex speech-to-speech, ~450ms turn-taking, #2 on open VoiceBench, and the first open full-duplex model with live tool-calling mid-conversation. 🛡️ Safety → Mistral Shieldstral-1.0-3B: 3B multimodal guardrail that takes your safety policy as plain text instead of fixed categories. Beats LlamaGuard-4-12B and ShieldGemma-9B on HarmBench (99.4) and ToxicChat (84.1) at a fraction of the size. Apache 2.0.show more

Victor M
54,264 görüntüleme • 21 gün önce
Introducing: OpenGranola 🔥 I built an open source meeting... copilot for macOS. It transcribes both sides of your call on-device, searches your own notes in real time, and hands you talking points right when the conversation needs them. No audio leaves your Mac. Point it at a folder of markdown files, pick any LLM through OpenRouter (Claude, GPT-4o, Gemini, Llama), and it just works. It's invisible to screen share too — nobody knows you have it. The whole thing is open source. Link belowshow more

yazin
293,178 görüntüleme • 5 ay önce
MiaoYan is a lightweight, native macOS Markdown editor for... writing and notes — one of my open-source macOS projects. It just shipped its biggest update to date, introducing Split View for side-by-side Markdown editing and live preview, with smooth bidirectional scroll sync. This release also brings a full visual refinement, layout cycling between sidebar, list, and focus modes, faster startup with native window animations, and broad performance improvements across editing, preview, and export. Built entirely in Swift, not Electron. MiaoYan is local-first, requires no account, collects no telemetry, and avoids AI features by design. I built it for clean Markdown, performance, and a calm writing environment.show more

Tw93
34,484 görüntüleme • 7 ay önce
Fuck yeah! MaskGCT - New open SoTA Text to... Speech model! 🔥 > Zero-shot voice cloning > Emotional TTS > Trained on 100K hours of data > Long form synthesis > Variable speed synthesis > Bilingual - Chinese & English > Available on Hugging Face Fully non-autoregressive architecture: > Stage 1: Predicts semantic tokens from text, using tokens extracted from a speech self-supervised learning (SSL) model > Stage 2: Predicts acoustic tokens conditioned on the semantic tokens. Synthesised: "Would you guys personally like to have a fake fireplace, an electric one, in your house? Or would you rather have a real fireplace? Let me know down below. Okay everybody, that's all for today's video and I hope you guys learned a bunch of furniture vocabulary!" TTS scene keeps getting lit! 🐐show more

Vaibhav (VB) Srivastav
139,105 görüntüleme • 1 yıl önce
llama.cpp isn't just for text LLMs anymore. Pure C++... zero shot voice cloning just officially landed in mainline. Text generation was only step one. If you’re building autonomous local AI agents, real time voice assistants, or edge workflows, instant low latency audio is the missing piece. Thanks to PR #26254, Alibaba’s state of the art Qwen3 TTS model family is now natively supported directly inside the llama.cpp repository under the multimodal (mtmd) framework. No Python bloat. No massive PyTorch CUDA overhead. Just raw, hyper optimized C++ running GGUF voice weights. Here is why this native update is a massive deal for the open source local AI stack: # Multimodal Architecture (.gguf + mmproj) Qwen3-TTS splits the workload between the base language model backbone and a multimodal projection adapter. llama.cpp handles this using the llama-tts binary, mapping the text model alongside its --mmproj projector to process audio tokens seamlessly. # Zero Shot Voice Cloning in Seconds You don't need fine tuning or massive dataset training. Feed the C++ engine a single 5 to 10 second .wav audio sample using the --tts-speaker-file flag, and it accurately clones the exact timbre, tone, and accent on the fly. # Real World T4 GPU Benchmark & Resource FootprintRunning the 1.7B Base model in 8-bit quantization (Q8_0): - VRAM Footprint: ~7 GB peak VRAM during active zero-shot cloning. - Audio Quality: Studio grade, natural-sounding voice output in seconds. • - Execution: Direct execution via native compiled binaries or sub process calls. # Coming Next to llama-server (PR #26603) Beyond CLI execution, a native POST /tts HTTP endpoint is currently being added to llama-server, which will soon allow you to trigger voice generation directly via standard REST API requests! # quick note on Colab compilation: Because this code was merged into mainline very recently, pre-built third-party binaries haven't fully caught up yet. Compiling llama-tts directly from source on Google Colab's free CPU instance can take about 1 hour (or ~1-2 minutes if targeting single GPU arch like -DCMAKE_CUDA_ARCHITECTURES=75). Be patient during the build step, or compile it locally on your own rig for instant execution! To test this out yourself, I built a zero config Google Colab notebook that compiles llama.cpp, downloads the Q8_0 GGUF files from HuggingFace, and spins up an interactive Gradio Studio UI so you can record/upload 3 second clips and clone voices in real time. Stop sleeping on native C++ audio. The era of bulky Python audio pipelines is officially over. Links to the free Google Colab notebook and the official ggml org GGUF HuggingFace model repository are in the replies below! available in q4 and q8 both variants, 1 GB and 1.85 GBs respectively (requires additional ~500MB mmproj gguf) Are you building local voice agents yet? What does your current audio stack look like? Drop your setups below!show more

Alok
47,881 görüntüleme • 28 gün önce
It's a wrap for Solana Consumer Day— A heartfelt... thanks to 260+ builders who gathered to define the future of consumer apps, incl. leaders from Nansen, Sky, Audius, Layer3, and Bonkbot s/o to Yash SEND for curating the event with us, and to chomp Jupiter @heliuslabs SOON - Solana Optimistic Network (Mainnet Arc) Sonic SVM tapestry Jambo @ReoffGenaud 1kx Ellipsis Labs Reown IRLs Cleverse Ventures for making it an unforgettable day! let this be the first in our many next community-led events 🔥show more

Superteam Vietnam
13,005 görüntüleme • 1 yıl önce
Is there a market for an intelligently designed TWIN... 60" 🌽 system? 60" strip till = 1/2 HP 1/2 the units of everything Could be more aggressive with depth of tillage in zone + CC biomass simultaneously Bigger tires Manure integration More intensive continuous cropping It has been proven from many growers that a simple modification of 52"8"52"8"52"8" rows not only fills the yield gap from 30's but could provide an uptick with the right hybrids able to capture more light and CO2 in more of its leaves and roots Banding nutrients deeper in the soil WITH big Covers that stay alive a lot longer. This adds resilience and keeps soil in place. Video is at Zack Smith plot this year where twin 60's outperformed 30's The plot was supposed to just compare 60's with 30's but we both knew static 60" rows will yield from 88-95% ... just like wide row cereals a slight modification to ensure root and plant spacing quickly narrows the gap while creating the bandwidth for more activities and synergies.show more

Jason Mauck
30,747 görüntüleme • 1 yıl önce