Today we’re open-sourcing Stable Audio Open Small, a 341M-parameter... text-to-audio model optimized to run entirely on Arm CPUs. This means 99% of smartphones can now generate music-production samples in seconds, right on-device with no internet required. Built for fast, on-the-go creation, it turns your next quick idea into up to 11 seconds of audio. Generate drum loops, foley, riffs, and textures right where you are. No cords 🔌 just chords 🎹 You can learn more here:show more

Stability AI
94,796 views • 1 year ago
HappyHorse 1.0 is now on OpenArt 🐎 It's the... #1 ranked text-to-video model on Artificial Analysis right now. Generate up to 15 seconds 1080p video with synchronised audio in a single pass and support for 7 languages.show more

OpenArt
2,310,185 views • 4 months ago
MiniMax H3 is now 50% OFF on Magnific for... 2K video, only until September 1. 🔥 I’ve been trying MiniMax H3 on Magnific, and it feels like a big upgrade for AI video creation. It’s not just about turning text into videos. You can use text, images, videos, and audio together in one prompt, giving you more control over the final video. Here’s what makes it stand out: - Multimodal: Use text, images, video, and audio in one prompt. - Multiple references: Add up to 9 images, 3 videos, and 3 audio files. - 2K video: Create videos up to 15 seconds long. - Built-in sound: Generate voice, music, and sound effects with the video. - Easy editing: Remove objects or transfer motion easily. - More control: Control the camera, characters, and voice. You can use it to turn posters into videos, moodboards into short films, and product images into ads. It also helps bring your ideas to life with realistic movement, lighting, reflections, and sound. The workflow is simple: give it your references → generate → edit → refine. Try MiniMax H3 on Magnific:show more

Markandey Sharma
96,964 views • 18 days ago
Today, we are adding Stable Video Diffusion, our foundation... model for generative video to the Stability AI Developer Platform API. The model can generate 2 seconds of video, comprising of 25 generated frames and 24 frames of FILM interpolation, within an average time of 41 seconds. Developers interested in utilizing Stable Video Diffusion through an API can access it now on the Stability AI Developer Platform. Learn more here:show more

Stability AI
175,837 views • 2 years ago
Meet Stable Audio 3.0, the open-weight model family built... for artistic experimentation. This is our open invitation to experiment with generative audio. We believe the best innovations are still waiting to be built. The 4-1-1 on 3.0: 📣 You own your outputs, and can distribute and commercialize them under the Stability AI Community License (up to $1 million in revenue). 🎵 New and improved capabilities include variable-length generation up to six minutes, and full song composition on portable devices, no GPU required. ✅ Trained on a fully licensed dataset. 🎨 You can customize the models on your own library with support for LoRa training, which we’ve documented for the first time. More on the models 👇show more

Stability AI
166,625 views • 3 months ago
MiniMax H3 is now on Magnific and honestly, there’s... a lot you can do with it. You can mix text, images, videos and audio in a single prompt up to 9 images, 3 videos and 3 audio references. Start creating now: It can generate up to 15s of 2K video with synced sound, including voice, music and effects. And you can go beyond generation too: edit clips, remove objects, transfer motion, and control the camera, character and voice. But the multi-reference workflow is probably my favorite. Give it your product, character, environment and motion references, and H3 pulls everything together. It feels like a much easier way to go from an idea to an actual finished video.show more

Kalsoom (ghotai )
47,888 views • 13 days ago
NVIDIA’s ARDY is a glimpse of where AI animation... is heading: real time, open source and you can play with it right now on Hugging Face Spaces Type what the character should do and generate a motion sequence in seconds. Give a sentence a body.show more

Hugging Apps
50,180 views • 1 month ago
Someone built an AI you power with a hand... crank it is called CrankGPT and its running without any battery, internet or data centre just you turning a handle like it is 1900 SqueezLabs built it using a Raspberry Pi 5 with 8GB of RAM an audio card and a 20 watt hand crank generator it takes about 30 seconds of cranking to boot into a working voice assistant and the onboard capacitor gives you roughly 20 seconds of runtime before you have to start cranking again they have used it to generate small images even write code 🤯 this proves you can run a real AI model on almost no power while everyone is building billion dollar data centres two guys put AI in a boxshow more

Sweep
11,610 views • 2 months ago
MiniMax H3 is a serious upgrade for AI video... creation. Instead of relying on just text prompts, you can combine images, videos, and audio references to control the character, motion, camera, and sound. What stands out: → Up to 9 image + 3 video + 3 audio references → Native synced voice, music & sound effects → 2K video generation up to 15 seconds → Instruction-based video editing → Motion transfer → Better character, camera & voice control The interesting part is that you can build a full scene from references instead of endlessly regenerating until the result feels right. 🔗 Now is a great time to try MiniMax H3 on Magnific, especially with the current 2K offer.show more

Tanvir Anjum
32,271 views • 17 days ago
🔴 SOME CHINESE DEVELOPERS JUST HUMILIATED THE ENTIRE PAID... AI VIDEO INDUSTRY WITH A FREE TOOL they released LongCat-Avatar, an open-source AI that turns a photo and audio file into a realistic talking video with synchronized lip movements. you can generate videos that run for minutes, completely free. no camera or studio needed. upload the image, add the audio, and let the model do the rest. it’s open source, FREE to use, and the repo is public. I’ll leave the repo in the comments.show more

MIKE
74,617 views • 13 days ago
no way this is not disturbing ad industry this... node based AI canvas can generate hundreds of for any product in seconds.. you just need to upload product photo and click Run on OpenCreator like, comment & repost -> dm canvas for free here's how it works:show more

el.cine
207,379 views • 9 months ago
HyENA is now live and it is here to... redefine Perpetuals trading from the ground up. Brought to you by Based, powered by , and built entirely on Hyperliquid HIP-3. HyENA introduces a new standard for on-chain trading: an internet trading engine with native yield. With HyENA, you can trade any asset on earth 24/7, all while your collateral continues to work for you in the background. No idle capital. No stale liquidity. Just a seamless, hyper-efficient trading experience. HyENA is not just an upgrade. We are presenting a new model for how global markets can operate. Hyperliquidshow more

Based
34,109 views • 8 months ago
This AI just turned me into a film director…... No editing skills. No timeline headaches. Just one prompt. This is Seedance 2.0 🎬 You can literally combine: → Text → Images → Videos → Audio And it understands everything. Even crazier? You can control it like this: Image → character Video → camera movement audio1 → music/voice It doesn’t just generate clips… It builds full cinematic scenes with: → Consistent characters → Smooth transitions → Realistic motion → Built-in lip sync Basically… From a single prompt → you get a multi-shot story. Not AI video. AI filmmaking. Go try it before everyone catches on 👇show more

Kshitij Mishra | AI & Tech
60,393 views • 4 months ago
Today, we're shipping MLX support for TADA, our open-source... text-to-speech model, which means the entire pipeline (LLM, flow-matching, and decoder) can now run locally on any Apple Silicon device. We're seeing a 45% reduction in memory usage and a 10x speed-up when using it quantized. With these improvements, you can use TADA on-device for OpenClaw or any personal chatbot. If you own a MacBook, Mac Mini, or Mac Studio, record a 10-second clip of any voice, type any text, and get high-quality, natural and expressive speech in real-time. Completely offline, completely free.show more

Hume AI
24,684 views • 5 months ago
Today, we are releasing Stable Video Diffusion, our first... foundation model for generative AI video based on the image model, Stable Diffusion. As part of this research preview, the code, weights, and research paper are now available. Additionally, today you can sign up for our waitlist to access a new upcoming web experience featuring a Text-To-Video interface. To access the model & sign up for our waitlist, visit our website here:show more

Stability AI
1,024,682 views • 2 years ago
Our open-source RAG app, Verba, can now run Ollama... models locally to embed documents into Weaviate and generate answers without your data leaving your device! I tested it with Llama3, and it works great + super quickly on my M2; In this recording, I'm doing a little Verba-ception, asking Verba what Verba is 🤤 You can also use open-source models like Mistral, Claude, or your own custom LLMs. This will make doing RAG with sensitive data possible since you won't depend on APIs anymore. Looking forward to everyone trying it out 🤗 Update's coming next week!show more

Edward
125,582 views • 2 years ago
I just open-sourced my /learn skill. Learn anything with... agents and HTML artifacts. I have been learning about all kinds of topics with it. Install the skill and interact with any agent to help you through any topic. Ask it to generate visual and interactive artifacts and help you go deeper or generate knowledge checks (e.g., quizzes). Upskilling myself on any topic is one of the most impactful ways I have been able to use AI agents. If you are a DAIR Academy pro member, you can use it with our AI Builder. Skill: Try now:show more

elvis
35,136 views • 2 months ago
You don't need a GPU for fast studio grade... voice cloning anymore. Qwen3 TTS (1.7B Q4_K_M) + mainline llama.cpp is officially the fastest way to generate zero shot voice clones using 100% pure CPU execution. Following up on my last post where we ran the Q8 model on a GPU, we just took local C++ voice synthesis a massive step further. The open source community quantized Alibaba's SOTA Qwen3 TTS model down to Q4_K_M GGUF, completely freeing local audio pipelines from dedicated graphics hardware. Here is the real world benchmark and hardware breakdown of running SOTA voice cloning on CPU: # Architecture & Model Setup Using Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf paired with the 8 bit multimodal projector (mmproj-Q8_0.gguf), llama.cpp executes the entire pipeline in pure C++. No PyTorch, no CUDA dependencies, and no VRAM bottlenecks. # Real-World Memory Footprint - Baseline RAM: 1.6 GB system idle. - Peak Generation RAM: 8 GB RAM during active voice synthesis. - Requirement: Any basic machine with at least 8 GB of system RAM can run this easily. # Real World CPU Benchmarks - Google Colab Free Tier (Throttled 2 Core CPU): Synthesizes a 5 sec studio quality audio clip (~8 words) in 45 seconds. - Modern Consumer CPU (Intel i5/i7 13th/14th Gen or AMD Ryzen 7000/9000): generation should drop to 5 to 20 seconds (nearly 1:1 real-time generation speed!). # Zero Shot Voice Cloning Quality Pass any 5 to 20 second .wav audio sample to the C++ engine using the --tts-speaker-file flag. It yields clean, natural sounding cloned speech with virtually zero quality loss compared to unquantized FP16 weights. To make testing seamless, I built an updated zero config Google Colab notebook. It pulls the official pre built llama.cpp CPU binaries (zero compilation time!) launches a live Gradio web app right in your browser. Record a 5 second clip from your mic (or drop a .mp3, .wav file), type text, and generate cloned audio on CPU. Native C++ audio models are making edge based, offline AI voice agents a reality. Links to the free Q4 CPU Colab notebook and the Q4_K_M GGUF HuggingFace repository are in the replies below! Which models have you been running on your CPUs? What CPU hardware are you using for local inference?show more

Alok
60,514 views • 25 days ago
Major major props to aroha for solving a major... limitation with Seedance 2.0. If you've tried uploading audio files directly, you know that Seedance 2.0 actually changes the lyrics and even the song itself BUT if you save out a blank video with the song and upload that video as a video omni reference it gives you perfect lipsync and no style drifting of the audio. I personally haven't notice style drift but it's 100% better than a straight audio upload which makes it unusable. This is actually MAJOR!!! Those music videos you've been wanting to generate with Seedance 2.0 are now solved. I tested this myself with some audio from Suno 5.5 which is the example you're watching right now and it works! I just wish Seedance 2.0 would actually fix the audio upload problem but this is an immediate fix. You don't need to type out the lyrics like I did. I think the best use case for using just an audio file is only if you want to keep a characters voice consistent.show more

Travis Davids
52,581 views • 4 months ago
here's a unique AI UGC format that you can... use to sell your products... you can use an image generator like GPT-image-2 to create a still image of a person in the bottom right corner & a green-screen behind them (make sure to prompt it to look organic/hand-held & not perfect quality) then send that to Seedance 2.5 with what you want him to say - you can get unlimited Seedance 2.5 for 33 days on Higgsfield right now, so you can generate these in bulk then in a video editor, you can do a chroma key on the greenscreen background to make it see-through & put a screenshot of whatever the topic of the video is behind the subject there are endless ways you can use this to sell stuff, let your creativity run wildshow more

EP
28,912 views • 25 days ago