Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

New TTS banger: Chatterbox Turbo 🤯 Zero-shot model that matches any reference voice with native paralinguistic tags, optimized for low-latency voice agents. ⬇️ Demo available on Hugging Face

41,631 görüntüleme • 7 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

llama.cpp isn't just for text LLMs anymore. Pure C++ zero shot voice cloning just officially landed in mainline. Text generation was only step one. If you’re building autonomous local AI agents, real time voice assistants, or edge workflows, instant low latency audio is the missing piece. Thanks to PR #26254, Alibaba’s state of the art Qwen3 TTS model family is now natively supported directly inside the llama.cpp repository under the multimodal (mtmd) framework. No Python bloat. No massive PyTorch CUDA overhead. Just raw, hyper optimized C++ running GGUF voice weights. Here is why this native update is a massive deal for the open source local AI stack: # Multimodal Architecture (.gguf + mmproj) Qwen3-TTS splits the workload between the base language model backbone and a multimodal projection adapter. llama.cpp handles this using the llama-tts binary, mapping the text model alongside its --mmproj projector to process audio tokens seamlessly. # Zero Shot Voice Cloning in Seconds You don't need fine tuning or massive dataset training. Feed the C++ engine a single 5 to 10 second .wav audio sample using the --tts-speaker-file flag, and it accurately clones the exact timbre, tone, and accent on the fly. # Real World T4 GPU Benchmark & Resource FootprintRunning the 1.7B Base model in 8-bit quantization (Q8_0): - VRAM Footprint: ~7 GB peak VRAM during active zero-shot cloning. - Audio Quality: Studio grade, natural-sounding voice output in seconds. • - Execution: Direct execution via native compiled binaries or sub process calls. # Coming Next to llama-server (PR #26603) Beyond CLI execution, a native POST /tts HTTP endpoint is currently being added to llama-server, which will soon allow you to trigger voice generation directly via standard REST API requests! # quick note on Colab compilation: Because this code was merged into mainline very recently, pre-built third-party binaries haven't fully caught up yet. Compiling llama-tts directly from source on Google Colab's free CPU instance can take about 1 hour (or ~1-2 minutes if targeting single GPU arch like -DCMAKE_CUDA_ARCHITECTURES=75). Be patient during the build step, or compile it locally on your own rig for instant execution! To test this out yourself, I built a zero config Google Colab notebook that compiles llama.cpp, downloads the Q8_0 GGUF files from HuggingFace, and spins up an interactive Gradio Studio UI so you can record/upload 3 second clips and clone voices in real time. Stop sleeping on native C++ audio. The era of bulky Python audio pipelines is officially over. Links to the free Google Colab notebook and the official ggml org GGUF HuggingFace model repository are in the replies below! available in q4 and q8 both variants, 1 GB and 1.85 GBs respectively (requires additional ~500MB mmproj gguf) Are you building local voice agents yet? What does your current audio stack look like? Drop your setups below!

Alok

47,881 görüntüleme • 5 gün önce

You don't need a GPU for fast studio grade voice cloning anymore. Qwen3 TTS (1.7B Q4_K_M) + mainline llama.cpp is officially the fastest way to generate zero shot voice clones using 100% pure CPU execution. Following up on my last post where we ran the Q8 model on a GPU, we just took local C++ voice synthesis a massive step further. The open source community quantized Alibaba's SOTA Qwen3 TTS model down to Q4_K_M GGUF, completely freeing local audio pipelines from dedicated graphics hardware. Here is the real world benchmark and hardware breakdown of running SOTA voice cloning on CPU: # Architecture & Model Setup Using Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf paired with the 8 bit multimodal projector (mmproj-Q8_0.gguf), llama.cpp executes the entire pipeline in pure C++. No PyTorch, no CUDA dependencies, and no VRAM bottlenecks. # Real-World Memory Footprint - Baseline RAM: 1.6 GB system idle. - Peak Generation RAM: 8 GB RAM during active voice synthesis. - Requirement: Any basic machine with at least 8 GB of system RAM can run this easily. # Real World CPU Benchmarks - Google Colab Free Tier (Throttled 2 Core CPU): Synthesizes a 5 sec studio quality audio clip (~8 words) in 45 seconds. - Modern Consumer CPU (Intel i5/i7 13th/14th Gen or AMD Ryzen 7000/9000): generation should drop to 5 to 20 seconds (nearly 1:1 real-time generation speed!). # Zero Shot Voice Cloning Quality Pass any 5 to 20 second .wav audio sample to the C++ engine using the --tts-speaker-file flag. It yields clean, natural sounding cloned speech with virtually zero quality loss compared to unquantized FP16 weights. To make testing seamless, I built an updated zero config Google Colab notebook. It pulls the official pre built llama.cpp CPU binaries (zero compilation time!) launches a live Gradio web app right in your browser. Record a 5 second clip from your mic (or drop a .mp3, .wav file), type text, and generate cloned audio on CPU. Native C++ audio models are making edge based, offline AI voice agents a reality. Links to the free Q4 CPU Colab notebook and the Q4_K_M GGUF HuggingFace repository are in the replies below! Which models have you been running on your CPUs? What CPU hardware are you using for local inference?

Alok

57,362 görüntüleme • 2 gün önce

Seedance 2.0 Valence-Arousal + FACS Created an example scene to show how valence-arousal and FACS can be used together. Prompt: 15s, cinematic emotional confrontation. Two characters @[chracter sheet ref] stand face-to-face inside a small apartment kitchen late at night. The room is dimly lit by a single warm overhead light and soft city lights leaking through the window. The atmosphere feels emotionally exhausted, tense and painfully intimate, like an argument that has been building for years. Modern cinematic realism, subtle handheld camera movement, shallow depth of field, soft film grain, emotionally restrained acting, realistic silence between dialogue lines. Beat 1: The emotional state remains at high arousal and medium-low valence. Camera: slow handheld side shot circling both characters tight over-the-shoulder close-ups brief eye-level two-shot showing emotional distance FACS Character A: AU4 + AU7 + AU17 Dialogue A: /juː ˈnev.ɚ ˈriː.ə .li lʊkt æt miː/ /juː wɚ ɔːlˌweɪz ˈsʌmˌwɛɹ ɛls/ Voice: tight restrained voice, controlled anger, uneven breathing Character A tries to stay calm while suppressing years of resentment. Beat 2: The emotional state gradually shifts toward very low valence and medium-high arousal. Camera: slow push-in toward Character B extreme close-up on trembling eyes and mouth wide static shot showing silence after the argument peaks FACS Character B: AU1 + AU4 + AU15 + AU25 Dialogue B: /aɪ wəz ˈtɹaɪ.ɪŋ maɪ bɛst/ /aɪ dɪdnt noʊ haʊ tə fɪks ˈɛv.ɹiˌθɪŋ/ Voice: breaking voice, unstable breath support, emotionally collapsing delivery Beat 3: The emotional state remains at very low valence and medium-low arousal. Camera: locked wide shot with silence between them slow close-up on both characters avoiding eye contact subtle rack focus between faces FACS Character A: AU1 + AU15 + AU17 Dialogue A: /ˈmeɪ.bi wiː stɑpt ˈlɪs.ən.ɪŋ ə lɔŋ taɪm əˈgoʊ/ Voice: emotionally exhausted, quieter delivery, fading anger replaced by sadness No exaggerated screaming, no violence, no comedy, no text overlay, no watermark.

Kōda

23,696 görüntüleme • 3 ay önce

For new followers: - I'm a long-time investor and builder in this space. - Founding Contributor of Realms.World ☁️. - Co-founder of Dojo. - Builder with the kings at Cartridge. - Starknet (Privacy Arc) class of '21. - Founder and Game Director of ETERNUM HAS MOVED. - Founder of Daydreams.Systems (x402, 8004 agents) My prime purpose for the past three years has been to build onchain infrastructure to enable the next generation of onchain experiences. This is done Starknet (Privacy Arc) as it is the superior VM for building complex applications—this will become clear soon enough. I work up and down the entire stack, from low-level indexing and contracts to GUI design. Nothing is out of scope. I have been pushing on agents for two years, mostly using existing frameworks like , until I came across @ElizaOS_ai in October. As I focused on building agents for ETERNUM HAS MOVED, it became clear that agents playing games require infinite paths to achieve goals. Thus, it's not scalable to hardcode functions—agents need to have total fluidity to take any action or call anything the game requires in any order. And ironically onchain infra is perfect for agent playgrounds because of its open nature. This exploration led me to create Daydreams.Systems (x402, 8004 agents), which focuses on the hardest problems of agents: long time-horizon goals using Hierarchical task networks (HTN). Daydreams agents don't require custom code—they work entirely based on 'sleeves'—which are just markdown files that explain how the agent can interact with the service (API docs, game guides, etc.) My thesis is simple. By focusing on the hardest problem (games), the design of the library will naturally lean towards an optimal structure for any problem an agent could face. We are early in this path and iterating with speed. If you are an onchain app developer or game builder—DM me, I want to know the architecture of your game so we can build sleeves together.

loaf

43,320 görüntüleme • 1 yıl önce

This guy built an AI pipeline that generates hyperrealistic fashion models in 47 minutes and now dropshippers pay him $1,400 to clone the entire system. He got tired of watching e-com brands lose $8K per photoshoot when a single product angle changed so he built a 9-node workflow that generates 127 product videos from one Pinterest photo without hiring a single model. Here's the exact breakdown: → Claude writes a 34-parameter JSON brand DNA before any image is touched target psychographics, price anchor, vibe matrix, anti-inspiration blacklist → Pinterest becomes the model source library but you can't just download and animate → Kling 2.6 takes that static JPG and turns it into 5-second video but only after the prompt architecture is locked → Negative prompt node runs 41 exclusion terms: no plastic skin, no CGI glow, no symmetry artifacts, no doll face, no synthetic lighting → That one step kills the "AI look" that tanks engagement by 67% in the first 3 seconds → TikTok Studio uploads 19 videos in one batch with zero manual captioning because the brand voice was pre-programmed in step one → Atlas scrapes Amazon product links and auto-generates a Shopify store with hero images, pricing tiers, scarcity copy, and mobile-optimized checkout in 90 seconds → The store goes live before the first TikTok video finishes processing The key move 94% of people skip: you can't animate the photo before you inject the negative prompt. If you send a raw Pinterest image straight into image-to-video the face morphs into a wax figure. The fabric loses texture. The hands grow extra fingers. The whole thing screams "AI" and your CTR dies. His system runs the exclusion filter first so the model moves like she's shot on an iPhone 15 Pro in natural light. One brand hit 2.6M views on TikTok in 11 days with zero paid ads and converted at 3.7% because the videos looked like organic UGC not polished studio content. Brands now pay him $1,400 for the full pipeline setup + $340/month to keep the store synced with new product drops and seasonal video batches. The entire system runs on $23/month in API costs and one laptop. No photographer. No model agency. No product samples. Just a prompt template, a Pinterest account, and the discipline to filter out the AI artifacts before you render movement.

Shade

536,735 görüntüleme • 2 ay önce

Claude Code + ChatGPT Images 2.0 is f*cking cracked 🤯 I rebuilt my static ad system inside Claude Code on the new ChatGPT Images 2.0 model. One brand name + one URL = 40 production-ready static ads. All inside Claude Code. Perfect for DTC brands and agencies who need high-volume ad creative without briefing a designer or spending hours in Canva. If you're finding winning ad concepts on Meta and manually recreating them one at a time — copying prompts, pasting product details, tweaking aspect ratios, downloading, organizing... This system eliminates the entire loop: → Give Claude a brand name and URL → It researches the brand's fonts, colors, packaging, and photography style → Builds a Brand DNA document from scratch → Fills in 40 proven ad templates (headline, us vs them, testimonial, UGC, review cards, stat callouts) with brand-specific details → Fires every prompt to ChatGPT Images 2.0 with your product photos as reference → Downloads finished ads into organized folders with an HTML gallery No manual prompt filling. No Canva templates. No copy-pasting between tools. What you get: → 40 ad formats filled with your exact brand colors, fonts, and copy → Text that actually renders correctly (the new model handles dense copy, logos, and multi-language callouts cleanly) → Product photos passed as reference so the model matches your real packaging → A reusable system — new brand, new folder, same pipeline Built 100% in Claude Code with ChatGPT Images 2.0. I put together a DIY playbook showing the exact architecture so you can build this yourself in Claude Code. Want it for free? > Like this post > Comment "CHAT" And I'll send it over (must be following so I can DM)

Mike Futia

189,312 görüntüleme • 3 ay önce

Claude Code + ChatGPT Images 2.0 is f*cking cracked 🤯 I rebuilt my static ad system inside Claude Code on the new ChatGPT Images 2.0 model. One brand name + one URL = 40 production-ready static ads. All inside Claude Code. Perfect for DTC brands and agencies who need high-volume ad creative without briefing a designer or spending hours in Canva. If you're finding winning ad concepts on Meta and manually recreating them one at a time — copying prompts, pasting product details, tweaking aspect ratios, downloading, organizing... This system eliminates the entire loop: → Give Claude a brand name and URL → It researches the brand's fonts, colors, packaging, and photography style → Builds a Brand DNA document from scratch → Fills in 40 proven ad templates (headline, us vs them, testimonial, UGC, review cards, stat callouts) with brand-specific details → Fires every prompt to ChatGPT Images 2.0 with your product photos as reference → Downloads finished ads into organized folders with an HTML gallery No manual prompt filling. No Canva templates. No copy-pasting between tools. What you get: → 40 ad formats filled with your exact brand colors, fonts, and copy → Text that actually renders correctly (the new model handles dense copy, logos, and multi-language callouts cleanly) → Product photos passed as reference so the model matches your real packaging → A reusable system — new brand, new folder, same pipeline Built 100% in Claude Code with ChatGPT Images 2.0. I put together a DIY playbook showing the exact architecture so you can build this yourself in Claude Code. Want it for free? > Like this post > Comment "CHAT" And I'll send it over (must be following @learnwithella so I can DM)

Ismail Khan

19,800 görüntüleme • 3 ay önce

Day 11/90 of Inference Engineering How does vLLM work and how is it used in production? Before we discuss how vLLM works internally, it helps to understand what vLLM is. At a high level, vLLM is an inference engine that is designed to serve LLMs to thousands of concurrent users efficiently while managing scarce compute and memory. The goal for vLLM is to maximize throughput and minimize latency; optimizing for the best inference economics and experience for end users. With every request from the end user, it eventually ends up in the engine core, gets scheduled alongside other requests from other concurrent users, executes on the GPU, and updates the KV cache with the new key and value vectors, and streams the tokens back to the user. The Scheduler decides what requests should execute next while continuously batching requests together to maximize GPU utilization. Continuous batching is an inference optimization that allows new requests to join a running batch as other requests finish generating tokens. This helps with keeping the GPU utilization high instead of letting it sit idle waiting for an entire batch to complete generating. After the scheduler dispatches the selected batch to the Model Executor, the Model Executor prepares the tensors and metadata required for inference, retrieves each request’s block table from KV Cache Manager, launches the optimized transformer forward pass on the GPU, computes the logits, updates the KV cache with the new key and value vectors, and finally returns the results for sampling and streaming. The KV Cache Manager uses the PagedAttention memory layout to allocate fixed-size cache blocks on demand and maintains a Free Block Queue on the CPU that tracks which blocks in the GPU’s Paged KV Cache are currently free. When a request needs additional KV cache space, the KV Cache manager takes a free block from the queue and assigns it to that request, thus avoiding an expensive search through GPU memory for available cache blocks. All of these components form the core of vLLM’s inference engine. The Scheduler determines what requests are executed, the Model Executor determines how those requests are executed, the KV Cache Manager determines where each request’s KV cache lives using the PagedAttention Memory Layout. This architecture enables vLLM to serve thousands of concurrent requests with high throughput, low latency, and efficient GPU memory utilization. Heres a little animation that visualizes everything! - I've also completed the forward pass for my mnist.c project. I had a nice chat with shrey birmiwal, such a knowledgeable guy. Excited to learn more about vLLM and implement a tiny-vLLM one day.

max fu

70,238 görüntüleme • 24 gün önce