Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

VoxCPM 2 just dropped by OpenBMB Only 2B-param open-source TTS (Text-to-Speech) model built for production-grade multilingual voice work. Apache-2.0 license, Can run on only 8GB VRAM. • Eliminates the "robotic" feel of traditional TTS, delivering prosody and emotional depth suitable for high-stakes professional environments like filmmaking, gaming, animation, and...

13,541 Aufrufe • vor 4 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Type a sentence, get any sound - from talking cats to singing saxophones. Brilliant release by NVIDIA ✨ NVIDIA just unveiled Fugatto, a groundbreaking 2.5B parameter audio AI model that can generate and transform any combination of music, voices, and sounds using text prompts and audio inputs Fugatto could ultimately allow developers and creators to bring sounds to life simply by inputting text prompts, → The model demonstrates unique capabilities like creating hybrid sounds (trumpet barking), changing accents/emotions in voices, and allowing fine-grained control over sound transitions - trained on millions of audio samples using 32 NVIDIA H100 GPUs 👨‍🔧 Architecture Built as a foundational generative transformer model leveraging NVIDIA's previous work in speech modeling and audio understanding. The training process involved creating a specialized blended dataset containing millions of audio samples → ComposableART's Innovation in Audio Control Introduces a novel technique allowing combination of instructions that were only seen separately during training. Users can blend different audio attributes and control their intensity → Temporal Interpolation Capabilities Enables generation of evolving soundscapes with precise control over transitions. Can create dynamic audio sequences like rainstorms fading into birdsong at dawn → Processes both text and audio inputs flexibly, enabling tasks like removing instruments from songs or modifying specific audio characteristics while preserving others → Shows capabilities beyond its training data, creating entirely new sound combinations through interaction between different trained abilities 🔍 Real-world Applications → Allows rapid prototyping of musical ideas, style experimentation, and real-time sound creation during studio sessions → Enables dynamic audio asset generation matching gameplay situations, reducing pre-recorded audio requirements → Can modify voice characteristics for language learning applications, allowing content delivery in familiar voices NVIDIA AI Developer

Rohan Paul

96,354 Aufrufe • vor 1 Jahr

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

221,558 Aufrufe • vor 5 Monaten

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 Aufrufe • vor 3 Monaten

Seedance 2.5 just got an upgrade with 1080p only on Higgsfield... you can now ONE SHOT commercials like this here's exactly how to do it (with a gift at then end): 1/ write the prompt as a timestamped breakdown > split the spot second by second > put the exact dialogue inside each window in quotes, the model speaks it word for word with lip sync > match the generation length to your timestamps, a 21 second script generated at 15 compresses and the delivery desyncs 2/ composition is named - never hoped for > name the camera: "handheld front camera, chest-up framing, natural micro shakes" for UGC, "35mm, slow push-in, product centered" for a produced spot > name the light: "golden hour through the windshield" or "ring light with slight reflections in the eyes" > name the grade: "high contrast, cool tones, warm skin" > whatever you leave unstated gets invented and locked for the whole clip 3/ characters hold when you anchor them > generate a first frame image before any video: real skin texture, visible pores, one fixed imperfection like freckles so drift becomes instantly visible > feed it in as the reference and open every prompt with "the same person as the reference, identical face, hair and outfit" > the @ system locks it harder: @.character for the face, @.style for the look, @.audio for the voice > then write the micro behaviors: a glance away and back, a pre-line breath, fingers adjusting grip... small involuntary movement is what reads human, and you get it by naming it 4/ sound is written - not defaulted > end every prompt with an audio block: "clear phone-mic voice with light room tone" for UGC, "clean studio voice, no echo" for a produced spot > name the music under the dialogue: "soft upbeat synth instrumental running quietly underneath" > one delivery word for the read: "delivery: fed up" produces a performance, an adjective stack produces nothing 5/ sharp text is quoted text > write the exact label or on-screen text in quotes with its style > unquoted text gets invented typography that garbles between frames > print brand names big, a bold label survives every shot while a small tag melts 6/ the consistency laws > repeat the product description verbatim in every prompt, faces anchor but products drift > count objects scene-wide: "exactly one bottle in the entire scene, no duplicate on any surface" > close the wardrobe: "small gold studs, no other jewellery, no rings, no watch" > every state change happens across a cut: the swatch on her hand in shot one, blended in shot two, no clip contains the transition... the viewer's brain supplies it and the shortcut for UGC: take an ad that already converted, ask gemini for a 1:1 timestamped breakdown of everything on screen, swap in your product and script -> that breakdown is your prompt set 9:16 for shortform, 16:9 for the spot, generate straight at 1080p and ship RT + reply to this post and i'll send you my full guide to make your own creatives

Machina

12,785 Aufrufe • vor 25 Tagen

I just vibe coded a static ad generator in Claude Code that creates 100+ Meta ads in minutes. All using the new, insane ChatGPT Images 2.0 model. One competitor ad + your product photo + your brand kit = dozens of on-brand variations, each targeting a different customer persona. Built 100% in Claude Code on the new ChatGPT Images 2.0. Perfect for DTC brands and agencies who need more statics at scale. Here's how it works: → Upload any competitor ad as your reference template → Add your product photos and brand kit (colors, fonts, logos) → AI generates 10 customer profiles from your brand research → Pick how many variations you want (10, 20, 30) → Tool fires every prompt to ChatGPT Images 2.0 with persona-specific copy for each one No designer back-and-forth. No Canva templates. No generic "Shop Now" on everything. What you get: → Ads that mirror winning concepts in your brand's voice → Text that actually renders correctly (the new model handles dense copy, logos, and multi-language callouts cleanly) → Copy targeted to specific customer pain points and personas → Multi-brand/client support with saved brand kits → Reusable customer profiles you build once and generate from forever I recorded a full walkthrough showing exactly how this works, including ALL the prompts I used so you can build it yourself. Want access to all the prompts for free? > Like this post > Comment "STATICS" And I'll send it over (must be following so I can DM)

Mike Futia

52,663 Aufrufe • vor 4 Monaten

Alibaba just released a coding model that hits 82 percent on SWE-Bench Verified. That is the highest score ever published for an open-source model. The weights are free. The license is Apache 2.0. You can run it today. The model is Qwen 4 Coder 32B. Here is what 82 percent on SWE-Bench Verified actually means. SWE-Bench Verified tests whether an AI can autonomously resolve real bugs pulled from real production GitHub repositories. Not synthetic exercises. Real open-source projects that real teams depend on. A model gets a bug report, reads the code, writes a fix, and either passes the test suite or it does not. At 82 percent, Qwen 4 Coder 32B resolves 82 out of every 100 real production bugs it is given. Without a human guiding it. On code it has never seen before. For comparison: Qwen 4 Coder 32B: 82 percent SWE-Bench Verified. Open source. Apache 2.0. Claude Fable 5: 80.3 percent SWE-Bench Pro. $10 input / $50 output per million tokens. Currently suspended. GPT-5.6 Sol: Competitive on Terminal-Bench. $5 input / $30 output per million tokens. An open-weight model that you can download and run for free just beat both of them on the benchmark designed to measure real software engineering capability. Here is the architecture. Qwen 4 Coder 32B is a 32 billion parameter dense model. Not a Mixture-of-Experts. Every parameter is active on every request. This matters for inference: a dense 32B model runs on 22 gigabytes of VRAM, which fits on a single high-end consumer GPU or a MacBook Pro with 64GB of unified memory. The smaller variant, Qwen 4 Coder 4B, runs at approximately 135 tokens per second on an M5 Max and fits inside 8 gigabytes of RAM. For a model with usable coding capability, that is a new bar for what fits in a single laptop. The training methodology continued Alibaba's approach of reinforcement learning on verifiable coding tasks. The model gets rewarded when its code passes tests. It gets penalized when it fails. Over millions of training steps, the model learns to write code that actually runs rather than code that looks plausible. License: Apache 2.0. Full commercial use. No attribution requirement. No revenue threshold. No monthly active user ceiling. Weights: Hugging Face, available today. Runs on: vLLM, Ollama, SGLang, and any standard GGUF-compatible inference engine. Qwen 4 32B also runs at approximately 135 tokens per second on an M5 Max chip, setting a new bar for what a sub-8GB model can do on Apple Silicon. The open-source coding model just beat the best closed-source model in the world on the benchmark designed to test whether AI can actually do software engineering. The weights are free. The subscription is optional. Source: Autom8Labs AI Insight July 2026, State of Open Source LLMs June 2026, Kunal Ganglani blog June 2026.

Harman

41,278 Aufrufe • vor 2 Monaten

Nobody's talking about the Claude trick that fixes every Seedance 2.0 video mistake. A cinematic action scene, an animated sequence, a product ad, and a dialogue scene are completely different types of content. Different camera language, different pacing, different reference logic. Most people try to force all of it through one universal prompt. The fix turns out to be simpler than expected: Claude Skills. It's not some secret hack. It's a file of instructions you load into Claude for one specific type of task. Once it's loaded, Claude stops acting like a general assistant and starts thinking like an expert in that one niche — with its own set of rules for that scene type. Here's what that looks like in practice: ▪ Cinematic skill — generates real camera language: dolly moves, crane shots, rack focus, shot discipline across multiple cuts. Test: a medieval battle between two armies, 15 seconds, dark fantasy. Result — three shots with the shot logic of an actual film, not just "epic battle" typed as text. ▪ Animation skill — locks in style and physics before a single shot is built. One test: a rain-soaked shonen fight in the style of Tokyo Revengers. Another: a quiet Ghibli-style farm scene. Same skill, completely different visual output — because the style gets fixed at the very top of the prompt. ▪ Product ad skill — keeps the product front and center in every shot: clean hero framing, commercial lighting, a full product description built from the reference image before any shots are generated. ▪ Dialogue skill — this is where most people fall apart. Lip sync, emotional direction, shot structure, and audio cues all have to work together. Test: an interrogation scene with a fourth-wall break — and the moment landed exactly as written, down to the pause and tone. The core idea is simple: instead of writing a prompt from scratch every time and hoping for the best, you build one skill per content type — and Claude asks the right setup questions (genre, tone, shot count, camera energy) before generating a fully structured Seedance 2.0 prompt. One prompt template for everything is exactly why 90% of AI videos look the same. This 13-minute video is free, and it's more useful than a $500 course. YouTube: "Skai Generated" - Thank you for sharing this invaluable material with us.

Zentrix⌚️

45,636 Aufrufe • vor 1 Monat

save this post to get the most out of unlimited Seedance 2.5 for up to 33 days on Higgsfield i'm going to show you how to use loops to produce ANY video format: ads, cinema, vlogs, UGC, music videos... with one system idea > vault > agent > references > images > script > video > montage > upscaling Seedance 2.5 one-shots a full 30 second video, audio generated in the same pass, carrying up to 30 image, 10 video and 10 audio references into a single generation here's a full breakdown of the setup: > idea: steal taste from work that already worked: - frameset․app and shotdeck․com for film stills - savee․com and cosmos․so for boards - eyecannndy․com for transitions then have a vision model name the lens, light, palette and grain of your picks in one locked paragraph you paste into every prompt > vault: an obsidian folder as your reference bible, one page per asset (idea, locked style, character sheets, reference images, the exact prompts that worked) plus one index page, reviewed after every session so it never rots into dead files > agent: three commands make every model callable from Claude Code: - npm install -g @ higgsfield/cli - higgsfield auth login - npx skills add higgsfield-ai/skills and your agent now submits, polls, retries and logs every job > references: build reference images by hand first, midjourney for cinema and stylized shots, nanobanana pro or gpt images 2 for realism use one locked style across the whole project, recurring characters turned into full sheets (front, side, back, blank background), and once locked you never regenerate them, you fix the motion prompt instead > images: frames before motion, always, a frame costs seconds and a clip costs minutes, so exploration happens at the cheap layer and only winners get animated > script: every shot gets the same six details, subject, action, place, camera, style, rules, and the 30 seconds splits into four timed beats inside one prompt, 0-6 set the scene, 6-14 build it out, 14-24 the turn, 24-30 the end > video: every reference gets a job and a boundary, "Video 1 defines motion and pacing" is half the instruction, "do not use the person's identity, clothing or scene" is the half that stops one reference leaking into shots it was never meant to touch > montage: the cut is a text file, one line per clip with its duration and an audio flag, ffmpeg renders the film from it, so the whole edit reruns in seconds > upscaling: once, at the end, on the finished cut, 720p while exploring, 1080p for keepers, 4K only for the master (use Topaz) for UGC ads, the same loop with two changes render the hook clip alone first, approve the face and the voice before anything else inherits them, then anchor every later clip with the approved hook's audio so one voice carries the whole ad and the script math is fixed, about 3.5 words per second, a 30 second ad is roughly 105 words, counted before anything renders unlimited means every loop above costs nothing to run... start one tonight

Machina

39,015 Aufrufe • vor 29 Tagen

Anthropic released Claude Design TODAY and it's now accessible at I spent the last hour giving it a first look, and shared my thoughts and results in the video below. This is a BIG drop. This is a new design surface from Anthropic, and it changes what "AI design" means. Short version: Claude can now design. Not "describe a design." Not "generate an image of a design." Actual production work — prototypes, wireframes, high-fidelity mocks, slide decks, landing pages — editable, on-brand, and ready to hand off. Here's what stood out on first look: → Real design surfaces Prototypes, wireframes, hi-fi, and slide decks — each with templates and proper structure, not just pretty screenshots. → Comment-based edits Leave a comment on any element and Claude revises it. This is the Figma-style review loop, with the designer replaced by a model that works at 3am. → Brand design systems You can feed it your system — colors, type, components — and it actually respects it. On-brand output, not generic AI slop. → Export anywhere PDF, PowerPoint, Canva, standalone HTML. Plus a built-in handoff straight to Claude Code for engineers to implement. → Import from real tools Figma, GitHub, and captured web elements come in as inputs. Your existing work is the starting line, not the discard pile. → Collaboration Share links for view / comment / edit — the exact tier system teams already expect. What I tested on Opus 4.7: • A 5-slide deck generated from a single screenshot. Claude asked clarifying questions BEFORE generating and shipped speaker notes by default. • A landing page build. Solid first pass, real components, real layout logic. • Multiple chats running concurrently. You can parallelize design work across threads like a small team. Why this matters: PMs, founders, marketers, and non-engineers can now create designs that engineers can actually ship with production-ready output and a claude code handoff built in. The gap between "I have an idea" and "here's a working prototype with my brand applied" just collapsed to minutes. Full walkthrough, live demos, exports, and honest takes on where it breaks below. P.S. • This is an Anthropic Labs product — NOT GA yet. • Claude Design is currently webapp only (no API), and does not yet support the Analytics API, Compliance API, or cost/usage reporting. • Availability: – Default ON for Pro / Max / Team – Default OFF for Enterprise Enterprise admins can toggle it on via RBAC in console (comes with a ~$20/user initial credit).

JJ Englert

32,445 Aufrufe • vor 4 Monaten

You have to really give it to OpenAI because Sora 2 is very impressive on a lot of fronts: - high quality video model with great physics - high quality audio in each video - high character consistency - multiple characters in one scene - accurate characters voice - social platform attached to it Before today the best AI video models were dominated by Chinese companies like ByteDance and Kuaishou and Google with Veo3. ByteDance makes TikTok, Kuaishou makes Kwai (similar app) and Google has YouTube to train on But none of these models had great character consistency, if it was a feature at all, let alone multiple characters in one scene. Generally you'd make a video and the face would slowly change into someone else, just not good On top of that Google was struggling with allowing people to upload characters scared it'd get abused for deep fakes, and just generally nerfing their model so you can't really use it for anything OpenAI solved that by re-thinking ownership over your characters smartly with Cameo, which is essentially "train yourself as a AI model" which we've all been doing in our apps for years, but in a more smart way, where you can control if only you make content with your appearance, or others too They've also added voice training to it immediately, which people would have to do separate on for ex ElevenLabs before On top of that the social platform aspect: Google's Veo 3 didn't have ANY community at all, while the Chinese video models did, but it was all more like weekly themed contests to win free credits, they never really managed to make it more than that, and it kinda stayed in this nerdy AI hacking vibe This vibe fits how hard it was/is to simply make a video featuring you or your friends with proper voice and audio and everything that Sora 2 does for you. You'd have to go to ElevenLabs to train your voices, then go to for ex Photo AI to train yourself as a person, then make videos, then add audio and voices, then edit them together, a lot of work! We don't know if Sora 2's social platform features will actually be used or take off, but it's a real cool experiment in trying to find a way to build a community around AI in a more Instagram-like way Being able to tag your friends and then add them as multiple characters is innovative in both the social and technical aspect So TL;DR OpenAI essentially took a lot of stuff that was already technically possible, then added new things that weren't possible yet, and then put it all together in a very friendly interface that even my mom can use, with generation times of just a few minutes which is extremely fast if you think of the pipeline behind it (multiple video generation + voice + audio etc.) And also importantly, it doesn't look like they nerfed it much for safety which is also very cool considering the legal risks So yes very very very impressive

@levelsio

178,100 Aufrufe • vor 11 Monaten

great to see more people generating 3d avatars with our new text-to-3d feature in forge. this marks a step in the right direction in putting powerful creation tools directly in the hands of everyone. we built forge entirely from the ground up over the past months as one of the key releases on our roadmap. having full proprietary ownership of the technology gives us complete control to shape its direction without depending on external platforms or third-party licenses. building on this foundation, our upcoming studio feature will let users generate high-quality accessories, clothing, and environments simply by typing natural language prompts. a single description can produce fully textured, production-ready 3d assets in seconds, with options to create multiple variations and refine them through follow-up instructions.every asset created in studio integrates seamlessly with the 3d ai agents made in forge on users can instantly apply clothing and accessories with automatic fitting, layer multiple items, and place their agents inside custom-generated environments. all clothing and accessories come pre-rigged and optimized, while environments include proper lighting and geometry for immediate use in games, animation, virtual worlds, and more. we have spent months thoughtfully designing how can deliver real, sustainable value back to the community. we will continue to share more details on token utility use cases and the economic flywheel we have built. the goal is to create a self-reinforcing system where creators earn through royalties, autonomous agents drive on-chain activity, and platform growth directly benefits active community members and token holders. together, these features enable a complete creative flow. from a simple idea, anyone can quickly build fully realized 3d ai agents standing in rich, custom scenes. we are excited to see what the community builds next.

nich

19,611 Aufrufe • vor 2 Monaten

🚨 The Next Evolution of AI Music is Here 🚨 We haven’t been standing still. We’ve been building at an incredible pace, with laser-sharp focus, pushing the boundaries of AI-powered music creation like never before. Our latest upgrade isn’t just more powerful—it’s more versatile, precise, and deeply creative than anything before. 🎶 Proof is in the sound: This song was generated from a simple prompt—“Blues with slight Arabian influence about a man lost in the desert searching for his bride.” Listen to the end and hear how $SUEDE AI captures emotion, style, and storytelling like never before. But this is just the beginning. Our new features are built for artists who want total creative control. Get extremely granular with how you craft and shape your sound: 🎛️ Full Production Control – Download an entire pack of every isolated instrument. 🎤 Use Your Own Voice – Or someone else’s. 📝 Exact Lyrics, Your Way – Have your words set to music seamlessly. 🎶 Reference Songs – Upload one for style analysis, extraction, and modeling—or simply note a publicly available track. 🔊 Text-to-Speech & AI Vocalists – Shape voices like never before. 🎼 Melody Collaboration – Upload a melody idea and let others build around it—or vice versa. However, due to cost considerations, we’ve capped it at 3 free songs per trial until subscription payments roll out in the next day or two. We’ll be launching a new payment gateway soon, so stay tuned for more details. And remember, all of this is powered by the $SUEDE token. It fuels the entire ecosystem, allowing artists to generate, own, and monetize their work like never before. We’re still working out a few kinks—like image generation—but prepare to be impressed. A major post is coming soon, breaking down these game-changing features and the revenue model behind them. Thread dropping soon. Turn notifications on. $SUEDE powers the future of culture. #SuedeAI #Web3Music

Suede Labs

17,425 Aufrufe • vor 1 Jahr