正在加载视频...

视频加载失败

Huge drop from Tencent! AuK - turns speech gen/editing into a text-prompt playground. You can control literally anything with natural language. - zero-shot voice cloning, no reference transcript - swap an accent, inject anger, fix background noise, or change a single spoken word in a track. - full TTS,...

16,216 次观看 • 2 天前 •via X (Twitter)

1 条评论

VOID 的头像
VOID2 天前

@realrebelai

相关视频

VoxCPM 2 just dropped by OpenBMB Only 2B-param open-source TTS (Text-to-Speech) model built for production-grade multilingual voice work. Apache-2.0 license, Can run on only 8GB VRAM. • Eliminates the "robotic" feel of traditional TTS, delivering prosody and emotional depth suitable for high-stakes professional environments like filmmaking, gaming, animation, and audiobooks. • 30-language multilingual: no language tag needed, just type in a supported language and generate directly. • Voice design: create a brand-new voice from a text description alone, like age, tone, pace, or emotion. No reference audio required. Describe the desired voice characteristics (gender, age, tone, emotion, pace …) in Control Instruction, and VoxCPM2 will craft a unique voice from your description alone. • Controllable cloning: clone from a short clip, then steer delivery style without losing the speaker’s core voice. • Ultimate cloning: use reference audio + transcript for continuation-style cloning that keeps the tiny vocal details. • 48kHz output: takes 16kHz reference audio and produces studio-quality speech without an external upsampler. • Real-time ready: around 0.3 RTF on RTX 4090, even lower with Nano-VLLM. • Commercial use: Apache-2.0 licensed. Developer-Friendly Infrastructure: - Native Torch Inference: Direct support for PyTorch-based workflows. - Training Flexibility: Supports both full-parameter and LoRA fine-tuning for specific domain adaptation. - Production Readiness: Compatible with voxcpm-nanovllm for large-scale, high-concurrency deployment.

Rohan Paul

13,541 次观看 • 5 个月前

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

221,558 次观看 • 5 个月前

Seedance 2.5 just got an upgrade with 1080p only on Higgsfield... you can now ONE SHOT commercials like this here's exactly how to do it (with a gift at then end): 1/ write the prompt as a timestamped breakdown > split the spot second by second > put the exact dialogue inside each window in quotes, the model speaks it word for word with lip sync > match the generation length to your timestamps, a 21 second script generated at 15 compresses and the delivery desyncs 2/ composition is named - never hoped for > name the camera: "handheld front camera, chest-up framing, natural micro shakes" for UGC, "35mm, slow push-in, product centered" for a produced spot > name the light: "golden hour through the windshield" or "ring light with slight reflections in the eyes" > name the grade: "high contrast, cool tones, warm skin" > whatever you leave unstated gets invented and locked for the whole clip 3/ characters hold when you anchor them > generate a first frame image before any video: real skin texture, visible pores, one fixed imperfection like freckles so drift becomes instantly visible > feed it in as the reference and open every prompt with "the same person as the reference, identical face, hair and outfit" > the @ system locks it harder: @.character for the face, @.style for the look, @.audio for the voice > then write the micro behaviors: a glance away and back, a pre-line breath, fingers adjusting grip... small involuntary movement is what reads human, and you get it by naming it 4/ sound is written - not defaulted > end every prompt with an audio block: "clear phone-mic voice with light room tone" for UGC, "clean studio voice, no echo" for a produced spot > name the music under the dialogue: "soft upbeat synth instrumental running quietly underneath" > one delivery word for the read: "delivery: fed up" produces a performance, an adjective stack produces nothing 5/ sharp text is quoted text > write the exact label or on-screen text in quotes with its style > unquoted text gets invented typography that garbles between frames > print brand names big, a bold label survives every shot while a small tag melts 6/ the consistency laws > repeat the product description verbatim in every prompt, faces anchor but products drift > count objects scene-wide: "exactly one bottle in the entire scene, no duplicate on any surface" > close the wardrobe: "small gold studs, no other jewellery, no rings, no watch" > every state change happens across a cut: the swatch on her hand in shot one, blended in shot two, no clip contains the transition... the viewer's brain supplies it and the shortcut for UGC: take an ad that already converted, ask gemini for a 1:1 timestamped breakdown of everything on screen, swap in your product and script -> that breakdown is your prompt set 9:16 for shortform, 16:9 for the spot, generate straight at 1080p and ship RT + reply to this post and i'll send you my full guide to make your own creatives

Machina

12,785 次观看 • 28 天前

🚀 Announcing Echo — our new frontier model for 3D world generation. Echo turns a simple text prompt or image into a fully explorable, 3D-consistent world. Instead of disconnected views, the result is a single, coherent spatial representation you can move through freely. This is part of a bigger shift in AI: from generating pixels and tokens to generating spaces. Echo predicts a geometry-grounded 3D scene at metric scale, meaning every novel view, depth map, and interaction comes from the same underlying world — not independent hallucinations. Once generated, the world is interactive in real time. You control the camera, explore from any angle, and render instantly — even on low-end hardware, directly in the browser. High-quality 3D world exploration is no longer gated by expensive equipment. Under the hood, Echo infers a physically grounded 3D representation and converts it into a renderable format. For our web demo, we use 3D Gaussian Splatting (3DGS) for fast, GPU-friendly rendering — but the representation itself is flexible and can be easily adapted. Why this matters: consistent 3D worlds unlock real workflows — digital twins, 3D design, game environments, robotics simulation, and more. From a single photo or a line of text, Echo builds worlds that are reliable, editable, and spatially faithful. Echo also enables scene editing and restyling. Change materials, remove or add objects, explore design variations — all while preserving global 3D consistency. Editing no longer breaks the world. This is only the beginning. Echo is the foundation for future world models with dynamics, physical reasoning, and richer interaction — environments that don’t just look right, but behave right. Explore the generated worlds on our website and sign up for the closed beta. The era of spatial intelligence starts here. 🌍 #Echo #WorldModels #SpatialAI #3DFoundationModels Check it out:

SpAItial AI

176,524 次观看 • 9 个月前

DRONE VIDEOGRAPHERS CHARGE $10K FOR THIS SHOT. HE PULLS IT FROM GOOGLE EARTH AND A PROMPT You never buy a drone, book a pilot, or leave the house. You pick any city on Earth, trace the flight path you want, and let Gemini render it as real-looking FPV footage. Clients pay thousands for this shot. You make it from a screenshot Here is the exact process: 1. Open Google Earth. Find the city or building you want. Frame the angle you'd want a drone to start from and take a screenshot 2. Draw the path. On that screenshot, draw a red line showing exactly where the drone should fly through the scene. This line is what the AI follows 3. Open Gemini and drop in the screenshot. Use the video generation in the Gemini app, the part that animates a still image into motion. Nano Banana handles images, the video engine is what turns your shot into footage 4. Paste the prompt. Tell it to follow the red flight path through the city, fast smooth motion, banking around buildings, golden-hour light, motion blur, 9:16 vertical, real FPV drone look. Full prompt is in the comments 5. Generate and clean it up. One clip is a few seconds. Stitch a couple together for a full flythrough and you have a reel Set the prompt once and you can re-run it for any location on the planet Who pays for this: Real estate agents, hotels, restaurants and event venues all need aerial b-roll and almost none can afford a real drone shoot Pull listings or venues with flat, ground-level photos and zero aerial footage. Send a free sample flythrough of their own location, then charge per clip or a monthly rate for ongoing reels One agent with ten listings is a recurring client, fully online Full prompt in the comments Bookmark this

Yarchi

53,401 次观看 • 3 个月前

NVIDIA just unleashed SANA-WM and it’s an absolute MONSTER for the future of open source AI! A blazing-fast 2.6B-parameter open-source world model that doesn’t just generate video… it creates controllable, physics-rich, high-fidelity worlds on demand. Why this is insanely powerful: • One image + text prompt + 6-DoF camera trajectory → generates 720p videos up to 60 seconds long with buttery-smooth, precisely controlled camera movement. You’re not just watching, you’re piloting the simulation. • Runs locally on a single consumer GPU (RTX 5090 level) thanks to heavy distillation + NVFP4 quantization. Full 60-second clip denoised in ~34 seconds. No massive clusters required. • 36× higher throughput than previous open models while rivaling (or beating) closed industrial giants in visual quality and consistency. • Trained lightning-fast: ~213K public videos in just 15 days on 64 H100s. • Built with next-level tech: Hybrid Linear Attention, dual-branch camera control, two-stage pipeline, and rock-solid metric-scale pose understanding. This is a true open world model, the foundation for embodied AI, robotics, autonomous systems, and hyper-realistic simulations that can run anywhere. Project: At our Zero-Human Company, we’re already running SANA-WM live in our core pipelines. It’s supercharging autonomous agent training, generating unlimited synthetic training data, and powering full end-to-end simulation loops, zero humans in the loop. The speed and control let us test thousands of edge-case scenarios overnight, iterate at lightspeed, and push our fully autonomous operations further than ever before. This is the kind of breakthrough that turns science fiction into daily reality. World models just leveled up — hard. The age of personal, local, controllable universes is here.

Brian Roemmele

619,329 次观看 • 3 个月前

If you are running local LLMs without N-gram speculative decoding, you are wasting massive amounts of compute. Whether your AI is editing a document, outputting structured JSON, or rewriting boilerplate templates, a huge chunk of the text it generates is highly repetitive or already exists right there in the prompt. Standard decoding wastes expensive GPU compute cycles "re thinking" every single token. By adding one hidden flag in llama.cpp, you can instantly fast forward through the repetition. Zero draft models. Zero extra VRAM. And virtually zero compute overhead. Google Colab hands you an enterprise grade NVIDIA Tesla T4 GPU with 16GB of VRAM for free. It’s the perfect Ubuntu Linux sandbox to build a bleeding edge inference engine from scratch. Recently, I showed you how to double your local speeds using MTP (Multi Token Prediction). But MTP requires a secondary neural network draft model. That eats into your precious VRAM (slightly though) and burns extra compute for every guess it makes. N-gram Speculative Decoding gives you a massive speed boost for exactly 0 memory cost and minimal compute. And it's faster than MTP when it works. Here is how it actually works under the hood: Standard autoregressive decoding is slow because it predicts one token at a time. If you ask an agent to format a long JSON object or update one line in an HTML file, it runs heavy matrix multiplications to calculate the probability of every single bracket, space, and letter from scratch. N-gram changes the game. It acts as a lightweight caching system. Instead of running heavy neural network math to guess the next word, it uses a simple hash table. Whenever the LLM starts outputting a sequence of tokens that already exists anywhere in its context window, N-gram instantly recognizes the pattern. Because it is just doing lightning fast string matching, the compute cost is practically zero. It "fast forwards" through the text, drafting the boilerplate instantly from memory, and the main model just verifies it in parallel. Pure speed. Using quantized GGUFs from Unsloth via HuggingFace, I spun up DeepMind’s massive Gemma 4 26B A4B QAT MoE on a free Colab instance to test this. Just look at the raw benchmark data on code editing task: Without N-gram: [ Prompt: 638.6 t/s | Generation: 45.9 t/s ] With N-gram: [ Prompt: 601.9 t/s | Generation: 107.1 t/s ] Here is the exact llama.cpp CLI command to activate it. Notice we don't even need the --model-draft flag: ./llama-cli -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -cnv -n 6000 -c 12000 -ngl 99 -fa on --spec-type ngram-mod Stop waiting for your GPU to re calculate words it already knows. I’ve built a free, interactive, cell by cell Google Colab notebook that lets you test this live in your browser. You can literally chat with the model and watch the text generation speed absolutely fly on the second turn when you ask it to edit a file. There are additional parameters for ngram-mod that you can tune once you get it working with the single flag. Link to the free Colab Notebook is in the comments below. It walks you through the entire stack: pulling pre built llama.cpp CUDA binaries for Linux, fetching GGUFs from HuggingFace, and spinning up the inference engine with ngram-mod from scratch. Let me know if you have already tried ngram-mod

Alok

31,765 次观看 • 1 个月前

THIS GUY TURNED 5 PROMPTING TIPS INTO A FREE AI CEO CHALLENGE The useful part is treating every prompt like you are briefing a very fast employee who has zero context. Most people open ChatGPT and type a wish. Pros give it a job. Try this instead: 1. Give it a role Not “help me with marketing.” Say: “Act as a B2B SaaS growth operator reviewing a landing page.” 2. Give it the real context Who is the customer? What are they buying? What have you already tried? What does success look like? 3. Give it constraints Length, tone, format, audience, banned words, examples to copy, examples to avoid. A vague prompt gets a vague answer. A constrained prompt gets something you can edit. 4. Ask for options before answers “Give me 5 angles, rank them, then explain the tradeoff.” This turns AI from an autocomplete box into a thinking partner. 5. Force it to show assumptions Before it writes, ask: “What are you assuming, what info is missing, and what would change your answer?” That one line saves a lot of fake confidence. Dan Martell’s video works because the promise is simple: 5 prompting habits that make AI feel less random. The reusable move is even simpler: Stop prompting for outputs. Start prompting for decisions. Bad: “Write me a post.” Better: “Here is the source, here is the reader, here is the angle, give me 3 hooks, choose the strongest, then draft in this style.” That is the difference between getting content-shaped noise and getting work you can actually ship. Caveat: prompts do not fix weak taste, bad data, or unclear strategy. But they do expose those problems faster. If your AI answers are generic, your prompt probably has no job, no context, no constraints, and no standard for what “good” means.

kocer

25,573 次观看 • 2 个月前

🔬 Exciting News! Our manuscript, "scGPT: toward building a foundation model for single-cell multi-omics using generative AI" is now finally published in Nature Methods (Nature Methods) 🎉 !!! (Re-)Introducing scGPT: A transformative foundation model engineered for single-cell omics analysis. Developed through the analysis of over 33 million human cells, scGPT sets a new benchmark for application versatility, offering both fine-tuning and zero-shot capabilities. Since its preprint in May 2023, scGPT has significantly impacted the field, evidenced by 13K+ installations, 600+ GitHub stars 🌟, and 40+ citations before its official publication! scGPT has been validated by numerous benchmark studies as a leading foundation model in single-cell analysis. Its pre-trained embeddings extend its utility beyond single-cell studies, enhancing a variety of downstream tasks including protein enrichment and genetic perturbation predictions. Some key updates lately: ---Expanded zero-shot applications for efficient reference mapping and integration, now with CellXGene census integration. ---Advanced perturbation analysis capabilities, including genome-scale perturb-seq data analysis and bulk sequencing data generalization. ---Upgraded scGPT package, offering versatile model loading compatible with PyTorch and flash-attn, for both GPU and CPU. ---Cloud-based scGPT applications for reference mapping, cell annotation, and gene regulatory network inference are available on ---Integration with Hugging Face for easier model training. Limitations: scGPT is an early foray into foundation models for single-cell omics, facing challenges like limited zero-shot learning in some tasks, pretraining constraints, data quality issues, and evaluation limitations. See our Supplementary Notes for details. 🚀 Future Work? Short-Term Goals: 1. Releasing a Mouse Model for broader analysis. 2. Developing a comprehensive evaluation suite for foundation models in single-cell analysis. 3. Creating a foundation model for single-cell spatial omics. 4. Enhancing zero-shot capacity by integrating scGPT with RAG (e.g., knowledge graphs). Long-Term Goals: 1. Expanding scGPT for comprehensive single-cell multi-omics analysis. 2. Developing an in-silico perturbation model for predicting genetic perturbation effects. 3. Merging scGPT with multi-modal genomic sequence models for a deeper understanding of cell biology. 📚 Access the paper on Nature Methods: 🔬Preprint in Bioarixv: 💻 All our codes/data/weights are open source: Wholehearted congratulations to all the authors, especially the two co-first authors, Haotian (Haotian Cui ) and Chloe (ChloeXWang), who are really the emerging superstars in AI and biology! Vector Institute Peter Munk Cardiac Centre AI U of T Department of Computer Science Department of Laboratory Medicine & Pathobiology University Health Network University of Toronto #scGPT #GenerativeAI #AI4Science #Combio #opensource

Bo Wang

199,776 次观看 • 2 年前

She's selling a $20,000 a month plan from a pool chair, one AirPod in. She says all you need is a phone. She's tapping a full iPad in landscape, silver nails flashing. Search cartoon kids shows. Point at Doggyland, 6.6 million views. Copy the ABC song transcript from the description. Paste it into ChatGPT. "Make me a prompt for a show just like this." Copy what it gives back. Open Picsart Flow. Paste. Pause at 0:38. The model in the dropdown is "Sora 2." Sora 2 isn't in Picsart. Sora 2 isn't anywhere. The resolution under it reads 1280x720. The clip is 12 seconds. Two seconds later the screen fills with bouncing letters, an elephant, a parrot, and a fully mixed song with the lyrics synced underneath. "And this is all in 4K." The prompt only asked for animation. It never asked for a song. A 12 second 720p clip became a finished music video between two taps. That didn't happen on screen. None of this is the point. The point is the workflow is a prop. There is no Sora 2 button that prints a 4K music video. The real thing is four cheap parts. Claude breaks the transcript into scenes and writes the visual prompt for each. Suno turns the lyrics into the song. Kling or Luma renders the clips. ffmpeg stitches them and Whisper syncs the captions. Pennies per call. One evening of Python. She wants you to comment FLOW so she can dm you a template that doesn't run. I rebuilt the real factory, Claude to Suno to Kling. Comment 720 and I'll dm it free.

Aeron

254,852 次观看 • 2 个月前

Introducing KausaCompute. Running AI models privately is expensive and complicated. Traditional cloud providers require credit cards, KYC verification, and complex setup processes. For many developers, especially in emerging markets, accessing GPU compute remains out of reach. KausaCompute changes that. Deploy any Docker container with NVIDIA GPU, pay with USDC from Maze Pocket, no KYC, no credit card, no cloud provider account needed. GPU pricing starts at $0.47/hr with competitive rates across all tiers. KausaCompute lives inside KausaLayer Pocket. Open a pocket, swap SOL to USDC using the built-in swap feature, and start deploying GPU containers right away. Everything stays within one ecosystem. What it actually does: Private LLM endpoints. Deploy Llama 3, Mistral, CodeLlama, or any open-source model as a personal API. No one logs the prompts. No one reads the data. Full control. AI coding assistants that never see external servers. Speech-to-text processing for thousands of audio files in minutes instead of hours. Sentiment analysis across millions of data points. Document summarization at scale. Multi-modal AI that processes images and text together. All running on dedicated NVIDIA GPUs. A6000, T4, L4, L40, A100, H100, and more. Pick the hardware, pick the duration, deploy in one click. The billing is straightforward. USDC is deducted from Maze Pocket before deployment. Pro tip: Maze Pocket supports multiple pockets per wallet. Create a dedicated pocket just for KausaCompute to keep GPU spending separate from other activities like trading or transfers. Clean separation, easy tracking. KausaCompute is not another cloud provider. It is the fastest path from USDC to a running GPU, with zero identity requirements. Live now at

KausaLayer

10,945 次观看 • 3 个月前

We’re excited to announce the release and open-source of HunyuanImage 3.0 — the largest and most powerful open-source text-to-image model to date, with over 80 billion total parameters, of which 13 billion are activated per token during inference.The effect is completely comparable to the industry’s flagship closed-source model.🚀🚀🚀 HunyuanImage 3.0 originates from our internally developed native multimodal large language model, with fine-tuning and post-training focused on text-to-image generation. This unique foundation gives the model a powerful set of capabilities: ✅Reason with world knowledge ✅Understand complex, thousand-word prompts ✅Generate precise text within images Different from traditional DiT architecture image generation models, HunyuanImage 3.0’s MoE architecture uses a Transfusion-based approach to deeply couple Diffusion and LLM training for a single, powerful system. Built on Hunyuan-A13B, HunyuanImage 3.0 was trained on a massive dataset: 5 billion image-text pairs, video frames, interleaved image-text data, and 6 trillion tokens of text corpora. This hybrid training across multimodal generation, understanding, and LLM capabilities allows the model to seamlessly integrate multiple tasks. Whether you're an illustrator, designer, or creator, this is built to slash your workflow from hours to minutes. HunyuanImage 3.0 can generate intricate text, detailed comics, expressive emojis, and lively, engaging illustrations for educational content. The current release focuses solely on text-to-image generation and future updates will include image-to-image, image editing, multi-turn interaction, and more. 👉🏻Try it now: 🔗GitHub: 🤗Hugging Face:

Tencent Hy

413,058 次观看 • 11 个月前