Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

We released Qwen3-Omni-Flash (2025-12-01 version) API Service. Smarter interaction, more human expression: · A/V Interaction: Significant boost in instruction following. Solves "dumbing down" in casual chats with rock-solid stability. · Precise Control: Enhanced System Prompt adherence for specific personas, styles, and lengths. · Multilingual Mastery: Solved language switching instability....

1,637,124 Aufrufe • vor 7 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

VoxCPM 2 just dropped by OpenBMB Only 2B-param open-source TTS (Text-to-Speech) model built for production-grade multilingual voice work. Apache-2.0 license, Can run on only 8GB VRAM. • Eliminates the "robotic" feel of traditional TTS, delivering prosody and emotional depth suitable for high-stakes professional environments like filmmaking, gaming, animation, and audiobooks. • 30-language multilingual: no language tag needed, just type in a supported language and generate directly. • Voice design: create a brand-new voice from a text description alone, like age, tone, pace, or emotion. No reference audio required. Describe the desired voice characteristics (gender, age, tone, emotion, pace …) in Control Instruction, and VoxCPM2 will craft a unique voice from your description alone. • Controllable cloning: clone from a short clip, then steer delivery style without losing the speaker’s core voice. • Ultimate cloning: use reference audio + transcript for continuation-style cloning that keeps the tiny vocal details. • 48kHz output: takes 16kHz reference audio and produces studio-quality speech without an external upsampler. • Real-time ready: around 0.3 RTF on RTX 4090, even lower with Nano-VLLM. • Commercial use: Apache-2.0 licensed. Developer-Friendly Infrastructure: - Native Torch Inference: Direct support for PyTorch-based workflows. - Training Flexibility: Supports both full-parameter and LoRA fine-tuning for specific domain adaptation. - Production Readiness: Compatible with voxcpm-nanovllm for large-scale, high-concurrency deployment.

Rohan Paul

13,541 Aufrufe • vor 3 Monaten

v4.5 just dropped for Pro & Premier subscribers 🔥 A wider range of genres, richer vocals, & enhanced prompt understanding for songs that match your vision. What’s New: 🙌 Expanded genres & smarter mashups: More genre options — Blends like midwest emo + neosoul or EDM + folk come together seamlessly. 🎤 Enhanced voices: Vocals now hit harder — with more depth, emotion, and range. From intimate whispers to full-on power hooks, v4.5 delivers with feeling. 🔊 More complex, textured sound: v4.5 picks up the subtleties that make your music shine — layered instruments, tone shifts, and sonic details with depth. Prompts like “leaf textures” or “melodic whistling” now come through with clarity and dimension. 🧠 Better prompt adherence: Your words hit harder. Mood, vibe, instruments, and detail are captured with precision—so what you imagine is what you hear. ✍️ Prompt enhancement helper: Drop in a few tags or a rough idea, hit Enhance, and get a rich, fully-formed style prompt you can roll with or remix. 🎭 Upgraded Covers + Personas: Covers hold onto more melodic detail. Genre switching feels seamless. Personas better preserve the vibe and character of your track — and now… 🤝🏽 Covers + Personas can be combined: Remix voice, structure, and style all at once. It’s a whole new way to create. 📈 Extended song length: Previously 4 minutes, now create up to 8 minutes without using Extend. 🏆 Improved audio: Fuller, more balanced mixes with reduced shimmer and degradation — everything sounds better.

Suno

439,775 Aufrufe • vor 1 Jahr

China unveils humanoid robot with lifelike skin and blinking eyes built for daily life | Prabhat Ranjan Mishra, Interesting Engineering Large Language Models (LLMs) and Vision-Language Models (VLMs) help process and interpret complex data from human interactions. A Shanghai-based company has developed humanoid robots that appear as real as humans. The advanced bionic humanoid robot is integrated with self-supervised AI algorithms. Named Elf V1, the robot can perceive the world, communicate, learn, and interact intelligently with its surroundings. Developed by AheadForm Technology, the robot offers up to 30 degrees of freedom, powered by a precise control system and an advanced AI learning algorithm. Robot offers expressive facial features The robot offers expressive facial features, moving eyes, and synchronized speech. It can also convey emotions and understand human non-verbal cues, making interactions more natural and engaging. The robot has highly interactive capabilities and lifelike appearances. AheadForm expects that its robots could soon seamlessly integrate into daily life, providing assistance, companionship, and support across various industries. “We believe that by developing realistic and expressive robot heads, we can bridge the gap between humans and machines, fostering a new era of interactive and intelligent robotics,” said the company in a statement. Reports revealed that to avoid the “uncanny valley” effect and be able to interact with us, they are given lifelike skin and capabilities to read our emotions and respond appropriately using dynamic expression simulation and emotion generation tech. Bionic skin and high-precision control system The Elf V1 series of humanoids features 30 facial muscles animated by brushless micro-motors and managed by a high-precision control system. Paired with an ability to detect their users’ emotions with low latency and bionic skin, their facial expressions are nearly identical to those of humans, reported CGTN. The company claims it’s pioneering the development of realistic humanoid robots designed to revolutionize human-robot interaction. It’s enhancing sophisticated humanoid robot heads that can express emotions, perceive their environment, and interact seamlessly with humans. By combining cutting-edge AI and advanced robotics, AheadForm aims to bring life to machines and transform how humans engage with technology. AI models boost robots’ responsiveness Seamless integration of Large Language Models (LLMs) and Vision-Language Models (VLMs) into the humanoid robots can help them process and interpret complex data from human interactions, enabling the robot to learn and adapt in real-time, achieving human-level understanding and responsiveness. AheadForm uses Brushless Motors that deliver ultra-quiet operation and high responsiveness, specifically designed for precision facial movements in humanoid robots. With its compact size, lightweight design, and energy efficiency, this motor is the ideal choice for next-generation robots that require precise, subtle facial control to create a truly human-like experience. Previously, the company unveiled the Lan Series that features realistic humanoid robots with soft skin and 10 degrees of freedom, offering a lifelike appearance and intuitive movements. This series is designed for cost-efficiency, for applications prioritizing mobility and manipulation.

Owen Gregorian

179,005 Aufrufe • vor 9 Monaten

Type a sentence, get any sound - from talking cats to singing saxophones. Brilliant release by NVIDIA ✨ NVIDIA just unveiled Fugatto, a groundbreaking 2.5B parameter audio AI model that can generate and transform any combination of music, voices, and sounds using text prompts and audio inputs Fugatto could ultimately allow developers and creators to bring sounds to life simply by inputting text prompts, → The model demonstrates unique capabilities like creating hybrid sounds (trumpet barking), changing accents/emotions in voices, and allowing fine-grained control over sound transitions - trained on millions of audio samples using 32 NVIDIA H100 GPUs 👨‍🔧 Architecture Built as a foundational generative transformer model leveraging NVIDIA's previous work in speech modeling and audio understanding. The training process involved creating a specialized blended dataset containing millions of audio samples → ComposableART's Innovation in Audio Control Introduces a novel technique allowing combination of instructions that were only seen separately during training. Users can blend different audio attributes and control their intensity → Temporal Interpolation Capabilities Enables generation of evolving soundscapes with precise control over transitions. Can create dynamic audio sequences like rainstorms fading into birdsong at dawn → Processes both text and audio inputs flexibly, enabling tasks like removing instruments from songs or modifying specific audio characteristics while preserving others → Shows capabilities beyond its training data, creating entirely new sound combinations through interaction between different trained abilities 🔍 Real-world Applications → Allows rapid prototyping of musical ideas, style experimentation, and real-time sound creation during studio sessions → Enables dynamic audio asset generation matching gameplay situations, reducing pre-recorded audio requirements → Can modify voice characteristics for language learning applications, allowing content delivery in familiar voices NVIDIA AI Developer

Rohan Paul

96,354 Aufrufe • vor 1 Jahr

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 Aufrufe • vor 11 Monaten

🚨SPIELBERG'S "DISCLOSURE DAY" TRAILER DROPS: IS HOLLYWOOD PREPARING US FOR SOMETHING REAL? The timing is not a coincidence. Steven Spielberg just released the first trailer for "Disclosure Day," arriving June 12, 2026. Emily Blunt plays what appears to be a news anchor whose voice transforms into an eerie clicking noise during a live weather report. People hooked to machines. Eyes changing. Crop circles forming in real time. One woman morphing into another. The tagline: "If you found out we weren't alone, if someone showed you, proved it to you, would that frighten you? This summer, the truth belongs to seven billion people." Congressional hearings on UAPs. Whistleblowers testifying about "non-human biologics." Pentagon officials using words like "disclosure" in official briefings. And now the director of Close Encounters, War of the Worlds, and E.T. returns with a film literally called "Disclosure Day." Is this entertainment, or the first step in preparing the public for ontological shock? One fascinating observation circulating on X: The alien speech in the trailer uses "sine-wave speech" technique. Sine-wave speech can be conditioned into listeners with the right primer, and subliminal induction has been proven capable of impacting subconscious choices when hidden within audio. The speech in Spielberg's trailer is obscured in sine-wave plus clicking sounds. Conditioning the masses? Or just brilliant sound design? Source: IGN / Brian Roemmele

Mario Nawfal

527,426 Aufrufe • vor 7 Monaten

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

221,026 Aufrufe • vor 3 Monaten

Reinforcement Learning from Human Feedback (RLHF) is gaining traction. This field aims to make AI more responsible by including human values and preferences. In this video, Nathan Lambert, a research scientist and RLHF team lead at Hugging Face explores its inner workings, applications and industry impact. RLHF has gained the spotlight in recent years. The growth of language models like Anthropic’s Claude and OpenAI's ChatGPT have increased interest in human-feedback integration. "There are some rumors that Open AI had two teams; one was doing RLHF and the other instruction fine-tuning. And the RLHF team kept getting more and more performance." Understanding RLHF The RLHF process has three main steps: Pre-training: Much like with GPT models, the journey starts with pre-training on a large corpus of data. This can range from text data, web scrapes, to specialized datasets. Reward Modeling: This is the RLHF counterpart of supervised fine-tuning in large language models. This stage involves creating a reward model that resonates with human values and preferences. RL Optimization: This stage parallels reward modeling and reinforcement learning in traditional AI models. The AI system fine-tunes itself based on the reward model, employing reinforcement learning algorithms for that extra layer of optimization. The Data Challenge Data collection and curation in RLHF closely resemble the challenges you'd encounter in large language model training. Datasets from organizations like OpenAI can serve as a useful foundation. However, the need for high-quality, task-specific data cannot be overstated. Implementing RLHF: A Practical Guide If you’re someone who loves getting hands-on with AI libraries like Hugging Face, implementing RLHF is right way to do. It’s essential to understand its limitations. Think about model stability, over-optimization, and exploration strategies, much like you would when prompt engineering. Ongoing Research and Next Steps While he suggests that some basics figured out, there are layers of complexity that still need to be unraveled: 1. New Benchmarks: How do we measure the effectiveness of RLHF? 2. Preference Modeling: How can the model be made to understand human preferences better? 3. Interpreting RLHF: Much like explainability in traditional models, how do we make RLHF more interpretable? 4. System-Wide Evaluation: Going beyond individual performance, how does RLHF affect an entire system? The Transformative Power of RLHF Whether you're an AI developer, a business analyst, or a marketer, RLHF promises to revolutionize your domain. Imagine customer service chatbots that understand human emotions better, or content generators that align more closely with human values. RLHF is an emerging field that focuses on enhancing machine learning models through human feedback. While it tackles important issues like bias and ethics, its broader goal is to improve system performance across various applications. Whether you're deeply invested in the ethics of AI or simply curious about advancements in machine learning, RLHF offers valuable insights. If you're interested in the next wave of AI development, this area is definitely one to watch.

Muratcan Koylan

27,005 Aufrufe • vor 2 Jahren

Tucker UNLEASHES on AG Pam Bondi with a blistering 3-minute rebuttal to her remarks on “hate speech.” The monologue version of Tucker is back. This is EXACTLY what we need from him right now. “There’s free speech and then there’s hate speech. This is the Attorney General of the United States, the chief law enforcement officer of the United States, telling you that there is this other category called hate speech. And of course, the implication is that’s a crime. There’s almost no sentence that Charlie Kirk. “And I, I’m not running the risk of appropriating his memory for my own ends by saying this. It’s provable. There’s no sentence that Charlie Kirk would have objected to more than that. And you’ve got to think the Attorney General didn’t think it through and was not attempting to desecrate the memory of the person she was purporting to celebrate that she just threw that out there, that she hadn’t thought about it. “You hope that. You hope that Charlie Kirk’s death won’t be used by a group we now call bad actors to create a society that was the opposite of the one he worked to build. You hope that. “You hope that a year from now, the turmoil we’re seeing in the aftermath of his murder won’t be leveraged to bring hate speech laws to this country. And trust me, if it is, if that does happen. There is never a more justified moment for civil disobedience than that, ever. “And there never will be. Because if they can tell you what to say, they’re telling you what to think. There is nothing they can’t do to you because they don’t consider you human. They don’t believe you have a soul. A human being with a soul. A free man has a right to say what he believes, not to hurt other people, but to express human his views. “And by the way, that thinking, and not to pile on the Attorney General, who’s a very nice person, but that thinking that she just articulated on camera there is exactly what got us to a place where some huge and horrifying percentage of young people think it’s okay to shoot people you disagree with to kill Nazis for saying things they don’t like. “Why do they believe that? How did we get here? Is it the video games? Is it the SSRIs? Yeah, probably. But what it really is is 12 and then 16 years of indoctrination in our schools at the hands of people who tell them that, who say exactly what the Attorney General just said. “Well, there’s free speech, which of course, we all acknowledge is important, so, so important. But then there’s this thing called hate speech. Hate speech, of course, is any speech that the people in power hate. But they don’t define it that way. They define it as speech that hurts people, speech that is tantamount to violence. And we punish violence, don’t we? “Of course we do. They’ve been taught that every year of their lives. And so naturally, most of them believe it. When Charlie Kirk is shot in the throat with a .36 on camera. I doubt very many young Americans want to see something like that or actually applaud the death of a man, a father, a husband. “But they’ve been told for their entire lives in schools exactly what Pam Bondi just told them. Well, there’s free speech, but then there’s also hate speech. And woe to those who engage in it because it’s a crime.”

The Vigilant Fox 🦊

262,867 Aufrufe • vor 10 Monaten

NEWS: Figure has unveiling its 3rd generation humanoid robot. • Features a completely redesigned sensory suite and hand system. Wireless charging in feet. • Next-generation vision system engineered for high-frequency visuomotor control. Its new camera architecture delivers twice the frame rate, one-quarter the latency, and a 60% wider field of view per camera within a more compact form factor • Soft goods, wireless charging, improved audio system for voice reasoning, and battery safety advancements • "Engineered from the ground-up for high-volume manufacturing. In order to scale, we established a new supply chain and entirely new process for manufacturing humanoid robots at BotQ." • Lower manufacturing cost • Softer, more adaptive fingertips increase surface contact area, enabling more stable grasps across objects of varied shapes and sizes. Each fingertip sensor can detect forces as small as three grams of pressure - sensitive enough to register the weight of a paperclip resting on your finger. • Multi-density foam to protect against pinch points, and is covered in soft textiles rather than hard machined parts. 9% less mass and less volume than Figure 02 • 10 Gbps mmWave data offload capability, allowing the entire fleet to upload terabytes of data • Upgraded audio hardware system for better real time speech-to-speech. Its speaker is twice the size and nearly four times more powerful, while the microphone has been repositioned for improved performance and clarity. • Charging coils in the robot’s feet allow it to simply step onto a wireless stand and charge at 2 kW Figure: "BotQ is Figure’s dedicated manufacturing facility designed to scale robot production. BotQ’s first-generation manufacturing line will initially be capable of producing up to 12,000 humanoid robots per year, with the goal of producing a total of 100,000 robots over the next four years. Instead of relying on contract manufacturers, Figure brought production of its most critical systems in-house to maintain tight control over quality, iteration, and speed."

Sawyer Merritt

140,241 Aufrufe • vor 10 Monaten

China’s pretty humanoid robot stuns by opening a car door in a ‘world’s first’ | Jijo Malayil, Interesting Engineering Mornine used onboard sensors and full-body control to locate the handle, adjust posture, and open a car door—no human input needed. AiMOGA Robotics has claimed to have reached a significant milestone in embodied AI with its humanoid robot, Mornine, autonomously opening a car door inside a functioning Chery dealership in China. Relying solely on onboard sensors, full-body motion control, and end-to-end reinforcement learning, Mornine performed the task without any human input. Unlike scripted or teleoperated robots, Mornie identified the door handle, adjusted its posture, and used coordinated force across its limbs and torso to complete the action—demonstrating advanced autonomy in a real-world setting. “The deployment marks one of the first instances of a service robot executing such a high-friction, physical interaction in a live commercial setting,” said the firm in a statement. In April, at the Shanghai Auto Show, automotive brands Omoda and Jaecoo, subsidiaries of Chery Automobile, introduced Mornine, designed for use in car dealerships. From sim to service Opening a car door may seem like a simple task, but AiMOGA Robotics views it as a pivotal moment in robotics—signaling a shift from simulation to real-world service, and from basic command execution to autonomous capability. Using only onboard sensors and full-body motion control, Mornine identified the door handle, adjusted her posture, and applied coordinated force across her limbs to open the door—entirely without human intervention. Mornine’s advanced sensor suite includes 3D LiDAR, depth and wide-angle cameras, and a visual-language model (VLM), enabling real-time perception of door position and opening status. Uniquely, Mornine wasn’t explicitly programmed to recognize door handles. Instead, she learned through reinforcement learning, undergoing millions of simulated cycles to focus on the right region and perform the task independently. “We never explicitly told the robot what a door handle is. It learned to focus on that region by itself,” said the engineering team at AiMOGA Robotics in a statement. The learned model was transferred to the real world using Sim2Real methods. Mornine continuously gathers live sensor data during operation, which feeds into a cloud-based training loop, allowing her to improve through continuous learning in real-world settings, reports Robotics Tomorrow. Now active in multiple Chery 4S dealerships in China, Mornine not only opens car doors but also assists with customer greetings, vehicle introductions, and item delivery—marking a step forward in humanoid robotics for commercial retail environments. AI meets retail Originally introduced as the AiMOGA Robot, Mornine was developed to support dealership sales by performing tasks such as explaining vehicle specifications, leading showroom tours, serving refreshments, and engaging with customers in multiple languages. First conceived by Chery as a virtual character to appeal to Generation Z using metaverse and virtual human technologies, Mornine gradually evolved into a real-world interactive humanoid. After multiple iterations of character and model design, Mornine debuted as a digital persona in animations, livestreams, and promotional content, gaining brand recognition. Chery later expanded the concept beyond the virtual space, resulting in the creation of the AiMOGA humanoid robot. Leveraging Chery’s expertise in autonomous driving, environmental sensing, and control systems, AiMOGA features full-stack capabilities in perception, cognition, decision-making, and execution. It uses multimodal sensing—combining speech, vision, and environmental data—to interpret user gestures, commands, and showroom dynamics. A bionic motion system and automotive-grade hardware enable dexterous movement and upright mobility, while multi-robot collaboration allows for coordinated tasks like guided tours. At the decision-making layer, Deepseek’s large language models enable natural language understanding and personalized interaction. In April 2025, Mornine officially began commercial service as an “Intelligent Sales Consultant” at the OMODA C5 JOYSTAR 4S dealership in Kuala Lumpur, Malaysia—marking her full transition from a virtual concept to a real-world humanoid sales assistant.

Owen Gregorian

67,975 Aufrufe • vor 1 Jahr