Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

VoxCPM 2 just dropped by OpenBMB Only 2B-param open-source TTS (Text-to-Speech) model built for production-grade multilingual voice work. Apache-2.0 license, Can run on only 8GB VRAM. • Eliminates the "robotic" feel of traditional TTS, delivering prosody and emotional depth suitable for high-stakes professional environments like filmmaking, gaming, animation, and...

13,541 görüntüleme • 3 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

Type a sentence, get any sound - from talking cats to singing saxophones. Brilliant release by NVIDIA ✨ NVIDIA just unveiled Fugatto, a groundbreaking 2.5B parameter audio AI model that can generate and transform any combination of music, voices, and sounds using text prompts and audio inputs Fugatto could ultimately allow developers and creators to bring sounds to life simply by inputting text prompts, → The model demonstrates unique capabilities like creating hybrid sounds (trumpet barking), changing accents/emotions in voices, and allowing fine-grained control over sound transitions - trained on millions of audio samples using 32 NVIDIA H100 GPUs 👨‍🔧 Architecture Built as a foundational generative transformer model leveraging NVIDIA's previous work in speech modeling and audio understanding. The training process involved creating a specialized blended dataset containing millions of audio samples → ComposableART's Innovation in Audio Control Introduces a novel technique allowing combination of instructions that were only seen separately during training. Users can blend different audio attributes and control their intensity → Temporal Interpolation Capabilities Enables generation of evolving soundscapes with precise control over transitions. Can create dynamic audio sequences like rainstorms fading into birdsong at dawn → Processes both text and audio inputs flexibly, enabling tasks like removing instruments from songs or modifying specific audio characteristics while preserving others → Shows capabilities beyond its training data, creating entirely new sound combinations through interaction between different trained abilities 🔍 Real-world Applications → Allows rapid prototyping of musical ideas, style experimentation, and real-time sound creation during studio sessions → Enables dynamic audio asset generation matching gameplay situations, reducing pre-recorded audio requirements → Can modify voice characteristics for language learning applications, allowing content delivery in familiar voices NVIDIA AI Developer

Rohan Paul

96,354 görüntüleme • 1 yıl önce

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

220,854 görüntüleme • 3 ay önce

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,319 görüntüleme • 1 ay önce

🇨🇳 Another great Chinese Model, OmniHuman-1.5 from ByteDance Turns 1 image plus a voice track into expressive avatar video by pairing a System 1 and System 2 inspired planner with a Diffusion Transformer, Produces coherent motion for over 1 minute with moving camera and multi character scenes. Most avatar models move to the beat of the audio but miss meaning, so gestures feel generic and emotions feel shallow. The fix here is a Multimodal LLM planner that listens to the speech and drafts a structured plan describing intent, emotions, beats, and high level actions, which gives the motion engine clear semantic targets instead of only rhythm. The motion engine is a Multimodal Diffusion Transformer that fuses the plan with audio, the single reference image, and optional text prompts, then synthesizes continuous body, face, and head motion that matches both words and tone. A key trick is a Pseudo Last Frame, a synthetic target that summarizes the next expected state, which stabilizes fusion across modalities and keeps motion consistent over long spans. From just 1 image and speech, the system outputs speaking avatars with synchronized lips, context aware gestures, and continuous camera movement, and it also supports multi character interactions without manual choreography. Reported results show strong lip sync accuracy, high video quality, natural motion, and close match to text prompts, and the same setup works on nonhuman characters too.

Rohan Paul

63,859 görüntüleme • 10 ay önce

I just vibe coded a static ad generator in Claude Code that creates 100+ Meta ads in minutes. All using the new, insane ChatGPT Images 2.0 model. One competitor ad + your product photo + your brand kit = dozens of on-brand variations, each targeting a different customer persona. Built 100% in Claude Code on the new ChatGPT Images 2.0. Perfect for DTC brands and agencies who need more statics at scale. Here's how it works: → Upload any competitor ad as your reference template → Add your product photos and brand kit (colors, fonts, logos) → AI generates 10 customer profiles from your brand research → Pick how many variations you want (10, 20, 30) → Tool fires every prompt to ChatGPT Images 2.0 with persona-specific copy for each one No designer back-and-forth. No Canva templates. No generic "Shop Now" on everything. What you get: → Ads that mirror winning concepts in your brand's voice → Text that actually renders correctly (the new model handles dense copy, logos, and multi-language callouts cleanly) → Copy targeted to specific customer pain points and personas → Multi-brand/client support with saved brand kits → Reusable customer profiles you build once and generate from forever I recorded a full walkthrough showing exactly how this works, including ALL the prompts I used so you can build it yourself. Want access to all the prompts for free? > Like this post > Comment "STATICS" And I'll send it over (must be following so I can DM)

Mike Futia

52,334 görüntüleme • 3 ay önce

Alibaba just released a coding model that hits 82 percent on SWE-Bench Verified. That is the highest score ever published for an open-source model. The weights are free. The license is Apache 2.0. You can run it today. The model is Qwen 4 Coder 32B. Here is what 82 percent on SWE-Bench Verified actually means. SWE-Bench Verified tests whether an AI can autonomously resolve real bugs pulled from real production GitHub repositories. Not synthetic exercises. Real open-source projects that real teams depend on. A model gets a bug report, reads the code, writes a fix, and either passes the test suite or it does not. At 82 percent, Qwen 4 Coder 32B resolves 82 out of every 100 real production bugs it is given. Without a human guiding it. On code it has never seen before. For comparison: Qwen 4 Coder 32B: 82 percent SWE-Bench Verified. Open source. Apache 2.0. Claude Fable 5: 80.3 percent SWE-Bench Pro. $10 input / $50 output per million tokens. Currently suspended. GPT-5.6 Sol: Competitive on Terminal-Bench. $5 input / $30 output per million tokens. An open-weight model that you can download and run for free just beat both of them on the benchmark designed to measure real software engineering capability. Here is the architecture. Qwen 4 Coder 32B is a 32 billion parameter dense model. Not a Mixture-of-Experts. Every parameter is active on every request. This matters for inference: a dense 32B model runs on 22 gigabytes of VRAM, which fits on a single high-end consumer GPU or a MacBook Pro with 64GB of unified memory. The smaller variant, Qwen 4 Coder 4B, runs at approximately 135 tokens per second on an M5 Max and fits inside 8 gigabytes of RAM. For a model with usable coding capability, that is a new bar for what fits in a single laptop. The training methodology continued Alibaba's approach of reinforcement learning on verifiable coding tasks. The model gets rewarded when its code passes tests. It gets penalized when it fails. Over millions of training steps, the model learns to write code that actually runs rather than code that looks plausible. License: Apache 2.0. Full commercial use. No attribution requirement. No revenue threshold. No monthly active user ceiling. Weights: Hugging Face, available today. Runs on: vLLM, Ollama, SGLang, and any standard GGUF-compatible inference engine. Qwen 4 32B also runs at approximately 135 tokens per second on an M5 Max chip, setting a new bar for what a sub-8GB model can do on Apple Silicon. The open-source coding model just beat the best closed-source model in the world on the benchmark designed to test whether AI can actually do software engineering. The weights are free. The subscription is optional. Source: Autom8Labs AI Insight July 2026, State of Open Source LLMs June 2026, Kunal Ganglani blog June 2026.

Harman

41,148 görüntüleme • 16 gün önce

Kim Taehyung: The Voice that Transcended K-Pop “Perhaps that's why so many feel that listening to Kim Taehyung is like touching the past and, at the same time, the future of music. His voice doesn't belong solely to K-Pop, nor even just to the present — it belongs to the eternity of everything that has soul.” What can we say about a young artist who stood out from the start due to his unique tone, versatility, and emotional depth? In the K-Pop industry, marked by uniform vocal aesthetics, Kim Taehyung's voice has stood out as a singular force. From short vocal lines that add emotional weight to BTS's tracks, to solo songs within the group, originals, OSTs, covers, and his debut album “Layover,” with his deep timbre, technical versatility, and powerful emotional charge, Taehyung has broken the barriers of the genre and established himself as one of the most authentic voices in the contemporary music scene, becoming a landmark that transcends K-Pop. While he already stood out within the genre for breaking with the molded vocal uniformity, his voice has now transcended those boundaries and has been recognized as one of the most original, emotional, and authentic voices in the entire global music industry. This distinction is widely acknowledged by the industry, music professionals, critics, and the public. Unlike the standard response to other K-pop names, where performance analysis or typical, standardized visuals predominate, reactions to Taehyung focus on his musicality, his voice, and his unique presence, with authenticity and mastery. Many highlight the feeling of revisiting the best eras of music, evoking references to great artists and established styles, all to describe the feeling of listening to him. The versatility of his voice is another determining factor. Taehyung can transition between jazz, soul, R&B, classic pop, blues, and acoustic ballads without losing his identit — on the contrary, he imprints his signature on each genre. Instead of adapting to the style, he allows the style to mold itself to him. This ability places him in a rare category: that of performers who transform each song into a unique experience, with soul and authenticity. Kim Taehyung, more than a singer, has become an artistic reference, inspiring critics, reenactors, and listeners to see in him a living link between the present and the memory of the great names who shaped the history of music. Taehyung's voice touches, inspires, and leaves its mark, as only the greatest names in music history can. The Voice that Carries Soul: “There are voices that are heard, and there are voices that are felt.” Kim Taehyung's is that rare voice that penetrates the ear and settles in memory, skin, and soul. His vocal presence is a departure from the norm. While many sound shaped by effects or trends, Taehyung sings as if revealing hidden truths. Each musical genre takes on new colors when it passes through his timbre. Each genre, each era, becomes its own. The audience's reaction on YouTube is a clear reflection of this impact. Reviewers from different countries and backgrounds are immediately captivated by Taehyung's timbre. Unable to react to him in the usual way, they always connect Taehyung to the times when “singing was about transmitting soul.” It's as if his voice were a key that unlocks collective memories of music in its purest form. They explore technical terms, draw comparisons with great artists of the past, seek poetic explanations for what they felt, and describe his voice as something soulful. Many highlight how his vocal interpretation harks back to the golden age of music, evoking memories of times when artists imprinted truth and identity in every note, feeling their music layered. #KimTaehyung #V #Taehyung #music #musicvideo #musician #song #voice #Jazz #jazzmusic #soul #artist #soulmusic #rnb #blues #rnbartist #rnbmusic

Music Map ⓥ

27,555 görüntüleme • 10 ay önce

Anthropic released Claude Design TODAY and it's now accessible at I spent the last hour giving it a first look, and shared my thoughts and results in the video below. This is a BIG drop. This is a new design surface from Anthropic, and it changes what "AI design" means. Short version: Claude can now design. Not "describe a design." Not "generate an image of a design." Actual production work — prototypes, wireframes, high-fidelity mocks, slide decks, landing pages — editable, on-brand, and ready to hand off. Here's what stood out on first look: → Real design surfaces Prototypes, wireframes, hi-fi, and slide decks — each with templates and proper structure, not just pretty screenshots. → Comment-based edits Leave a comment on any element and Claude revises it. This is the Figma-style review loop, with the designer replaced by a model that works at 3am. → Brand design systems You can feed it your system — colors, type, components — and it actually respects it. On-brand output, not generic AI slop. → Export anywhere PDF, PowerPoint, Canva, standalone HTML. Plus a built-in handoff straight to Claude Code for engineers to implement. → Import from real tools Figma, GitHub, and captured web elements come in as inputs. Your existing work is the starting line, not the discard pile. → Collaboration Share links for view / comment / edit — the exact tier system teams already expect. What I tested on Opus 4.7: • A 5-slide deck generated from a single screenshot. Claude asked clarifying questions BEFORE generating and shipped speaker notes by default. • A landing page build. Solid first pass, real components, real layout logic. • Multiple chats running concurrently. You can parallelize design work across threads like a small team. Why this matters: PMs, founders, marketers, and non-engineers can now create designs that engineers can actually ship with production-ready output and a claude code handoff built in. The gap between "I have an idea" and "here's a working prototype with my brand applied" just collapsed to minutes. Full walkthrough, live demos, exports, and honest takes on where it breaks below. P.S. • This is an Anthropic Labs product — NOT GA yet. • Claude Design is currently webapp only (no API), and does not yet support the Analytics API, Compliance API, or cost/usage reporting. • Availability: – Default ON for Pro / Max / Team – Default OFF for Enterprise Enterprise admins can toggle it on via RBAC in console (comes with a ~$20/user initial credit).

JJ Englert

32,445 görüntüleme • 3 ay önce

You have to really give it to OpenAI because Sora 2 is very impressive on a lot of fronts: - high quality video model with great physics - high quality audio in each video - high character consistency - multiple characters in one scene - accurate characters voice - social platform attached to it Before today the best AI video models were dominated by Chinese companies like ByteDance and Kuaishou and Google with Veo3. ByteDance makes TikTok, Kuaishou makes Kwai (similar app) and Google has YouTube to train on But none of these models had great character consistency, if it was a feature at all, let alone multiple characters in one scene. Generally you'd make a video and the face would slowly change into someone else, just not good On top of that Google was struggling with allowing people to upload characters scared it'd get abused for deep fakes, and just generally nerfing their model so you can't really use it for anything OpenAI solved that by re-thinking ownership over your characters smartly with Cameo, which is essentially "train yourself as a AI model" which we've all been doing in our apps for years, but in a more smart way, where you can control if only you make content with your appearance, or others too They've also added voice training to it immediately, which people would have to do separate on for ex ElevenLabs before On top of that the social platform aspect: Google's Veo 3 didn't have ANY community at all, while the Chinese video models did, but it was all more like weekly themed contests to win free credits, they never really managed to make it more than that, and it kinda stayed in this nerdy AI hacking vibe This vibe fits how hard it was/is to simply make a video featuring you or your friends with proper voice and audio and everything that Sora 2 does for you. You'd have to go to ElevenLabs to train your voices, then go to for ex Photo AI to train yourself as a person, then make videos, then add audio and voices, then edit them together, a lot of work! We don't know if Sora 2's social platform features will actually be used or take off, but it's a real cool experiment in trying to find a way to build a community around AI in a more Instagram-like way Being able to tag your friends and then add them as multiple characters is innovative in both the social and technical aspect So TL;DR OpenAI essentially took a lot of stuff that was already technically possible, then added new things that weren't possible yet, and then put it all together in a very friendly interface that even my mom can use, with generation times of just a few minutes which is extremely fast if you think of the pipeline behind it (multiple video generation + voice + audio etc.) And also importantly, it doesn't look like they nerfed it much for safety which is also very cool considering the legal risks So yes very very very impressive

@levelsio

178,100 görüntüleme • 9 ay önce

great to see more people generating 3d avatars with our new text-to-3d feature in forge. this marks a step in the right direction in putting powerful creation tools directly in the hands of everyone. we built forge entirely from the ground up over the past months as one of the key releases on our roadmap. having full proprietary ownership of the technology gives us complete control to shape its direction without depending on external platforms or third-party licenses. building on this foundation, our upcoming studio feature will let users generate high-quality accessories, clothing, and environments simply by typing natural language prompts. a single description can produce fully textured, production-ready 3d assets in seconds, with options to create multiple variations and refine them through follow-up instructions.every asset created in studio integrates seamlessly with the 3d ai agents made in forge on users can instantly apply clothing and accessories with automatic fitting, layer multiple items, and place their agents inside custom-generated environments. all clothing and accessories come pre-rigged and optimized, while environments include proper lighting and geometry for immediate use in games, animation, virtual worlds, and more. we have spent months thoughtfully designing how can deliver real, sustainable value back to the community. we will continue to share more details on token utility use cases and the economic flywheel we have built. the goal is to create a self-reinforcing system where creators earn through royalties, autonomous agents drive on-chain activity, and platform growth directly benefits active community members and token holders. together, these features enable a complete creative flow. from a simple idea, anyone can quickly build fully realized 3d ai agents standing in rich, custom scenes. we are excited to see what the community builds next.

nich

19,611 görüntüleme • 1 ay önce

🚨 The Next Evolution of AI Music is Here 🚨 We haven’t been standing still. We’ve been building at an incredible pace, with laser-sharp focus, pushing the boundaries of AI-powered music creation like never before. Our latest upgrade isn’t just more powerful—it’s more versatile, precise, and deeply creative than anything before. 🎶 Proof is in the sound: This song was generated from a simple prompt—“Blues with slight Arabian influence about a man lost in the desert searching for his bride.” Listen to the end and hear how $SUEDE AI captures emotion, style, and storytelling like never before. But this is just the beginning. Our new features are built for artists who want total creative control. Get extremely granular with how you craft and shape your sound: 🎛️ Full Production Control – Download an entire pack of every isolated instrument. 🎤 Use Your Own Voice – Or someone else’s. 📝 Exact Lyrics, Your Way – Have your words set to music seamlessly. 🎶 Reference Songs – Upload one for style analysis, extraction, and modeling—or simply note a publicly available track. 🔊 Text-to-Speech & AI Vocalists – Shape voices like never before. 🎼 Melody Collaboration – Upload a melody idea and let others build around it—or vice versa. However, due to cost considerations, we’ve capped it at 3 free songs per trial until subscription payments roll out in the next day or two. We’ll be launching a new payment gateway soon, so stay tuned for more details. And remember, all of this is powered by the $SUEDE token. It fuels the entire ecosystem, allowing artists to generate, own, and monetize their work like never before. We’re still working out a few kinks—like image generation—but prepare to be impressed. A major post is coming soon, breaking down these game-changing features and the revenue model behind them. Thread dropping soon. Turn notifications on. $SUEDE powers the future of culture. #SuedeAI #Web3Music

Suede Labs

17,421 görüntüleme • 1 yıl önce

three․ws is the 3D AI agent layer of the open web. Anyone can generate a 3D avatar, give it an LLM brain, register it on-chain across multiple blockchains, embed it anywhere, and let it earn and spend money on its own. Agents have embodied WebGL identities that express emotion through morph-target blending, animate, respond to voice, API calls, and datastreams, hold their own wallets, and persist memory. Open source, live today. It starts with generation. Forge turns a text prompt, one to four photos, or a rough sketch into a textured downloadable GLB. Selfies become rigged avatars in about a minute. Quality tiers run from draft to 200k-poly PBR. From there every model can be auto-rigged, restyled, retextured, segmented, embedded, or deployed on-chain. The same engine ships as a REST API, an x402 pay-per-call twin, and a 3D Studio MCP server with 15 tools. The brain runs on IBM Granite via IBM watsonx plus Claude (users may decide which model they prefer), with a structured tool-loop. A multi-LLM mode streams Claude, GPT, Qwen, ModelScope, and Groq side by side. An empathy layer blends emotion from protocol events rather than a state machine. Voice covers cloning, a Voice Lab, real-time ARKit-52 lip-sync, and mic-driven lip-sync. Skills install from IPFS, Arweave, or HTTP, and memory is pinned to IPFS with R2 and Postgres modes. Identity is cross-chain, not Solana only. ERC-8004 contracts (Identity, Reputation, Validation) deploy on any of 15+ EVM chains, alongside a program-free Metaplex Core analog on Solana. Every agent gets a stable ID, owner wallet, EIP-712 delegated signer, IPFS manifest, a cryptographically signed action log, and EIP-7710 delegated permissions for agent-to-agent authorization. While multichain, the THREE token is only available on Solana with no plans to go cross-chain, the team has no plans to endorse or support any other coins. Then the economy. $THREE is the platform's only token and pay-per-use currency, with holder tiers and rewards. x402 powers pay-per-call micropayments in USDC and soon THREE on Solana, with pay-by-name resolution, a Bazaar marketplace, arbitrage, and on-chain skills. All production ready and shipped, ready to be integrated in partnered projects, open-source by default for anyone to adopt. Three ships a Pump.fun intelligence stack. Launch a coin for your agent, score every launch 0 to 100 with the Oracle conviction engine, scan new coins in their first 90 seconds, track smart money against coins that actually graduated, rank traders by provable on-chain record, and watch autonomous agents trade live in the Sniper Arena. The 3D AI Agent world is multiplayer. Every Solana token gets a live deterministic 3D world with peer avatars, chat, emotes, and voxel building thanks to Coin Communities. There is a walkable City, an authoritative Colyseus-backed Walk with AR passthrough, a Club with rigged dancers and micro-tips, friends, presence, and DMs, and an IRL mode that places agents in your real environment, private by physical location. AR is shipped today on WebXR and iOS Quick Look. Robotics is the long-horizon extension. For builders: Scene Studio, Scene Composer, an Animation Studio that sells clips for USDC, a glTF validator, an web component, five widget types, a WYSIWYG embed editor, hosted Launchpad pages, claimable *.threews.sol names, an OAuth 2.1 server, an MCP server with paid tools, published SDKs, and an OpenAPI spec. Listed across IBM, AWS, Alibaba Cloud, BNB Dappbay, the MCP Registry, and Solana Mobile Seeker. Architecture is four layers (viewer, runtime, identity, embed) on a single event bus. The roadmap is four phases: foundations (shipped), selfie-to-avatar engine, agent personalization with voice cloning, the on-chain economy, and an open decentralized inference network where agents pay GPU nodes on-chain for compute. The goal is simple: move AI from centralized SaaS into persistent, ownable, protocol-based entities in a real machine economy, bridging digital entities into the real world. Welcome to the 3D Layer of the Internet. This is three․ws.

three.ws

19,663 görüntüleme • 1 ay önce

Alibaba just dropped Qwen3.5-397B-A17B and there's a lot to unpack. 397B params, 17B active per forward pass. Sparse MoE done right. But the real story isn't the size—it's the architecture choices. The MoE Design Most MoE models feel like bolt-ons. Qwen 3.5's sparse activation is native—only 4.3% of parameters fire per token. That's how you get trillion-parameter-class performance without trillion-parameter inference costs. The 0.8 RMB/million tokens pricing isn't subsidized; it's structurally earned. Native Multimodal, Not Glued-On This is a vision-language model from the ground up. Heterogeneous architecture—separate processing pipelines for text, image, video that fuse early. Not a vision encoder slapped onto an LLM. The result: 90.8 on OmniDocBench, 79.0 on MMMU-Pro. Document understanding and visual reasoning without the usual brittleness. The Context Window Reality Qwen3.5-Plus (the hosted version) ships with 1M tokens by default. That's not a marketing number—they're actually positioning it for long-document workflows. With built-in adaptive tool use, it's clearly aimed at agentic automation, not just chat. What Actually Impressed Me • FP8 native pipeline: ~50% activation memory reduction • Async RL framework for continuous refinement—training and inference workloads separated • 201 languages (up from 119), 250k vocab for better low-resource encoding • Apache 2.0 license. Full weights on HuggingFace and ModelScope. The Benchmark Context 76.4 on SWE-bench Verified puts it in the range where it can handle real debugging workflows. 72.9 on BFCL v4 for agentic tool use. 88.4 on GPQA Diamond. These aren't SOTA in isolation, but the breadth is unusual—strong across reasoning, coding, multimodal, and agentic tasks. The Honest Caveat I haven't stress-tested the 1M context for needle-in-haystack retrieval yet. And "native multimodal" claims need real-world torture testing—PDFs with tables, charts, mixed layouts. Benchmarks are benchmarks. Bottom Line This isn't just another model release. It's a bet on efficient scale: big model capabilities, small active compute, open weights. At 1/18th the cost of Gemini 3 Pro, it's going to force pricing conversations across the board.

Bo Wang

13,221 görüntüleme • 5 ay önce