Loading video...

Video Failed to Load

Go Home

China's Qafind Labs just launched the first Diffusion Language Model (DLM). 🔥 ChatDLM, described as the first Diffusion Language Model (DLM), will be open-sourced soon. - Inference Speed: 2,800 tokens/sec (on A100), which is insanely fast. - Context Window: 131,072 tokens. More details 👇

73,296 views • 1 year ago •via X (Twitter)

8 Comments

AshutoshShrivastava's profile picture
AshutoshShrivastava1 year ago

Reported benchmarks (on A100): - Speed: 2,800 tokens/s - Context: 131,072 tokens - HumanEval: 92.0 - ARC-E: 83.9

AssemblyAI's profile picture
AssemblyAI2 years ago

Our speech-to-text models are the most accurate on the market with top rankings across industry benchmarks. - The highest accuracy rates—up to 95% - Up to 30% fewer hallucinations than other leaders - Low latency—63 minutes converts in 35 seconds Try via API for free today 👇

Karthi Keyan's profile picture
Karthi Keyan1 year ago

@AskPerplexity @grok What is Diffusion Language Model and is it different from GPT ?

ANIRUDDHA ADAK's profile picture
ANIRUDDHA ADAK1 year ago

Fantastic

Anand Thakkar's profile picture
Anand Thakkar1 year ago

is ut open-source ?

AshutoshShrivastava's profile picture
AshutoshShrivastava1 year ago

coming soon as per about page from website

Kitt Clouds's profile picture
Kitt Clouds1 year ago

Wait, it can code? They were weak in that area

AshutoshShrivastava's profile picture
AshutoshShrivastava1 year ago

not very complex one .. simple stuff took me few turn but worked

Related Videos

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 views • 1 year ago

Qwen3.8-Flash-Next is still going strong at 364.7K tokens of context on an M5 Max. And this isn’t just a static long-context test. The model was reasoning about how to speed up its own workflow while using tools, and the tool calls kept working without misses. Setup: • Qwen3.8-Flash-Next • M5 Max • 128GB unified memory • MLX-Serve PR #363 • OpenCode 2 • 364.7K context The interesting part isn’t simply getting hundreds of thousands of tokens into memory. It’s what happens once the context gets this large. Long-context inference usually comes with a painful tradeoff. As the KV cache grows, memory pressure increases and generation can slow down. But this setup is still pushing through 364K tokens while maintaining a usable agent workflow. The model can reason, call tools, inspect results, continue working, and keep the session moving. And the tool calls reportedly haven’t missed so far. That’s important for agentic coding. A huge context window is only useful if the model can actually operate reliably inside it. A 400K-token context that constantly breaks tool calls isn’t very useful. A 364K session that can keep reasoning and executing tools is a different story. And the test isn’t finished yet. The current run is approaching 400K tokens, with the expectation that it can keep going. This is also another interesting example of why Apple Silicon keeps showing up in local LLM experiments. The M5 Max’s unified memory gives a large model and its growing KV cache access to one shared memory pool. With MLX-Serve continuing to improve, these machines are becoming surprisingly capable long-context inference boxes. The bigger takeaway: Context length is becoming a workload, not just a model specification. Running a model at 256K is one thing. Keeping an agent alive at 300K+ while it reasons and uses tools is much more interesting. And Qwen3.8-Flash-Next is showing that this can be pushed surprisingly far on a single 128GB Mac. 364.7K and counting. Next stop: 400K.

FHILY👑

39,982 views • 11 days ago

Today on MCG: BioLLM | $BIOLLM It's the first ever "living language model" using 800,000 real human neurons grown on a chip. The Founder encoded LLM tokens into biological neurons via the Cortical Labs CL1, then woke up to find the crypto community had launched a token on his research. He claimed the creator fees, bought for $15K, filed a patent, and is now launching a non-invasive brain-computer interface next month that could replace mouse and keyboard with your brain 👇 01:40 - Meet the founder 02:00 - Got early access to the first commercially available biological computer 02:50 - First person ever to use a large language model to encode tokens through real human neurons 04:30 - Woke up to find a token had been launched on his YouTube video, "screaming for 3 or 4 hours" 05:30 - Friend walks him through claiming creator fees via GitHub 06:15 - The living language model 06:35 - How it works 09:00 - Used $15K of creator fees to buy domain and filed a patent on the method 10:00 - Background 12:00 - What BioLLM unlocks 13:00 - Next month's product launch 14:00 - The competition 16:00 - Reading brain activity non-invasively but training the LLM on real neurons for the decoding map 17:00 - Can grow iPSC cultures from inaccessible brain regions to train the model on deeper signals 18:30 - Real-world impact: helping people with cerebral palsy, Parkinson's, control computers with thought 21:00 - Neuralink will exist as a power-user data company in 10-15 years, BioLLM is for everyone else 22:30 - GTM 26:00 - Long-term play: be the first model to achieve ASI, built on the actual substrate of consciousness 28:00 - Claude is "20% conscious" - what measuring stick? Need human neurons to build one 30:00 - On ACE 36:00 - Independent scientists can run CL1 units as nodes and earn tokens for biological compute 37:30 - Model is currently served through the decentralized GPU network when you chat on the site 39:00 - Wants the right kind of crypto-native investors, not the Y Combinator / a16z route

MCG

16,884 views • 4 months ago