Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I’ve been testing Hy4 preview in WorkBuddy, and the most interesting part is not simply the model size, it’s how much practical work it can handle with a relatively focused active parameter count. Hy4 preview brings together stronger code understanding, generation, and editing; improved document and information processing; workflow...

56,152 Aufrufe • vor 7 Tagen •via X (Twitter)

25 Kommentare

Profilbild von MAUREEN
MAUREENvor 6 Tagen

A capable workflow should support iteration without losing the original goal.

Profilbild von Stones
Stonesvor 6 Tagen

Practical coding support is about solving real problems, not only producing output.

Profilbild von Bonita🤎
Bonita🤎vor 6 Tagen

Code generation becomes more useful when editing is part of the same process.

Profilbild von RICHIE..😎
RICHIE..😎vor 6 Tagen

Developers need tools that can follow the context of an evolving project.

Profilbild von SHERIFF™🤠
SHERIFF™🤠vor 6 Tagen

Real workflow testing reveals more than a benchmark summary ever can.

Profilbild von George
Georgevor 6 Tagen

It helps when a model can contribute across planning, writing, and revision.

Profilbild von JA KISII™ 🇰🇪
JA KISII™ 🇰🇪vor 6 Tagen

Focused active capacity can be valuable for everyday coding tasks.

Profilbild von Tyla
Tylavor 6 Tagen

Developers need systems that can work with context instead of ignoring it.

Profilbild von JOMBA
JOMBAvor 6 Tagen

A useful coding assistant should help maintain momentum across tasks.

Profilbild von CATHEY
CATHEYvor 6 Tagen

Better code support lets teams spend more time on the harder decisions.

Profilbild von Betty
Bettyvor 6 Tagen

The workflow matters as much as the model behind it.

Profilbild von Ochora🇰🇪☆, KC
Ochora🇰🇪☆, KCvor 6 Tagen

Practical usefulness matters more than scale when work needs to get done.

Profilbild von CATHEY
CATHEYvor 6 Tagen

Real productivity comes from fewer interruptions during the building process.

Profilbild von Kiage Clinton, KC
Kiage Clinton, KCvor 6 Tagen

Editing support matters because first drafts rarely stay unchanged.

Profilbild von iano
ianovor 6 Tagen

The best tools help turn a rough idea into something workable.

Profilbild von chenly🥹
chenly🥹vor 6 Tagen

Generating and refining should feel like parts of one process.

Profilbild von Jayden Cole
Jayden Colevor 6 Tagen

Practical performance is what determines whether a model earns repeat use.

Profilbild von Amazing
Amazingvor 6 Tagen

Strong understanding can reduce the back and forth during technical work.

Profilbild von Fatuma
Fatumavor 6 Tagen

Good tools reduce the effort required to move from draft to improvement.

Profilbild von Weigo
Weigovor 6 Tagen

A model should be judged by the quality of work it helps complete.

Profilbild von JADUONG
JADUONGvor 6 Tagen

Focused capability can make a big difference in regular product work. .

Profilbild von JADUONG
JADUONGvor 6 Tagen

Code understanding matters when a task involves more than writing new files.

Profilbild von Alex Ryan
Alex Ryanvor 6 Tagen

Model size only matters if it translates into useful results.

Profilbild von REX JAMES
REX JAMESvor 6 Tagen

A model becomes more helpful when it can revise work thoughtfully.

Profilbild von Saeed Anwar
Saeed Anwarvor 6 Tagen

Fewer rounds and lower token usage matter more than benchmark scores in production agent cost. What tasks did you actually run it on?

Ähnliche Videos

🚨 I just built a game with an open-source AI model. And honestly… I didn’t expect it to be this capable. Tencent Hunyuan just released Hy4 preview, and it’s already pushing into the top tier of open-source models. Three major releases in six months. That pace is crazy. Here’s what Hy4 preview brings: → 770B total parameters → 49B active parameters → 1M+ token context window → Fully open-source But the numbers aren’t even the most interesting part. Hy4 preview was built around one goal: real-world productivity. Coding. Engineering. Office work. Science. Gaming. Finance. Security. And Tencent didn’t build it in isolation. Hy4 preview was co-designed alongside real products like WorkBuddy, using expertise and real-world data from across Tencent’s ecosystem. So I decided to test it the way I actually like testing AI models: I gave it a game idea and let WorkBuddy help turn it into a playable experience. 🎮 From the initial concept to the actual game logic, it was surprisingly smooth. And the benchmark results back up the hype: 163 internal experts 203 engineering tasks Hy4 preview — 2.99/4 Kimi K3 — 2.94/4 GLM 5.3 — 2.92/4 It also beats GLM 5.2 on benchmarks and comes remarkably close to GLM 5.3. Then comes the part I really like: 💰 ¥6/M input tokens 💰 ¥18/M output tokens 💰 ¥0.30/M cache hits Flagship-level capability without the flagship-level price. And right now, you can try Hy4 preview FREE through WorkBuddy for the next two weeks. If you’re curious what it can actually do, don’t just read the benchmarks. Build something with it. 🔗 Tencent Hy Tencent AI

Aryan Rakib

63,055 Aufrufe • vor 6 Tagen

I wanted to see what Hy4 preview could actually do when I gave it a real creative project instead of another “write me some code” prompt. So I went with something I’d genuinely want to play: 🐬 a cute dolphin as the hero 🌊 colorful underwater worlds 💎 glowing pearls to collect 🐠 sea creatures and obstacles 🪸 coral reefs, bubbles & hidden treasure 🎮 smoother animations and progressively harder levels The result is this underwater platformer 👇 And this is where Hy4 preview gets interesting to me. Behind the scenes, we’re talking about 770B parameters, 49B active parameters and a 1M+ token context window. It’s also fully open-sourced. But big numbers only matter if they translate into something useful. What I’m really testing is how much actual work the model can take off my plate understanding the idea, handling the implementation, keeping all the moving parts connected, and getting from a blank canvas to something that feels like a real experience. And there’s another update that makes this even more interesting: Hy4 preview was upgraded yesterday to significantly reduce both conversation turns and token consumption. That means less back-and-forth, faster thinking, and a smoother overall experience when working through complex tasks. That’s also why the cost-effectiveness angle stands out to me. If a model can handle more of the workflow while using fewer tokens and requiring fewer iterations, the value isn’t just in the benchmark score. It’s in how much more you can actually build with it. This dolphin has me wanting to test Hy4 preview with some much crazier ideas next. 🐬👀 Tencent Hy Tencent AI ☞

Md Riyazuddin

63,562 Aufrufe • vor 8 Tagen

I wanted to see where Hy4 preview would stop, so I gave it a deliberately demanding prompt: “Build a working Cyber Incident Command Center.” Not a mockup. A real, responsive app with a live threat map, animated attack routes, incident simulation, filters, investigation panels, event logs, ownership controls, and working isolate/resolve actions. Hy4 preview planned the build, wrote the application, ran it locally, tested the interactions, checked the mobile layout, fixed issues, and opened the finished product for review. Then I clicked “Simulate.” A new incident appeared, the dashboard updated, I filtered it by severity, opened the evidence trail, isolated the affected system, and resolved it. The KPI counters changed with the action. That’s a much more useful coding demo to me than watching code appear in an editor. Hy4 preview is Hunyuan’s third major release in six months. It’s fully open-sourced, with 770B total parameters, 49B activated, and a 1M+ token context window. Tencent built it around productivity scenarios through close work with experts in software engineering, gaming, finance, and security, plus direct co-design with tools like WorkBuddy. In an internal blind test covering 203 engineering tasks and 163 Tencent experts, it scored 2.99/4—slightly ahead of Kimi K3 at 2.94 and GLM 5.3 at 2.92. And the pricing stays practical: $0.834/M input $2.501/M output $0.042/M cache hits If you want to test it yourself, Hy4 preview is free inside WorkBuddy for the next two weeks, for a limited time. Tencent Hy Tencent AI WorkBuddy

Clara Bennett

12,864 Aufrufe • vor 17 Tagen

Qwen3.8-Flash-Next is still going strong at 364.7K tokens of context on an M5 Max. And this isn’t just a static long-context test. The model was reasoning about how to speed up its own workflow while using tools, and the tool calls kept working without misses. Setup: • Qwen3.8-Flash-Next • M5 Max • 128GB unified memory • MLX-Serve PR #363 • OpenCode 2 • 364.7K context The interesting part isn’t simply getting hundreds of thousands of tokens into memory. It’s what happens once the context gets this large. Long-context inference usually comes with a painful tradeoff. As the KV cache grows, memory pressure increases and generation can slow down. But this setup is still pushing through 364K tokens while maintaining a usable agent workflow. The model can reason, call tools, inspect results, continue working, and keep the session moving. And the tool calls reportedly haven’t missed so far. That’s important for agentic coding. A huge context window is only useful if the model can actually operate reliably inside it. A 400K-token context that constantly breaks tool calls isn’t very useful. A 364K session that can keep reasoning and executing tools is a different story. And the test isn’t finished yet. The current run is approaching 400K tokens, with the expectation that it can keep going. This is also another interesting example of why Apple Silicon keeps showing up in local LLM experiments. The M5 Max’s unified memory gives a large model and its growing KV cache access to one shared memory pool. With MLX-Serve continuing to improve, these machines are becoming surprisingly capable long-context inference boxes. The bigger takeaway: Context length is becoming a workload, not just a model specification. Running a model at 256K is one thing. Keeping an agent alive at 300K+ while it reasons and uses tools is much more interesting. And Qwen3.8-Flash-Next is showing that this can be pushed surprisingly far on a single 128GB Mac. 364.7K and counting. Next stop: 400K.

FHILY👑

39,982 Aufrufe • vor 10 Tagen

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 Aufrufe • vor 3 Monaten

ChatGPT o1 is the most "intelligent" AI model and it's not even close! Full o1 generates thinking steps ~50 faster than preview. It's more accurate, reliable, and got better on harder tasks that require advanced reasoning and knowledge. I ran a few tests on it already. Here are my observations: Full video with examples & explanations: Strengths - impressive at math, code, and knowledge-intensive tasks. Weakness - it only failed on a cross-word puzzle but I think it might be solvable when a web search becomes available. In the end, while very efficient with complex knowledge use, it's still constrained by data it's trained on. Speed - the thinking steps are generated a lot faster! Not a fair comparison with the open alternatives but I think this improves the overall user experience. "Knowledgeable and highly intelligent" - as mentioned in the demo by OpenAI researchers, o1 is great at dealing with ambiguity and filling in knowledge gaps. I was impressed by how it implemented an agentic solution (with lots of details) from a basic diagram of architecture (with minimal details). Check out the sample video. Better Task Coverage - Due to the speed and the ability to make sense of instructions and intent (i.e., know when to response fast and when to "think" deeply) much better, it feels like it might be more useful for a broader range of tasks. Image understanding - the image understanding capability is mysterious (often leads to faster responses but no thinking) but impressive. More experiments and notes soon. Stay tuned!

elvis

102,643 Aufrufe • vor 1 Jahr