Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

~ 6.5 - 6.7 t/s for GLM 5.1 on M5 Max 128GB Added “Dense” model export, now model load is only 5s ! Experts are streaming from SSD, so we do not pre-load it. Added direct SSD->Slot memory path, removed prefetch... Many dead end experiments. See Export a “dense-only...

29,421 Aufrufe • vor 4 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Alibaba just released a coding model that hits 82 percent on SWE-Bench Verified. That is the highest score ever published for an open-source model. The weights are free. The license is Apache 2.0. You can run it today. The model is Qwen 4 Coder 32B. Here is what 82 percent on SWE-Bench Verified actually means. SWE-Bench Verified tests whether an AI can autonomously resolve real bugs pulled from real production GitHub repositories. Not synthetic exercises. Real open-source projects that real teams depend on. A model gets a bug report, reads the code, writes a fix, and either passes the test suite or it does not. At 82 percent, Qwen 4 Coder 32B resolves 82 out of every 100 real production bugs it is given. Without a human guiding it. On code it has never seen before. For comparison: Qwen 4 Coder 32B: 82 percent SWE-Bench Verified. Open source. Apache 2.0. Claude Fable 5: 80.3 percent SWE-Bench Pro. $10 input / $50 output per million tokens. Currently suspended. GPT-5.6 Sol: Competitive on Terminal-Bench. $5 input / $30 output per million tokens. An open-weight model that you can download and run for free just beat both of them on the benchmark designed to measure real software engineering capability. Here is the architecture. Qwen 4 Coder 32B is a 32 billion parameter dense model. Not a Mixture-of-Experts. Every parameter is active on every request. This matters for inference: a dense 32B model runs on 22 gigabytes of VRAM, which fits on a single high-end consumer GPU or a MacBook Pro with 64GB of unified memory. The smaller variant, Qwen 4 Coder 4B, runs at approximately 135 tokens per second on an M5 Max and fits inside 8 gigabytes of RAM. For a model with usable coding capability, that is a new bar for what fits in a single laptop. The training methodology continued Alibaba's approach of reinforcement learning on verifiable coding tasks. The model gets rewarded when its code passes tests. It gets penalized when it fails. Over millions of training steps, the model learns to write code that actually runs rather than code that looks plausible. License: Apache 2.0. Full commercial use. No attribution requirement. No revenue threshold. No monthly active user ceiling. Weights: Hugging Face, available today. Runs on: vLLM, Ollama, SGLang, and any standard GGUF-compatible inference engine. Qwen 4 32B also runs at approximately 135 tokens per second on an M5 Max chip, setting a new bar for what a sub-8GB model can do on Apple Silicon. The open-source coding model just beat the best closed-source model in the world on the benchmark designed to test whether AI can actually do software engineering. The weights are free. The subscription is optional. Source: Autom8Labs AI Insight July 2026, State of Open Source LLMs June 2026, Kunal Ganglani blog June 2026.

Harman

41,278 Aufrufe • vor 1 Monat

First impressions on Muse Glimmer! It's incredibly fast for a dense model, currently running an average of 208tps with a max of 274tps on a single 5090 with their DFLASH config. Comparatively, though, both using Open Code, Qwopus Coder (with thinking off) produced a much better shark survival game than the one I got from Glimmer. Meta's new dense model is currently just lacking some HTML canvas taste, but this is something that can be added via SFT as long as the model is stable and capable from a back-end programming perspective. And it seems to be, without a doubt. The big kicker here is that I ran this at extra high thinking, and it did not take long at all to run. Our current local leader, Qwen 27B 3.6, has a tendency to overthink, but with glimmer, that is not the case. Right now, my recommendation for general local programming (Apps, Games, Websites, Visual Tools) in this class is still Qwopus Coder with thinking disabled, or Qwopus Fusion with thinking enabled. Of course Shark Survival is a very basic domain-specific test, but I find that the result scales very well across many domains. If we're going to be shipping apps generated entirely locally, visual taste is somewhat of a bare minimum requirement, solely in my opinion, and Qwen's models in this class offer significantly more at the moment. That's actually why I initially started getting into finetuning with Qwen 3.5, they were the first base that was able to do really good front-end with some opus-trace fine-tuning. Qwen 3.6 has taste even in the base model, and we know Qwen 3.8 is going to blow us all away! Regardless, this looks like a very tempting new base model. As a first offering from Meta in this class for a long time, I am incredibly impressed and elated to have it. We now finally have a proper Single GPU frontier race, instead of us just begging Qwen for more releases. Single GPU open frontier model race is a VERY good thing. Please keep pushing Meta

Kyle Hessling

17,399 Aufrufe • vor 13 Tagen

Responsible AI should not feel like the path of most resistance. That was the strongest idea from my SAS Innovate conversation with Reggie Townsend, leading Data Ethics, Governance, and Social Impact at SAS Software. Reggie framed governance not as a compliance layer, but as a way to scale human judgment. Too often, we innovate first and govern later. Model selected. Agent deployed. Process built. Then governance arrives at the end and feels like friction. Bolt it on after the fact, and resistance is guaranteed. The opportunity: design responsible AI, so it becomes intuitive, action-oriented, and useful in the flow of work. That is what Reggie meant by making responsibility "irresistible." Second point: Use cases must lead. When everyone can access the same models, differentiation will not come from the technology. It will come from how leaders define outcomes, govern applications, and connect business value to institutional trust. The risk does not live in the model. The value does not live in the model. Both live in the use case. For CEOs and boards, this is the shift: from model-first oversight to outcome-first accountability. Better questions before scaling AI: - What human decision are we shaping? - What business outcome are we improving? - What risk are we containing? - What judgment are we extending? Responsible AI becomes strategic when it helps people make better decisions, faster, with greater confidence. Most leaders can't see where governance sits inside their AI operating model. SAS AI Navigator makes it visible: Design for the human. Not only the technology.

Sabine VanderLinden

835,838 Aufrufe • vor 3 Monaten

Here is the first look of the Financial Model for Pakistan Stock Exchange we have built. At this moment, I don't know if it will be successful or not. Yet it's backtesting has given promising results. I believe it is real time and future testing that matters the most. To make it transparent, I will keep on updating on "X" the top picks from the Model. I request you not to invest a single money based on it's results because it is in the early stage. Every feedback from your side would be highly encouraged. The problem was that investors only see MARI, FFC, ,OGDC PPL, MEBL, Luck like companies because they are mainstream. But remember that there are 457 listed companies and there are some hidden gems that are usually missed by retail investors like us. One such example was SSOM that just exploded in few weeks. Thats why, the purpose of the model was to find those companies that have a lot of potential but are not covered by any brokerage house and yet there results are excellent. Based on this model, some time back, I added positions in companies like AGIL (Agri Autos) SPEL (Synthetic Papers ltd), NATF (National Foods), KSB Pumps, Treet. Because scores of all these companies is excellent. They remained stagnant for long time so my conviction on the model was also declining. But it were last two weeks that made a difference and so was my own conviction on this model. Big names like Wealth ⚡️ Wise Value Investor Doctor in PSX Tayyab 🇵🇰 are already aware of it. On Thursday I added HCAR. I will keep you updated just for testing purpose I again repeat not to invest. I will risk my own money. If it is successful, we intend to make a website where you all have access to this data. How it works and how to read it, I will explain it to you in a separate thread.

Shayan Baig

54,727 Aufrufe • vor 1 Jahr