Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Local models just got a free intelligence upgrade. Someone took the new Qwen3.8-27B (the one that already runs in 22GB) and only changed the chat template. No fine-tune. No new training. Just a smarter system prompt that kills all the filler and forces the model to stay sharp. Result:...

75,251 Aufrufe • vor 1 Monat •via X (Twitter)

34 Kommentare

Profilbild von Harman
Harmanvor 1 Monat

Link is here

Profilbild von Systems Pimp 🎩🦯🕶️
Systems Pimp 🎩🦯🕶️vor 1 Monat

There is a lot on the table harness side generally speaking in the same vain.

Profilbild von Harman
Harmanvor 1 Monat

Seems great!

Profilbild von Camaleón Raro
Camaleón Rarovor 1 Monat

prompt engineering on local quants without structural constraints still leaves reasoning loops open-ended. we added strict AST schema validation in local agent prompts, dropping syntax error retries ~70% instantly. are you enforcing prompt templates in CI or directly inside local agent runners?

Profilbild von Harman
Harmanvor 1 Monat

Nice, AST validation is solid we enforce Sharp template straight in the agent runners.

Profilbild von Aria Tech
Aria Techvor 1 Monat

no training just a better prompt and now it beats opus

Profilbild von Harman
Harmanvor 1 Monat

Exactly!

Profilbild von BayesChat
BayesChatvor 1 Monat

Qwen3.8-27B’s potential is massive, a simple template tweak can unlock its full capability without retraining, that's impressive.👍

Profilbild von Harman
Harmanvor 1 Monat

Its incredible!

Profilbild von Johannes Steen
Johannes Steenvor 1 Monat

For my RTX 3060:

Profilbild von covfefe ali
covfefe alivor 1 Monat

your post sounds like ai

Profilbild von 0xE  @enol.app
0xE  @enol.appvor 1 Monat

It is a bad template and is baking in some opinionated system instructions in.

Profilbild von Buswe
Buswevor 1 Monat

Cambiar solo la plantilla de chat guía mejor el razonamiento sin tocar los pesos, por eso el salto de calidad sale gratis y lo aplicas al instante.

Profilbild von Harman
Harmanvor 1 Monat

Free quality upgrade.

Profilbild von MJ 🇯🇲 🇺🇸 🇯🇵
MJ 🇯🇲 🇺🇸 🇯🇵vor 1 Monat

@grok what’s the reviews from real-world usage regarding this technique?

Profilbild von Internet Labs
Internet Labsvor 1 Monat

Closed-source labs spending $100M on fine-tuning, only for a guy on LocalLLaMA to match Opus performance by editing a Jinja file.

Profilbild von Bughunter Geek
Bughunter Geekvor 1 Monat

I can't get 3.8-27B to run on my MacBook Pro with standard M5 and 24 GB. So I feel saying 22 GB of RAM is enough to run this model is a bit misleading unless you mean additional RAM on top of the RAM used by the system and other services that may be running on your PC.

Profilbild von Peace
Peacevor 1 Monat

Results from Sharp are impressive

Profilbild von Seb
Sebvor 1 Monat

Qwen 3.8 was already stupidly good locally. Now we’re optimizing the weights without touching the weights

Profilbild von Nic Wienandt - mtecnic
Nic Wienandt - mtecnicvor 1 Monat

ok but what about us poor vllm users

Profilbild von Forlais
Forlaisvor 1 Monat

Nice, same weights asked better is the cheapest upgrade going, a bit like squeezing them 9x lossless with the contents still exact and nothing about the model changing.

Profilbild von Shaun Thomas
Shaun Thomasvor 1 Monat

I just tested this and not only did I lose 10-15% tok/s, but it took about 50-100% more tokens for the same task. The froggerinc template it's based on is almost as bad. It seems to play a lot more nicely with dflash2 for some reason, though.

Profilbild von Lindsay Rex
Lindsay Rexvor 1 Monat

how about dropping it in to claude then. lols.

Profilbild von Saman Ahmed
Saman Ahmedvor 1 Monat

This is why I keep saying the harness matters almost as much as the model.

Profilbild von Harman
Harmanvor 1 Monat

Exactly!

Profilbild von BentleyPC
BentleyPCvor 1 Monat

It did make a huge performance improvement for me so kudos to the author(s)

Profilbild von Elias Sandburg
Elias Sandburgvor 1 Monat

@grok cooked?

Profilbild von Modelplane
Modelplanevor 1 Monat

Chat template changes are underrated—they can shift output distribution more than people expect. Curious if you tested it across different tasks or just one benchmark?

Profilbild von Glenn Hernandez
Glenn Hernandezvor 1 Monat

wonder if this would work for any Gemma Model as well. Or Glimmer, etc

Profilbild von Sophia Data Queen
Sophia Data Queenvor 1 Monat

Harness and prompting still move the needle hard.

Profilbild von 0xagonally
0xagonallyvor 1 Monat

hey, remember when certain CEOs got mocked to high heaven for getting all excited about publishing a TEXT FILE? no, I didn't think so..

Profilbild von Mitchell Agoma
Mitchell Agomavor 1 Monat

Prompt and template changes can move behaviour without changing weights. Version the template, benchmark capability and latency, then probe for regressions hidden behind a stronger average response. Related QA analysis:

Profilbild von NullFoundry
NullFoundryvor 1 Monat

Its kind of outdated since the 2.2 is out, regardless thank you!

Profilbild von AI Mastery Guide
AI Mastery Guidevor 1 Monat

free upgrade with zero training is huge if real

Ähnliche Videos

There is no best model. There's a lot of noise about models right now. Who is training them, who owns them, where legal intelligence should live. One question actually matters: what produces the best outcome for the legal task in front of you? That's how we decide things at Legora. We optimize for the end-to-end outcome on a legal task. The model is one layer of that system, not the system. Models are uneven and the frontier changes almost weekly. One model plans a long job well, another runs deep analysis across thousands of documents. Some have to be told exactly what to do, and some are fine with a vague brief. They all break in different ways. So our lawyers write evals and we test them with the Legora BAR, our benchmark for agentic reasoning. Every model takes every test, and the model that wins gets the work. We post-train when we know it buys our customers better performance on a specialized task. Training is a tool we reach for when it helps, nothing more than that. The intelligence that compounds sits in the orchestration layer. Precedents, review standards, client requirements. That knowledge has to stay editable, auditable and portable. In our system, a changed review standard is an edit that takes effect the same day, with no new model training required. No lawyer should have to worry about which model did the work, any more than they think about which chip is in their laptop. They should only care about the quality of the work. That's what we are focused on. If you want the engineering version of this argument rather than the CEO version, our CPO, Bryan Tsao, and CTO, Jacob Lauritzen, take it apart in the video below.

Max Junestrand

49,515 Aufrufe • vor 22 Tagen

Chinese AI models are wiping billions off Big Tech right now. Google just lost $200 billion in a single day, and the model it needed to fight back still isn't ready. Gemini 3.5 Pro, Google's most powerful model, is months behind schedule. Alphabet stock dropped 4.4% that same day. The Deepseek moment is happening again, and the new model is FAR bigger. On the same day Google's delay leaked, a Beijing lab called Moonshot released Kimi K3. It is the largest open model ever built, with 2.8 trillion parameters. It took the number one spot on the Frontend Code Arena, a live coding leaderboard, passing Anthropic's best model. And Moonshot is giving it away for free on July 27. The genius part: Anyone with enough computers can download it and run a frontier level AI without paying a cent to a US company. A single task on Kimi K3 costs about 94 cents. The same work on some American models costs nearly double. So why would a company keep paying premium prices for a model it can now get for free? The entire US AI business is built on selling access to models that cost billions to train. If a free Chinese version does most of the same work, that pricing power starts to crack. And Kimi is close to the best. On one closely watched intelligence ranking it scored 57, just behind the top American models GPT-5.6 Sol and Fable 5, and ahead of Claude Opus 4.8. Bank of America told clients that Kimi proves Chinese labs can keep making big leaps even with limited chips. And the founder of Moonshot, Yang Zhilin, learned to build AI as a researcher INSIDE Google. Google literally wrote the 2017 paper that made all of these models possible. Now the people who studied its work are using it to destroy Google, and handing it out for free. What happens next: Kimi K3's weights go public on July 27. Google reports earnings on July 22, and everyone will be asking the same question about Gemini. If free models keep topping the charts, every valuation built on paid AI access has to be rewritten. What do you think?

Ricardo

47,790 Aufrufe • vor 2 Monaten

elon musk grabbed the source code openai open-sourced by accident, rewrote it in rust over a weekend, and shipped it as a free coding agent that does everything $200/mo chatgpt pro does. why pay $200 to openai and $200 to claude when this runs for $8 the swarm above is one weekend of exactly that: thousands of agents pouring through four endpoints, three paid seats billing $1.80 a task while the free fork bills $0. musk co-founded openai, walked out, and when they left codex on github under a permissive license, he forked it, stamped grok on it, and gave it away what the free version does that the $200 seat charges for: the agent · openai's own engine -> it reads your repo, writes patches, runs your tests, and loops until they pass, exactly like codex -> because under the hood it is codex, just faster and free. you are paying $200 for the paid skin of a tool now sitting on github the license · apache-2.0, un-revocable -> free to use, free to fork, free to ship inside your own product with zero strings -> openai cannot pull it back. musk made sure the license is the kind that never expires the switch · one line, no new tools -> point it at any openai-compatible or claude-compatible endpoint, including an $8 kimi backend -> same terminal, same workflow, gpt-5.6 and opus 5 just quietly lose the seat the bill · $400 down to $8 -> chatgpt pro plus claude max is $400 a month. the free agent plus an $8 kimi key does the same daily work -> that is a 98% cut, built out of openai's own source code, handed to you by the guy suing them here is the part they will fight me on: openai did not lose this to a better model, they lost it to their own license and an enemy with a weekend free. the $200 was never the tool, it was the toll, and musk just put openai's own logo on the road around it drop your $400/mo ai stack to $8. the run above is openai's own agent, rewritten free, doing the job it bills $200 a month for. the full breakdown is in the article below

starmex

111,684 Aufrufe • vor 1 Monat

no money for grok or midjourney? this tool is for you. there's a FREE tool created by an anon dev. open-source. runs locally. 117k stars on github. it generates: > images & video > 3d models > audio > 20+ models here's how to set it up in under 5 minutes: 1️⃣download ComfyUI Desktop go to and grab the desktop app for your system. windows 10+, mac (apple silicon), or linux. it installs like any normal app, it sets up python and every dependency for you in the background. no terminal, no config files. 2️⃣open it first launch, it spins up its own environment automatically. you just wait a few seconds and you're in. you'll land on a node canvas, that's the whole interface. 3️⃣load a starter workflow top menu → Workflow → Browse Templates → Image Generation. click it. this drops a ready-made setup onto your canvas so you don't build anything from scratch. 4️⃣grab a model comfyui ships empty on purpose, the model is the brain, and you pick it. in the template, the "Load Checkpoint" node has a Download button when no model is installed. click it. it pulls one in for you (a few GB, this is the only real wait). 5️⃣install ComfyUI Manager this is the one add-on you don't skip. it lets you install models, custom nodes, and updates with a click instead of the command line. grab it from github (link in comments). it's the difference between fighting comfyui and flying in it. one honest note: an NVIDIA gpu makes this fast, apple silicon works great too, and a weak machine still runs it just slower. that's the whole setup. you now own an image, video, and 3D studio that costs you nothing per month. save this. and the next time grok or midjourney asks for your card. you won't need it. disclaimer: comfyui itself is 100% free. so are the local models (sdxl, flux, wan 2.2, ltx-2). some premium models like seedance are pay-per-use api models, only if you want top-tier quality. the free local ones cover most of what you need. (github link in the comments) follow and turn on post notification for daily AI contents.

m0h

14,542 Aufrufe • vor 3 Monaten

-> someone cloned claude -> design interface and -> made it completely free -> it's work on YouTube -> and also suitable for kids -> it’s called open design -> and it’s live on github -> same clean split-screen ui -> you get in claude artifacts -> prompt on the left, live -> design/code preview on -> the right, type what you -> want to build and it -> generates the ui in real -> time, but here’s the twist -> you pick the ai model -> not locked into one -> company, want to use -> gemini, mistral, llama, -> deepseek any model -> with an api work -> if you’re running local -> models with ollama -> that works too -> no subscription walls -> the big difference -> vs claude artifacts -> works with any free -> ai model you’re not -> paying $20/mo just to -> design, use free tiers -> local models, or whatever -> you already have access to -> fully local, your prompts -> and code never leave -> your machine unless -> you want them to -> no data training -> no cloud storage -> privacy by default -> no usage limits -> claude cuts you off -> after a few designs -> here you can generate, -> iterate, break things -> and rebuild all day -> the only limit is your -> don’t like how a button -> works, change it -> want to add your own -> components, go ahead -> you own the tool -> so if you’ve been gatekept -> by paywalls or worried -> about sensitive prompts -> going to some company’s -> servers, this fixes that. -> same workflow, more -> control, zero monthly fee

BeingInvested

12,134 Aufrufe • vor 4 Monaten