Video wird geladen...
Video konnte nicht geladen werden
Local models just got a free intelligence upgrade. Someone took the new Qwen3.8-27B (the one that already runs in 22GB) and only changed the chat template. No fine-tune. No new training. Just a smarter system prompt that kills all the filler and forces the model to stay sharp. Result:... show more
75,251 Aufrufe • vor 1 Monat •via X (Twitter)
34 Kommentare

Link is here

There is a lot on the table harness side generally speaking in the same vain.

Seems great!

prompt engineering on local quants without structural constraints still leaves reasoning loops open-ended. we added strict AST schema validation in local agent prompts, dropping syntax error retries ~70% instantly. are you enforcing prompt templates in CI or directly inside local agent runners?

Nice, AST validation is solid we enforce Sharp template straight in the agent runners.

no training just a better prompt and now it beats opus

Exactly!

Qwen3.8-27B’s potential is massive, a simple template tweak can unlock its full capability without retraining, that's impressive.👍

Its incredible!

For my RTX 3060:

your post sounds like ai

It is a bad template and is baking in some opinionated system instructions in.

Cambiar solo la plantilla de chat guía mejor el razonamiento sin tocar los pesos, por eso el salto de calidad sale gratis y lo aplicas al instante.

Free quality upgrade.

@grok what’s the reviews from real-world usage regarding this technique?

Closed-source labs spending $100M on fine-tuning, only for a guy on LocalLLaMA to match Opus performance by editing a Jinja file.

I can't get 3.8-27B to run on my MacBook Pro with standard M5 and 24 GB. So I feel saying 22 GB of RAM is enough to run this model is a bit misleading unless you mean additional RAM on top of the RAM used by the system and other services that may be running on your PC.

Results from Sharp are impressive

Qwen 3.8 was already stupidly good locally. Now we’re optimizing the weights without touching the weights

ok but what about us poor vllm users

Nice, same weights asked better is the cheapest upgrade going, a bit like squeezing them 9x lossless with the contents still exact and nothing about the model changing.

I just tested this and not only did I lose 10-15% tok/s, but it took about 50-100% more tokens for the same task. The froggerinc template it's based on is almost as bad. It seems to play a lot more nicely with dflash2 for some reason, though.

how about dropping it in to claude then. lols.

This is why I keep saying the harness matters almost as much as the model.

Exactly!

It did make a huge performance improvement for me so kudos to the author(s)

@grok cooked?

Chat template changes are underrated—they can shift output distribution more than people expect. Curious if you tested it across different tasks or just one benchmark?

wonder if this would work for any Gemma Model as well. Or Glimmer, etc

Harness and prompting still move the needle hard.

hey, remember when certain CEOs got mocked to high heaven for getting all excited about publishing a TEXT FILE? no, I didn't think so..

Prompt and template changes can move behaviour without changing weights. Version the template, benchmark capability and latency, then probe for regressions hidden behind a stronger average response. Related QA analysis:

Its kind of outdated since the 2.2 is out, regardless thank you!

free upgrade with zero training is huge if real
