Video wird geladen...
Video konnte nicht geladen werden
Prior to release, we shared a version of Cohere North Mini Code with AI engineers and answered some questions. Here's a quick illustrated walkthrough of the model's architecture and training process. Small models fill an important niche. They: 1. run on more widely available hardware 2. handle tasks within... show more
89,469 Aufrufe • vor 3 Monaten •via X (Twitter)
14 Kommentare

The 30 billion parameter mixture of experts model stacks 49 Transformer blocks, the first of which is dense. The MoE layers have 128 experts, and activate 8 for each token. Leading to 3 billion active parameters. The self-attention setup interleaves sliding window attention and full attention in 3:1 ratio. Our team describes this choice in "Rope to Nope and Back Again: A New Hybrid Attention Strategy"

GGUF quantized version now out kudos to @UnslothAI:

@cohere Love that you're sharing details on the internals like this. This level of detail is something that helps the whole community advance.

@cohere It is great but I wish it was better than qwen 3.6 27b or gemma 4 31b.

@cohere Light models working on sub tasks is smart.

@cohere small models hitting the sweet spot rn. curious how north mini handles tool calling vs something like haiku, that tradeoff usually gets glossed over

@cohere ملهم👏

@cohere love a clean visual breakdown... small models handling specific sub-tasks instead of giant bloated ones is just good system design

@cohere Small models are quietly becoming the backbone of scalable AI systems. 🚀

@cohere I mean you didn't beat Qwen, and I think you should probably focus on that goal, rather than release to release, or try a smaller target and beat the likes of LFM2.5 and their 8B-A1B.

@cohere 话是对的但有一种不会成的宿感在

@cohere dm me :)

@cohere small models, big cap 🥱

@cohere small models really do have their perks super flexible too




