正在加载视频...
视频加载失败
Prior to release, we shared a version of Cohere North Mini Code with AI engineers and answered some questions. Here's a quick illustrated walkthrough of the model's architecture and training process. Small models fill an important niche. They: 1. run on more widely available hardware 2. handle tasks within... show more
14 条评论

The 30 billion parameter mixture of experts model stacks 49 Transformer blocks, the first of which is dense. The MoE layers have 128 experts, and activate 8 for each token. Leading to 3 billion active parameters. The self-attention setup interleaves sliding window attention and full attention in 3:1 ratio. Our team describes this choice in "Rope to Nope and Back Again: A New Hybrid Attention Strategy"

GGUF quantized version now out kudos to @UnslothAI:

@cohere Love that you're sharing details on the internals like this. This level of detail is something that helps the whole community advance.

@cohere It is great but I wish it was better than qwen 3.6 27b or gemma 4 31b.

@cohere Light models working on sub tasks is smart.

@cohere small models hitting the sweet spot rn. curious how north mini handles tool calling vs something like haiku, that tradeoff usually gets glossed over

@cohere ملهم👏

@cohere love a clean visual breakdown... small models handling specific sub-tasks instead of giant bloated ones is just good system design

@cohere Small models are quietly becoming the backbone of scalable AI systems. 🚀

@cohere I mean you didn't beat Qwen, and I think you should probably focus on that goal, rather than release to release, or try a smaller target and beat the likes of LFM2.5 and their 8B-A1B.

@cohere 话是对的但有一种不会成的宿感在

@cohere dm me :)

@cohere small models, big cap 🥱

@cohere small models really do have their perks super flexible too




