
Mike Gannotti
@MichaelGannotti • 51,514 subscribers
💻 Michael Gannotti, North Carolina, Founder the SMF Works Project, Principal AI Solutions Engineer Microsoft. Army Vet, 🚫NO DMs! Opinions=MY OWN
Videos

My Hermes Install Step by Step: 1. Install Ubuntu Linux on a machine – Once installed make sure you have the latest updates 2. Install both Google Chrome and Microsoft Edge browser and log in to your accounts to synch bookmarks/favorites 3. Set up an Ollama account at – I have the annual Pro Plan (if I can ever come up with the funds I will probably spring for the Max plan but Pro is simply awesome) 4. Install Ollama and then run “ollama run glm-5.1:cloud” – It will then have you authenticate to your account and add your machine 5. Install Hermes (watch my video for explanation around this as you may get interrupted during install) 6. Once Hermes is installed you will be prompted to configure. I chose the default quick configure. For model provider select Ollama, provide your Ollama API key, then for model choose desired model. At this point in time I recommend glm-5.1 7. Download Obsidian as your Hermes second brain, set up a vault then tell Hermes to integrate it and provide Hermes the vault location 8. Start building and have fun!
Mike Gannotti108,834 views • 3 months ago

Qwen3.6-35B-A3B-NVFP4 Benchmark: MoE vs. Dense Performance Report This report details a performance comparisonbetween two artificial intelligence models, focusing on the Qwen3.6-35B-A3B-NVFP4 and its predecessor. While both models achieved flawless accuracy across a comprehensive 65-test benchmark covering vision, video, and reasoning, the 35B version demonstrated significant speed advantages. Utilizing a Mixture of Experts architecture, the 35B model achieved nearly four times higher throughput and drastically reduced latency for initial responses. It also showed superior efficiency in speculative decoding and successfully managed massive context lengths up to 128,000 tokens. Ultimately, the source recommends the 35B model for high-speed production environments, while noting the 27B model remains more stable for specific complex reasoning modes.
Mike Gannotti10,457 views • 24 days ago
No more content to load