
Mike Gannotti
@MichaelGannotti • 51,514 subscribers
💻 Michael Gannotti, North Carolina, Founder the SMF Works Project, Principal AI Solutions Engineer Microsoft. Army Vet, 🚫NO DMs! Opinions=MY OWN
Videos

Do yourself a favor make your IT admin enable you for Frontier in Microsoft 365 and as soon as it’s available… get Scout!! Trust me on this. Scout is going to be not only your go to for AI. Scout will become your go to for … everything! Microsoft Microsoft365
Mike Gannotti887,458 Aufrufe • vor 1 Monat

My Hermes Install Step by Step: 1. Install Ubuntu Linux on a machine – Once installed make sure you have the latest updates 2. Install both Google Chrome and Microsoft Edge browser and log in to your accounts to synch bookmarks/favorites 3. Set up an Ollama account at – I have the annual Pro Plan (if I can ever come up with the funds I will probably spring for the Max plan but Pro is simply awesome) 4. Install Ollama and then run “ollama run glm-5.1:cloud” – It will then have you authenticate to your account and add your machine 5. Install Hermes (watch my video for explanation around this as you may get interrupted during install) 6. Once Hermes is installed you will be prompted to configure. I chose the default quick configure. For model provider select Ollama, provide your Ollama API key, then for model choose desired model. At this point in time I recommend glm-5.1 7. Download Obsidian as your Hermes second brain, set up a vault then tell Hermes to integrate it and provide Hermes the vault location 8. Start building and have fun!
Mike Gannotti108,834 Aufrufe • vor 3 Monaten

Qwen3.6-35B-A3B-NVFP4 Benchmark: MoE vs. Dense Performance Report This report details a performance comparisonbetween two artificial intelligence models, focusing on the Qwen3.6-35B-A3B-NVFP4 and its predecessor. While both models achieved flawless accuracy across a comprehensive 65-test benchmark covering vision, video, and reasoning, the 35B version demonstrated significant speed advantages. Utilizing a Mixture of Experts architecture, the 35B model achieved nearly four times higher throughput and drastically reduced latency for initial responses. It also showed superior efficiency in speculative decoding and successfully managed massive context lengths up to 128,000 tokens. Ultimately, the source recommends the 35B model for high-speed production environments, while noting the 27B model remains more stable for specific complex reasoning modes.
Mike Gannotti10,457 Aufrufe • vor 24 Tagen
Keine weiteren Inhalte verfügbar