Video wird geladen...
Video konnte nicht geladen werden
I got a fully maxed out MacBook, mostly so I can run local models fast. And omg here is mixtral running on ollama... almost cannot believe how fast it is. This is with no internet!! A model that beats GPT-3.5 running locally! What!
196,086 Aufrufe • vor 2 Jahren •via X (Twitter)
11 Kommentare

@ollama Will it keep my lap nice and warm

@ollama this is actually how @anotherjesse heats his office in the winter (not even joking)

Get superior performance with Unplugged Performance's UP-05 forged wheels. Available for all S, 3, X, and Y Tesla Models.

@ollama Now plug that into AnythingLLM and get the speed + vector database, and a lot more in a single package. Supports any @ollama model! I need to get an M3, my intel with Ollama just ain't cutting it.

@ollama could you do /set verbose i want know how much tokens/s did M3 got

@ollama 36 t/s!

@ollama That just means robots running local neural networks will be here far faster than people expect! Especially since AI chips are advancing faster than printer tech did in the 90’s!

@DanielleFong @ollama I feel like this is one of the only reasons anyone is getting a fully maxed out MacBook these days.

@ollama That’s why I maxed out my MacBook too

@ollama How does it do at ingesting longer prompts? If you put in half the context window as prompt, is it much slower?

@ollama interesting question! just tested with ~2000 token input and it still was very fast, 31 tokens/second:
