Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I got a fully maxed out MacBook, mostly so I can run local models fast. And omg here is mixtral running on ollama... almost cannot believe how fast it is. This is with no internet!! A model that beats GPT-3.5 running locally! What!

196,086 Aufrufe • vor 2 Jahren •via X (Twitter)

11 Kommentare

Profilbild von sanket patel
sanket patelvor 2 Jahren

@ollama Will it keep my lap nice and warm

Profilbild von Charlie Holtz
Charlie Holtzvor 2 Jahren

@ollama this is actually how @anotherjesse heats his office in the winter (not even joking)

Profilbild von UNPLUGGED PERFORMANCE
UNPLUGGED PERFORMANCEvor 2 Jahren

Get superior performance with Unplugged Performance's UP-05 forged wheels. Available for all S, 3, X, and Y Tesla Models.

Profilbild von Tim Carambat
Tim Carambatvor 2 Jahren

@ollama Now plug that into AnythingLLM and get the speed + vector database, and a lot more in a single package. Supports any @ollama model! I need to get an M3, my intel with Ollama just ain't cutting it.

Profilbild von Vikr
Vikrvor 2 Jahren

@ollama could you do /set verbose i want know how much tokens/s did M3 got

Profilbild von Charlie Holtz
Charlie Holtzvor 2 Jahren

@ollama 36 t/s!

Profilbild von Josh Marino
Josh Marinovor 2 Jahren

@ollama That just means robots running local neural networks will be here far faster than people expect! Especially since AI chips are advancing faster than printer tech did in the 90’s!

Profilbild von Josh Whiton
Josh Whitonvor 2 Jahren

@DanielleFong @ollama I feel like this is one of the only reasons anyone is getting a fully maxed out MacBook these days.

Profilbild von Mbongeni Ndlovu
Mbongeni Ndlovuvor 2 Jahren

@ollama That’s why I maxed out my MacBook too

Profilbild von Unemployed Capital Allocator
Unemployed Capital Allocatorvor 2 Jahren

@ollama How does it do at ingesting longer prompts? If you put in half the context window as prompt, is it much slower?

Profilbild von Charlie Holtz
Charlie Holtzvor 2 Jahren

@ollama interesting question! just tested with ~2000 token input and it still was very fast, 31 tokens/second:

Ähnliche Videos