Загрузка видео...

Не удалось загрузить видео

На главную

Someone got DeepSeek-R1-0528-Qwen3-8B running on an iPhone 16 Pro using MLX. It runs but takes ages to respond, and the phone gets hot fast. 8B models on phones aren't sci-fi anymore. via u/adrgrondin on r/LocalLLaMA

82,742 просмотров • 1 год назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

i wish this wasn’t real, but bro!! this was my experience just this morning, this was someone i took on a date last week of december, we were meant to meet up again, but she had to leave for service(in the east) on her arrival we couldn’t talk, i texted her on snap(see first media) but she didn’t reply, two days later i texted her on whatsapp and the message didn’t tick twice, it was then i realized something is wrong, i had to call her, she explained to that her phone got bad on her arrival, she asked me for financial assistance, i sent her the money to fix the phone, days later she didn’t mention anything about the phone, then i had to ask her during one of our conversation, it was then she told me that “it’s a panel issue, it cant be fixed anymore” i sympathized with her then continued our conversation, days later she asked for another money for the same phone fixing, i explained my current situation to her, i told her about the expenses on me and i can’t spend anymore money out of budget, she grumbled and sighed, i had to apologize to her lol, just two days back i noticed she reposted a post on her ticktock and the message on whatsapp finally ticked twice, i asked her on call if she has fixed the phone, she said no, saying that the money i sent won’t be enough since it’s panel issue, i said okay,( though i was suspicious but i didn’t make a big deal out of it) lo and behold i came across her tiktok page with a new post(content) i asked her about the phone and this was her response(last media), i deleted her number and blocked her.💔💔💔
0:10

Sensitive content

i wish this wasn’t real, but bro!! this was my experience just this morning, this was someone i took on a date last week of december, we were meant to meet up again, but she had to leave for service(in the east) on her arrival we couldn’t talk, i texted her on snap(see first media) but she didn’t reply, two days later i texted her on whatsapp and the message didn’t tick twice, it was then i realized something is wrong, i had to call her, she explained to that her phone got bad on her arrival, she asked me for financial assistance, i sent her the money to fix the phone, days later she didn’t mention anything about the phone, then i had to ask her during one of our conversation, it was then she told me that “it’s a panel issue, it cant be fixed anymore” i sympathized with her then continued our conversation, days later she asked for another money for the same phone fixing, i explained my current situation to her, i told her about the expenses on me and i can’t spend anymore money out of budget, she grumbled and sighed, i had to apologize to her lol, just two days back i noticed she reposted a post on her ticktock and the message on whatsapp finally ticked twice, i asked her on call if she has fixed the phone, she said no, saying that the money i sent won’t be enough since it’s panel issue, i said okay,( though i was suspicious but i didn’t make a big deal out of it) lo and behold i came across her tiktok page with a new post(content) i asked her about the phone and this was her response(last media), i deleted her number and blocked her.💔💔💔

Ishaaka

706,530 просмотров • 7 месяцев назад

Big win for open-source LLMs! DeepSeek V4 Pro holds the top open-weights score on SWE-bench Verified, in the GPT-5.5 range. GLM 5.2 leads the open-weight intelligence index and sits near the closed frontier on long-horizon coding. But this leaderboard number is a weak proxy for real performance. It comes from one task set, run through one harness, served at one precision. The same weights can even score differently across providers, since many hosts quantize activations to fp8 and drift the model off its reference weights. Real performance is determined based on whether a model can read a repo, make coordinated edits across files, run the tests, and recover when one breaks. By that measure, the top open models hold up, but only inside the right harness. The teams that actually put DeepSeek V4 into production pipelines as a frontier substitute got there through the harness they built around the model, not by picking a stronger model. If you want to see this in practice, Cline (64k+ stars) has actually built that harness around open models, tuned so they run at production quality. And it's tuned so that these LLMs can run at production quality, with plan and act modes, checkpoints, and terminal feedback. ClinePass is the new access layer on top of it. It runs a curated set of those models inside Cline, narrowed to the ones tested for coding-agent use, with 2 to 5x the standard rate limits and no separate provider accounts, keys, or billing to track. The video below shows the setup, and I worked with the team to put this together. It runs alongside custom keys and local models as well, not in place of them.

Avi Chawla

44,124 просмотров • 1 месяц назад

Holy shit... Microsoft open sourced an inference framework that runs a 100B parameter LLM on a single CPU. It's called BitNet. And it does what was supposed to be impossible. No GPU. No cloud. No $10K hardware setup. Just your laptop running a 100-billion parameter model at human reading speed. Here's how it works: Every other LLM stores weights in 32-bit or 16-bit floats. BitNet uses 1.58 bits. Weights are ternary just -1, 0, or +1. That's it. No floats. No expensive matrix math. Pure integer operations your CPU was already built for. The result: - 100B model runs on a single CPU at 5-7 tokens/second - 2.37x to 6.17x faster than llama.cpp on x86 - 82% lower energy consumption on x86 CPUs - 1.37x to 5.07x speedup on ARM (your MacBook) - Memory drops by 16-32x vs full-precision models The wildest part: Accuracy barely moves. BitNet b1.58 2B4T their flagship model was trained on 4 trillion tokens and benchmarks competitively against full-precision models of the same size. The quantization isn't destroying quality. It's just removing the bloat. What this actually means: - Run AI completely offline. Your data never leaves your machine - Deploy LLMs on phones, IoT devices, edge hardware - No more cloud API bills for inference - AI in regions with no reliable internet The model supports ARM and x86. Works on your MacBook, your Linux box, your Windows machine. 27.4K GitHub stars. 2.2K forks. Built by Microsoft Research. 100% Open Source. MIT License.

Guri Singh

2,180,357 просмотров • 5 месяцев назад