Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I got Llama 3 running in my browser using only my GPU with my Wi-Fi switched OFF completely client-side WebGPU is a new feature in browsers where JS can use the GPU of the device and apparently you can run LLMs on it too, and it's fast! My end...

261,455 Aufrufe • vor 2 Jahren •via X (Twitter)

9 Kommentare

Profilbild von @levelsio
@levelsiovor 2 Jahren

One way to deal with the limitations of having to download model - let user use cloud model but show an indicator ☁️ Cloud Based Conversation or smth then once it's locally ran switch to ✅ 100% private local conversation - if user GPU or connection doesn't allow for locally run LLM give them the cloud one

Profilbild von Mario Hachemer
Mario Hachemervor 2 Jahren

I don't think users will be satisfied with the quality of even 7b models for that use case. Local compute is techie-brain thinking, no b2c product has ever won because of privacy advantages.

Profilbild von @levelsio
@levelsiovor 2 Jahren

Yep but there's 70B too, issue is how do you get that 40GB file to the user

Profilbild von @levelsio
@levelsiovor 2 Jahren

LM Studio is great but unusable for non-techies, it's still very hard You have to download it, install it, go to the right tab, search the model, choose the right quantization (this is insane btw, why doesn't it choose the best for me) It's very engineer minded and not user friendly

Profilbild von Xenova
Xenovavor 2 Jahren

While WebGPU-powered LLMs are amazing for demos (and to show what's possible in a browser context), I think the most exciting applications will use smaller task-specific models (especially for vision and multimodal tasks). For example: - Background removal: - Image segmentation (SAM):

Profilbild von Francisco Cordoba - Francisco Money
Francisco Cordoba - Francisco Moneyvor 2 Jahren

Unless your users are tech savy and super concern about privacy they are not going to download a file.

Profilbild von @levelsio
@levelsiovor 2 Jahren

It'll be automatic ofc

Profilbild von hugo alves
hugo alvesvor 2 Jahren

Still don’t understand why you named your product “the rapist ai”

Profilbild von Fleetwood
Fleetwoodvor 2 Jahren

On-device is 100% the future - particularly for private use cases like this. Combine this with TTS e.g so that users can speak & build the rapport. HF has been working hard on web + on-device with Transformers.JS & my project:

Ähnliche Videos

BREAKING: GPT-5.6 Sol is out—AND Codex has been merged into ChatGPT Desktop as ChatGPT Codex. This combo model and desktop app harness are the gold-standard for knowledge work in AI. 5.6 is powerful, fast, half the price of Fable, and my default for almost everything. We’ve been testing it internally Every 🪨 for about a month across coding, writing, design, and knowledge work. Here’s our day-zero vibe check: - An A-tier coder—but it’s not Fable. Sol scored 56/100 on our Senior Engineer benchmark compared to a 91 for Fable. I think the 56/100 undersells it, it's an excellent implementor, and very smart. But Fable just writes conceptually cleaner code and works better at the top end of task complexity. PRO-TIP: Use GPT-5.6 as Fable's subagent for the most goated combo in AI coding. - The best writer of the frontier models. It’s clearer and more concise than Fable or Opus 4.8, without the overexplaining or weird private language. It can one-shot marketing emails, help you workshop taglines, and explain complex concepts clearly. It's also super fast, which makes it easy to collaborate with. - Design is better, but not top-tier. It has noticeably more taste than 5.5, but Fable and Opus 4.8 are still playing at a different level. See examples in the video and vibe check below. - The real leap is knowledge work. Sol is the first model I’ve trusted to run whole loops of knowledge work—not just help with individual tasks. I use it to process email, surface decisions from meetings and Slack, find job candidates, scan Facebook Marketplace for furniture, and log my meals. It has shifted my job from doing the work to tending the system that does it. - The merged app is fine. I was extremely worried about this because I love the Codex app. OpenAI was caught in an interesting position: How to make an agent orchestration app for regular ChatGPT consumers, coders, and businesses all in one app. They now split the interface between ChatGPT Work and ChatGPT Codex. They're basically the same except Work hides code. And "Chat" has been demoted to 2nd tier status for quick questions in either one. It's not a big leap, but it's not a huge setback either. And it remains my favorite of the desktop agent orchestration apps. Verdict: If I really had to put my finger on it, I'd say Fable has way more big model smell. But that means it's a skill in itself to get value out of it—99% of people are still not there yet. GPT-5.6 is almost as powerful, but is easy to use, fast, and relatively cheap. It should give you an early sense of where model work is going. Full Every 🪨 Vibe Check:

Dan Shipper

146,074 Aufrufe • vor 2 Monaten

David Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @jason: “Should they trust any of these LLMs with their proprietary knowledge for fear of having it cribbed into a core LLM?” david friedberg: “I have had experiences where we've asked some fairly novel scientific questions, and (the AI model) identifies it as a novel insight. It's like, ‘Oh, never thought about that, interesting, blah, blah, blah.’ And then using a different account, asking the next version (of the model) later, I've now experienced this. It's like, ‘Oh, well, you could do this,’ and it actually just describes this exact thing that we had in our chat in the previous version. Now, these are a handful of anecdotal experiences, but I know the domain that we work in, and the niche of it, and the ideation of this stuff, and the novelty of this stuff, and the lack of papers being published, and so on. So I know that there isn't some new corpus of information out there that's training the new model. So all I can say at that point is that my conversation or our analyses have been used for training.” David Sacks: “Okay, this does raise a really good question. What does it mean that the model is allowed to train on unidentifiable data?” Friedberg: “Well, that's my point. So it doesn't use any of my personal information, but it can use an insight derived from our chat, which it can then say is some training data that is unrelated. But the truth is, it's actually a piece of IP that's our organization’s IP, and our engagement back and forth. We don't have any NDA or confidentiality provisions or protections with them being a service provider back to us. This is why I care a lot about open source because I don't want them having my chat logs because they can use it for training to create an IP advantage that is now diffused to the rest of the market.”

The All-In Podcast

54,050 Aufrufe • vor 4 Tagen