Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing LlamaEdge — lightweight & portable LLM tools for your local, edge & server devices. Based on WasmEdge, LlamaEdge apps are <5 MB, self-contained (no complex dependencies), sandboxed (cloud ready), and can be orchestrated by container tools. LlamaEdge apps are portable across diverse CPUs, GPUs, and OSes. It is...

18,094 Aufrufe • vor 2 Jahren •via X (Twitter)

9 Kommentare

Profilbild von Rohan Paul
Rohan Paulvor 2 Jahren

Looks awesome, especially with a tech stack of Rust + Wasm + llama.cpp

Profilbild von wasmedge
wasmedgevor 2 Jahren

And more to come! We are supporting inference backends beyond llama.cpp. TensorRT-LLM and Intel Transformer Library. Stay tuned!

Profilbild von Sujānt
Sujāntvor 2 Jahren

the video @_rajkhare @iamgingertrash

Profilbild von will (exo/acc) 🦙
will (exo/acc) 🦙vor 2 Jahren

fyi @0x440x46 @AdonAlternative

Profilbild von josephs.tez
josephs.tezvor 2 Jahren

This is awesome, I can’t wait to try it out!

Profilbild von Aditya Poddar ⚛️
Aditya Poddar ⚛️vor 1 Jahr

I am new to wasmEdge, and I am already excited. I think it can/will make running AI models more accessible. Thus reducing barrier for entry. The plugin ecosystem will boost development in the space.

Profilbild von Joy Pullela
Joy Pullelavor 2 Jahren

I really don't want to refute it. But this inconsistency is too obvious. Who should I trust? You say AI doesn't make people unemployed. But Lee Kai-fu said that AI will quickly and destructively make people unemployed, even larger than all of human tech revolutions added together

Profilbild von Joy Pullela
Joy Pullelavor 2 Jahren

The screenshots I took before were all in Chinese. Now I found some English ones. You can also go to Google and search for it. Video and screenshot evidence of himself. The interviewer can also attest.

Profilbild von Joy Pullela
Joy Pullelavor 2 Jahren

After all, he wore a suit and went to such a high-level interview. He said it so confidently. The facts are there. Or is he in a suit ,full of lies. I agree with you above. So what kind of lie is Kai-Fu Lee telling?

Ähnliche Videos

Introducing Pods Hyperspace Pods lets a small group of people - a family, a startup, a few friends, to pool their laptops and desktops into one AI cluster. Everyone installs the CLI, someone creates a pod, shares an invite link, and the machines form a mesh. Models like Qwen 3.5 32B or GLM-5 Turbo that need more memory than any single laptop has get automatically sharded across the group's devices - layers split proportionally, inference pipelined through the ring. From the outside it looks like one OpenAI-compatible API endpoint with a pk_* key that drops straight into your AI tools and products. No configuration beyond pasting the key and changing the base URL. A team of five paying for cloud AI burns $500–2,000 a month on API calls. The same team's existing machines can serve Qwen 3.5 (competitive on SWE-bench) and GLM-5 Turbo (#1 on BrowseComp for tool-calling and web research) for free - the hardware is already on their desks. When a query genuinely needs a frontier model nobody has locally, the pod falls back to cloud at wholesale rates from a shared treasury. But for the daily work - code reviews, refactors, research, drafting - local models handle it and nobody gets billed. And when it is idle, you can rent out your pod on the compute marketplace, with fine-grained permissions for access management. There's no central server involved in inference. Prompts go from your machine to your pod members' machines and back: all of this enabled by the fully peer-to-peer Hyperspace network. Pod state - who's a member, which API keys are valid, how much treasury is left - is replicated across members with consensus, so the whole thing works on a local network. Members behind home routers don't need port forwarding either. The practical setup for most pods is three models covering different jobs: Qwen 3.5 32B for code and reasoning, GLM-5 Turbo for browsing and research, Gemma 4 for fast lightweight tasks. All running on hardware you already own. Pods ship today in Hyperspace v5.19. Model sharding, API keys, treasury, and Raft coordinator are all live. What Makes This Different - No middleman. Your prompts travel from your IDE to your pod members' hardware and back. There is no server in between reading your data. - No vendor lock-in. Pod membership, API keys, and treasury are replicated across your own machines using Raft consensus. If the internet goes down, your local network keeps working. There is no database in someone else's cloud that your pod depends on. - Automatic sharding. You don't configure layer ranges or calculate VRAM budgets. Tell the pod which model you want. It figures out how to split it across whatever hardware is online. - Real NAT traversal. Your friend behind a home router with a dynamic IP? Works. No VPN, no Tailscale, no port forwarding. The nodes handle it. - Free when local. This is the part that matters most. Cloud AI bills scale with usage. Pod inference on local hardware scales with nothing. The marginal cost of your 10,000th prompt is the electricity your laptop was already using. Coming soon: - Pod federation: pods form alliances with other pods. - Marketplace: pods with spare capacity can sell inference to other pods.

Varun

308,492 Aufrufe • vor 3 Monaten

My Prediction Market Tools list + Trading Workflow A) Tools I use: - Betmoar: Filtering + Analysing specific smart money wallets - HashDive - Prediction Market Analytics: Spotting volume and OI changes + whale flows - Polymarket Analytics: Comparing markets across different Prediction Markets - Polycule: Trading on the go - Polysights: Advanced Metrics related to volatility and trends B) My workflow: 1) I usually trade sports + crypto price prediction markets since I have some kind of expertise there 2) I use Betmoar to filter and find interesting markets. Try to enter new markets as early as possible since there are significant mispricings 3) I then use Hashdive to see what OI changes and whale flows are like for a particular market 4) Polymarket Analytics to compare markets across PM and Kalshi to see where I can get the best price 5) Trade directly through PM and Kalshi for now to execute if im on my laptop and Polycule when I'm not 6) I also check in daily on smart money wallets I follow on BetMoar to see if they've made any interesting trades recently I don't focus on Arbitrage, LPing or delta-neutral trades at all. Prefer to either take directional trades where I have an information edge or I'm early Sometimes I'll 'gamble' on sports markets because it's fun C) Takeaways for you: - Only trade in markets you have in-depth knowledge/edge - Being early is lucrative in prediction markets - Pro tools can save you hours trying to find alpha - Don't use too many tools. Use a few and master them instead - Everyone has their own process. Don't copy mine it won't work for you. The point of this tweet was to educate you on how you can build your own process LMK if there's any other useful tools I should check out that could optimise my workflow further!

Yoshi

31,700 Aufrufe • vor 10 Monaten

Holy shit... Microsoft open sourced an inference framework that runs a 100B parameter LLM on a single CPU. It's called BitNet. And it does what was supposed to be impossible. No GPU. No cloud. No $10K hardware setup. Just your laptop running a 100-billion parameter model at human reading speed. Here's how it works: Every other LLM stores weights in 32-bit or 16-bit floats. BitNet uses 1.58 bits. Weights are ternary just -1, 0, or +1. That's it. No floats. No expensive matrix math. Pure integer operations your CPU was already built for. The result: - 100B model runs on a single CPU at 5-7 tokens/second - 2.37x to 6.17x faster than llama.cpp on x86 - 82% lower energy consumption on x86 CPUs - 1.37x to 5.07x speedup on ARM (your MacBook) - Memory drops by 16-32x vs full-precision models The wildest part: Accuracy barely moves. BitNet b1.58 2B4T their flagship model was trained on 4 trillion tokens and benchmarks competitively against full-precision models of the same size. The quantization isn't destroying quality. It's just removing the bloat. What this actually means: - Run AI completely offline. Your data never leaves your machine - Deploy LLMs on phones, IoT devices, edge hardware - No more cloud API bills for inference - AI in regions with no reliable internet The model supports ARM and x86. Works on your MacBook, your Linux box, your Windows machine. 27.4K GitHub stars. 2.2K forks. Built by Microsoft Research. 100% Open Source. MIT License.

Guri Singh

2,180,357 Aufrufe • vor 4 Monaten

Proud to announce the in-depth collaboration between Kingnet and Alibaba Cloud in AI Gaming. Alibaba Cloud provides world-leading cloud computing, big data, and AI services, with disclosed revenue exceeding $15 billion in 2024, which is one of the most renowned global server providers. When two superpowers collide, the game changes. 🌊AI Gaming R&D By integrating Qwen 's LLM and Alibaba Cloud 's PAI platform (including PAI-iTAG, PAI-Designer, PAI-DSW, PAI-DLC, and PAI-EAS), Kingnet has emerged as one of the gaming industry's pioneers in AIGC-powered content generation and AI rendering. Together, we are accelerating the realization of no-code game development. 🌊GPU Computing Resources Alibaba Cloud delivers GPU-accelerated elastic computing services with exceptional processing power, supporting diverse workloads including deep learning, scientific computing, graphics visualization, and video processing - providing robust GPU computing capabilities for KingnetAI's demanding requirements. 🌊Cloud Service Optimization Cloud server deployment has become the mainstream choice for small and mid-sized game studios in global operations. Leveraging Alibaba Cloud server advantages, we will develop and deploy more cloud-native games to meet user demands. The disruptive innovation we're bringing to the industry: 🔸Minute-scale game asset production replaces traditional week/month-long cycles 🔸Single-digit dollar development costs VS traditional four-figure entry thresholds 🔸AI-powered NPCs with behavioral engines deliver dynamic player interactions, breaking static story constraints, etc. 🔜Kingnet AI V2 is approaching launch. The Agent system and game generation engine will be officially deployed across 3 chains: 🔹Leveraging Solana high throughput and low gas fee , Solana has consistently been a developer favorite, latest product will be deployed on Solana - with users paying $SOL for on-demand asset creation fees. 🔹Another key partner is BNB Chain ,We are actively participating in both the #BNBAIHack and the latest MVB 10. Powered by BNB Chain long-standing support for AI innovation. Kingnet V2 and NFT drop will be deployed on BNB Chain, providing developers and the community with comprehensive game-generation tools and support. 🔹As an early strategic partner of Kingnet, TON 💎 @TONEastAsia was one of the earliest chain to connect Web2 and Web3, Kingnet V2 will be deployed on TON, providing TON game developers with low-cost, high-efficiency asset generation, and supporting users to use $TON as an asset generation cost. The Future of AI Gaming is coming.

Kingnet AI

149,774 Aufrufe • vor 1 Jahr