Loading video...

Video Failed to Load

Go Home

The NVIDIA Ada Lovelace architecture is ultra efficient and provides an incredible performance boost: 🟢 GeForce RTX 4070 vs RTX 3070 Ti at 1440p in Warhammer 40,000: Darktide ⬆️ More frames ⚡️ Fewer watts #BeyondFast

541,692 views • 3 years ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Run Gemma 4 26b MTP on 8 GB VRAM GPUs at 25+ tokens/second. Flags included! local llm space is moving at terminal velocity. only 3 days ago google released gemma 4 26b a4b qat quants. more efficient than before, ran on 8gb vram at 20 tok/sec. and now just a few hours ago, mainline llama.cpp merged a massive update and we just shattered our own record. decode throughput went 25-40% up on the same 8 GB VRAM setup! Before MTP: 20 tps -> After MTP: 28 tps! llama.cpp just officially merged PR #23398 ("add Gemma4 MTP"), bringing native Multi-Token Prediction (MTP) support to Gemma 4 models. By running speculative drafting on the same 8GB VRAM RTX 4060 setup, my decode throughput on a 64k context instantly leaped to a blistering 25–27 tokens/sec thats 25-30% increase with the same hardware. Here is the architectural catch you need to know: Unlike the Qwen 3.5 and 3.6 series, which bake the MTP heads directly into the base GGUF, the Gemma 4 MTP head is not built in. You must download a separate, specialized MTP drafter GGUF (the assistant model) to act as the speculator. (I've dropped the download link in the replies). copy and try the exact flags: -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --spec-type draft-mtp --spec-draft-n-max 6 --spec-draft-p-min 0.7 --spec-draft-model gemma-4-26b-A4B-it-assistant-Q4_0.gguf -c 64000 -v n-max 4 and p-min 0.7 is also worth checking out. benchmark on your setup and workflow. if you have a single 8 gb vram nvidia rtx 4060, 3060, 3070, 2080, 2070, grab the MTP drafter GGUF link in the comments and try it yourself. Check it out even if you have asmaller or a larger gpu, such as a single rtx 3090, 4090, 3060, 2060. MTP works for all gemma 4 sizes such as gemma 4 12b, gemma 4 31b etc. but remember to grab the correct mtp draft assistant models respectively. what are you benchmarking today

Alok

200,913 views • 3 months ago

Samsung Galaxy S25 Ultra has completely targeted its competitors at the iPhone, and no longer competes with Chinese brand phones. This is because young people in South Korea are almost completely occupied by Apple, and Samsung’s strategy is to pull these young people back, so it strives to make Galaxy look like the Apple iPhone. There are several reasons for not competing with Chinese brands: 1. On the surface, although the Chinese Ultra models have powerful cameras and are suitable for those who pursue the ultimate in images, they are only limited to these people. Overall, the sales of Ultra models are not high, and even the sum of all brands of Ultra models cannot be compared with the sales of S24Ultra. 2. Moreover, these models with powerful cameras are relatively thick and heavy, with serious camera bulges, and the design is not perfect. It cannot be perfect. This design may not be suitable for everyone. Samsung will not easily take the risk to adopt this design in the global market. 3. The infinitely enhanced camera configuration will greatly increase the cost, which is difficult for Samsung to accept, and the production of this frequently updated camera may not be able to support Samsung's sales demand of tens of millions. In short, the Chinese brand Ultra model is more like a special non-popular model suitable for geeks. The Galaxy S25 series is defined as a popular model, and it must consider everyone's feelings and sufficient supply. This is why Samsung will not design the S25 Ultra as a Chinese brand Ultra. It remains consistent with the iPhone 16 Pro Max, but subtly. It is always slightly better than it For example, it is a little thinner (8.2mm vs 8.25mm), a little lighter (219g vs 227g), a little narrower bezel, a little stronger performance, a little more camera (retain 3x), a little more ultra-wide-angle pixels (50MP vs 48MP), a little bigger battery, and a little faster charging. Even the most incredible improvement: One UI 7.1 is a little smoother than iOS18 software. This is happening.

Ice Universe

347,352 views • 2 years ago

Run Gemma 4 26B MoE on 8GB VRAM with 250k context at 20+ tokens/sec If you own any 8GB VRAM graphics card, stop what you are doing. Local AI just had its absolute "Holy Shit" moment for budget hardware. Yesterday, I benchmarked Unsloth Gemma 4 12B Q4_K_XL on an 8GB card. The community went wild but immediately demanded more: "Can we run a 25B+ model on budget GPUs?" Today, I’m delivering exactly that. I am running a massive 26B parameter Mixture of Experts (MoE) model locally on a standard 8GB VRAM setup with 250k full native context!. If you own an RTX 3060, 3070, 4060, or any budget GPU with 8GB of VRAM, the local AI paradigm has completely changed. The performance metrics are astonishing: - 20 tokens/sec flat decode throughput. - Stable, flat decode speed even with massive prompts. - I threw a 60k token prompt at it, and it still clocked in at 20 TPS without dropping a single frame. # What about prefill? Yes, Time To First Token (TTFT) is slightly high when swallowing massive contexts. But with a solid 200 tokens/sec prefill speed, the wait is barely noticeable and highly usable. And this is running completely without Multi Token Prediction (MTP) active. How is this possible? It’s the magic of Google's new QAT (Quantization Aware Training) quants for Gemma 4. The model weight file (unsloth gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf) is only 13.2 GB, making it the ultimate local powerhouse. # The Test Setup: CPU: Intel Core i7 RAM: 16GB System RAM GPU: NVIDIA GeForce RTX 4060 Laptop GPU (8GB VRAM) # The Secret Sauce (The -cmoe Flag) To make this work properly on any 8GB card, you must use the -cmoe (CPU MoE) flag in llama.cpp. This flag isolates the heavy MoE expert weights directly to system memory (CPU/RAM) while letting your GPU focus strictly on the Attention layers and the KV Cache. It prevents VRAM spillage and holds the throughput rock solid. # The flags: -m "gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf" -cmoe -c 248000 -v Once running, just open the UI on localhost and toggle the new reasoning lightbulb icon in the text input box to watch the model perform multi step thinking. Are you still running smaller models, or are you ready to scale up your budget local setups? Let's discuss in the replies

Alok

292,770 views • 3 months ago

The "I don't have enough VRAM" excuse just died. I’m running Meta’s new 30B Muse Glimmer Q6_K_XL with a massive 130k context window on just 26GB VRAM FREE compute on Kaggle. Kaggle provides you free 2x Nvidia T4 GPUs. 30 hours usage each week! Yesterday, I showed you the violent throughput of Muse Glimmer on a single RTX 4090. Today, we are securing a Dual NVIDIA T4 GPU cluster with 32GB of total VRAM for exactly $0 and dropping the massive 24.5GB Q6_K_XL GGUF onto it. Here is the exact Kaggle workflow and benchmarking breakdown: # 1. The Storage Bypass & Setup I built a clean cell by cell script in the file. We dynamically fetch the CUDA accelerated llama.cpp binaries and use wget to stream the model directly into Kaggle's /kaggle/tmp scratch storage, which cleanly bypasses their 19.5GB output directory limit. # 2. The Multi GPU Performance With the -ngl 99 flag offloading all model layers across both T4 GPUs (32GB VRAM combined), we pushed a massive 131,072 token context window (-c 131072). The benchmark numbers: Prefill: 265.9 t/s Decode: 9.0 t/s VRAM Total: 26.5 GB # 3. The Architecture Insight The Q6_K_XL model itself is 24.5 GB. Because of Muse Glimmer's aggressive 16:1 GQA, the unquantized KV cache for a massive 130k context window only takes up 2 GB of memory. No heavily degraded Q4 KV quantization required. It just works. No compiling from source. No credit card. No OOM crashes. Zero excuses. If you’re running a single RTX 3090, 4090, or 5090, you need to experience this hyper efficient KV cache right now before the upcoming Qwen 3.8 27B drop completely steals your VRAM tomorrow. pick the Q4 or Q5 quants for 24 GB VRAM rigs. I'm dropping the Unsloth huggingface GGUF links and the free Kaggle notebook link in the replies. spin up your own instance, and show me your multi GPU benchmarks.

Alok

19,370 views • 1 month ago

my 8 GB VRAM gaming laptop is absolutely going to hate me for this. but I still did it. ran a 31b dense model (Gemma 4 31b Q4) with only 8 GB VRAM last week I ran Gemma 4 26B A4B a mixture of experts model on my RTX 4060 and hit 25–28 tokens/sec using llama.cpp's new MTP support. smooth. snappy. but MoE has a secret: it only activates 4B parameters per token despite having 26B total. that's why it flies. so the real question started haunting me. what if I throw a full, no tricks, every parameter fires on every token, 31B DENSE model at the same machine? # Hardware: GPU: NVIDIA RTX 4060, 8 GB VRAM RAM: 16 GB CPU: Intel Core i7 H Laptop. Gaming. Modest. The model: gemma-4-31B-it-qat-UD-Q4_K_XL.gguf (model's unsloth huggingface link in the comments) This is Google DeepMind's flagship dense model in the Gemma 4 family that can run on single consumer GPU. It packs a hybrid attention architecture, supports up to 256K context natively, and is QAT (Quantization Aware Training) optimized, meaning it retains far more quality than standard post training quants at the same bit depth. This is NOT the MoE. This is 31 BILLION dense parameters, every single one of them loaded. # the flags I used: -m gemma-4-31B-it-qat-UD-Q4_K_XL.gguf -cnv --spec-type draft-mtp --spec-draft-model mtp-gemma-4-31B-it.gguf --spec-draft-n-max 8 --spec-draft-p-min 0.6 -c 6000 -v Multi Token Prediction (MTP) is still active here. Separate draft GGUF required, same as the 26B setup. # Results: → Decode: ~3 tokens/sec → Prefill: ~2 tokens/sec → Context: 6000 tokens → Hardware crying quietly in the corner: yes so is 3 tps actually usable? For real time back and forth chat? Not ideal. You're not having a fluid conversation at 3 tps. but slow ≠ useless. And this is where it gets genuinely interesting. think about how senior devs actually work in a real team. But when something is architectural, deeply complex, or needs serious reasoning? they walk down the hall and escalate to the senior. That's exactly the local AI agent architecture this unlocks: → Fast orchestrator model (Gemma 4 26B MoE at 25+ tps) handles routing, simple queries, tool calls, memory. The junior dev. → Gemma 4 31B dense is the senior, called only when the fast model genuinely hits a wall. Hard multi step reasoning. Complex code generation. Deep architectural decisions. The agentic loop stays fast. Only the hard hops touch the 31B. That's a legitimate production grade local AI architecture on a budget hardware. (requires 2 8gb gpus) other workflows where 3 tps is completely fine: - overnight batch jobs. summarize documents, extract structured data, review code. Fire it off. Sleep. wake up to results. - One shot deep reasoning - Silent code audit loops, you write and test, the 31B reviews diffs and flags issues in the background between your sprints - Any workflow where output quality > output speed A few weeks ago, nobody was running a 30B+ dense model on a single consumer GPU with 8 GB VRAM. At all. Now we're doing it on an Intel i7-H gaming laptop with a NVIDIA RTX 4060, thanks to llama.cpp + QAT quants + MTP speculative drafting. Google DeepMind said the Gemma 4 31B targets "consumer GPUs and workstations." They were not exaggerating. The hardware bar to run serious frontier class models locally keeps dropping. the tools are here. the models are here. you just have to be willing to abuse your laptop a little. what workflows would you actually run on a local 3 tps 31B dense model? genuinely curious. drop it below.

Alok

63,689 views • 3 months ago

50% more context unlocked for Qwen 3.8 27b Q4_K_XL dflash 2 on a single RTX 4090 (24 GB VRAM) I found a hidden VRAM tax in llama.cpp. By combining my custom 2 bit DFlash 2 drafter with one overlooked server flag, I just unlocked another +80,000 tokens of context. Qwen3.8-27B is now running a massive 250,000 context at 75 tokens/s on a single RTX 4090. Here is the secret: By default, `llama-server` reserves massive chunks of your VRAM to handle multiple concurrent users (batching). If you are running a single user session, you are bleeding memory for features you aren't using. By passing the `--parallel 1` flag, you force the engine to dedicate 100% of your 24GB VRAM buffer to a single user. When we combine the VRAM saved by our Q2_K 2-bit drafter with the VRAM saved by `--parallel 1`, the context ceilings absolutely explode: Note: all benchmarks carried out with a massive 28k prompt. Ubuntu 22. ### THE NEW 24GB PHYSICAL LIMITS (Single RTX 4090): # 1. The "Repo Swallower" (Q4 KV Cache): - Context: 250,000 tokens (Up from 170k!) - Speed: 73.66 t/s decode | 1,608 t/s prefill - Peak VRAM: 23.8 GB # 2. The "High-Precision SWE" (Q8 KV Cache): - Context: 150,000 tokens (Up from 100k!) - Speed: 75.01 t/s decode | 1,667 t/s prefill - Peak VRAM: 23.9 GB # 3. The "Pristine Attention" (Unquantized FP16 KV): - Context: 90,000 tokens - Speed: 80.58 t/s decode | 1,699 t/s prefill - Peak VRAM: 23.92 GB ### HOW TO RUN THE 250K GOD STACK TODAY: (Requires PR #27342 + my Q2_K Hugging Face drafter) llama.cpp flags: ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q2_K.gguf --spec-type draft-dflash --spec-draft-n-max 3 -c 250000 -ngl 99 --parallel 1 --port 8080 -ctv q4_0 -ctk q4_0 We are pushing a quarter million tokens of context with speculative DFlash 2 decoding at 73 tokens/second on a single consumer gaming GPU. I dropped my custom 2 bit Hugging Face GGUF links, visual performance graphs, and the PR #27342 build instructions in the replies below. If you own a single RTX 3090 or 4090, it is officially time to cancel your API subscriptions and let local silicon eat the cloud. how much monthly API spend does an optimized 4090 rig like this actually replace for you?

Alok

39,189 views • 1 month ago

If you thought the Gemma 4 31B (dense) model was fast, sit down. I just benched the updated Gemma 4 26B A4B MoE on a single RTX 4090 (24 GB VRAM) 9,200 t/s prefill. 160 t/s decode. 250,000 context window. All on a single consumer RTX 4090. The numbers are completely unhinged. The 31B is a dense behemoth. But the 26B is a Mixture of Experts (MoE), specifically an Active 4 Billion (A4B). It holds 26B parameters of knowledge but only activates 4B per token. Because its inference memory footprint is so light, I didn’t even need KV cache quantization to hit a quarter million context. Compiled the latest llama.cpp from source on Ubuntu 22 (CUDA 13). Fed it a 28k token prompt, and manually cranked the batch sizes (-b 2048 -ub 2048) to absolutely redline the Tensor Cores. Here is the benchmarking breakdown: # 1. The Baseline (No MTP) Even without speculative decoding, the A4B architecture flies. llama.cpp flags: ./build/bin/llama-server -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -c 250000 -ngl 99 -fa on -b 2048 -ub 2048 --port 8080 -v Context Ceiling: 250,000 tokens (21.5 GB VRAM) Prefill: 9,200 t/s (Absurd) Decode: 124 t/s # 2. The MTP Overdrive Injected the new MTP draft model to enable Speculative Decoding. llama.cpp flags: ./build/bin/llama-server -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --spec-type draft-mtp --spec-draft-model mtp-gemma-4-26B-A4B-it.gguf --spec-draft-n-max 4 --spec-draft-p-min 0.7 -c 250000 -ngl 99 -fa on -b 2048 -ub 2048 --port 8080 -v Context Ceiling: 250,000 tokens (22.96 GB VRAM) Prefill: 7,054 t/s (MTP draft overhead slightly caps prefill) Decode: 156 t/s # The Agentic Architecture Insight Why does this matter? Because you can now build a killer local agentic loop on a consumer desktop. Use the 31B dense model (from the previous post) as your heavy, deliberate Orchestrator / Verifier / Planner. Pass the actual execution tasks to this 26B MoE. At 160 t/s, this MoE can chew through code generation, tool calling, and massive RAG document retrieval over a 250k context window almost instantly, drastically speeding up your agentic loop. If you own a single RTX 3090 or 4090 and haven't tried this specific stack yet, you need to pull these latest updates and run it. Local inference just leveled up. Hugging Face links to the Unsloth 26B QAT quants and MTP drafters are in the replies. performance graphs also available in the replies.

Alok

40,993 views • 2 months ago

Qwen 3.8 27B (dense) running on a single RTX 4090 (24GB VRAM) at 65 tokens/sec decode with MTP! 260,000 context window or 65 tokens/sec decode with native MTP. The API cartel should be terrified. We are officially running frontier tier agentic AI (benchmarks comparable to claude opus 4.6 max) on a single consumer gaming GPU. I benchmarked Qwen3.8-27B on a single NVIDIA RTX 4090 (24GB VRAM, Ubuntu 22) using Unsloth’s Dynamic Q4_K_XL GGUF on the latest llama.cpp. Here is the complete benchmark breakdown across both Context Scaling and MTP Overdrive (28k prompt baseline): ### PART 1: The Context Scaling Matrix (Pure Throughput) # 1. Standard FP16 KV Cache (Unquantized): - 80k Context: 2,664.7 t/s prefill | 40.68 t/s decode | 22.36 GB VRAM - 100k Context: 2,678.7 t/s prefill | 40.89 t/s decode | 23.59 GB VRAM (100k is the hard ceiling for unquantized f16 KV in 24GB VRAM) # 2. Q8 Quantized KV Cache (-ctv q8_0 -ctk q8_0): - 130k Context: 2,639.1 t/s prefill | 40.96 t/s decode | 22.18 GB VRAM - 170k Context: 2,653.9 t/s prefill | 40.70 t/s decode | 23.68 GB VRAM (170k is the sweet spot for heavy agentic coding workflows) # 3. Q4 Quantized KV Cache (-ctv q4_0 -ctk q4_0): - 260k Context: 2,659.8 t/s prefill | 40.70 t/s decode | 23.00 GB VRAM Full 262k native context residing entirely in 24GB VRAM. Zero system RAM offload. Stress test with a monster 142k real-world prompt (-c 170000, Q8 KV): - Prefill: 1,829.50 tokens/s - Decode: 31.3 tokens/s - VRAM: 23.7 GB rock solid ### PART 2: Native MTP Overdrive (Trading Context for Speed) Since MTP heads are baked into the architecture, enabling native speculative drafting pushes decode speeds straight to 60 t/s with zero external draft model: # 1. MTP + Q8 KV Cache: - 80k Context: 2,370.66 t/s prefill | 59.25 t/s decode | 23.4 GB VRAM (MTP state buffers eat slightly more memory, making 80k the ceiling for Q8) # 2. MTP + Q4 KV Cache: - 130k Context: 2,391.09 t/s prefill | 60.10 t/s decode | 23.5 GB VRAM (Sweet spot: 130,000 context running at a screaming 60 tps decode) ### Qwen3.8-27B vs Muse Glimmer 30B Two days ago I benched Meta's Muse Glimmer 30B hitting 130k context unquantized (19.3 GB VRAM) pulling 50-75 t/s decode. If you own a single RTX 3090 or RTX 4090, you have zero excuse to burn API credits. ### The Reproduction llama.cpp flags: 1. Max Context Stack (260,000 Context @ 41 tps): ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -c 260000 -ngl 99 --port 8080 -ctv q4_0 -ctk q4_0 2. MTP Overdrive Stack (130,000 Context @ 60 tps): ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -c 130000 -ngl 99 --port 8080 -ctv q4_0 -ctk q4_0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 Unsloth's Hugging Face GGUF links, intelligence/agentic benchmark details, and performance charts are posted in the replies. Local compute is eating the cloud alive. How much monthly API spend does a 24GB setup like this actually replace for you?

Alok

384,271 views • 1 month ago

Free NVIDIA GPU with 16 GB VRAM GPU for Running Local LLMs! If you want to master local LLMs but you're waiting until you can afford a $1,500 GPU, you're honestly not going to make it. The open source AI ecosystem is moving way too fast for you to wait on your budget to catch up. Especially when you can build a bleeding edge inference engine from scratch right now, completely for free. You don't need a heavy local rig to start. Google is literally letting you use an enterprise grade NVIDIA Tesla T4 GPU for $0/hour. At standard cloud computing rates (~$0.20/hr), Google Colab’s 4 hour daily free tier hands you roughly $24 worth of data center tier GPU compute every single month. And most people just waste it. Let’s talk about the hardware you get access to for free. The NVIDIA Tesla T4 is an absolute workhorse: - Architecture: NVIDIA Turing (TU104) - VRAM: 16GB GDDR6 (320 GB/s bandwidth) - Compute: 320 Tensor Cores | 2560 CUDA Cores - Performance: 130 TOPS INT8 | 8.1 TFLOPS FP32 - Power: Sipping energy at a max 70W TDP This is the exact same hardware I used to run DeepMind's Gemma 4 26B A4B QAT MoE at a 250,000 context window without a single Out Of Memory (OOM) crash. If you have a web browser and 10 minutes, you have everything you need. I’ve put together a fully documented, cell by cell Google Colab notebook that teaches you exactly how to do this. Here is what the notebook actually teaches you: - How to provision an Ubuntu Linux environment with CUDA 13.0 and verify your driver stack. - How to pull the source code and compile the latest llama.cpp C++ binaries from scratch, specifically optimizing the build for your exact GPU using the -DCMAKE_CUDA_ARCHITECTURES=native flag. - How to directly download quantized local LLMs (GGUF format) straight from HuggingFace using the CLI. - How to manage 16GB VRAM limits, offload neural network layers to the GPU, and push massive context windows. Compile raw llama.cpp, ollama run a model, or spin up the LM Studio CLI. Pick whatever stack you are comfortable with. just start building. No hardware. No credit card. No excuses. Bookmark this post right now so you don't lose the tutorial. Even if you don't have time to run it today, you are going to want this workflow in your engineering toolkit. The link to the free Colab Notebook is in the comments below. Lemme know if you need more tutorials like this.

Alok

182,483 views • 2 months ago

Muse Glimmer, A 30B parameter dense model swallowing a 130,000 token context window using only 19.3 GB of VRAM (extreme efficiency). No KV cache quantization required. I just benched the new Muse Glimmer 30B (dense) on a single RTX 4090. We are pulling 3,100+ t/s prefill and 75 tokens/second decode. The throughput is violent. Meta superintelligence lab just open sourced this agentic beast, explicitly engineered to dominate 24GB consumer cards. I pulled the latest llama.cpp source on Ubuntu 22 (CUDA 13) to see if the specs were real. Fed it a 28k token prompt. Here is the exact llama.cpp God Stack and benchmarking breakdown: # 1. The Deep Context Run (No Speculative Decoding) The architecture uses a massive 16:1 GQA (Grouped Query Attention) ratio. This means the KV cache footprint is practically non existent. ./build/bin/llama-server -m Muse-Glimmer-30B-UD-Q4_K_XL.gguf -c 130000 -b 4096 -ub 4096 -ngl 99 --port 8080 Prefill: 3134.95 t/s Decode: 50.00 t/s VRAM: 19.34 GB (I hit 130k context on pristine, unquantized f16 cache and still had 4.5 GB of VRAM left over. Absolute witchcraft). # 2. The DFlash Speculative Overdrive Meta shipped this with a DFlash block diffusion drafter. Let's trade that extra VRAM for pure speed. ./build/bin/llama-server -m Muse-Glimmer-30B-UD-Q4_K_XL.gguf -md dflash-kquant.gguf --spec-type draft-dflash --spec-draft-n-max 3 -c 80000 -b 4096 -ub 4096 -ngl 99 --port 8080 Prefill: 1293.69 t/s Decode: 75.00 t/s VRAM: 23.93 GB (Maxed out on card) the dflash gguf is additional 1.6 GBs # The Architecture Insight (Muse Glimmer vs. Gemma 4 31B) If you look at my Gemma 4 31B tests from last week, getting 140k context required heavily degrading the memory with Q4 KV quantization (gemma 31b q4 can do only about 40k context with unquantized kv on a 24gb card). That "unzipping" overhead bottlenecked Gemma's MTP decode speeds down to 65 t/s. Muse Glimmer completely sidesteps this bottleneck. By using aggressive 16:1 GQA, it keeps the KV cache in native f16 format at massive context lengths. Flash Attention gets to run at maximum uncompressed speed, letting the DFlash drafter push decode safely to 75 t/s without compute lag. With a 76% on SWE Bench Verified and seamless local tool calling, this model looks promising. Unsloth's Hugging Face GGUF links, intelligence/agentic benchmark details, and inference throughput performance graphs are posted in the replies. For 24GB rig, what’s your current go to model?

Alok

65,480 views • 1 month ago

i spent $26,600 on cloud GPU rentals over 14 months before i found a NVIDIA DGX Spark at $2,999 (founder's edition) or $3,999 (shipping price) it paid for itself in 6 weeks i run 200B parameter models locally now and my old cloud provider keeps sending me loyalty discount emails the math on that $26,600 is embarrassing to type out loud $1,900/month for 14 months, H100 instances on a specialist cloud provider, because anything bigger than a 70B model simply would not fit anywhere else i paid the invoices like they were a utility bill and told myself it was just the cost of doing serious AI work it took me over a year to find out it wasn't 14 months, broken down: → months 1-4: $1,400-1,600/month - felt like manageable infrastructure overhead → months 5-9: crept to $1,900-2,100 as i started running DeepSeek-class experiments, costs tracking directly with model size → months 10-12: one agent loop ran for 36 hours against a 130B model while i slept, that month hit $2,400 → month 13: ran the cumulative total for the first time, saw $23,800, felt physically sick → month 14: another $2,800 month while i waited for the hardware to ship the box is the NVIDIA DGX Spark - roughly the footprint of a large mac mini, powered by a GB10 Grace Blackwell chip with 128GB of unified LPDDR5X memory that unified memory is the whole thing an RTX 4090 has 24GB of VRAM, which means a 70B model in full BF16 precision physically does not fit, you're quantizing down or you're renting cloud, those are your options this box loads a 200B parameter model quantized and serves it through vLLM over localhost, same API interface the cloud endpoint used the migration took one line of code - i changed the base URL from the provider's endpoint to 127.0.0.1:8000 and everything just worked electricity to run continuous 200B inference locally comes out to about $12/month the payback arithmetic is almost too clean: $2,999 hardware cost against $1,900/month saved, the box paid for itself before i'd owned it two months what i didn't account for was how completely the cost model changes your behavior when there's no hourly meter running, you greenlight experiments you'd never approve on cloud - agent loops that churn for hours, running 10,000 documents through a reasoning pass at 3am, speculative fine-tuning jobs you'd normally skip because the cost felt unjustifiable i ran more experiments in the first 30 days after the box arrived than in the four months before it the loyalty discount email landed about 8 weeks after i cancelled the cloud subscription 15% off my next three months, valued customer, we'd love to have you back i didn't reply the box was already running

Argona

22,355 views • 3 months ago

Exciting News from HotkeySwap: TAOevm is Officially Live! We’re thrilled to announce the launch of TAOevm—a powerful addition to the HotkeySwap ecosystem that marks a significant step in our mission to make DeFi smarter, more accessible, and truly cross-chain. This milestone represents months of dedicated work and innovation, bringing you an AI-powered, user-friendly experience that connects seamlessly across blockchain networks. With a fresh look and expanded features, HotkeySwap and TAOevm are here to transform your DeFi journey. Whether you’re looking to explore, trade, or invest, HotkeySwap’s ecosystem is built to make it easy and rewarding. What’s New in the Hotkey Ecosystem? Hotkey’s latest updates bring you seamless cross-chain tools, analytics, and a collaborative environment to explore, trade, and engage. Here’s what’s now live and ready for you: • TAOevm Blockchain: TAOevm introduces full EVM compatibility to expand decentralized finance and smart contract capabilities. While Bittensor remains focused on its core mission as a decentralized AI and machine learning network, TAOevm provides an independent platform for DeFi innovation. This separate chain approach allows Bittensor’s community to access EVM benefits without compromising Bittensor’s AI-first vision. With a capacity of 200,000 transactions per block and blocks validated every 2 to 3 seconds, TAOevm delivers 60k to 100k TPS at peak, Hybrid Proof of Work security, and ultra-low fees—empowering developers to create new possibilities within an ecosystem that respects and reinforces the unique strengths of both platforms. • Hotcurves is the pump fun for TAOevm: As part of our commitment to community-driven innovation, we’re introducing Hotcurves—a vibrant platform on TAOevm where you can explore new projects, invest early, and connect with a community of like-minded enthusiasts who love to innovate and have fun. Hotcurves simplifies the token deployment process, allowing anyone to create a token in minutes with built-in liquidity pools for instant trading. With secure features like automatic LP burns and contract renouncement, Hotcurves offers a safer, more accessible environment for developers and investors alike. New Websites & Resources Our new branding and websites bring a unified, intuitive experience to the entire Hotkey ecosystem. Discover everything you need to navigate, learn, and grow with Hotkey: • Official Site: Your central hub for all Hotkey news, resources, and network updates. • TAOevm Mainnet Explorer: Track live transactions within the TAOevm mainnet, a powerful blockchain addition to Hotkey’s infrastructure. • TAOevm Testnet Explorer: Experiment with the TAOevm testnet and test your strategies with access to the Testnet Faucet. • TAOevm Bridge: Seamlessly transfer assets across chains within the Hotkey ecosystem. • TAOevm DEX: Engage in secure and efficient decentralized trading on TAOevm. • TAOevm Charts: Get real-time insights with charts tailored to TAOevm, offering analytics similar to DeScreener or DEXTools, integrated seamlessly with Hotkey. • Hotcurves: Discover and invest in early-stage projects on TAOevm’s very own community launchpad. • Documentation: Access comprehensive guides for HotkeySwap, TAOevm, and Hotcurves, providing everything you need to maximize your experience across the ecosystem. A New Brand, A Bold Vision With our new branding, Hotkey’s mission to make DeFi inclusive, intelligent, and community-centered is stronger than ever. The new look symbolizes our dedication to building an ecosystem where every user, investor, and trader can confidently explore and create within DeFi. This rebrand isn’t just about appearances—it’s a pledge to continuously evolve and support our growing community. We are forever thankfully for the continued support up to this point and we hope by showcasing our EVM you will be able to gain a glimpse into the capabilities of the team as we now gear up for the launch of the v2 dex in the coming weeks.

hotkeyswap

40,408 views • 1 year ago

You Can't Vibe-Code Trust Avishai Abrahami, Co-Founder & CEO of Wix , interviewed by Harry Stebbings (kevin andres) Summary: Wix trades at a $2.8B market cap on $2.1B of revenue while the market ascribes roughly zero value to a business throwing off $400M a year in free cash flow. Wix CEO Avishai Abrahami's argument is that the market can't yet price what AI actually threatens: the moat is trust and business logic, and neither gets vibe-coded away. His response is to own the disruptor (Base44), train his own narrow models, and stay committed through a storm he insists always arrives on a random Wednesday. 1. Trust is the moat. The real value of Salesforce is trust: JP Morgan and huge banks let it hold all their customer data, and the CRM itself is a small part of that. "What other platform will JP Morgan trust for their customers' data? None." That trust took years to build and can't be reconstructed by an agent scraping a database, so the companies whose value lives in trust survive the SaaS apocalypse while the ones reduced to piping get commoditized. 2. The business-logic wall. "You're not going to vibe-code Shopify no matter how good you are. The business logic is too hard." Wix tested this directly: they asked a team of professional developers to build the operating logic for a single hairdresser in Base44, gave up after a week, brought in a stronger team, and still failed two weeks later. Complex operational software is far harder than a demo suggests, which is why the pizza shop and the hairdresser stay Wix customers rather than build their own stack. 3. Own the disruptor. Wix bought Base44, a one-person company, for $80M, and it now does over $150M in ARR, roughly double what they paid. Abrahami frames the future as three buckets: owners who never want to build, owners who vibe-code everything themselves, and a mix in the middle over the next five or six years. Rather than bet on which wins, Wix owns the tool customers would defect to, so a customer who switches platforms still switches to Wix. 4. Trading on someone else's news. "Today we are trading on other companies' news. We're not trading on Wix news. We're trading on what OpenAI or Anthropic or Google are saying." Base44 alone, valued on vibe-coding peer multiples, should be worth around $8B, which means the market assigns less than zero to Wix's core. Abrahami's response is to detach: he doesn't wake up checking whether the stock moved 20%, because the only thing he can influence is the business. 5. The narrow model. Wix fine-tuned and combined its own models and now matches top-tier frontier quality on Base44 tasks at far lower cost. The logic: they sit on a huge stream of training data from watching what users try and where they fail, so a model built for Base44 can skip what frontier models carry, like knowledge of Chinese poetry, and go deep on what someone means when they say "build me a task manager to tell my boyfriend where he's wrong." A narrow target is easier to hit than a frontier model, and Wix already runs a trained model on website generation that's faster, cheaper, and makes fewer errors, retrained weekly on a live feedback loop. 6. Quality before cost. When Harry cites Chamath's claim that open source runs 14-16x cheaper, Abrahami pushes back: that holds for small tasks, but for something as complex as Base44 the savings land at 5-10%, and his own model runs 1-30% cheaper than frontier, not the order of magnitude people assume. More to the point, this is the wrong time to chase cost: "20% more quality, 20% less cost, I'll go for the quality." It's a brand-new market that's just starting, and the job now is to make the product better. 7. The but is very big. "We all give too much credit for AI. It's amazing, it's incredible, it's super powerful, but the but is pretty big." He asked Claude to write a safety protocol and got six mandatory gates, then pushed back on each one and watched the model cave until only one survived, downgrading the rest from "must test" to "might want to look at later." We over-trust these systems, and that reflex, treating a Reddit post as equivalent to research published in Nature, is where the danger lives. 8. Customer support still breaks. Wix has 3,500 people and its single biggest department is customer support, serving 192 countries. They tried hard not to build their own AI support agent, tested many off-the-shelf products, and concluded flatly: "It doesn't work. We tried, we tried again, it didn't work." The gap between hyped AI support startups and what actually ships in production is the tell that the technology is earlier than the marketing, maybe five years from being different. 9. Buybacks as dividends. Wix had $1.5B sitting in the bank it couldn't put into a major acquisition because it was focused on the new product and Base44, so it bought back stock at a low price, with admittedly terrible short-term timing. Abrahami is unbothered: "The big question is where it's going to be in three years, not what happened in the last three months." He argues buybacks are a fantastic, underused tool, essentially a dividend to every shareholder, and companies should lean on them to balance stock-based compensation instead of endlessly diluting. 10. Execution, not finance. A low stock price makes M&A currency less valuable, but Abrahami says that's not his real constraint. Base44 was a one-person company; Wix had to build an entire company around it, staffing it with people pulled from the core. "I don't know how to do another one of those at the same time and have the same quality." The bottleneck on the next acquisition is execution capacity, not the balance sheet. 11. Chosen to be here. The one thing money buys beyond food security is freedom, and the deepest form of that freedom is knowing you're here by choice. "I'm here because I've chosen to be here. Nobody made me." He could move to Costa Rica or dance carnival in Brazil, and choosing to stay and run a public company through a crashing stock is where he finds his power. Money also made him more impatient and a bit lazier, and more rational because he's no longer deciding from fear. 12. The random Wednesday. Resilience starts with accepting the storm will come, because we assume that if yesterday was easy tomorrow will be too, and reality doesn't move in gentle slopes. "The worst thing that happens is probably some random thing on some random Wednesday. It's not something you get a lot of warning for." His anchor, borrowed from Babylon 5, is that you get there when you get there and the weapons you have are the weapons you have, so the only real question is whether you're doing the best you can with what you control.

Gokul Rajaram

22,843 views • 1 month ago

The wait is finally over — Spartan Fuel has arrived. Imagine Maximum muscle growth Skin-splitting pumps Limitless energy Laser-like focus What’s the secret? An all-encompassing, comprehensive intra-workout blend. Everybody knows about the importance of pre and post workout nutrition, but intra-workout nutrition is often neglected. The optimal time to fuel your body is when you are working out, breaking down muscle fibers, expending stored glycogen, and depleting electrolytes. In order to stimulate maximum muscle growth and achieve a vicious, skin-splitting pump, we need to fuel our bodies. Spartan Fuel contains a comprehensive blend of EAAs, fast-acting carbohydrates, electrolytes, and mitochondrial boosters that provides your body with EXACTLY what it needs to perform at it’s highest level. Many of you are probably like myself; you put on your best Walter White impression and whip up a concoction of supplements to try and find that edge. Why should we leave gains on the table, right? Well, I had enough of guessing and watching my supplement cabinet (and monthly bill) continuously grow. And then it struck me. There’s no product that properly combines EAAs, carbs, and electrolytes. After lifting for years, absorbing everything I could on X, I figured it was time to make my own mark. That’s when I decided to connect with the man himself, BowTied Biohacker . Leveraging his expertise and my vision, we created a formula that would change the way we lift forever. We created a product that would be FELT IMMEDIATELY. After sending samples out to a bunch of bros here on X, the feedback was overwhelming — we had struck gold. Other-worldly pumps Gas tanks that were always running on full People crushing their log books, pumping out more reps than ever before… DURING EVERY SET Everyone felt like a million bucks. Once you try it, it you will never want to workout without it. It’s THAT good. We all push ourselves hard. Many of us train to failure. We want to get jacked. We want to get shredded. We all want to unleash our inner warrior in the gym. Now there’s a way to totally lock-in and dominate with intensity during very workout. We all know the feeling. Once in a while we have a workout that just blows us away. We feel stronger than ever, locked-in. We leave the gym with a high that has us feeling on top of the world. Now picture every single workout being that amazing. No other product can deliver the boost you need to consistently perform at the highest level. It’s basically a PED. We didn’t skimp out on quality. We didn’t cut corners. We included EVERYTHING needed to maximize results and boost performance. What many people don’t know is that in order for a supplement to truly be effective, the ingredients need to be dosed in the proper ratios. And that’s just what we did. And just when you think it can’t get any better (there has to be some catch, right…right?) We kept it natural — no artificial ingredients or sweeteners. This is something that digest like a dream, hits the bloodstream instantaneously, and fuels your muscles. Our competitors don’t do this. They sell a bunch of ingredients separately, trying to sell more products. Or they sell proprietary blends and junk loaded with fillers. Spartan Fuel makes your life easier. One tub. One scoop (or 2 if you’re like me and want to go hard). No more wasting money purchasing the entire supplement store. Whether you start drinking it on your way to the gym or as you begin your workout, you will quickly feel the difference. This fall, you can dominate every workout, and supercharge your winter bulk. And this post would not be complete without giving a huge thank you to @Thomas_Salamus_ TJ was instrumental to the birth of Spartan Fuel since Day 1. If you love his products, you’ll love Spartan Fuel. Rest assured, the quality is unmatched. The first batch is limited, so act now and don’t miss this opportunity to unlock your true potential.

Spartan

80,629 views • 11 months ago

**Doris Yin Speech at China Guizhou Zunyi GCV Barter Conference** Hello to the community leaders, GCV ambassadors, merchants, and pioneers of GCV Guizhou in China! Today is January 12, 2025, marking the first GCV Barter Conference in China in New Year and the 13th Barter Conference overall. I would like to extend my sincere gratitude to the organizers of this conference, the Guizhou Zunyi GCV Community, and the co-organizers, Barter Huishang (Guizhou) Digital Economy Industry Group Co., Ltd. I also want to acknowledge the following GCV ambassadors for their active dedication and contributions to this conference: **GCV Ambassador of China:** - Yang Zhizhong - Cai Zaiqiao **Ambassadors of Guizhou Province GCV:** - Wang Shiqiong - Cai Weisheng - Guo Jiaqing **Zunyi GCV Ambassadors:** - Wang Jianbo - Luo Nanlu **GCV ambassadors at the district and county level in Zunyi City** Additionally, I would like to express my heartfelt thanks to our numerous GCV merchants and sponsors. Without your support, we would not have been able to hold such a grand and large-scale event. Today's gathering in Zunyi, a sacred site of the revolution, reminds me of the Red Army's 25,000-mile Long March. Their perseverance and sacrifice continue to inspire us. The Zunyi Conference took place from January 15 to 17, 1935, and exactly 90 years later, we are gathered here today. The defining characteristics of the Zunyi Conference included the commitment to uphold the truth, correct mistakes, establish the correct leadership of the Party Central Committee, and creatively develop and implement strategies that fit the nature of the Chinese revolution. Today, our Zunyi Conference will also be recorded in the history of blockchain, as every effort you have put in has contributed to building a strong network ecosystem. Our partial fiat and partial distribution policy serves as a solution for the rapid development of the ecosystem during the closed mainnet of the Pi Network. As we all know, the first quarter of this year will bring about the successful mainnet launch of Pi Network. After six long years of challenges and perseverance, all of our pioneers will have the opportunity to witness this significant historical moment. What an exciting and proud day this will be! It has not been easy for everyone to persist through these six years; it requires great blessings, unwavering faith, and the courage to overcome difficulties. Today, our pioneers in Zunyi, Guizhou Province, gathering for this GCV barter conference holds great significance. I see that ten companies are providing products for barter, with nine companies, including Guizhou Meitan County Daoqin Hospital and Barter Huishang (Guizhou) Digital Economy Industry Group Co., Ltd., sponsoring this event. Once again, I extend my heartfelt thanks to all of you. The GCV Barter Conference serves multiple purposes. It is not only about creating GCV data or demonstrating the strength of our China region to CT, but also about showing how closely we align with their vision and mission. Additionally, it provides robust evidence for a substantial number of KYC and migration initiatives in China. More importantly, what we do today aims to boost China’s future economic development. Once the main network of the Pi Network is launched, we anticipate a significant demand for Chinese products from numerous international pioneers, which will in turn generate a large volume of export orders. At the same time, there will be international merchants looking to export their products to China. Once OM, import and export transactions will be conducted using the new currency, facilitating the vision of a stable currency and enabling seamless and reliable exchanges with fiat currency. Therefore, the merchants who engage now will have the advantage of being early adopters. The Pi Network offers a partner program and a MapofPi program. To participate in the partnership, businesses are required to have a company website. We invite businesses with websites to join us. However, if you do not have a company website, you can still join the Mapofpi program, which encompasses a wide range of industries, allowing participation from both large companies and small traders. Registration for the Mapofpi does not require a business license or website; various entities including shops, hospitals, schools, hair salons, accounting firms, law firms, restaurants, and hotels are welcome to register. Please select an active merchant and support GCV at $314,159. Prior to the OM launch, it is advisable to use partial fiat currency and partial Pi to ensure that merchants can cover their costs and fulfill their tax obligations. Recently, on January 9, we established the China GCV Industry Chain Alliance, which aims to create an industrial chain that facilitates the circulation of Pi among merchants, thereby reducing the burden of exchanging fiat currency after OM. During the enclosed mainnet, you can assist merchants in registering as Pi Network Partners and Mapofpi . Ms. Lumari is our Global GCV CT executive director and her goal is to have 200,000 registered Mapofpi merchants worldwide. My personal target is to reach 100,000 registered merchants in China alone. This goal is achievable given the over 58 million enterprises and more than 20 million pioneers in China. If we can effectively convey that Pi Network WEB 3.0 blockchain technology will significantly enhance human productivity and that the business opportunities from accepting partial Pi and partial FIAT during the 60 days before OM will present numerous benefits and minimal risks to merchants, then it is likely that no merchant will be unfavorably surprised by the initiative. This strategy offers a multitude of advantages with virtually no downsides. Furthermore, it benefits pioneers by allowing them to transfer purchasing power to the community and minimize fiat currency expenses in their daily life. Consequently, the GCV data we generate will significantly benefit the Chinese pioneers, as a large number of registered merchants can transform the China region from a high-risk area to a safe zone. Not only can this region be promoted to a VIP area, which would enjoy expedited KYC and mapping processes, but it will also allow pioneers and merchants to thrive together in our ecosystem. This collaboration will enhance the prosperity of our country and empower the China region to contribute to the welfare of communities worldwide. Once OM, it will play a crucial role in the economic development of both China and the world. If you pay attention to our migrartion speed, you might have noticed that it has slowed down recently. From December 17th to around the 30th, the migrating speed was over 50,000 to 100,000 per day, but now it has dropped to just over 10,000. What is the reason for this decline? If it was previously possible to migrate over 100,000 per day, why has it changed? The CT has stated that they will OM until the first quarter of this year to bring the migratiion in line with KYC amounts. However, if it's technically feasible to achieve a higher migration speed, why isn’t it being done? The answer is quite simple: it depends on what everyone does with the Pi after such large migration numbers. If pioneers rush to buy and sell, hold onto their Pi coins, or trade at low value, it will impact the speed and efficiency of the next migration in these regions. This principle is not only theoretically valid but has proven true in practice. For instance, countries like the Philippines, Indonesia, and Malaysia have a solid educational foundation in GCV. Most pioneers there are highly aware of the risks involved in participating in the black market, which allows them to generate a substantial amount of GCV data. As a result, their migration speed is notably fast, and there are many large wallet migrated. To help the CT regain momentum, we all need to cooperate. Engage with the migrated Pi and participate in the GCV barter ecosystem. Be cautious of individuals who aim to deceive you for personal gain; devaluing the Pi often serves as a tactic to exchange something small for your valuable treasure. It's crucial to educate pioneers about the true value of what they hold and encourage them to avoid dishonest practices. I urge everyone to actively participate in partial Pi and partial FIAT barter. The more GCV data we generate, the more secure our wallets will be. Therefore, it's important for everyone to read and share the Pioneer Handbook I wrote which has been translated into 30 languages to raise awareness among pioneers. By learning from the Pioneer Handbook and participating in GCV bartering, we can improve China's migration efforts and foster ecological development. This stability can ensure that the value of our Pi endures for future generations, rather than becoming worthless in a few years. Wouldn't that be something we want to preserve for our children and grandchildren? Today's message is lengthy but very important, and I hope you take the time to understand it. I wish our Guizhou Zunyi Conference great success! Thank you to all GCV Ambassadors, Merchants, and Pioneers for your incredible support! Your efforts today are planting the seeds for a prosperous future, and I hope you find safety and fulfillment in the days to come. May your wishes come true! Wishing you health and happiness! Let’s work together to create a better future! I also hope you all have a joyful Chinese New Year! Doris Yin 🪷🪷🪷 Founder, Global GCV Movement January 12, 2025

Doris Yin 东方紫莲🪷

18,340 views • 1 year ago

🚨 Protocol Update #9 It's incredible how time flies when you’re laser-focused on building and delivering the essential products that form the backbone of decentralized finance. Hatom has now been live on the Mainnet for over a year, and we're proud to say that this entire period has been free of issues or downtime. Our platform has been battle-tested during volatile market conditions, and each of our products has performed exactly as expected—solidifying our place as a cornerstone in the #MultiversX ecosystem. Describing last year as “incredible” feels like an understatement. We’ve witnessed unprecedented growth across the entire #MultiversX ecosystem, particularly in terms of TVL and yield opportunities. The day before Hatom launched its Lending Protocol and Liquid Staking on Mainnet, #MultiversX had a total TVL of $95 million. Within two weeks, the ecosystem surpassed $200 million in TVL, with Hatom driving over 50% of that growth. At its peak, Hatom reached over $280 million in TVL, accounting for more than 70% of the chain’s total TVL. What's even more remarkable is that, after initially using Treasury funds to incentivize users, Hatom has shifted to distributing rewards solely from protocol revenue. This marks the start of a fully sustainable, real-yield model, proving our products' rapid product-market fit and long-term viability. A Recap of the Past Year Here’s a quick overview of what we’ve accomplished in the past year: • Launched the first Lending Protocol in the #MultiversX ecosystem, along with the Liquid Staking Protocol on Mainnet. • Surpassed $100 million in TVL within just five days of the launch. • Deployed the HTM Booster Module and Accumulator. • Launched the Tao Bridge and Tao Liquid Staking, bringing over 33k $TAO into the #MultiversX ecosystem in just two weeks. • Implemented multiple upgrades to core infrastructure. • $HTM became the second-largest ESDT token after $EGLD. • Distributed over $3.85 million in rewards to our users. We are happy to announce that Hatom V2 is now live! After an incredible year of growth, we’re excited to take the next step toward becoming the leading liquidity hub across multiple chains. We invite you to explore our newly rebranded website at marking the beginning of our omni-chain journey. This rebranding reflects our bold vision and sets the stage for a full overhaul of our dApps, delivering a fresh and enhanced experience for all users. Achieving self-sustainability in such a short time, we now focus on research and development. Instead of pursuing many ideas, we’re committed to building high-impact products that create perfect synergies within our ecosystem. With that said, let’s dive into the key topics of this update: USH and Booster V2. Hatom USD (USH) We’ve highlighted USH in several updates, and it’s great to see the community recognizing its potential. USH is set to be one of the most impactful products on #MultiversX, providing a key revenue stream for Hatom while helping us maintain competitive rates and long-term sustainability. USH is the result of extensive research and careful development, designed to seamlessly fit into the Hatom ecosystem. While many DeFi projects are raising millions for new stablecoins, USH stands as another powerful product within our hub. The time has finally come for USH to be unveiled to the public, and we are excited to announce that USH will officially launch on Devnet on 28th October. While we’ve thoroughly tested for bugs internally, we’re excited to engage the community in this critical phase. To encourage participation, we’ll offer incentives for those testing USH on the Devnet, with more details to be shared at launch. Understanding USH's architecture is key to how it functions within our ecosystem. Let’s break it down step by step, starting with an explanation of each component. Facilitators USH’s minting process is driven by Facilitators—smart contracts responsible for the controlled minting and burning of USH. At launch, two primary facilitators will handle these tasks, each with distinct functionality: 1. Lending Protocol Facilitator The Lending Protocol Facilitator allows users to mint USH using a variety of supported collateral assets directly into the Hatom Lending Protocol. Unlike traditional lending mechanisms, where interest rates fluctuate based on the utilization rate, the minting of USH has fixed interest rates, thanks to Hatom's unique role as the entity managing the minting process. In a scenario where a user is minting USH through this facilitator using multiple assets as collateral, the protocol automatically prioritizes collateral with the lowest Minting APY. Let’s consider an example where a user deposits: - $1,000 in USDC (with a collateral factor of 80% and a 2% Minting APY) - $1,000 in BTC (with a collateral factor of 75% and a 3% Minting APY) - $1,000 in HTM (with a collateral factor of 70% and a 4% Minting APY) Based on these parameters, the user can mint a maximum of $2,250 worth of USH, distributed as follows: - $800 from $USDC (80% of $1,000) at 2% Minting APY - $750 from $BTC (75% of $1,000) at 3% Minting APY - $700 from $HTM (70% of $1,000) at 4% Minting APY The overall Minting APY will be a weighted average of these individual APYs, calculated based on the proportion of USH minted from each collateral type. Now, if the user decides to borrow only $1,000 worth of USH, the APY is determined as follows: - The first $800 will be borrowed from $USDC at 2% APY - The remaining $200 will be borrowed from $BTC at 3% APY This results in an effective Minting APY of 2.2%, reflecting a weighted average of the APYs across the borrowed amounts. It’s important to note that EGLD and wTAO, along with their liquid staking derivatives such as sEGLD and swTAO, can only be used as collateral in the Isolated Pools (which will be explained in the next section), not in the Lending Protocol 2. Isolated Pools Facilitator The Isolated Pools Facilitator allows users to mint $USH at zero interest using $EGLD, $wTAO, or their liquid staking derivatives ( $sEGLD or $swTAO) as collateral. Here’s how it works: When depositing EGLD or wTAO • These assets are staked through the Hatom Liquid Staking Protocol, generating the staking APY. • The staked assets are then deposited into the Lending Protocol, earning a supply APY, but are not activated as collateral. When depositing sEGLD or swTAO • When users deposit staking derivatives into the Isolated Pools, the protocol holds the staking derivatives, but the user's exposure is immediately shifted to the underlying asset ( $EGLD or $wTAO). This means the user no longer benefits from the staking rewards of the derivative, and instead, their exposure is entirely tied to the value and price movements of the underlying asset. • The staked assets are deposited into the Hatom Lending Protocol, earning the supply APY, but again not being activated as collateral. Since the protocol generates revenue from staking and supplying assets in the Lending Protocol, this income is used to incentivize the USH Staking Module. The protocol buys HTM tokens from the open market and distributes them, along with all fees generated by other facilitators, as rewards to stakers. We believe that the Isolated Pools Facilitator is one of the most important pieces of the USH ecosystem. Its potential impact on the TVL within both the Hatom ecosystem and the broader #MultiversX blockchain is immense and the revenue generated by this facilitator through fees will significantly bolster the overall growth of the protocol. To illustrate the potential of Isolated Pools, let’s use the following example: • $50 million worth of $EGLD is deposited into the Isolated Pools, generating a 6% staking APY • $50 million worth of $wTAO is also deposited, earning a 15% staking APY The total staking rewards generated from these assets would be: • $EGLD staking rewards: $50 million × 6% = $3 million annually • $wTAO staking rewards: $50 million × 15% = $7.5 million annually In total, the protocol generates $10.5 million in staking rewards annually. These rewards are then used to buy back HTM tokens from the open market, driving significant buying pressure on the HTM token itself. The purchased HTM tokens are distributed to USH LP stakers in the USH Staking Module, alongside the revenue generated by the Lending Protocol Facilitator. TVL and Yield Impact As we explore the broader impact of USH and the Isolated Pools, it becomes evident how these mechanisms contribute to the overall growth of the Hatom ecosystem, particularly in terms of TVL and potential yield generation. Based on the above numbers, if $50 million worth of $EGLD and $50 million worth of $wTAO are deposited into the Isolated Pools with a 75% collateral factor, we could mint up to $75 million worth of $USH. However, to prioritize safety, we’ll mint only 50% of the maximum, resulting in $37.5 million worth of $USH. In an ideal scenario, but also very unlikely, the $37.5 million $USH would be deposited in the Staking Module to generate rewards. In order for $USH to be deposited in the Staking Module, it is paired with another token (e.g., $USDC or $EGLD) to form Liquidity Pool (LP) position, contributing $75 million to the USH Staking Module. Additionally, the $100 million deposited in the Isolated Pools cycles through Liquid Staking and into the Lending Protocol, contributing a total of $300 million in TVL. Total TVL Breakdown: • $300 million from assets flowing through Isolated Pools ($100m) → Liquid Staking ($100m) → Lending Protocol ($100m) • $75 million from LP positions in the USH Staking Module Total TVL = $375 million As mentioned above, the $100 million deposited in Isolated Pools generates approximately $10.5 million annually in staking rewards (6% APY from $sEGLD and 15% APY from $swTAO). If all minted $USH is deposited into the Staking Module, the $75 million staked would benefit from these rewards, resulting in a 14% APY for USH LP stakers. On top of the protocol’s rewards, liquidity providers earn additional fees from their LP positions on decentralized exchanges, creating the perfect opportunity for all the participants in the USH Staking Module looking for attractive yields. USH Stability: The Peg Mechanism Ensuring the stability of USH is paramount, and to maintain its value close to $1 under all market conditions, we’ve implemented a robust dual peg mechanism. This system consists of two key layers of protection—Soft Peg and Hard Peg—designed to keep USH stable through both market-driven incentives and other mechanisms for scenarios where the Soft Peg mechanism can’t reclaim the peg. 1. Soft Peg Mechanism The Soft Peg Mechanism helps keep USH stable around its $1 value by encouraging market participants to act when USH trades above or below $1. When USH trades below $1 Users can buy USH at a discount, on a DEX, and repay their USH loans on Hatom, as USH is always valued at $1 on the protocol. This action removes $USH from circulation, helping to restore its price. When USH trades above $1 Users can borrow USH from the protocol at $1 and sell it on the open market at the higher price, increasing the circulating supply of USH and pushing its price back down to $1. 2. Hard Peg Mechanism (Redemption Mode) In cases where the Soft Peg alone cannot restore USH to $1 and its price drops significantly below the peg, the Hard Peg Mechanism is triggered through Redemption Mode. This mechanism allows any market participant to step in and help restore the peg by repaying USH loans for other borrowers, seizing their collateral at the full $1 value. It's important to note that Redemption Mode is only activated in the Isolated Pools and does not impact users minting USH through the Lending Protocol. Here’s how Redemption Mode works: When USH trades below $1 and the Redemption Mode is activated, redeemers can buy USH at the lower market price (e.g., $0.95), and use it to repay borrowers' debts at the full $1 value within the protocol. The redeemer receives collateral in the form of liquid staked tokens(such as $sEGLD or $swTAO) equivalent to the USH they repaid at its full $1 value, profiting from the difference between the discounted purchase price and the redemption value. The borrower being redeemed also benefits by receiving a redemption bonus, which allows them to keep a portion of their collateral after part of it is seized after loan was repaid. This system ensures that borrowers are not penalized during redemption, creating a balanced mechanism where both the redeemer and the borrower have something to gain. Redemption Mode differs from Liquidation in several ways: Redemption is triggered by USH falling below $1 and involves repaying borrower accounts to restore the peg. Both the redeemer and the borrower benefit, with the redeemer profiting from the price difference, and the borrower receiving a bonus from their collateral. Liquidation occurs when a borrower’s collateral falls below a certain threshold, making them risky. During liquidation, a portion of the borrower’s loan is repaid, and the collateral is seized, while also incurring a liquidation penalty. Redemption Mode uses a data structure known as a Red-Black Tree to efficiently monitor and rank all borrower positions within the protocol smart contract itself. This structure dynamically tracks borrowers based on their Borrow Limit Used, which is the percentage of collateral they have utilized relative to their borrowing capacity. The system prioritizes borrowers with the highest Borrow Limit Used, meaning those who have borrowed the most relative to their collateral are considered first for redemption. USH Airdrop Regarding the USH Airdrop, we would like to inform you that snapshots will end once USH is deployed on the Public Mainnet. The airdrop will be concluded shortly after, once all liquidity pools are stable and we determine the optimal moment to distribute the rewards to the community. USH Staking Module & Booster V2 The USH Staking Module will play a critical role in maintaining deep liquidity for USH while offering users high-yield opportunities. By staking USH LP tokens, such as USH/USDC and USH/EGLD, users can earn rewards generated by USH facilitators. This approach strengthens USH’s liquidity pools, making them robust enough to handle significant trades without destabilizing its price, thus reinforcing USH’s peg and overall stability. Beyond creating robust liquidity, the USH Staking Module serves as the key utility module within the USH ecosystem, designed to provide users with an opportunity to earn high yields on their USH holdings in a sustainable and organic way. All rewards distributed through the module are generated by various products across the Hatom ecosystem, ensuring long-term sustainability. For users seeking a more stable yield, the USH/USDC LP provides lower risk and steady returns. Those looking to leverage their EGLD holdings can opt for the USH/EGLD LP, which can be staked in the USH Staking Module. A key advantage of staking in the USH Staking Module is that rewards are based on the full value of the LP, not just the USH portion, maximizing your yield potential. As we continue to grow, we’ll be adding more LPs, providing users with even greater flexibility and options for staking their USH in the module. While our current focus is on LP tokens, we’re also exploring the possibility of allowing direct USH staking in the future, expanding the staking opportunities across the ecosystem. The Integration of Booster V2 with the Staking Module Booster V2 will be available for testing with the USH Devnet release, and with its introduction, we’ve strengthened the relationship between the HTM token and USH. Our ecosystem now features two independent boosters: one for the Lending Protocol and one for the USH Staking Module, each operating with the goal of maximizing yields for users. Key Improvements in Booster V2 Booster V2 brings several enhancements that elevate the functionality and user experience: Support for Multiple Token Types: Users will be able to deposit Pool Tokens, Farm Tokens, Dual Farm Tokens, or Staked HTM Tokens (via xExchange). Only the HTM portion will be considered for boosting. Unlimited Staking: The cap on HTM deposits will be removed, allowing users to stake without limits. This will foster a competitive environment where the more HTM you stake, the higher your potential APY. Integrated xExchange Management: Users will be able to manage their xExchange positions directly from the Booster dashboard. This will include creating pools, farming, dual farming, and staking HTM tokens, all from one convenient dashboard. Energy Management Integration: Booster V2 will allow users to manage their xExchange Energy directly from the dashboard, providing an additional way to boost rewards even further. Seamless Migration: Users will be able to migrate HTM between the Lending Protocol Booster and the USH Staking Module Booster without any cooldown periods, making it easier to optimize strategies across both modules. How the Yields Work Booster V2 will introduce a more structured and competitive approach to yield distribution across both the Lending Protocol and the Staking Module. HTM Booster in the Lending Protocol Base APY (First Batch): This is available to all users who stake a specific percentage of HTM relative to their collateral value. Any user can achieve this Base APY by staking the required amount of HTM. Boosted APY (Second Batch): After achieving the base level, users can boost their returns further by staking additional HTM, competing for the second batch of rewards. The more HTM staked beyond the base threshold, the higher the potential yield. USH Staking Module Yields Staking APY: Users who deposit USH-related LP tokens without boosting through the HTM Booster will still receive a Staking APY. This ensures that even passive participants which are not looking to stake their HTM in the Booster can take advantage of the USH Ecosystem to generate yields. Booster APY: Similar to the system in the Lending Protocol, users can stake HTM to unlock a Base APY. Beyond this threshold, any additional HTM staked will increase their APY in a competitive manner, allowing users to maximize their returns based on the amount of HTM they commit to boosting their positions. Rollout Plan for USH USH will be deployed in a phased rollout to ensure smooth implementation: Public Devnet: Open for testing, with incentives for participants to explore and stress-test the platform. Private Mainnet: A limited launch with partners to mint USH, bootstrap USH liquidity and generate initial protocol revenue. Public Mainnet: A full-scale launch, enabling all users to mint, stake, and trade USH. We know DeFi can be complex, which is why we’re committed to providing the tools and resources needed to navigate our ecosystem. With the USH Public Devnet launch, we’ll release updated documentation offering clear guidance on Hatom’s products. Developer documentation is also in the works, and we’re exploring the idea of a Hatom Academy for educational resources. Plus, we’ll soon roll out content focused on USH, helping users fully tap into its potential within Hatom and the MultiversX ecosystem. What’s Next? Hatom Pulse As Hatom grows, our focus remains on pushing DeFi boundaries while expanding across multiple ecosystems. Although this update doesn’t include a full roadmap—that will come later—our priority is clear: expanding Hatom across chains. To stand out in the competitive DeFi landscape, we’re committed to developing standout products. With that in mind, we’re excited to give you an exclusive preview of one of our most innovative products in development: Hatom Pulse. Over-collateralized non-custodial lending protocols, liquid staking, and over-collateralized stablecoins already exist on #Ethereum. What sets us apart is the synergy between these components within a unified ecosystem. By integrating these pillars, we tackle capital inefficiencies, allowing one protocol to enhance strategies that benefit the others, maximizing returns across the board. For example, when USH is minted, it means that EGLD is deposited, liquid-staked, and supplied in the lending protocol—all three protocols working in harmony. Hatom Pulse will elevate this synergy to another level, solving key issues faced by Aave, Compound Labs , and other leading protocols. We believe this innovation will be pivotal as we work to gain market share while expanding cross-chain. Our proof of concept will be deployed and battle-tested on #MultiversX, but the real growth will come when we scale this to markets that are thousands of times larger. This will be a turning point for Hatom. So, what is Hatom Pulse? On Hatom, like on Aave and other leading lending protocols, the largest assets used as collateral are often not borrowed, leading to substantial revenue loss for the protocol. This also results in very low income on the supply side, as borrowing fees depend on utilization rates, which only increase when borrowing activity rises. Generally, lending protocols are used to provide assets for borrowing stablecoins or for leveraging liquid staking strategies. This inefficiency locks up billions of dollars in dormant assets, and users earn very low supply rates on their collateral, which doesn’t help offset their loan interest. Hatom Pulse is designed to address these inefficiencies by leveraging the synergy between our existing products. It creates sophisticated vaults that activate dormant assets, unlocking advanced yield opportunities through a delta-neutral strategy. By utilizing assets like $EGLD, $sEGLD, $wTAO, and $swTAO, Hatom Pulse enables users to engage in delta-neutral strategies, where we long and short these assets on (CEXs), earning funding rates and staking rewards while keeping their assets intact. (The exact strategy, along with all the details, will be shared once USH is fully established). Initially, these vaults will operate on CEXs, where liquidity is highest, and will be managed through custodians like Copper.co to mitigate counterparty risks. Later, we plan to extend this to DEXs where all operations will be governed by smart contracts, ensuring full decentralization. serves as a strong proof of concept for us in this regard. However, our strategy will differ, as our focus will be on protecting the unit value, rather than the dollar value. Although Hatom Pulse is still in its research phase, early estimates suggest that this product alone could generate over 18% annual returns on $EGLD and more than 35% on $wTAO, with what we believe to be minimal risk. It’s important to note that these figures reflect current metrics based on internal calculations and may slightly differ upon product launch. But imagine reaching this on #Ethereum, while allowing users to borrow using their assets—this could be a disruptive protocol. We believe Hatom Pulse has the potential to become a cornerstone product as we transition into an omni-chain future. In a competitive DeFi landscape, it could give us a significant edge by offering something truly groundbreaking, capable of competing with well-established protocols across various chains. This strategy represents immense untapped potential. Hatom Pulse is being developed for risk-averse users who seek higher returns without excessive risk. By addressing inefficiencies in current DeFi strategies, we aim to offer a secure, robust option for yield generation that could rival established protocols. It's been an intense year for our team, and we sincerely thank the community for their patience, trust, and unwavering support as we've worked hard to build and deliver these groundbreaking products. As Hatom's omni-chain expansion nears, we remain focused on improving our existing products and researching new innovations to stay ahead in this competitive market. Our goal is to build a comprehensive DeFi ecosystem, accessible across all blockchains. With USH approaching its Mainnet release, we're proud of how our products have reshaped the DeFi landscape on MultiversX. By filling key gaps in the on-chain economy, we've created opportunities for users to generate yield, unlock the potential of decentralized finance, and provide strong utility for EGLD. In just over a year, we’ve built a strong ecosystem, but this is only the beginning. We’re ready to go even further, developing better products and unlocking new opportunities for our users. We’ll share more about our expansion plans in a dedicated post, staying focused on what matters most. Rest assured, what’s coming will be truly impressive for Hatom and our growing community!

Hatom Labs

182,997 views • 1 year ago