Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Speed & power start from the ground Want to get better? Do this routine 2-3x per wk! 1. Bilateral extensive pogo 2x30 seconds 2. Unilateral extensive pogo 1x30 seconds 3. Bilateral intensive pogo 2x8 seconds 4. Unilateral intensive pogo 1x8 seconds Build big time feet/ankles!

36,313 Aufrufe • vor 1 Jahr •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

深藏不漏的本地视频生成大师 Gary 用黑白深度+精准分镜提示词复刻了陶阿狗君的创意换装视频 我用 Minimax H3 + gary的黑白深度,微调了一下提示词,就用新的模特复现了 复刻的具体步骤: 1. 准备好模特的人物三视图 2. 生成四个场景下的人物参考图,可以用 gpt image 2.5 换人 3. 稍微调整提示词里对原来模特的特征锁定,按照新模特重写 4. 用新的提示词 + 参考图片 + 黑白深度视频交给 Minimax H3 生成 微调后的提示词👇 defines Rei's opening outfit only: a fitted black satin mini dress with thin spaghetti straps, a lace-trimmed V-neckline, a small front bow with trailing ribbons, off-the-shoulder lace ruffle sleeves, a ruched bodice, a ruffled lace hem, bare legs and black pointed-toe stiletto pumps. Preserve the dress's silhouette, materials and front/back construction as shown. Use this outfit for the white-studio opening only; do not reproduce the character-sheet layout. defines her Paris outfit, front-facing presentation pose, drink, sunglasses, city map and UI. defines her Rome outfit, front/three-quarter presentation pose, props, city map and UI. defines her Cairo outfit, presentation pose, shopping bags, city map and UI. defines her Sydney outfit, front-facing finishing pose, city map and UI. is an unmodified slightly turned standard face crop from the original Rei character card. It is the highest-priority facial identity reference for ALL outfits: preserve its eye shape and spacing, eyelids, nose, lips and natural proportions without beautification. Pictures 2–5 supply new city outfits and environments, never a substitute face. This is the same adult woman in different clothes throughout the entire video. is a black-and-white relative-depth reference for body movement, camera pullback and chronological timing only. Depth brightness represents distance, not skin color or face detail. Render natural full color using the pictures. A featureless face in the depth map must not become the back of Rei's head. Use the scene pictures to resolve which way her face and chest point after each turn. Generate ONE continuous 12-second vertical video. The timed Shot labels below are successive phases of this one output, not instructions to insert camera cuts. Keep the reference timeline at its original speed. Do not compress the opening to make room for the cities. Only Rei rotates during transformation; the camera must not orbit her. A back view is a brief passing orientation during a spin, never a held destination pose. Each completed transformation reveals her face and the FRONT of the new outfit, as in its corresponding picture. Change clothing, handheld props, city map and city-name UI together; do not change only the background while leaving the previous outfit. Shot 1 [0.00–2.50 seconds]: White studio and opening black sailor uniform from . Rei looks at the phone held in front of her chest. Follow the opening camera pullback in . Her torso remains generally facing the camera; allow the original small phone and head gestures. This interval contains no full-body spin, no location pin yet, no city background and no wardrobe change. Shot 2 [2.50–3.15 seconds]: The red location pin appears above Rei. This event ONLY adds the pin. She continues looking at her phone in the white studio, still wearing the opening black sailor uniform. The pin does not trigger a spin or a scene change. Do not begin the transformation here. Shot 3 [3.15–4.05 seconds]: Follow the small preparatory side pivot in , then return toward the camera-facing starting orientation while holding the phone. This is the preparatory movement, not the rapid transformation spin. Retain the opening black sailor uniform, white background and red pin throughout this interval. No Paris map or summer outfit before 4.05 seconds. Shot 4 [4.05–6.00 seconds]: At 4.05 seconds begin the rapid transformation rotation following . Pass continuously through the side and back orientations without pausing. Around 4.25 seconds, during this rotation, change together into the Paris outfit, drink, sunglasses, map and navigation UI from . By 4.50 seconds complete the rotation and reveal Rei FACING THE CAMERA, with her face and the front of her navy-and-ivory striped top visible. From 4.50 to 6.00 seconds remain in this front/three-quarter FRONT presentation, with the drink and sunglasses gesture from . The cream full-length trousers, striped top and red neck scarf replace the opening black sailor uniform completely. No back-facing hold, no phone, no opening black sailor uniform during the Paris presentation. The top search bar and bottom location card both read 巴黎. Shot 5 [6.00–8.00 seconds]: Begin the next reference turn at 6.00 seconds. Around 6.25 seconds change outfit, props, city and UI to Rome from . Complete the turn by 6.55 seconds into a front/three-quarter FRONT view showing Rei's face, sage-green midi dress with brown belt. Until 8.00 seconds use the Rome pose with the paper guide and warm Trajan column and Roman Forum map. Both city labels read 罗马. Back-facing frames occur only while the turn is in progress; no stopped back view after the turn. Shot 6 [8.00–10.00 seconds]: Begin the next reference turn at 8.00 seconds. Around 8.25 seconds change together to the ivory linen shirt, beige full-length trousers, ochre scarf, shopping bags, Cairo map and UI from . By 8.55 seconds reveal her face and the front of the ivory linen shirt. Hold the front/three-quarter presentation with natural small movements until 10.00 seconds. Both city labels read 开罗. Keep the Cairo outfit and map together; do not introduce previous clothes or Sydney early. Shot 7 [10.00–12.00 seconds]: Begin the final reference turn at 10.00 seconds. Around 10.25 seconds replace the Cairo clothes and shopping bags, city and labels with the Sydney sky-blue overshirt, white T-shirt, navy knee-length shorts and white sneakers, Sydney Tower map and UI from . By 10.55 seconds face the camera with the white T-shirt beneath the sky-blue overshirt visible, hands raised and one leg bent backward as in . Keep this front-facing Sydney presentation with subtle natural movement through 12.00 seconds. Both city labels read 悉尼. No additional spin, destination or wardrobe change after this final reveal. Keep Rei's facial identity, natural body proportions, hair and bangs throughout. Clothing changes must not alter her face, age, body shape, hair color or hair length. Do not reproduce character-sheet panels, extra people, face closeups or depth-map gray colors. Do not reproduce the blank faces in Picture 1's full-body clothing views; Rei must have her normal face whenever her face is visible. Keep the city UI style consistent; update both location labels at each city change. No generated speech or music. The angle in Picture 6 describes identity only, NOT a fixed head pose. Face direction, expression, head tilt, hand gestures and full body movement follow Video 1, with each completed city presentation matching its corresponding scene picture. In Paris the sunglasses are opaque black and fully cover both eyes and eye sockets. Keep the sunglasses on throughout the Paris presentation. At the Rome transformation, switch to the accessories shown in Picture 3.

Hoody

34,819 Aufrufe • vor 12 Tagen

Gemma 4 26B A4B MoE - 500+ t/s decode - Single RTX 4090 (24 GB VRAM) - Llama.cpp concurrency 24 - q8 kv cache How many API users can you simultaneously host on a single RTX 4090 (24 GB VRAM) before it crashes? Yesterday, I proved you can host 14 active users using unquantized memory. Today, I used 8 bit KV Cache Quantization to hack the VRAM footprint. I successfully scaled to 24 concurrent users without a single dropped connection. A 71% server capacity boost for free. By adding the -ctk q8_0 -ctv q8_0 flags to llama.cpp, you compress the KV cache context memory from 16 bit to 8 bit. This unlocks massive concurrency limits on Gemma 4 26B (MoE) on a single 24GB consumer GPU. Here is the exact telemetry from pushing 8 bit quantization to its absolute physical edge: # TEST 1: The 24 User Concurrency Max Server Config: 24 slots (np 24) | 4,096 context per slot | 98,304 Total Context Client Load: 24 simultaneous requests (2,000 token prompt per user) Unquantized KV cache for this load requires 28GB+ VRAM (Instant OOM). Quantized to Q8, it allocated safely at 23.35 GB. The C++ engine crunched the entire batch in 28.5 seconds. Decode Speed: 21 t/s (Per User) | 500 t/s (Agg) # TEST 2: The 48 User Queue Overload What happens to a compressed cache during a traffic spike? Server Config: 24 slots (np 24) | 4,096 context per slot | 98,304 Total Context Client Load: 48 simultaneous requests (2k token prompt per user) Zero queue drops. The scheduler flushed and hot swapped the 8 bit memory flawlessly on the fly, completing all 48 users in 66.0 seconds (a perfect 2.3x queue scaling multiplier). Decode Speed: 18 t/s (Per User) | 430 t/s (Agg) # TEST 3: The 8 User RAG Slam Server Config: 8 slots (np 8) | 60,000 context per slot | 480,000 Total Context Client Load: 8 simultaneous requests (30k token prompt per user) It allocated 23.83 GB VRAM and chewed through ~240,000 prefill tokens in 46 seconds under massive memory pressure. Prefill Speed: 6,200 t/s (Agg) Decode Speed: 22 t/s (Per User) | 175 t/s (Agg) # The Engineering Alpha (The Quantization Tradeoff): You gain a massive 71% increase in server capacity, but what do you lose? Compute latency. Because the cache is stored in 8 bit, the GPU's cores have to dequantize the memory back to 16 bit on the fly during every single prefill step. In my unquantized tests yesterday, single slot prefill was hitting ~1,500+ t/s. Today, under the heavy 48-user Q8 load, prefill dropped as low as ~750 t/s. You trade a few seconds of initial prefill latency to essentially double your API hosting capacity. For production high volume SaaS, this is the ultimate unit economics cheat code. Here is the exact command to run a 24 user Q8 continuous batching server on your own single 4090, single 3090 or any 24gb vram rig: ./build/bin/llama-server -m gemma-4-26B-A4B-it.gguf -c 98304 -np 24 -b 2048 -ub 2048 -ngl 99 -fa on -ctk q8_0 -ctv q8_0 --port 8080 (Note: -c 98304 allocates exactly 4,096 tokens of context per user across 24 slots). Hugging Face links to the Unsloth Gemma 4 26B QAT quants along with performance graphs available in the replies. Would you trade 3 seconds of Time To First Token latency to double your active user capacity?

Alok

17,465 Aufrufe • vor 2 Monaten

iPhone Takes a Photo of You Every 5 Seconds While you’re using your iPhone, it takes a photo of you every 5 seconds. More precisely - when you pick up or put down your iPhone, and also when it’s unlocked and sits motionless for a long time, the Face ID sensors take pictures every 5 seconds. You can see this through the infrared lens: it looks like the iPhone is taking a photo in the dark with a flash. Even if you cover the sensor, it will still keep trying to take a picture. In reality, of course, the iPhone isn’t taking photos. It scans your face to figure out whether you’re looking at it. In Face ID settings this mode is called Attention Awareness. You’re working on your iPhone, it’s charging, you get a call - and the iPhone “takes pictures” with the Face ID camera. As soon as it realizes you’re looking at it, the ringtone gets quieter. The same thing works with alarms and other notifications. This is not actually taking photos - it’s the TrueDepth infrared sensors scanning your face for the Attention Aware Features. How to turn this off on iPhone. You need to disable Attention Aware Features. Method 1 (main one): 1. Open Settings. 2. Tap Face ID & Passcode. 3. Enter your passcode. 4. Find Attention Aware Features and turn it off. Method 2 (via Accessibility): 1. Open Settings → Accessibility. 2. Tap Face ID & Attention. 3. Turn off Attention Aware Features. After that, your iPhone will stop automatically lowering the volume of calls, alarms, and notifications when you’re looking at the screen, and it won’t scan your face as actively for “attention.”

Officer's Notes

18,831 Aufrufe • vor 26 Tagen

I just ran Gemma 4 31B on @CerebrasSystems at 1,800+ tokens/sec and it's multimodal. For context: that's 35x faster than a typical GPU endpoint, and the first token (reasoning included) lands in 1.5 seconds. This isn't a benchmark slide, I recorded the inference live. Prompt I used: "Create a simulation of an iPhone. Include at least one working dummy note taking app, a functional notification pulldown, high quality graphics, single HTML file, any libs via CDN." - Generation time: 3 seconds. - Notes app worked. - Notification panel worked. - Rendered first try. This is what wafer-scale inference unlocks, not just "faster," but a different category of product. When generation is this fast, you stop waiting and start iterating in real time. Why this matters: Gemma 4 31B is Google DeepMind's flagship open weight model, Apache 2.0 licensed, dense (not MoE), and built for efficiency over raw parameter count. It scores close to Claude Haiku 4.5 on the Artificial Analysis Intelligence Index (30 vs 29) but runs ~18x faster on Cerebras. It's also the first multimodal model on Cerebras's platform, meaning you can now feed it screenshots, documents, charts, and UI states at wafer scale speed. # Applications I'm most excited about: - Screenshot → Insight: Drop in a dashboard or document screenshot, get structured findings back instantly. no waiting, no batching. - Live UI generation: Full interactive interfaces (like my iPhone sim) generated and rendered in under 2 seconds. - Screenshot -> Patch: Feed it a broken UI + console error, get a minimal code fix and verification steps back. - Computer use & agentic loops: See -> reason -> act - verify, fast enough to keep a human in the loop instead of waiting on the model. - Long context summarization: Full research reports condensed into decision ready summaries you can read and requery in one sitting. The bigger unlock isn't the speed number itself, it's that agentic and multimodal loops (see -> reason -> output -> tool call -> verify -> retry) finally run in real time instead of feeling sluggish. As Logan Kilpatrick (Logan Kilpatrick) put it: "If every model was doing 2,000 tokens per second, you wouldn't build the same product and just have it be faster, you'd build different products." Gemma 4 31B is live now on Cerebras Inference Cloud in public preview. If you're building multimodal, agentic, or real time apps, this is worth testing today. What would you build with such insane inference throughput?

Alok

12,962 Aufrufe • vor 3 Monaten

Here are medically supported methods that can help. 1. Practice the start–stop technique This is one of the most recommended methods for premature ejaculation control. How it works: - During sexual activity, stop stimulation when you feel close to ejaculation. - Wait about 20–30 seconds until the urge reduces. - Start again. Repeating this trains the body to delay ejaculation. 2. Try the squeeze technique When you feel very close to climax: - Gently squeeze the head of the penis for a few seconds. - This reduces arousal and delays ejaculation. This technique is commonly used in behavioral therapy for premature ejaculation. 3. Strengthen pelvic floor muscles Stronger pelvic muscles improve ejaculation control. The best exercise is Kegels. How to do it: - Tighten the muscles you use to stop urine. - Hold for 5 seconds, then relax. - Repeat 10–15 times, 3 times daily. 4. Slow down the pace Fast thrusting increases stimulation and leads to quicker ejaculation. Try a slower rhythm, short pauses, and changing positions to control arousal. 5. Focus more on foreplay Longer foreplay reduces pressure to perform quickly during penetration and improves satisfaction for both partners. 6. Reduce anxiety and stress Performance anxiety can cause early ejaculation. Relaxation techniques, deep breathing, and confidence building can improve sexual endurance. 7. Use condoms They may reduce sensitivity slightly, helping some men last longer. 8. Maintain good physical health Regular exercise improves blood flow, stamina, and hormonal balance, which supports sexual performance and reduces the risk of erectile dysfunction. Helpful habits: regular workouts, good sleep, limiting alcohol, and avoiding smoking. Lasting longer during sex usually improves with practice, pelvic muscle strengthening, and better arousal control. Hope you learnt something new? Please share for others to learn.

declasiq_esmi

57,109 Aufrufe • vor 1 Monat

I spent 48 hours running AI from my phone. Here are 11 things that turned out to be possible and 3 that almost cost me money Forgot my laptop at home and thought the day was lost. Opened a terminal from my phone and decided to see how long I could last Lasted 2 days. Not just lasted but made $840 What works from a phone: 1. Set up Claude Code through SSH in 10 minutes while riding the subway 2. Get Telegram pushes every time a wallet enters a position 3. Copy a trade with 1 tap without taking out my earbuds 4. Launch scripts by voice through Shortcuts 5. Monitor 3 wallets simultaneously without a single lag 6. Get a morning report at 7 AM as a regular message 7. Rebuild the bot when it crashed while sitting in a cab 8. Check PnL without opening a browser 9. Add a new wallet to tracking in 30 seconds 10. Set up auto-copying without confirmation on verified wallets 11. Get a full strategy breakdown of a wallet through Claude Code in a regular chat And here is what almost killed the deposit: 1. My finger slipped and I entered at twice the planned size. Did not notice for 20 minutes. Got lucky that the position ended up in profit but it could have gone very differently 2. My phone died at 2 AM. Missed the exit signal and the position dropped $110 while I slept. By morning I realized that a power bank is just as much a part of the strategy as the bot itself 3. The delay when copying was 40 seconds. On a 15-minute market that is an eternity. The price moved from 8 cents to 23, and instead of a 12x return I got a 4x. Still profit but you feel the difference immediately Total for 48 hours: +$840. Screen time on the phone: 47 minutes. Never needed the laptop The entire time I was following the same wallet. That is the 1 that was sending me signals at 2 AM: The phone turned out to be a fully functional control panel. But this control panel has no safety switch. And that is worth remembering every time you are tapping with 1 hand in the coffee line

Blaze

93,477 Aufrufe • vor 6 Monaten

CRAZY 🚨 I built a scanner that shows you world's sentiment on social media in real-time Right now, it is pointed at Netanyahu's speech at the United Nations Imagine reading all posts on X and Instagram at the speed of light Meet How does it work? 1. It picks up posts from a set of keywords, topics, and accounts from X and Instagram 2. then uses the Sage decision model to filter if it's relevant to the UN speech, whether it's supportive or critical for Benjamin Netanyahu - בנימין נתניהו, and the post's narrative topic (ie., is the post about the Speech, Israel, or Iran?) 3. The image and text are classified separately, providing different weights to the sentiment and narrative rankings. This all produces short, mid, and long-term sentiment measurements. 1. Ben reacts every 5 seconds. 2. Narrative weights are updated every hour. 3. The 24h dial shows conglomerate sentiment measurement. I built this yesterday to test out Sage's new reasoning and image capabilities. It's the only AI model we used to build it. This is the power of decision models. Sage reasons when things are complicated and can interpret images, making this app possible. Disclaimer: This tool is politically neutral. It's difficult to touch anything related to geopolitics and be fairly neutral. We did our best on filter selection - if you think the method could have been better, please drop a comment below. I will use World Signal for future trends. If you'd like me to point it at a narrative next week, send me a comment below.

Big Iron Chris

32,250 Aufrufe • vor 9 Tagen

Turning a simple idea into a dreamy visual ✨ Made with Supercool.com Video.Prompt⤵️ Create an extremely funny, adorable cinematic 3D animated short film, exactly 10 seconds long. CHARACTER CONSISTENCY: Use the exact same tiny baby otter character throughout the entire video. The otter is extremely small and adorable, with fluffy brown wet fur, a round cute face, huge expressive eyes, tiny ears, small whiskers, chubby cheeks, and tiny paws. Keep the exact same facial features, fur color, proportions, anatomy, and appearance in every shot. SCENE 1 — INSTANT COMEDY START (0–2 seconds): Open with the baby otter sitting on a wet tree stump beside a fast-flowing rainy forest river. The otter is holding a tiny fish-shaped snack with both paws and happily eating it. Suddenly, a gigantic crocodile silently rises behind the otter. The crocodile opens its enormous mouth directly behind the tiny otter. The otter keeps eating completely unaware. Use a funny camera angle showing the enormous crocodile behind the tiny innocent otter. SCENE 2 — THE REALIZATION (2–4 seconds): The otter suddenly notices the crocodile's giant shadow. It slowly stops chewing. Its eyes move upward. The otter slowly turns its head. It sees the gigantic crocodile. Extreme close-up of the otter's face. Its eyes become ridiculously huge. Its mouth opens in complete shock. The crocodile looks extremely confident. SCENE 3 — UNEXPECTED REACTION (4–6 seconds): Instead of running away, the tiny otter suddenly stands on its little back legs. It raises both tiny paws toward the crocodile. The otter tries to look extremely scary. It puffs out its tiny chest. Its cheeks become puffed up. It makes the funniest possible "angry" face. The contrast between the tiny harmless otter and the enormous crocodile should be hilarious. SCENE 4 — CROCODILE PANICS (6–8 seconds): The crocodile suddenly gets confused. It leans backward. The otter takes one tiny step forward. The crocodile becomes increasingly nervous. Then the otter suddenly sneezes. A tiny puff of water shoots from its nose. The crocodile reacts as if it has been attacked by a monster. It dramatically jumps backward into the river with an enormous SPLASH. SCENE 5 — FINAL COMEDY PUNCHLINE (8–10 seconds): The otter looks at the camera with a completely confused expression. It looks at its own tiny paws. Then it looks back toward the river. The crocodile's head slowly appears from the water again. The crocodile looks embarrassed. The otter casually goes back to eating its fish snack as if absolutely nothing happened. End with the crocodile staring at the otter in disbelief while the otter happily continues eating. VISUAL STYLE: High-end cinematic 3D animated movie quality, adorable stylized characters, extremely expressive facial animation, exaggerated comedy, realistic wet fluffy fur, detailed rain droplets, realistic water simulation, dramatic splashes, lush rainy forest, misty atmosphere, cinematic lighting, volumetric fog, beautiful reflections, smooth animation, dynamic camera movement, cinematic depth of field, polished feature-film rendering. COMEDY: Make every reaction exaggerated and visually funny. The humor should come from the extreme size difference between the tiny otter and giant crocodile, unexpected pauses, facial expressions, sudden movement, exaggerated splash effects, and perfect comedic timing. The otter should never actually hurt the crocodile. Keep the tone playful, harmless, cute, and family-friendly. CAMERA: Fast cinematic close-ups, dramatic reveal, reaction close-up, low-angle crocodile shot, slow-motion splash, and final comedic close-up. FORMAT: Exactly 10 seconds. Vertical 9:16. Fast pacing. Strong hook in the first second. Clear visual storytelling. Big comedic payoff at the end. No dialogue. No text. No subtitles. No logos. No watermark.

Zarnab Ai

11,311 Aufrufe • vor 22 Tagen

how to start a faceless youtube channel in 2026 with affiliates that actually makes money i've written 5,000+ scripts across 50+ niches. here's the exact process i'd follow if i had to start from $0 tomorrow STEP 1 - pick a niche with money in it, not one you "like" a million views from teenagers and a million views from 55-year-olds are not the same paycheck. finance, retirement, sleep, and world politics pay 5-10x what entertainment does because the advertisers are different. boring + high CPM beats fun + broke. every time. STEP 2 - steal what already works don't reinvent anything. find 5 channels that went from 10k to 100k recently (not the giants the ones who just broke through) and study what they changed. their titles, their hooks, their first 30 seconds. STEP 3 - the script is 70% of everything a mediocre video with a great script beats a beautiful video with a boring one every single time. cold open on the most intense moment. new question every 30 seconds. never let them guess what's coming next. if you can delete a sentence and lose nothing, delete it. STEP 4 - package it or it dies nobody sees the video. they see the title and thumbnail first. one visual, one emotion, 3 words max on the thumbnail. a title that makes NOT clicking feel like a mistake. 40 good videos die every day because of 40 bad titles. STEP 5 - make money before youtube pays you affiliate links in the description from video 1. an email freebie by week 2. a $19 guide by month 2. you should earn before adsense ever turns on. waiting for monetization is the #1 beginner mistake. STEP 6 - read the retention graph like scripture it tells you the exact second people left and why. cliff in the first 30s = fix your intro. dip in the middle = that's the sentence that bored them. it's the only feedback that matters. that's it. that's the whole game. the niche is your ceiling. the script decides how close you get to it. i post this stuff daily. follow so you don't have to learn it the slow, expensive way like i did.

Tryahd

42,048 Aufrufe • vor 1 Monat

Qwen3.8-Flash-Next now reaches ~43 tok/s after a 122,902-token prompt on ONE DGX Spark. ⚡🚀 MTP k=2 won my draft-depth sweep, with +42.5% mean decode over no draft. The PLE table stays fully on-device. I promised the deeper MTP tests. Here are the results, and now you can explore them in an interactive benchmark page too. 𝗧𝗪𝗢 𝗗𝗥𝗔𝗙𝗧 𝗧𝗢𝗞𝗘𝗡𝗦 𝗪𝗢𝗡 Mean single-request decode with 32K context configured: MTP k=2: 39.21 tok/s MTP k=3: 36.42 tok/s MTP k=1: 35.18 tok/s No draft: 27.51 tok/s k=2 also produced the fastest individual sweep run: 41.34 tok/s. Four runs each for no draft, k=1 and k=2. Seven for k=3. Decode excludes time to first token. Here, k means speculative draft depth, not quantization bits. k=3 produced more tokens per step, but the extra drafting work did not pay off in throughput. k=2 is my current pick for this setup. 𝗧𝗛𝗘 𝟭𝟮𝟯𝗞-𝗧𝗢𝗞𝗘𝗡 𝗣𝗥𝗢𝗠𝗣𝗧 𝗧𝗘𝗦𝗧 I then ran a separate long-prompt comparison: Actual input: 122,902 tokens Configured context: 262,144 Requested output: 128 tokens One request at a time MTP k=2: ~43 tok/s No draft: 26.2 tok/s Time to first token: 110.6 seconds with MTP 107.0 seconds without it The win here is generation speed, not faster prefill. To keep the scope clear: 256K was the configured limit. This was a real ~123K input, not a completely filled 256K window or a full k sweep at that depth. 𝗣𝗟𝗘 𝗦𝗧𝗔𝗬𝗦 𝗢𝗡 𝗧𝗛𝗘 𝗦𝗣𝗔𝗥𝗞 Whole model on-device: 78.57 GiB Packed 5-bit PLE table: 30.4 GiB, included in that total No NVMe PLE offload in this build. This is still turboderp’s 3.05bpw_h5_ng5 EXL3 pack, served through my vllm-exl3 integration. My work here is the serving integration and testing. These are preliminary performance measurements, not a quality evaluation or a claim of bit-exact full-output parity. 𝗘𝗫𝗣𝗟𝗢𝗥𝗘 𝗧𝗛𝗘 𝗥𝗘𝗦𝗨𝗟𝗧𝗦 The benchmark page has the individual sweep values, long-prompt comparison, and measurement scope. You can play the animation, export the charts, or download the HTML and data to render them yourself. No Spark needed to view the results. Credit to turboderp / ExLlamaV3 for the pack and kernels, vLLM for the serving engine, and Qwen Qwen Developers for the model. Recipe + reproduction: Interactive benchmark:

Cruz

12,258 Aufrufe • vor 25 Tagen