Can LLMs keep correct medical judgment when the context... is misleading? LLMs ace medical exams. But real-world use cases rarely come with clean context the way exams do. So we investigated whether LLMs preserve correct medical judgment under misleading context. Our findings: > Across 11 frontier model configs, accuracy dropped from 71% to 38%. > A 14-member clinical panel from 7 countries identified serious potential harm in 38.2% of reviewed cases. 1/8show more

Hongjian Zhou
52,945 views • 3 months ago
LongWriter Unleashing 10,000+ Word Generation from Long Context LLMs... discuss: Current long context large language models (LLMs) can process inputs up to 100,000 tokens, yet struggle to generate outputs exceeding even a modest length of 2,000 words. Through controlled experiments, we find that the model's effective generation length is inherently bounded by the sample it has seen during supervised fine-tuning (SFT). In other words, their output limitation is due to the scarcity of long-output examples in existing SFT datasets. To address this, we introduce AgentWrite, an agent-based pipeline that decomposes ultra-long generation tasks into subtasks, enabling off-the-shelf LLMs to generate coherent outputs exceeding 20,000 words. Leveraging AgentWrite, we construct LongWriter-6k, a dataset containing 6,000 SFT data with output lengths ranging from 2k to 32k words. By incorporating this dataset into model training, we successfully scale the output length of existing models to over 10,000 words while maintaining output quality. We also develop LongBench-Write, a comprehensive benchmark for evaluating ultra-long generation capabilities. Our 9B parameter model, further improved through DPO, achieves state-of-the-art performance on this benchmark, surpassing even much larger proprietary models. In general, our work demonstrates that existing long context LLM already possesses the potential for a larger output window--all you need is data with extended output during model alignment to unlock this capability.show more

AK
50,995 views • 2 years ago
Big win for open-source LLMs! DeepSeek V4 Pro holds... the top open-weights score on SWE-bench Verified, in the GPT-5.5 range. GLM 5.2 leads the open-weight intelligence index and sits near the closed frontier on long-horizon coding. But this leaderboard number is a weak proxy for real performance. It comes from one task set, run through one harness, served at one precision. The same weights can even score differently across providers, since many hosts quantize activations to fp8 and drift the model off its reference weights. Real performance is determined based on whether a model can read a repo, make coordinated edits across files, run the tests, and recover when one breaks. By that measure, the top open models hold up, but only inside the right harness. The teams that actually put DeepSeek V4 into production pipelines as a frontier substitute got there through the harness they built around the model, not by picking a stronger model. If you want to see this in practice, Cline (64k+ stars) has actually built that harness around open models, tuned so they run at production quality. And it's tuned so that these LLMs can run at production quality, with plan and act modes, checkpoints, and terminal feedback. ClinePass is the new access layer on top of it. It runs a curated set of those models inside Cline, narrowed to the ones tested for coding-agent use, with 2 to 5x the standard rate limits and no separate provider accounts, keys, or billing to track. The video below shows the setup, and I worked with the team to put this together. It runs alongside custom keys and local models as well, not in place of them.show more

Avi Chawla
44,124 views • 2 months ago
Does LLM really need to be a helpful assistant... all the time? No. If you want to simulate people, “perfectly helpful” could be the wrong objective. Meet OdysSim, a journey toward LLMs beyond assistants, as behavioral foundation models (10B tokens of real human behavior; 23 sim benchmarks, finally in one place. new open models: outperform or on par with GPT-5.5, Gemini 3.1, or Claude Opus 4.7 in many behavior-sim dimensions). Human behavior simulation is becoming essential. Agent evaluation needs realistic users before real users show up. Medical and classroom training need realistic patients and students. Social science needs synthetic participants at scale. But real people are not ideal assistants. Real patients panic or ignore good advice. Real students misunderstand. Real customers are vague, picky, impatient, or simply leave. Human behavior is messy, diverse, and often imperfect. Frontier LLMs are getting better at math, code, and long-horizon tasks. They are NOT getting better at simulating human behavior. If anything, they drift the other way: more assistant-ish, more homogeneous, fewer of the errors and quirks real humans show. This is no accident. The whole pipeline is built for helpfulness and task success, not behavioral realism. And you can't prompt your way out of that. So we rethink the recipe from scratch and release: 🧠 The OdysSim corpus: 21.4M real human interactions (~10B tokens) from 62 sources, every conversation retrofitted with social grounding (who is talking, and why) 📏 SOUL-Index: 23 human-behavior benchmarks unified into one suite across 5 axes 🤖 OSim-8B: open weights; tops more SOUL-Index benchmarks than any frontier model, acts more like a real user than any of them on τ-bench (nearly matching real humans in the reaction dimension), and writes far more human-like text along the way.show more

Xuhui Zhou
143,103 views • 3 months ago
Call for Genuine Dialogue with Foreign Medical Graduates (FMGs)... We are deeply concerned by the recent actions of the National Medical Commission National Medical Commission regarding the new policies affecting Foreign Medical Graduates #FMGs . Despite multiple attempts by #FMGs to engage in constructive dialogue about these policies, the doors have been closed to them, leaving many feeling unheard and marginalized. Key Issues Highlighted: 1.Lack of Engagement with FMGs: •When FMGs approached the NMC to discuss the recent circulars impacting their education and career pathways, they were met with a closed door. This lack of openness contradicts the NMC’s stated commitment to inclusivity and transparency. •On other occasions, the NMC has publicly invited dialogue with medical professionals on various issues. The exclusion of FMGs from such conversations suggests a discrepancy in how different members of the medical community are treated. 2.The Need for Authentic Communication: •Effective policymaking requires genuine engagement with all stakeholders. FMGs, who represent a significant portion of the medical workforce, deserve to have their voices heard and their concerns addressed in a fair and transparent manner. •The perception that the NMC is avoiding meaningful dialogue with FMGs undermines trust and suggests a lack of willingness to address the real challenges faced by these dedicated professionals. Our Stance: We at @udfaindia believe that a body like the NMC, which is designed to represent and listen to the medical community, must be open to all its members. Ignoring or sidelining any group is counterproductive and contrary to the principles of equality and fairness. Why This Matters: •Equity and Representation: All medical professionals, regardless of where they received their education, should have the opportunity to engage with regulatory bodies and influence policies that affect their careers and lives. •Transparency and Trust: For the NMC to maintain its credibility and legitimacy, it must foster an environment of openness and responsiveness, especially towards those who are most impacted by its decisions. •Potential for Unrest: If the NMC continues to act unilaterally without genuine dialogue, it risks alienating a significant portion of the medical community. This could lead to increased discontent and potential protests, which is not in anyone’s interest. Our Call to Action: •Open the Doors to Dialogue: We urge the NMC to actively engage with FMGs and their representatives to discuss the implications of recent policies and seek mutually acceptable solutions. •Demonstrate True Commitment: Show that the commitment to inclusivity and open dialogue is not just for show, but a fundamental principle guiding the NMC’s actions. •Listen and Represent: The NMC must fulfill its mandate to listen to and represent all members of the medical community, including FMGs who have contributed significantly to healthcare in India. Let’s Work Together for a Fair and Inclusive Medical Community. We remain hopeful that the NMC will reconsider its approach and embrace a more inclusive and open stance towards FMGs. It is only through unity and collaboration that we can ensure a just and effective healthcare system for all. #OpenDialogue #FairTreatment #SupportFMGs #NMCAccountability #FMGS National Medical Commission ALL FMGs ASSOCIATION(AFA) Free Press Journal Anil Sharda ravish ndtv Karn Pratap Singhshow more

Dr.Arun Kumar
13,546 views • 2 years ago
Dear Tarun Chitra 1. We are the original creators... of DeSci back in 2016. What DeSci has become today is largely unrelated with its original model of producing rigorous peer-reviewed scientific studies published in reputable medical journals. The model we introduced. We are tirelessly fighting against pseudoscience, and we are showing the world that people can understand the difference between legit science and pseudoscience with the success of $INNBCV. Yes, meritocracy is possible in crypto. Even against all odds. 2. We are the only project in the entire crypto space that ever funded, performed, and published highly innovative HIV cure research ( We are the project that produced the first peer-reviewed study on blockchain-based biomedical data storage in the world’s most reputable scientific network, Springer Nature ( $INNBCV is not for the privileged few; it is for the many. We resisted all the pressure from those who wanted us to provide big allocations to VIPs of other DAOs “because it is good for the marketing” and put our users first, ensuring a fair launch, a launch for the people, and they turned $70k into $2,000,000. $INNBCV shows that you can have a sustainable model, provided you are backed by actual science. And thanks to the amazing guys at daos.fun baoskee and Solana community. Behind our project there is the sweat and blood of years of work to produce publications in the most reputable medical journals. Just to put things into perspective, it took us 3 years to publish our latest work in Springer Nature. 3. Unlike many other projects, we had no ICO/VCs, meaning we had to prove ourselves every single day because we are only supported by our community. If we deliver products, we survive; it is either publish or perish for us, and that’s why we have such a close connection to our community. $INNBCV is a struggler, $INNBCV is a survivor, $INNBCV is not for the privilege of the few but for the people. Our community makes it possible by supporting us. You guys are the real heroes.show more

InnovativeBioresearch🇮🇹
10,867 views • 1 year ago
After a few more hours, I think I've figured... out Opus 5. Opus 5 is trained to be more agentic than anything I've used. All Claude 5 models are like that. So what changes? The way to interact with Opus 5 or contextualize it won't work the same way as with other models. It loves exploring, so it doesn't need much guidance for it. Unique preferences, artifacts, and references compliment it well and enable cleaner and more effective exploration and execution. Now that it can explore more effectively on its own and understand intent better, the best thing to do is to get out of its way (e.g., it doesn't need examples of your preferences; a clear high-level description of it works best). It's truly agentic in that sense. A good first step to provide better context for Opus 5 is to distinguish between what's situational and what needs persistence. Regardless, persistent system prompts and CLAUDE.MD needs to stay lightweight. Remove memories and tool descriptions from these. CLAUDE.MD is also a great place to tap into progressive disclosure by linking command/skills to it. On the situational side, agent skills and auto-memory can leverage progressive disclosure and the improved ability of the model to use its external context/knowledge. Conflicting and unnecessary instructions, which are common at this layer (mainly to ensure reliability), are going to throw off this model easily. That's the biggest change I had to make. Simple, clean, and clear prompts and skills work best. I had to clean a lot of my skills and system prompts. The way I prompt remains the same (usually clear and well-scoped). MCP tool descriptions are also more descriptive and have been deduped from the system prompt. Anthropic released a guide on the new rules for context engineering, which was helpful here. I started to test the recommendations and created a little artifact with the things that worked along the way. This might feel like a lot of work. Believe me, it has been frustrating. But I think we can expect future frontier models to become more agentic and smarter at figuring out the right context/gaps. The best thing to do is to prepare for that now. Boris Cherny mentioned that Opus 5 is their least prompt-injectable model yet. I am not sure if that was something they intentionally trained for or if it emerged based on how it was trained, which is to be extremely agentic in nature and more direct in execution.show more

elvis
37,757 views • 1 month ago
Here's what The Browser Company's AI eng & ML... teams are working on for Dia right now: (This is a pitch to come work for us; info at end) 🤖 COMPUTER USE – we've built our own bespoke APIs on top of Chromium to optimize latency, accuracy, and cost of computer-using agents. Demo attached. Big breakthroughs here in recent weeks. 🛡️ ON-DEVICE MODELS – we've built our own custom infra to run everything from encoder-only models to full LLMs on device. It's cross-platform, supports LoRa adapters, and optimized for the GPU. This system preserves privacy and enables fast inference times. 🧠 MEMORY – with your permission, Dia automatically tailors your AI experiences to you, personally, based on the tabs you open while browsing normally every day. We're also bringing vertical memory to specific features. ♻️ DATA FLYWHEELS – our Fall/Winter P0 is to double-down on training custom models based on implicit signals from daily use of Dia. Dia should get smarter and more useful the more people use it. Whether via RL, auto-generated prompts, or otherwise. If this work sounds interesting to you please visit our jobs page or email [email protected]. Hiring nearly every related role -- from ML engineers to people prototyping with AI and context/prompt writers -- everyone encouraged to apply!!show more

Josh Miller
68,130 views • 1 year ago
Gemma 4 26B A4B MoE - 500+ t/s decode... - Single RTX 4090 (24 GB VRAM) - Llama.cpp concurrency 24 - q8 kv cache How many API users can you simultaneously host on a single RTX 4090 (24 GB VRAM) before it crashes? Yesterday, I proved you can host 14 active users using unquantized memory. Today, I used 8 bit KV Cache Quantization to hack the VRAM footprint. I successfully scaled to 24 concurrent users without a single dropped connection. A 71% server capacity boost for free. By adding the -ctk q8_0 -ctv q8_0 flags to llama.cpp, you compress the KV cache context memory from 16 bit to 8 bit. This unlocks massive concurrency limits on Gemma 4 26B (MoE) on a single 24GB consumer GPU. Here is the exact telemetry from pushing 8 bit quantization to its absolute physical edge: # TEST 1: The 24 User Concurrency Max Server Config: 24 slots (np 24) | 4,096 context per slot | 98,304 Total Context Client Load: 24 simultaneous requests (2,000 token prompt per user) Unquantized KV cache for this load requires 28GB+ VRAM (Instant OOM). Quantized to Q8, it allocated safely at 23.35 GB. The C++ engine crunched the entire batch in 28.5 seconds. Decode Speed: 21 t/s (Per User) | 500 t/s (Agg) # TEST 2: The 48 User Queue Overload What happens to a compressed cache during a traffic spike? Server Config: 24 slots (np 24) | 4,096 context per slot | 98,304 Total Context Client Load: 48 simultaneous requests (2k token prompt per user) Zero queue drops. The scheduler flushed and hot swapped the 8 bit memory flawlessly on the fly, completing all 48 users in 66.0 seconds (a perfect 2.3x queue scaling multiplier). Decode Speed: 18 t/s (Per User) | 430 t/s (Agg) # TEST 3: The 8 User RAG Slam Server Config: 8 slots (np 8) | 60,000 context per slot | 480,000 Total Context Client Load: 8 simultaneous requests (30k token prompt per user) It allocated 23.83 GB VRAM and chewed through ~240,000 prefill tokens in 46 seconds under massive memory pressure. Prefill Speed: 6,200 t/s (Agg) Decode Speed: 22 t/s (Per User) | 175 t/s (Agg) # The Engineering Alpha (The Quantization Tradeoff): You gain a massive 71% increase in server capacity, but what do you lose? Compute latency. Because the cache is stored in 8 bit, the GPU's cores have to dequantize the memory back to 16 bit on the fly during every single prefill step. In my unquantized tests yesterday, single slot prefill was hitting ~1,500+ t/s. Today, under the heavy 48-user Q8 load, prefill dropped as low as ~750 t/s. You trade a few seconds of initial prefill latency to essentially double your API hosting capacity. For production high volume SaaS, this is the ultimate unit economics cheat code. Here is the exact command to run a 24 user Q8 continuous batching server on your own single 4090, single 3090 or any 24gb vram rig: ./build/bin/llama-server -m gemma-4-26B-A4B-it.gguf -c 98304 -np 24 -b 2048 -ub 2048 -ngl 99 -fa on -ctk q8_0 -ctv q8_0 --port 8080 (Note: -c 98304 allocates exactly 4,096 tokens of context per user across 24 slots). Hugging Face links to the Unsloth Gemma 4 26B QAT quants along with performance graphs available in the replies. Would you trade 3 seconds of Time To First Token latency to double your active user capacity?show more

Alok
17,465 views • 1 month ago
This might be the greatest counting video ever. I... don’t want to spoil it, so watch it first, then come on back and let’s discuss… ____ There are so many fantastic things happening here as our hero (age: 2 years, 7 months) counts his plastic eggs. In this video we get a sense that he is well on his way to mastering two important early math skills: rote counting and one-to-one correspondence. Rote counting is the process of remembering and reciting numbers in their correct order. It’s a memorized skill - but an important one that helps children to understand that numbers exist in a fixed sequence. One to one correspondence is the process of matching each spoken number to one (and only one) physical object. This little guy pulls it off perfectly, counting with accuracy to 11. His inclusion of 8:30 between 8 and 9 was not only the best laugh I’ve had all week, but a telling stroke of genius as well. Where did it come from? I’d be willing to wager it’s his bedtime - which he’s learned comes between 8 and 9pm. So smart! 🧠 This little mathematician was shared to TT by smaxwell18.show more

Dan Wuori
115,189 views • 1 year ago
Free NVIDIA GPU with 16 GB VRAM GPU for... Running Local LLMs! If you want to master local LLMs but you're waiting until you can afford a $1,500 GPU, you're honestly not going to make it. The open source AI ecosystem is moving way too fast for you to wait on your budget to catch up. Especially when you can build a bleeding edge inference engine from scratch right now, completely for free. You don't need a heavy local rig to start. Google is literally letting you use an enterprise grade NVIDIA Tesla T4 GPU for $0/hour. At standard cloud computing rates (~$0.20/hr), Google Colab’s 4 hour daily free tier hands you roughly $24 worth of data center tier GPU compute every single month. And most people just waste it. Let’s talk about the hardware you get access to for free. The NVIDIA Tesla T4 is an absolute workhorse: - Architecture: NVIDIA Turing (TU104) - VRAM: 16GB GDDR6 (320 GB/s bandwidth) - Compute: 320 Tensor Cores | 2560 CUDA Cores - Performance: 130 TOPS INT8 | 8.1 TFLOPS FP32 - Power: Sipping energy at a max 70W TDP This is the exact same hardware I used to run DeepMind's Gemma 4 26B A4B QAT MoE at a 250,000 context window without a single Out Of Memory (OOM) crash. If you have a web browser and 10 minutes, you have everything you need. I’ve put together a fully documented, cell by cell Google Colab notebook that teaches you exactly how to do this. Here is what the notebook actually teaches you: - How to provision an Ubuntu Linux environment with CUDA 13.0 and verify your driver stack. - How to pull the source code and compile the latest llama.cpp C++ binaries from scratch, specifically optimizing the build for your exact GPU using the -DCMAKE_CUDA_ARCHITECTURES=native flag. - How to directly download quantized local LLMs (GGUF format) straight from HuggingFace using the CLI. - How to manage 16GB VRAM limits, offload neural network layers to the GPU, and push massive context windows. Compile raw llama.cpp, ollama run a model, or spin up the LM Studio CLI. Pick whatever stack you are comfortable with. just start building. No hardware. No credit card. No excuses. Bookmark this post right now so you don't lose the tutorial. Even if you don't have time to run it today, you are going to want this workflow in your engineering toolkit. The link to the free Colab Notebook is in the comments below. Lemme know if you need more tutorials like this.show more

Alok
182,483 views • 2 months ago
As always everyone is blind staring at the progress... of LLMs for coding and chat But meanwhile the new SOTA video model Seedance 2.5 has been slowly rolling out and it's really quite exceptional It's made by ByteDance (TikTok) who of course have lots of training data With just a few reference pics, it can get quite close to how you look IRL and you can do quite professional video shots with just a prompt I'd say it's the first video model that's now at the level of image models with the level of character likeness, cracking that in image models also took about 3 years (2022-2025) Generating 15 seconds takes about 4 minutes I put it live now on Photo AI, you can use it under [ Make video ] from the sidebar with just a prompt and your model selected! So you don't need to take an AI photo first and then turn that into a video! Saves lots of time :D It's more expensive than but I kept the credits the same (30 for 1 video) It also works inside the new video editor and you can make changes in your video with [ Magic edit ] in both the main app and the video editor Also a message for my server guy Daniel Lockyer (it can do voice too and you can even submit a voice sample of yourself, but I didn't here)show more

@levelsio
711,073 views • 1 month ago
If you have an RTX 3090 or 4090, Mia... just shipped you a free massive upgrade in both speed and intelligence. I will explain to you why this will make your Qwen 3.8 27B on your card, even better, and my flags for running it. Qwen3.8-27B, EXL3 3.5bpw, DFlash2 speculative decode, RTX 4090. Single stream. The kit is from MiaAI-Lab, EXL3 is turboderp's format. I re-measured everything on my own card because the my first benchmarks seemed off. It turns out it really does run much faster. WHY EXL3 IS A DIFFERENT ANIMAL The old way (Q4_K_M) rounds each weight to the nearest 4-bit value independently. Every weight introduces its own rounding error. Those errors accumulate across millions of weights and causes drift (Slightly dumber). EXL3 is a fundamentally different compression algorithm. Instead of rounding each weight on its own, it encodes the entire weight vector as a path through a constrained codebook and spreads the rounding error across dimensions using a Hadamard transform. The result is that at the same bits per weight, more of the original model's intelligence is preserved. The important part is this CAN ACTUALLY BE MEASURED. The cleanest way to see that is KL divergence against a high-precision teacher. Lower means the quantized model thinks more like the original. On the malaiwah independent teacher-logit panel for GLM-5.3-Flash: EXL3 4bpw: 0.0246 nats Official FP8: 0.0206 nats NVFP4: 0.0605 nats EXL3 sits 0.004 nats behind native FP8 at half the size. NVFP4 at higher bit width is 2.5x further from the teacher. That panel is GLM-5.3-Flash, not Qwen 3.8. Cited as the mechanism, not as this run's data. But the point stands: EXL3 is not just smaller, it is smarter per bit than the formats most people are running. WHAT I MEASURED I first measure 108 tok/s from a single run. After that number looked too good to be true. I reran it. It looks like after a warm up, the numbers are even better. Basically, like people long thought, the RTX 3090 and RTX 4090 are actually superb AI computer cards. Hence why NVIDIA stopped shipping them with NVLINK since the 4090. Short context ceiling (~2k in, 1016-token output, TTFT-separated): 135, 138, 153, 174, 133, 129 tok/s across 6 runs. Sustained longform (2040-token essay): 105.4, 94.5, 98.4 tok/s Short answer (504 tokens): 93.2 tok/s The honest shape: ~130-150 tok/s at short context is the ceiling, ~94-105 sustained on longform. The ceiling matters because that is what people feel in chat. The old dense Q4_K_M on llama.cpp ran ~37 tok/s on this same card. (No MTP), with MTP about 60 tok/s Sustained is roughly 2.5-3x. Ceiling is closer to 4x. Same model, different quantization and engine. The multiplier comes from EXL3, the ExLlamaV2 engine, and DFlash2 together. CONCURRENCY IS A RTX 4090 LANE. Just like the old config on the 4090, the 24gb vram, means a long context can only hold one stream, and running concurrency requires to lower context length, because it runs fast it sort of makes up for it by being faster than slower GPU chips. CONTEXT LADDER The recipe doc measured prefill. I re-ran it with TTFT separated from decode, because decode is what you actually feel after the first token. ~5k in: decode 140 tok/s (TTFT 0.5s), needle HIT ~18k in: decode 85 tok/s (TTFT 0.3s*), needle HIT ~73k in: decode 28 tok/s (TTFT 1.8s), needle HIT ~146k in: decode 16 tok/s (TTFT 2.3s), needle HIT (*0.3s at 18k is a prefix-cache hit from the paired pass. Cold prefill for reference: ~2,020 tok/s at 17k falling to ~508 at 153k.) Needle hit at every depth, mine and the original 7/7. Retrieval is intact at max context. Speed is not: decode falls ~9x from short to max. Past ~50k tokens this stops being a chat tool and becomes a batch tool. At 146k it works, but nobody is typing interactively against 16 tok/s. WHERE IT BROKE The model's native context is 262k. The README says DFlash2 fits ~220k on a 24GB card. My 200,704-token attempt failed with insufficient VRAM. Dropped to 168,960 and it booted. The real ceiling is somewhere between 168,960 and 200,704 and I never tested that gap. I jumped to a value that worked and called it done. That is ~32k tokens of context I left on the table. One thing the numbers taught me: JSON tokenizes at ~1.5 chars/token, prose at ~3.9. "150k tokens of JSON" needs ~2.7x more filler than the same estimate in prose. Size by real tokens, not estimates. THE UPGRADE If you own a 4090 and you are running Q4_K_M on llama.cpp, you are leaving a good bit of speed and measurable intelligence on the table. The same model, on the same card, with a better quantization and engine, goes from 60 tok/s to 130-150 at short context and 94-105 sustained. The model also thinks closer to the original because EXL3 preserves more of the output distribution per bit than the old rounding method. The recipe is in the first reply. Everything above came from one 4090 and one afternoon of re-measuring. The decode ladder especially needs independent numbers. If your card gets different falloff, that is worth knowing. Recipe and flags/ findings in reply 👇show more

Yume_X
38,685 views • 15 days ago
I told you to claim your free 16GB NVIDIA... GPU for learning Local LLMs. Now I’m going to show you how to double its inference speed without touching the hardware. Google Colab gives you an enterprise grade NVIDIA Tesla T4 GPU for free, roughly 4 hours every single day. It is the absolute perfect sandbox for learning AI engineering, testing inference flags, and pushing massive context windows. The local AI timeline is moving way too fast. If you aren't using Multi Token Prediction (MTP) yet, you are leaving massive performance on the table. I just pushed DeepMind’s Gemma 4 26B to 64.9 t/s on this exact free tier. Let's look at the raw benchmark data running on an Ubuntu Linux environment with the latest compiled llama.cpp binaries and quantized GGUFs from Unsloth via HuggingFace: # Qwen 3.5 9B (Dense): Base: [ Prompt: 626.7 t/s | Generation: 21.0 t/s ] With MTP: [ Prompt: 539.1 t/s | Generation: 24.8 t/s ] # Gemma 4 26B QAT (MoE): Base: [ Prompt: 634.2 t/s | Generation: 48.3 t/s ] With MTP: [ Prompt: 572.1 t/s | Generation: 64.9 t/s ] If you are paying attention, this single Colab notebook reveals 3 massive observations about the current state of local LLMs: # 1. The MTP Speedup (Software Overclocking) Standard autoregressive decoding guesses one token at a time. MTP acts like a highly optimized, built in speculative decoder. It predicts multiple future tokens at once and the main model verifies them in parallel. The result? Zero accuracy loss and a massive throughput increase. Gemma jumped from 48 to 65 t/s just by flipping a flag. # 2. The MoE Paradox (Bigger is Faster) How does a 26B parameter model absolutely destroy a 9B model in raw speed on the exact same hardware? Architecture. Qwen 3.5 9B is a dense model. it activates all 9 billion parameters for every single token. Gemma 4 26B is a Mixture of Experts (MoE) model. It routes data efficiently, activating only 4B parameters per token. You get the reasoning capabilities of a 26B model with the compute cost of a 4B model. 3. Thinking Efficiency When I ran the exact same complex prompt on both models, the larger MoE spent significantly fewer "thinking" tokens to arrive at the correct answer. A smarter model doesn't just give better answers; it gets to the point faster, saving you compute cycles and preserving your context window. # Want to run this yourself? Here are the exact llama.cpp CLI commands. For Qwen (MTP is baked into the main model): ./llama-cli -m Qwen3.5-9B-UD-Q4_K_XL.gguf -p "Explain quantum computing." -n 2000 -c 8000 -ngl 99 -fa on --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 For Gemma (Using a separate lightweight draft model): ./llama-cli -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --model-draft mtp-gemma-4-26B-A4B-it.gguf -p "Explain quantum computing." -n 2000 -c 8000 -ngl 99 -fa on --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 Stop waiting for a $3,000 rig. Boot up Colab, pull these models, and start building your stack. I’ve put together a completely free, cell by cell Google Colab notebook that automates this entire workflow so you can test it yourself in 5 minutes and learn. Link to the notebook is in the comments below. Experiemt with different MTP parameters, context windows and post your results in the comments.show more

Alok
170,442 views • 2 months ago
This Chinese developer launched Llama 70B locally on a... MacBook on a plane and for a full 11 hours without internet ran client projects. He was sitting by the window on a transatlantic flight with a MacBook Pro M4 with 64 GB of memory. WiFi on board cost $25 for the flight. He declined. No cloud API, no connection to Anthropic or OpenAI servers, no internet at all. Just a local Llama 3.3 70B on bf16 and his own orchestrator script. The model runs through llama.cpp. Generation speed, 71 tokens per second. Context around 60,000 tokens. Memory usage, 48.6 GiB out of 64. Battery at takeoff, 3 hours 21 minutes. And he gave the orchestrator this system prompt before takeoff: "You are an offline orchestrator running on a single MacBook. There is no network. The only resources you have are local files in /Users/dev/work, the Llama 70B inference server at localhost:8080, and a battery budget of 3 hours 21 minutes. Process the queue at /Users/dev/work/queue.jsonl (one client task per line). For each task: draft → run local evals → save artefact to /Users/dev/work/done/. Save context checkpoints every 12 tasks so you can resume after a battery swap. Stop only on empty queue or when battery drops below 5%." So the system knows exactly what resources it is running on. It knows it has no connection to the outside world for the next 11 hours. It knows it has finite memory and a finite battery. It knows the human will not intervene until the plane lands. The system runs in 1 loop. Takes a task from the queue, runs it through inference, saves the artifact, writes a checkpoint. Task after task, just like that. And only when the battery drops below 5% does the orchestrator automatically pause, waits for the laptop to switch to the backup power bank, and continues from the last checkpoint. Here is what the system actually writes in his log during the flight: "saved context checkpoint 8 of 12 (pos_min = 488, pos_max = 50118, size = 62.813 MiB)" "restored context checkpoint (pos_min = 488, pos_max = 50118)" "prompt processing progress: n_tokens = 50 / 60 818" "task 37016 done | tps = 71 s tokens text → /Users/dev/work/done/proposal_westside.md" Outside the window, clouds, blue sky, and no WiFi. On the tray, 1 MacBook, an open terminal on 2 screens, and an inference server on localhost. From what I have observed, this is the cleanest offline AI workflow I have seen in the past year: 11 hours of flight, $0 for WiFi, and the entire client queue closed before landing.show more

Blaze
1,843,280 views • 4 months ago
This is the second of two posts I’m making... about the teaching of Andy Stanley because I have a burden to warn people to be like the Bereans "examining the Scriptures daily to see if these things were so." (Acts 17:11). I have attached a short video clip where Stanley says we don’t know Jesus rose from the dead because the bible tells us so. He says it’s the other way around. Now how do we know Jesus bodily rose from the dead? Did any of us see it happen? No, it happened 2000 years ago. But haven’t people researched the resurrection and found all sorts of evidence consistent with Him rising from the dead. Yes, but such evidence, though powerful, and we can certainly use in our witnessing, does not ultimately prove He rose from the dead. In fact, no matter what evidence we point to in geology, biology, astronomy etc., none of this proves in an ultimate sense the bible is true. Now it is certainly true such evidence properly interpreted does confirm the bible’s history. For instance, the molecule of heredity DNA is a complex information system and language system. No one has seen matter produce information or a language from matter by natural processes. Our observations and experience show information and language have to come from an intelligence. Such evidence confirms an intelligence behind life. This certainly confirms the first verse of the bible, “In the beginning God created… .” But nonetheless, it’s not absolute proof. Think about it. We are finite beings living in the present. We don’t know everything. We don’t know how much we don’t know or do know in relation to whatever there is to know! When we try to interpret evidence of the present in relation to the past, how do we know we have all the relevant information to make the correct interpretation. Some information we don’t have could totally change our interpretation. That has certainly happened with scientists solving crimes using circumstantial evidence. When new evidence comes along, some of the interpretations change and certain people thought to be guilty were found they were innocent after-all. We need to have all the information needed. But we can’t know everything. However we have a book that claims over three thousand times to be the Word of God. This book claims that God moved people by His spirit to write His Word, what He wants revealed to us about life, the universe and history. This book, the bible, tells us that God knows everything, He is infinite in knowledge and wisdom. He has all information. “In whom are hidden all the treasures of wisdom and knowledge” (Colossians 2:3). If God’s Word is what it claims to be (and it is), then the infinite Creator God has revealed to us the key information we need to know to have the ability to correctly interpret this world in relation to the past, present and its purpose and meaning. “Therefore, we never stop thanking God that when you received his message from us, you didn’t think of our words as mere human ideas. You accepted what we said as the very word of God—which, of course, it is. And this word continues to work in you who believe” (1 Thessalonians 2:13). This means if we build our thinking on God’s Word, we build a Christian worldview to enable us to look at the world through “biblical glasses” and have the ability to correctly interpret and understand it. And Genesis 1-11 is the history that is foundational to the rest of the bible and thus our worldview. So, when we start with God’s Word we learn that the God revealed to us in the bible is the One Who created all things. Now by just looking at the world, for instance at DNA, we might deduce that there’s an intelligence behind life. But we would not know who that intelligence is unless revealed to us. The bible reveals who that intelligence is, God the father, Son, and Holy Spirit. But, if we just looked at the world with all its death, suffering and disease, we could assume that the intelligence behind life must be an ogre to make such a violent disease ridden suffering world. But when we start from God’s Word, we understand there was no death or disease to start with, but these entered the world because of sin. We also find man has a problem called sin which alienated him from God. We even find in Genesis that God promised someone would come to save us from our sin and restore our relationship with God (Genesis 3:15; 3:21). We learn later on that “someone” is Jesus--the one who became flesh for us, the babe in a manger 2000 years ago. It's important to understand we can’t know anything absolutely unless an absolute authority has revealed to us what we need to know. Now Andy Stanley claims we don’t know Jesus rose from the dead because the bible tells us so, but claims it’s the other way round, that because Jesus rose from the dead, we can believe what is says about this in the bible. In reality, Andy Stanley is actually claiming he knows all information about the resurrection to know it’s true so then he can proclaim what the bible says about the resurrection is true. This is not so. And remember from the previous post, Andy Stanley accepts man’s view of evolution and millions of years as true to declare what Genesis records about creation is not all true. So why shouldn't people take the word of others who claim Jesus didn't rise from the dead to then declare the account of the resurrection in the gospels can't be true. Stanley, as a finite fallen human being with very limited knowledge, starts outside the bible to go to the bible to make pronouncements over God’s written Word. No wonder he rejects a literal Genesis. Sadly, this is the case for the majority of our church leaders and Christian academics, particularly when it comes to Genesis. Think about it. Really Stanley is acting in accord with our sin nature because of what happened in Genesis 3. Part of our sin nature concerns us wanting to be our own god. Consider Genesis 3:1 and Genesis 3:5, the temptation by the devil. Satan tempted Adam and Eve to question God’s Word, and be want to be like God to decide good and evil, truth, etc., for themselves. In actuality, the statements Andy Stanley is making about Christianity and God’s Word are reflecting this sin problem we have. He is letting his sin nature master over him in this instance instead of letting God’s Word tell us clearly what we should believe. And I would say that about all Christians who reinterpret parts of the bible (like Genesis) because of beliefs from outside the bible. Now the whole bible is actually about Jesus, from the very first verse, that Jesus is the Creator (Colossians 1:16 “For by him all things were created,”), & and he is the Savior (Revelation 5:9), to the very last verse, “He who testifies to these things says, “Surely I am coming soon.” Amen. Come, Lord Jesus! The grace of the Lord Jesus be with all. Amen” (Revelation 22:20–21) God reveals all we need to know about Jesus in His Word. We know Jesus rose from the dead because the bible tells us so! We know we are sinners because the bible tells us so. We know we can be saved through faith in Christ because the bible tells us so. We know we need to repent of sin because the bible tells us so. As Christians we know we will spend eternity in Heaven because the bible tells us so. “So faith comes from hearing, and hearing through the word of God” (Romans 10:17). I still remember singing the chorus as a child Jesus loves me this I know, for the bible tells me so. Yes, the bible tells me so – that’s how know who Jesus is, that He is the Creator, that he died and rose from the dead, and that He is our Savior.show more

Ken Ham
233,976 views • 3 years ago
We’re excited to introduce ShinkaEvolve: An open-source framework that... evolves programs for scientific discovery with unprecedented sample-efficiency. Blog: Code: Like AlphaEvolve and its variants, our framework leverages LLMs to find state-of-the-art solutions to complex problems, but using orders of magnitude fewer resources! Many evolutionary AI systems are powerful but act like brute-force engines, burning thousands of samples to find good solutions. This makes discovery slow and expensive. We took inspiration from the efficiency of nature. ‘Shinka’ (進化) is Japanese for evolution, and we designed our system to be just as resourceful. On the classic circle packing optimization problem, ShinkaEvolve discovered a new state-of-the-art solution using only 150 samples. This is a big leap in efficiency compared to previous methods that required thousands of evaluations. We applied ShinkaEvolve to a diverse set of hard problems with real-world applications: 1/ AIME Math Reasoning: It evolved sophisticated agentic scaffolds that significantly outperform strong baselines, discovering an entire Pareto frontier of solutions trading performance for efficiency. 2/ Competitive Programming: On ALE-Bench (a benchmark for NP-Hard optimization problems), ShinkaEvolve took the best existing agent's solutions and improved them, turning a 5th place solution on one task into a 2nd place leaderboard rank in a competitive programming competition. 3/ LLM Training: We even turned ShinkaEvolve inward to improve LLMs themselves. It tackled the open challenge of designing load balancing losses for Mixture-of-Experts (MoE) models. It discovered a novel loss function that leads to better expert specialization and consistently improves model performance and perplexity. ShinkaEvolve achieves its remarkable sample-efficiency through three key innovations that work together: (1) an adaptive parent sampling strategy to balance exploration and exploitation, (2) novelty-based rejection filtering to avoid redundant work, and (3) a bandit-based LLM ensemble that dynamically picks the best model for the job. By making ShinkaEvolve open-source and highly sample-efficient, our goal is to democratize access to advanced, open-ended discovery tools. Our vision for ShinkaEvolve is to be an easy-to-use companion tool to help scientists and engineers with their daily work. We believe that building more efficient, nature-inspired systems is key to unlocking the future of AI-driven scientific research. We are excited to see what the community builds with it! Learn more in our technical report:show more

Sakana AI
360,407 views • 11 months ago
The human brain is truly a marvel of nature.... If you horribly reductive, and boiled it down to a language model, you'd be looking at roughly 100 trillon parameters running as a sparse MoE architecture Only about 1-5% of neurons fire at any given moment, meaning the brain "activates" maybe 1-5 trillion parameters per inference step. For context, the largest AI models we've built probably top out around 5 trillion parameters. The brain is roughly 100x larger. Even its active params at any given moment are larger than almost every model in existence today. Here's what melts my brain (pun intnended) though Your brain does all of this on about 20 watts of power, less than a dim light bulb. Training a frontier AI model consumes enough electricity to power small cities for months. Running inference across data centers pulls megawatts. Your brain runs 24/7 for 80+ years on the equivalent of a phone charger. We haven't come close to matching the brain's scale. And we're not even in the same universe when it comes to efficiency. Evolution spent 500 million yrs optimizing the most energy-efficient intelligence architecture ever known. we're trying to brute force our way there with compute and electricity. Nature is still the best engineer in the room.show more

am.will
130,883 views • 5 months ago
📍Theory of Space (accepted at #ICLR2026) Theory of Mind... → hidden mental states Theory of Space → hidden spatial beliefs from passive observers “What do I know?” to active explorers “What don’t I know, and how do I reduce that uncertainty?” Theory of Space is to evaluate if foundation models can actively construct, revise, and exploit internal spatial beliefs. We quantify Active-Passive Gap. Not just measure task accuracy, but how much uncertainty is reduced per step, and how many steps are needed in total for agents to build stable spatial beliefs. Exploration should prioritize information gain and reduce uncertainty per step. Instead, we observe LLMs/VLMs explore redundantly with stalled belief updates. Key findings: 1. Active agents perform worse than rule based programs 2. Cognitive Map Failures & Belief Drift (beliefs about previously observed objects degrades over time; new updates corrupt earlier correct perceptions) 3. Poor Vision Identification & Belief Inertia in Belief Revision Website: Code: Data: Theory of Space is a joint effort of Northwestern Engineering, Stanford AI Lab, Allen School, Cornell Computer Science. Led by the amazing WilliamZhang, jointly done with Zihan Huang, yue wang, Jieyu Zhang, Lester Xue, @wzihanw, Qineng Wang, Keshigeyan Chandrasegaran, Ruohan Zhang, Yejin Choi, Ranjay Krishna, Jiajun Wu, Fei-Fei Lishow more

Manling Li
53,079 views • 7 months ago
How long should my baby use a pacifier? I... get this question a lot and it’s a tough one - both because there isn’t a single correct answer and because (like feeding and sleep) the topic brings out lots of strong opinions. But if you ask me, the family in this video has the right idea. Infants are born with a strong sucking reflex and pacifiers can help them to soothe and sleep. There’s even some evidence to suggest that sleeping with pacifiers might reduce the risk of Sudden Infant Death Syndrome (SIDS). In short: for babies (up to a year), I’m a big fan. But it’s not uncommon to see children with pacifiers well into toddlerhood, and in some cases, even beyond. And here I’d raise some important cautions. Children who rely heavily on pacifiers may be more prone to middle ear infections. And dentists note that prolonged pacifier use can affect your child’s teeth and create bite issues. Perhaps most importantly is their potential to impact expressive language development. Your child’s ability to speak is an important one. After a point, language shapes not only the content of our thinking, but the very structure of our cognition. By otherwise occupying the mouth over long periods of time, pacifiers may slow language development by limiting opportunities for expression. Speaking with a pacifier in the mouth can also lead to distortion of speech sounds (even when they aren’t in the mouth). All told, I’m an advocate for beginning to wean off of pacifiers at around a year of age - which is why this video spoke to me. We see an infant appropriately using one and big sister demonstrating her expressive language, her mouth unencumbered and free to chatter away happily. The transition can be difficult - but not nearly as challenging as for a child who has become dependent over a period of years. Do/did you use pacifiers with your child? Why or why not? How did you help transition away from their use? This sweet siblings were shared to IG by _lullabye_luxuries_.show more

Dan Wuori
158,122 views • 2 years ago