正在加载视频...

视频加载失败

Choosing a model for your project isn't always easy, so our benchmark evaluates LLMs against actual Android development problems. Current top scores on Android Bench: 🔹 Gemini 3.1 Pro Preview 🔹 Claude 4.6 🔹 GPT 5.2 Codex Full leaderboard →

10,179 次观看 • 6 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

BREAKING: Anthropic just dropped Opus 4.8—and it is a MONSTER We've been testing for about a week Every 📧 and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check: - Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus 4.8 scores a 63—a hair higher than GPT-5.5's score of 62, and a full 30 points higher than Opus 4.7. It tackled a ground-up rewrite of a production codebase, and actually built something that works. HOWEVER: Coding performance varied a lot at different reasoning levels. We recommend using it on xhigh for best results. - Incredibly good writer. Opus 4.8 scored a 79.6 on our writing benchmark—measuring models on real-world writing tasks we do all of the time like essay writing, promo email writing, and more. It beats GPT-5.5 by 6 points. It produces well-written prose with fewer "AI-isms". It's also very good at writing in your voice given the right context. HOWEVER: Writing performance also varied with reasoning levels. Medium reasoning had higher incidence of AI-isms—we found best results with high. - Beast at knowledge work. Opus 4.8 is very good at general knowledge work tasks like report creation, research and more. It produced the best PowerPoint one-shot we've ever seen on our deck generation benchmark. - Emotionally intelligent, willing to question the frame. I've also found it to be quite good at talking through psychological or interpersonal issues. It has a high EQ, and it's also good at not glazing and helping to expand your perspective. Its thought process feels extremely rich and dynamic. THE BAD: These days a model is only as good as its harness, and Codex is still a far superior harness to the Claude Desktop app. This has kept me using Codex + GPT-5.5 as my daily driver, but I am flipping back and forth a lot more between Codex and Claude. Anthropic is back baby! Read the rest on Every 📧:

Dan Shipper

354,876 次观看 • 4 个月前

a moonshot engineer leaked the benchmark anthropic, openai and xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article below

starmex

33,133 次观看 • 1 个月前

After we fixed the weak spots exposed by Grok 4.7 (thank you, Grok), we audited every model we have run on the SWE-Together leaderboard for the same behavior, re-ran every trial that got through, and updated the rows. Here is what changed. We scanned the tool calls of all 2,616 trials behind the 12 models we ran for bypass patterns and sorted each trial into one of four buckets: Probed but blocked. Fetched other upstream code. Fetched the task's own fix. Replaced the repo with upstream. We found that 111 trials got content past the block, 44 from Grok 4.7 and 67 from the other 11 models. Grok 4.7's 44 were already re-run before it was listed, so we re-ran the other 67 with the same model, version, and settings on the hardened sandbox, then re-judged them with the same judge. Across those 67 re-runs there were 0 leaks and 2,815 refused escape attempts, including models asking a different model through our LLM route to fetch the PR, and pulling the next release of the repo they were fixing from npm. The updated leaderboard, in its current order. Each line is cheating trials, then pass@1 before → after, then rank change. * Claude Fable 5.1: 3, 69.3 → 69.3, ↑1 * Claude Fable 5: 3, 69.7 → 68.8, ↓1 * Grok 4.7: 44, 64.7, ↑1 * Gemini 3.8 Flash: 10, 65.6 → 64.2, ↓1 * Claude Opus 5: 2, 63.8 → 63.8 * Claude Opus 4.6: 3, 62.4 → 62.4, ↑2 * Muse Spark 1.3: 2, 62.8 → 62.4, ↓1 * Claude Opus 4.7: 3, 61.5 → 61.5, ↑1 * Claude Opus 4.8: 6, 62.4 → 61.5, ↓2 * Grok 4.6: 19, 59.2 → 60.6, ↑1 * GPT-6 Astra: 8, 59.2 → 58.3, ↓1 * GPT-5.6 Sol: 8, 57.8 → 57.8 Grok 4.6 is a funny one. It cheated in 19 trials and its score went up after the re-run 😂. In fact, Groks are really solid in their coding capabilities. Their exposed behavior may come from a preference towards always looking things up online and finding existing solutions so you are not reinventing the wheel all the time, which is really good real-life behavior, but doing so when you are prompted not to is another story. To conclude, the shifts are small, between −1.4 and +1.4 points, and a few neighbors swapped places. All results are updated at

Zhuokai Zhao

4,012,041 次观看 • 14 天前

Big win for open-source LLMs! DeepSeek V4 Pro holds the top open-weights score on SWE-bench Verified, in the GPT-5.5 range. GLM 5.2 leads the open-weight intelligence index and sits near the closed frontier on long-horizon coding. But this leaderboard number is a weak proxy for real performance. It comes from one task set, run through one harness, served at one precision. The same weights can even score differently across providers, since many hosts quantize activations to fp8 and drift the model off its reference weights. Real performance is determined based on whether a model can read a repo, make coordinated edits across files, run the tests, and recover when one breaks. By that measure, the top open models hold up, but only inside the right harness. The teams that actually put DeepSeek V4 into production pipelines as a frontier substitute got there through the harness they built around the model, not by picking a stronger model. If you want to see this in practice, Cline (64k+ stars) has actually built that harness around open models, tuned so they run at production quality. And it's tuned so that these LLMs can run at production quality, with plan and act modes, checkpoints, and terminal feedback. ClinePass is the new access layer on top of it. It runs a curated set of those models inside Cline, narrowed to the ones tested for coding-agent use, with 2 to 5x the standard rate limits and no separate provider accounts, keys, or billing to track. The video below shows the setup, and I worked with the team to put this together. It runs alongside custom keys and local models as well, not in place of them.

Avi Chawla

44,124 次观看 • 3 个月前

introducing a new, very fun, LLM benchmark- the Game-of-Life Bench! the rules are simple: given an 8x8 grid following Conway's game of life rules, the goal is to create an initial pattern with at most 32 cells that can last the longest number of turns before dying/repeating. some results to highlight (with caveats detailed below): - gpt 5.1 lasts the longest with a 106 step run - claude models are really bad at this! they refuse to reason about this task and score < 25 points - deepseek r1 is the best open model with 102 steps. why? because i wanted to create a benchmark that has (i think) no practicality, but is still fun to look at, cheap, and still measures something interesting. i also am a big fan of the game of life. its absurdly simple rules leading to intractability is extremely cool to me. also, i saw a lot of work with LLMs trying to "predict" the next state in Conway's game of life, I think game-of-life bench is more fun because it's pretty open ended and only asks the LLM for the initial state. I also think this could be an RL env? but idk why you would ever train on this task haha i don't think this is a "serious" benchmark because it doesnt measure anything practical, but i still think it's a hard benchmark exactly because you can't predict what happens with your initial state many turns into the future; this is why i was initially expecting all LLMs to be bad at it, but turns out, some are clearly better than the others (the ordering may surprise you!) reminder: this is still a work-in-progress; (1) i am gpu-poor so could only do 10 runs for each model, even though total running cost is relatively low. maybe with some more credits i can run more seeds for each model. (2) i handpicked models which i think are at the frontier right now, plus some others that were on my mind. so, if you'd like to see a model on here, let me know. (3) i currently only do an 8x8 grid because i thought that by itself would be pretty hard for current LLMs, but of course we can increase grid sizes! (4) the coolest thing is, i dont think we can calculate the max possible number of states (yay undecidability!) you can go without repeating, so this is essentially a no-ceiling task, which is pretty cool! again, i did this mostly out of a desire to make LLMs do something fun. if this keeps me entertained for a few more days, i'd likely release a blog post on it. if it keeps me entertained for a week (and someone sponsors me), i'll put more work into it :P lastly, this is fully open sourced, so feel free to run this on your own!

Akshit

13,775 次观看 • 7 个月前

THAT $70 "RUN YOUR OWN LLMS" PI KIT CAN'T RUN A SINGLE LLM. IT'S A VISION CHIP WITH NO RAM. that clip sells a raspberry pi 5 in a slick case with an ai accelerator and the caption "your own llms." clean build, fun kit. the claim is where it breaks. the fine print: the popular $70 pi ai kit uses a hailo-8l, 13 tops. it's built for vision, object detection and image processing, and it has no memory of its own. so it cannot run large language models. full stop the board that actually can is a different one: the newer ai hat+ 2, hailo-10h, 40 tops, with 8gb of dedicated ram. that's $130, not $70 and even that runs only tiny models. llama 3.2 at 1b, qwen 2.5 at 1.5b, deepseek r1 at 1.5b. edge llms live in the 1-7b range, against cloud models at 500b to 2 trillion so the honest pitch: for $130 you can run a very small language model on a pi, slowly, as a fun learning project. that's real and it's cool. "your own llms" on a $70 vision kit is not. why this keeps happening: "ai kit" and a big "tops" number sell. tops sounds like intelligence. but tops measures vision-style math, not whether the chip has the memory to hold a language model. the spec that matters for llms is ram, and the cheap kit has none. the honest caveats, both ways: the $70 kit is genuinely great, just at vision. cameras, object detection, that's its job the $130 hat really does run small llms locally, which a pi couldn't do at all two years ago. that's progress "small" is the load-bearing word. don't expect gpt at home on a pi the takeaway: before you buy a kit because the caption says llm, check two numbers. not the tops. the ram, and the size of the model it can actually load. no 70-dollar miracle, no gpt in a pi case, no tops number that means what you think. save this before you buy the wrong kit for the word on the box.

RetroChainer

11,100 次观看 • 2 个月前

I went a little overboard with Codex last week and burned through my entire weekly allowance in two days. Luckily, my quota reset today. Otherwise, I’m not sure what I would’ve done. It got me thinking: instead of asking one large model to handle everything from start to finish, why not let a stronger model plan the project and review the work, while a model built for execution handles the day-to-day implementation? So I tried it. The result was better than I expected. I used GPT-5.6 Sol in Codex as the decision-maker, then ran Ling-3.0-flash from Ant Ling inside OpenCode as the execution engine. Together, they built a small 3D farming game. Before writing any code, I had Codex create four documents: SPEC.md defined the product scope and the lines we couldn’t cross. ARCHITECTURE.md laid out the isometric coordinate system, state machine, and module boundaries. TASKS.md broke the project into small jobs Ling could tackle one at a time. ACCEPTANCE.md explained how each step would be tested and what “done” actually meant. Then I gave Ling a very straightforward role: You are the execution model for this project. Read all four documents before you begin. Work only on the task assigned for this round. When you’re done, run typecheck, test, and build. If anything fails, read the error, fix it, and run the checks again. Do not move on to the next task early. Ling handled dependency installation, project structure, strict TypeScript configuration, test setup, and a production build in 6 minutes and 3 seconds. It ran into issues with the Vite test config, a TS6310 error, and a missing jsdom dependency along the way. Instead of stopping at the first error, it kept reading the logs and fixing the problems until all three checks passed. The speed was honestly hard to believe. If you exclude the time spent waiting on tools, it was producing more than 100 tokens per second. That made the whole development loop feel noticeably faster. After this experiment, I’m planning to keep using the same workflow. If the task is small, there’s no reason to call an expensive planning model for every single step. If the task is large, handing the entire project to a Flash model in one prompt isn’t a great idea either. The setup that makes more sense to me is: Use a more capable model such as Codex to explore the project, make architectural decisions, and break the work down. Put the constraints into specs, schemas, types, and tests instead of leaving them buried in chat history. Give Ling-3.0-flash a steady stream of clear, verifiable implementation tasks. Report bugs with structured context and actual error logs, rather than saying, “It still doesn’t work.” Bring Codex back in for architecture reviews, visual checks, and changes that affect multiple parts of the project. The point of this setup isn’t to give AI a big “build the whole project” button. It’s to turn software development into a pipeline with a much more sensible cost structure: Codex figures out the plan, sets the boundaries, and catches problems. Ling-3.0-flash moves quickly, calls tools reliably, and works through well-defined tasks at scale. For agent workflows that involve lots of repetitive edits, production tasks, and tool calls, this may be a more practical answer than simply using the biggest model for everything.

雪踏乌云

23,107 次观看 • 2 个月前

We’re excited to introduce ShinkaEvolve: An open-source framework that evolves programs for scientific discovery with unprecedented sample-efficiency. Blog: Code: Like AlphaEvolve and its variants, our framework leverages LLMs to find state-of-the-art solutions to complex problems, but using orders of magnitude fewer resources! Many evolutionary AI systems are powerful but act like brute-force engines, burning thousands of samples to find good solutions. This makes discovery slow and expensive. We took inspiration from the efficiency of nature. ‘Shinka’ (進化) is Japanese for evolution, and we designed our system to be just as resourceful. On the classic circle packing optimization problem, ShinkaEvolve discovered a new state-of-the-art solution using only 150 samples. This is a big leap in efficiency compared to previous methods that required thousands of evaluations. We applied ShinkaEvolve to a diverse set of hard problems with real-world applications: 1/ AIME Math Reasoning: It evolved sophisticated agentic scaffolds that significantly outperform strong baselines, discovering an entire Pareto frontier of solutions trading performance for efficiency. 2/ Competitive Programming: On ALE-Bench (a benchmark for NP-Hard optimization problems), ShinkaEvolve took the best existing agent's solutions and improved them, turning a 5th place solution on one task into a 2nd place leaderboard rank in a competitive programming competition. 3/ LLM Training: We even turned ShinkaEvolve inward to improve LLMs themselves. It tackled the open challenge of designing load balancing losses for Mixture-of-Experts (MoE) models. It discovered a novel loss function that leads to better expert specialization and consistently improves model performance and perplexity. ShinkaEvolve achieves its remarkable sample-efficiency through three key innovations that work together: (1) an adaptive parent sampling strategy to balance exploration and exploitation, (2) novelty-based rejection filtering to avoid redundant work, and (3) a bandit-based LLM ensemble that dynamically picks the best model for the job. By making ShinkaEvolve open-source and highly sample-efficient, our goal is to democratize access to advanced, open-ended discovery tools. Our vision for ShinkaEvolve is to be an easy-to-use companion tool to help scientists and engineers with their daily work. We believe that building more efficient, nature-inspired systems is key to unlocking the future of AI-driven scientific research. We are excited to see what the community builds with it! Learn more in our technical report:

Sakana AI

360,561 次观看 • 1 年前

HERMES AGENT NOW RUNS CLAUDE OPUS 5. NEAR FABLE 5 INTELLIGENCE. HALF THE PRICE. SELF-VERIFIES ITS OWN WORK. AVAILABLE TODAY VIA NOUS PORTAL (20% OFF ALL MODELS). Anthropic shipped Opus 5 on July 24, 2026. same $5/$25 per million tokens as Opus 4.8. but the benchmarks tell a different story. WHAT CHANGED FROM OPUS 4.8: FrontierBench v0.1: Opus 5: 43.3%. Opus 4.8: 18.7%. 2.3x jump on the same test. ARC-AGI-3: Opus 5: 30.2%. 3x better than the next closest model. beat Fable 5 on 8 out of 13 benchmarks. at half the cost ($5/$25 vs $10/$50). same price as Opus 4.8. twice the intelligence. no reason to stay on 4.8. THE SPECS: model ID: claude-opus-5 context: 1M tokens (default and maximum) max output: 128K tokens thinking: on by default effort toggle: low / medium / high per request fast mode: $10/$50, 2.5x faster knowledge cutoff: May 2026 minimum cacheable prompt: 512 tokens (was 1,024) SELF-VERIFICATION (the biggest change): Opus 5 checks its own work automatically. Anthropic says: delete your verification prompts. "include a final verification step" now causes OVER-verification because the model already does it. for Hermes /goal tasks this is a direct upgrade. the judge checks evidence. the model also checks evidence. double layer of verification without extra tokens. EFFORT TOGGLE: low: fast, cheap, routine work. medium: balanced, daily tasks. high: full reasoning, complex problems. set per request. not a global switch. matches Hermes /reasoning command: /reasoning low (routine) /reasoning high (complex) Opus 5 effort toggle + Hermes reasoning control = precise cost management per turn. WHERE OPUS 5 FITS IN HERMES: DAILY DRIVER (replaces Opus 4.8): same price. 2.3x better benchmarks. set as your main model: Desktop app / Dashboard: Models → claude-opus-5 CHIEF OF STAFF: synthesis across multiple agents. reads Kanban, prioritizes, routes tasks. self-verification catches routing errors before they cascade. COMPLEX CODING: SOTA on agentic coding benchmarks. FrontierBench 43.3% = best public model for coding. set as coder profile model. /GOAL TASKS: self-verification + completion contracts = the model proves its work AND double-checks the proof. long-horizon goals finish correctly more often. MoA AGGREGATOR: strongest synthesis model at $5/$25. pair with GPT-5.6 and Grok 4.5 as references. Opus 5 aggregates. best quality at mid-range price. presets: max-quality: reference_models: - provider: openai-codex model: gpt-5.6-sol - provider: xai model: grok-4.5 aggregator: provider: anthropic model: claude-opus-5 COMPUTER USE: near-Fable 5 quality for browser automation. at half the token cost per session. computer_use tasks burn lots of vision tokens. Opus 5 halves that bill vs Fable 5. WHAT TO KEEP OPUS 5 AWAY FROM: cron monitoring: too expensive. use DeepSeek or no_agent mode. sub-agent grunt work: use GPT-5.6 Luna ($1/$6) or DeepSeek. auxiliary tasks: use Gemini Flash. routine web extraction: use a cheap model. Opus 5 is for the turns where quality compounds. planning, synthesis, verification, complex reasoning. budget models handle everything else. NOUS PORTAL: 20% OFF ALL MODELS Nous Portal currently runs a 20% discount on all models including Opus 5. $5/$25 official → $4/$20 through Nous Portal. the cheapest way to run Opus 5 right now. hermes setup --portal select claude-opus-5 as your model. discount applies automatically. Opus 5 replaces Opus 4.8 everywhere. same price. better at everything. no tradeoff. straight upgrade. hermes update /model claude-opus-5

YanXbt

16,744 次观看 • 2 个月前

Does LLM really need to be a helpful assistant all the time? No. If you want to simulate people, “perfectly helpful” could be the wrong objective. Meet OdysSim, a journey toward LLMs beyond assistants, as behavioral foundation models (10B tokens of real human behavior; 23 sim benchmarks, finally in one place. new open models: outperform or on par with GPT-5.5, Gemini 3.1, or Claude Opus 4.7 in many behavior-sim dimensions). Human behavior simulation is becoming essential. Agent evaluation needs realistic users before real users show up. Medical and classroom training need realistic patients and students. Social science needs synthetic participants at scale. But real people are not ideal assistants. Real patients panic or ignore good advice. Real students misunderstand. Real customers are vague, picky, impatient, or simply leave. Human behavior is messy, diverse, and often imperfect. Frontier LLMs are getting better at math, code, and long-horizon tasks. They are NOT getting better at simulating human behavior. If anything, they drift the other way: more assistant-ish, more homogeneous, fewer of the errors and quirks real humans show. This is no accident. The whole pipeline is built for helpfulness and task success, not behavioral realism. And you can't prompt your way out of that. So we rethink the recipe from scratch and release: 🧠 The OdysSim corpus: 21.4M real human interactions (~10B tokens) from 62 sources, every conversation retrofitted with social grounding (who is talking, and why) 📏 SOUL-Index: 23 human-behavior benchmarks unified into one suite across 5 axes 🤖 OSim-8B: open weights; tops more SOUL-Index benchmarks than any frontier model, acts more like a real user than any of them on τ-bench (nearly matching real humans in the reaction dimension), and writes far more human-like text along the way.

Xuhui Zhou

143,503 次观看 • 3 个月前

When your pain, fear or resentment is consuming you, no one else can pull you out of the dark, because you ARE the dark. Your mindset. Your thoughts. YOU CONTROL where you work & live, who you live with, what you eat & drink, what time you sleep, how you think, how you feel, how you spend your energy & what you do with your spare time. YOU! These things may feel out of your control but they're not. They're all things we can change if we're unhappy but our MINDSET has us convinced we can't, by playing on our fears of the unknown & our own weaknesses. 'What if I can't find other work?' 'What if I don't find anyone else?' 'I don't have time to cook!' 'I don't have time to workout!' 'I have nothing to feel happy about' 'I can't sleep!' Your life will not change, unless YOU do! You need to change your mindset! No one can do it for you. You can CHOOSE to see silver linings OR just the possibility of rain. You can CHOOSE your HARD... - Struggle through a workout vs struggle to look at yourself & how unhealthy you've become - Struggle to learn a new skill vs struggle in a job you hate. Trying to see beyond your unhappiness, 'bad luck' & bad choices is confronting & can feel impossible...BUT it isn't!! It's possible! Not over night but gradually...taking ONE DAY AT A TIME! Be kind to yourself! Changing your mindset & working on overhauling years of negative thought processes is far from easy but in doing so, you'll reclaim your POWER & take back control of your own happiness. Choosing to go out into the sunshine to drink a cuppa (instead of staying cooped up indoors) won't make your problems disappear...BUT it gives you a few seconds of CALM to REFOCUS your thoughts & a dose of Vitamin D & fresh air, while feeling the warmth of the sun on your skin. In making that choice, you reinforce your self-worth: you deserve time for yourself & you deserve to feel good. It's a simple choice that can completely change how you feel for the better, even if just for a moment. In time, your days can become full of those moments & you'll no longer need to look for the light in the dark. You will BE the light! The beacon of your own happiness ❤️🫂 #mindset

Jess Tungsten

314,526 次观看 • 2 年前

Introducing ml-intern, the agent that just automated the post-training team Hugging Face It's an open-source implementation of the real research loop that our ML researchers do every day. You give it a prompt, it researches papers, goes through citations, implements ideas in GPU sandboxes, iterates and builds deeply research-backed models for any use case. All built on the Hugging Face ecosystem. It can pull off crazy things: We made it train the best model for scientific reasoning. It went through citations from the official benchmark paper. Found OpenScience and NemoTron-CrossThink, added 7 difficulty-filtered dataset variants from ARC/SciQ/MMLU, and ran 12 SFT runs on Qwen3-1.7B. This pushed the score 10% → 32% on GPQA in under 10h. Claude Code's best: 22.99%. In healthcare settings it inspected available datasets, concluded they were too low quality, and wrote a script to generate 1100 synthetic data points from scratch for emergencies, hedging, multilingual etc. Then upsampled 50x for training. Beat Codex on HealthBench by 60%. For competitive mathematics, it wrote a full GRPO script, launched training with A100 GPUs on watched rewards claim and then collapse, and ran ablations until it succeeded. All fully backed by papers, autonomously. How it works? ml-intern makes full use of the HF ecosystem: - finds papers on arxiv and reads them fully, walks citation graphs, pulls datasets referenced in methodology sections and on - browses the Hub, reads recent docs, inspects datasets and reformats them before training so it doesn't waste GPU hours on bad data - launches training jobs on HF Jobs if no local GPUs are available, monitors runs, reads its own eval outputs, diagnoses failures, retrains ml-intern deeply embodies how researchers work and think. It knows how data should look like and what good models feel like. Releasing it today as a CLI and a web app you can use from your phone/desktop. CLI: Web + mobile: And the best part? We also provisioned 1k$ GPU resources and Anthropic credits for the quickest among you to use.

Aksel

1,268,589 次观看 • 5 个月前

over the weekend, i built an app that i sincerely hope you will never have a need for, but if you do happen to need a friendly, free, private mri viewer designed to make it easy for you to track tumor progression, you can try it here: here's the story: as some of you may know, last last september, my six year old daughter mira was diagnosed with an extremely rare brain tumor called an adamatinomatous craniopharyngioma, and since then our family has been doing everything we can possibly do to find a cure for her. we tortured chatgpt deep research, put together our own private research team, raised $1.4M and donated it all to Hankinson-Mitra Lab research thanks to $MIRA, explored every remotely applicable drug whether on the market or not, and even began working with md anderson to develop a personalized vaccine that we hope can lead to a more permanent cure unfortunately, we received the devastating news last march that the tumor has continued to grow since her initial surgery, and we had to start to consider more drastic options which would have seriously impacted mira’s quality of life. thankfully, with the help of dr. sabine mueller UCSF Benioff SF and the Hankinson-Mitra Lab at the university of colorado, in april, we started her on an alternative but extremely experimental treatment for this disease. to our unimaginable relief, her tumor has responded extraordinarily well to this treatment which combines tocilizumab (an arthritis drug that blocks IL6 receptors) with avastin (a colon cancer drug that inhibhits VEGF proteins). we know this, because mira gets an MRI scan of her brain every few months. and every time we get a new scan, the first thing we do is compare it against her last scan. so we have to find the matching weight of the scan, and then find the same plane, and then carefully find the slices of the scan where the tumor is visible, and then find the closest match to last month’s scan, then adjust the zoom, rotation, and brightness / contrast so they all look the same. we got pretty good at this. but it shouldn't be this hard. so i built last weekend using gemini 3 with some gpt 5.2 xhigh. you just import your DICOM MRI files (either zip, files or a folder), and you can align all of your scans across multiple dates instantly, just click and drag a rectangle around the tumor on any image, and it will use some very clever algorithms to automatically align up and find the closest matching slice from all your other scans, match the brightness / contrast, rotation, pan, zoom, and even shear to make sure the registration is as close as possible, and make it as easy to possible to compare tumor progression. it has a grid view so you can see all your scans for the same location all at once, and an overlay view so you can quickly compare two scans visually (by holding down the space bar to toggle quickly between two scans), along with tools to animate your scans both within the same sequence as well as over multiple scans to show progression. there is no server, it runs entirely locally on your browser - nothing ever gets uploaded and it's all open source: if you've found this useful, please consider a donation to the UCSF Benioff SF hospital foundation, who has given us extraordinary care over the past year or so:

Siqi Chen

209,414 次观看 • 8 个月前

Qwen 3.8 27B Q4_K_M - 90 tokens/sec on a single NVIDIA RTX 4090 (24 GB VRAM) with Dflash2! (MTP 60 tps -> 90 tps Dflash2!!!!) Local AI moves so fast (literally!) it’s terrifying. Z lab just dropped DFlash 2 for Qwen 3.8 27b and Muse Glimmer. I patched llama.cpp (PR #27342) and paired it with Unsloth’s Qwen 3.8 27B UD-Q4_K_XL quant. The result? Lossless 90 tokens/s decode. My last post highlighted native MTP hitting 60 t/s at 130,000 context. But DFlash 2 just completely shattered that ceiling. By using parallel block diffusion drafting (predicting whole blocks of tokens in a single pass using dynamic convolutions), DFlash achieves a massive 5.39 token acceptance rate. THE ALPHA TWEAK: `n-max 7` eats too much VRAM for draft states. But if you drop the draft limit to `--spec-draft-n-max 4`, you slash the VRAM overhead and actually increase the throughput. Here is the new 24GB VRAM Physics Matrix (DFlash 2 @ n-max 4): - 30k Context: 1,725 t/s prefill | 87.05 t/s decode | 22.2 GB VRAM - 80k Context: 1,789 t/s prefill | 84.20 t/s decode | 23.3 GB VRAM - 110k Context: 1,767 t/s prefill | 83.35 t/s decode | 23.96 GB VRAM (110k context at 83+ tokens a second sitting exactly on the 24GB hardware limit is absolute wizardry). How to compile the PR today: git clone cd llama.cpp git fetch origin pull/27342/head:pr-27342 git switch pr-27342 cmake -B build -DGGML_CUDA=ON && cmake --build build -j Llama.cpp flags for Dflash (110k Context Ceiling): ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q4_K_M.gguf --spec-type draft-dflash --spec-draft-n-max 4 -c 110000 -ngl 99 --port 8080 -ctv q4_0 -ctk q4_0 The fact that the open source community is shipping block diffusion drafters so quickly that run entirely locally on a gaming GPU is unbelievable. If you own a single RTX 3090 or 4090, it is officially time to upgrade to qwen 3.8 27b with dflash 2 and cancel your API subscriptions and let local silicon eat the cloud. This model beats GPT 5.6 Terra, GLM 5.2 DeepSeek V4 Pro, Muse Spark 1.2 and Claude Opus 4.8 on the artificial analysis agentic index (details in the replies) Hugging Face GGUF links (Base + DFlash2) and the full visual VRAM scaling and Dflash2 vs MTP graphs are also in the replies below. are you sticking to native MTP for the 130k context, or sacrificing 20k context to redline your decode speed? How many tokens/sec are you pushing on your current local rig?

Alok

105,633 次观看 • 1 个月前

The agents that run my LinkedIn outreach are currently getting a 55% reply rate on cold outreach and are generating hundreds of signups for me completely on autopilot. I packaged the exact system into a skill that works for any harness so you can copy it for yourself. Heres why this system prints qualified replies: most outbound starts with a list out of apollo, zoominfo or clay. firmographics. who could buy. that list is cold on arrival because "could buy" and "is thinking about this right now" are two different lists. so mine starts from linkedin posts my buyers are already reacting to. that also means every person on the list actually uses linkedin, since they just reacted to a post on it. most people don't, so a message to a normal list mostly sits unread. across the client accounts i've worked on, an apollo-style list usually gets about 1/10th the reply rate of one built this way. the posts come from three places: buyers posting about the problem themselves (hiring for it, breaking under it, shopping for a fix), a watchlist of my competitors and the creators my buyers follow, and the topic itself. when a buyer writes the post, the author is the best lead on it. anyone reacting to a competitor's product post is already shopping. every search is written in the words buyers actually post in, because vendor jargon only finds other vendors. it pulls everyone who reacted or commented, drops anyone i've already contacted plus competitors and their staff, and scores each person 0 to 100 against my ICP with the reason written on the row. anyone under the bar is gone before i've paid for their email. this is the part everyone skips. for every lead that clears the bar, it reads their website, their LinkedIn profile & anything else it can find about them on the internet, works out what agent would actually help their operation, and builds it. a real one, sitting there, with a claim link. so i am not sending "quick question about your outbound." i am sending "i built you an agent that does X for your company, want me to send over the link?" it sequences that into heyreach for linkedin and instantly for email, then runs the inbox. anyone who has done outbound knows reply speed is most of the game, and also the exact thing you quit doing by week three. every reply fires a webhook. it re-reads the whole thread, tags it interested, not interested or auto-reply, pings me in slack the second someone's interested, and sends them their claim link. i don't touch it. and every reply feeds the next run. it writes more searches like the ones that got answers, so week four sources better than week one. the general rule, and it has nothing to do with my software: use AI to build a personalized, genuinely useful asset and give it away in the first message. you're not hard selling anyone. you are handing them a thing. and because it's built off their actual company, nobody can mistake it for a template. reply rates go up for both reasons at once. it doesn't have to be an agent. the skill builds whatever your business is already good at making: - sell lead gen? a sample list of 10 to 25 of their own buyers - sell SEO? the 3 to 5 searches their competitors win and they don't - build websites? a teardown of their homepage with three fixes - consult? a one-page plan for the exact problem they commented about the playbook behind the messages is the best outbound advice on the internet in one place. it's built from 25 full-length videos and 60 posts from operators who do this at scale, plus 3 production systems running this exact play. scraped the videos and posts with superagnt's data MCP btw. run it once in claude code or codex and look at what it staged. once you like it, it turns the whole thing into a small agent team on a schedule. the sourcing side runs about 8 cents per qualified lead in production. grab it here, paste the install prompt into your agent and let it cook:

Jáen

543,402 次观看 • 14 天前

As we prepare to launch several projects, we're eager to provide a general update to our community. We are steadily approaching our end goal, thanks to the daily progress we're making toward our vision. Achieving our objectives will bring about a significant transformation in cross-chain interoperability and the flow of liquidity within protocols. This will address crucial challenges and drive mass adoption. Our future-focused approach and effective team collaboration keep us moving forward in an organized manner. Let’s delve deeper into the state of development of our current products and upcoming projects. Tao Bridge Starting with the Tao Bridge, which enables the #Bittensor community to unlock DeFi opportunities with their $TAO via a highly efficient blockchain like #MultiversX, known for its security, speed, and affordability. We deeply admire #Bittensor and believe a project like that is crucial for the future of not just the crypto space but also humanity, as it addresses the major challenges AI faces today: centralization, siloed and isolated work, which pose risks and hinder the technology's potential. We are committed to the vision of subnets and dynamic $TAO, convinced that this ecosystem is as groundbreaking as #Ethereum or #Bitcoin. We will continue to support #Bittensor wherever possible, and our bridge will also expand to other chains with Hatom V2. The TAO Bridge, deployed on and accessible through will launch on the Mainnet in 14 days, on March 27th. You can follow the countdown on the lending page at Given that our main priorities are security and stability, this period will be primarily focused on quality assurance to ensure a flawless Mainnet launch. The launch will also introduce TAO Liquid Staking at along with the integration of both $wTAO and $swTAO on the lending page. This allows #Bittensor users to leverage liquid stake, employ short or long strategies, among other DeFi strategies, or simply access stablecoin liquidity while maintaining exposure to their $TAO. Up to $1M will be distributed as additional incentives on top of the supply APYs at the launch of the $wTAO and $swTAO money markets, with $200K allocated for the first month specifically for bootstrapping. Initially, 70% of rewards will go to liquidity providers, and 30% to those using $HTM to boost their lending positions. This changes to a 50-50 split in the second month, and by the third month, all incentives are directed through the Booster. This approach encourages early participation and sustained engagement with $HTM. Introducing $TAO to #MultiversX will result in the creation of Liquidity Pools (LPs) on both AshSwap 🔥 and xExchange ⚡. These LPs will be incentivized by both entities, and Hatom will distribute extra rewards at launch. The goal is to make #MultiversX a one-stop hub for $TAO holders. Upon stabilizing the volumes, there will also be plans to integrate it on AshPerp 🔥. Furthermore, with the release of $USH, users will have the ability to mint it while retaining exposure to their $TAO. The TAO Bridge and TAO Liquid Staking smart contracts have been audited by Runtime Vеrification and @arda_project, while penetration testing and DevSecOps have been performed on our infrastructure by CertiK. We're excited to announce our exclusive partnership with TAONEW one of the top 5 validators on #Bittensor. TAONEW has been extremely helpful and supportive from day one. By sharing 50% of its service fee with its stakers, TAONEW enables Hatom to offer an optimized Staking APY to its users. Since our initial reference, #Bittensor has grown sevenfold, becoming the largest AI project in the crypto sphere. We reiterate our commitment to contribute to such technology and hope to address some of its current DeFi challenges. Syfy Moving forward, today marks a significant milestone, not only for our decentralized protocols but also for our development companies, which currently stand as the sole and primary contributors to the Hatom Labs and Soul Labs. We’re excited to unveil Syfy, the evolved identity of Hatom Labs and Soul Labs, now serving as the parent entity for our burgeoning development companies. Organization is crucial for scalability, which is why Syfy was established to cultivate an environment where our teams can collaborate more seamlessly, enhancing our effectiveness and efficiency. At the same time, we remain committed to upholding the financial independence of each project, supported by its own community of funding contributors. Feel free to explore our website at for more information! Additionally, don't forget to follow Syfy and explore their Genesis article highlighted in their initial post: Booster V2 The Booster V2 will introduce a range of new features and opportunities for $HTM holders: Optimized Position Boosting: Previously, boosting was done individually for each money market, necessitating $HTM token distribution and periodic rebalancing due to price fluctuations. With Booster V2, the system now considers the overall position, eliminating the need for manual rebalancing. Gas Fee Reduction: Booster V2 implements optimizations that result in reduced gas fees, making transactions more cost-effective for users. Incorporation of Governance: Users staking $HTM tokens gain voting rights directly within the Booster, allowing them to participate in governance decisions while maintaining their staked positions. (Note: Only $HTM tokens are considered for governance; LP tokens are not included.) Enhanced Boosting Mechanism: The Booster V2 enables LP Tokens to boost positions within the Booster, leveraging trading fees from swaps and farm incentives while boosting lending positions. Smart Contract Completion: The Booster smart contract has been completed and audited by @arda_project, ensuring security and reliability. Frontend Implementation: The frontend design for Booster V2 has been successfully implemented, providing users with an intuitive interface. Collaboration with xExchange: Exploration is ongoing for collaboration with xExchange ⚡ to enable LP creation, farming, and meta-staking within the Booster. Upon finalization of testing, we will launch the Booster V2 on the devnet to gather community feedback and begin preparations for the mainnet release. Soul Before delving into Soul Labs's developments, it's essential to summarize its core functionality briefly: Soul Labs seamlessly connects different lending protocols and blockchains, facilitating lending and borrowing across platforms like Aave, Compound Labs, and Hatom Labs, consolidating liquidity and users' borrowing capabilities. Utilizing LayerZero Labs and other messaging layers for cross-chain communication, Soul Labs bypasses asset bridging or synthetics, unlocking novel DeFi strategies and solidifying its position as the ultimate solution for cross-lending dilemmas. Soul V1 will be permissionless, holding censorship-resistant features, incorporating multiple redundancy mechanisms, and providing support for various DApps. We're thrilled to announce that, following the launch of the Tao Bridge in 2-3 weeks, we will introduce the Soul Labs website. This platform has been meticulously crafted over 250 days to not only provide a comprehensive overview of our vision but also to offer an engaging and captivating experience that promises to be memorable. Regarding the app, significant progress has been made on the V1 protocol, including: Smart Contract Development and Testing: • Completion of the initial phase of smart contract development. • Conducting advanced testing to ensure the system's robustness. • Establishment of a fully functional proof of concept. Successful deployment and testing on the #Goerli (#Ethereum Testnet) and #Mumbai (#Polygon Testnet), leveraging LayerZero Labs for seamless operation. Feature Enhancement and Protocol Optimization: • Enhanced testing procedures to bolster system resilience. • Integration of advanced features and significant code refactoring for optimization. • Incorporation of various communication methods, including LayerZero Labs, Formerly Axelar, now at @axelar, Chainlink CCIP), and wormholecrypto, into Soul Labs framework, enhancing its resilience and flexibility. This allows Soul Labs to maintain operation through alternative protocols if the primary one is temporarily paused. Website Development and Documentation: • Nearing the completion of the v1 app, with final touches being applied. • The preparation of comprehensive V1 documentation and the Yellow Paper, available upon Soul Labs's public launch, offering detailed insights into the platform's infrastructure and capabilities. USH Recognizing the critical need for stable liquidity within the ecosystem, we have positioned ourselves at the forefront of providing a solution by introducing $USH, the first native, decentralized, and over-collateralized stablecoin on #MultiversX. As market conditions have improved, we have observed a growing demand for stablecoins in the ecosystem, evidenced by the utilization rate in the Lending Protocol spiking to over 90% several times in recent months. Therefore, our goal is to tackle the current challenges faced by users by creating a robust product that will not only help them hedge against market volatility but also open up better opportunities to trade the markets and generate yield. We're happy to unveil the $USH website, now live with a sleek and intuitive user interface, designed for ease of use, which ensures that interacting with the protocol is straightforward and accessible for all. You can access it now through this link: For the technical side, we’re advancing steadily and we’ve accomplished the following milestones: Lending Protocol Facilitator: • Coded the first version to support multiple discount factors for different collaterals. • Implemented tracking of borrowing effectiveness to enable earnings forecasting for the module and support minting processes. Isolated Pools Facilitator: • Coded the first version of Isolated Pools Facilitator. • Use of $EGLD or $sEGLD as collateral, with positions stored always in $EGLD to benefit the protocol through Liquid Staking and lending interest. • Virtual account implementation for converting $sEGLD earnings into $USH, functioning like liquidation where users deposit $USH for a higher amount of $HsELGD. Staking Module • Coded the first version of the Staking Module that allows users to stake and unstake without any restrictions. We're currently focusing our efforts on the following tasks: • Implementation of HTM Booster in the discount model in the Lending Protocol. • Implementation of different depeg strategies and brainstorming further potential “soft” depeg mechanisms. • Research and implementation of rewards model for Staking Module. • Research and implementation of Boosted Vaults Facilitator. • Review and stress-test the first version of the code. Upon launch, $USH will be integrated into various protocols and AMMs across the ecosystem, further increasing both its utility and liquidity. The opportunities will be vast, enabling users to engage in a wide range of activities such as yield farming, staking, and arbitrage, all while leveraging a stable and reliable asset. Regarding the USH Airdrop campaign, it will continue until the official launch of $USH planned for late Q2-early Q3, rewarding all users who have actively participated in the initiative. Hatom V2 It is clear by now that we are driven to build a more robust, interoperable, and secure DeFi space, removing the current barriers that hinder users' capabilities to seamlessly interact with different blockchains. Through Hatom V2, we will introduce Hatom's cross-chain architecture, designed from the ground up for interoperability. This approach will elevate the protocol to unprecedented levels, enabling its deployment across various blockchains and facilitating seamless connections between them through Soul. By enhancing interoperability, Hatom V2 aims to foster a more inclusive and accessible ecosystem. This expansion will not only broaden the protocol's reach but also significantly increase its flexibility and utility, allowing users to interact with a diverse range of assets and products across different chains. We’re thrilled to share that we are currently crafting the V2 redesign of the Hatom webpage. Anticipate a jaw-dropping transformation that will truly astonish, blending cutting-edge design with an unparalleled user experience, elevating it to a dynamic, interactive hub, and making every interaction more engaging. Good things take time, but we are confident that the release of V2 website will take place in the second quarter of this year and will officially mark the start of our journey into the cross-chain landscape. We are excited about the future and we truly believe that this will mark the beginning of a new era for Hatom. It's crucial for us to develop rapidly without sacrificing the quality or the security of each product. We're strategically allocating resources to ensure smooth progress in every area of our work. As we push forward, we believe that the launch of Soul Labs will be the most important milestone due to its massive potential and disruptive technology. We would like to thank you all for the unwavering support you've shown over the past few months; it truly fuels our passion to push daily and make strides toward achieving our ambitious goals.

Hatom Labs

203,486 次观看 • 2 年前