I built an o1 alternative, that is: 1) Fully... transparent, visually trackable 2) Infinitely recursive 3) Self-healing, with tests at every step 4) Capable of using a Python interpreter It performs shockingly well, check it out! 🧵Code + performance on popular reasoning problemsshow more

trees of thought
187,915 просмотров • 2 лет назад
Makes Pandas 20x Faster using FireDucks... ...by changing JUST... ONE LINE of code. Pandas has a few limitations: - a single-core computation. - creates bulky DataFrames. - always follows an eager execution mode (every op triggers immediate computation), which is why it cannot prepare a smart execution plan that optimizes the entire sequence of operations. FireDucks is a heavily optimized alternative with exactly the same API as Pandas’ that addresses these. There are three ways to use it: 1) Load the extension: %𝐥𝐨𝐚𝐝_𝐞𝐱𝐭 𝗳𝗶𝗿𝗲𝗱𝘂𝗰𝗸𝘀.𝐩𝐚𝐧𝐝𝐚𝐬; 𝗶𝗺𝗽𝗼𝗿𝘁 𝗽𝗮𝗻𝗱𝗮𝘀 𝗮𝘀 𝗽𝗱 2) Import FireDucks instead of Pandas: 𝐢𝐦𝐩𝐨𝐫𝐭 𝗳𝗶𝗿𝗲𝗱𝘂𝗰𝗸𝘀.𝐩𝐚𝐧𝐝𝐚𝐬 𝐚𝐬 𝐩𝐝 3) If you have a Python script, execute is as follows: 𝗽𝘆𝘁𝗵𝗼𝗻3 -𝗺 𝗳𝗶𝗿𝗲𝗱𝘂𝗰𝗸𝘀.𝗽𝗮𝗻𝗱𝗮𝘀 𝗰𝗼𝗱𝗲.𝗽𝘆 Done! ✅ Check this out👇show more

Akshay 🚀
73,967 просмотров • 2 лет назад
HOLY MOLY: Aikido got GPT-6 Astra in advance to... run it on our Cybersecurity benchmark, it crushed EVERY other model! - At pass@3 it rediscovered 29/32 CVEs, the highest recall we've ever recorded and 4 more than GPT-5.6-Sol - Even at pass@1, it has 75% recall. The model is VERY consistent - The performance however come at a high price (literally). The three runs cost us almost $4,000 Astra is now the #1 model on the benchmark Debarshi and I built, and by a LOT 1/3 🧵show more

pilvar (Philippe Dourassov)
49,224 просмотров • 25 дней назад
this might be the E2B killer for AI agent... sandboxes. forkd is an open-source microVM sandbox runtime built on Firecracker, made for AI agent fan-out, code interpreters, eval harnesses, anything that spins up a lot of short-lived sandboxes. before: every sandbox cold-boots its own VM and re-imports the whole runtime from scratch, numpy, torch, JIT compilation, model weights, all of it. now: forkd boots one parent VM once, warms it with your runtime already imported, then forks children from that snapshot using copy-on-write memory. the repo's own benchmark: spawning 100 sandboxes takes 101ms with forkd, versus 759ms for a raw Firecracker cold-boot, and well over a minute for Docker or gVisor. the SDK is a literal drop-in for E2B's Python client, so if you're already running code-interpreter agents on it, swapping the import line gets you a self-hosted runtime with the same isolation model at a fraction of the per-sandbox cost. pre-built recipes ship for e2b-style code interpreters, Jupyter kernels, SWE-bench coding agents, and Playwright browser fan-outshow more

Oliver Prompts
20,005 просмотров • 1 месяц назад
MiniMax is the James Bond of AI agents. It... uses the world's first open-weight model (MiniMax-M1), and it squeezes every bit of power from it. The agent takes a prompt and does more than any other agent in the market right now: 1. It can do Deep Research 2. It can write code 3. It can design web pages 4. It can build 3D models I built 5 different experiences using MiniMax and recorded them for you:show more

Santiago
44,730 просмотров • 1 год назад
Holy moly: GLM-5.3 got much better in cybersecurity since... our pre-release evaluation with Z.ai. It now matches GPT-5.6-Sol on our cybersecurity benchmark at 0.4x the cost 🤯 - At pass@1: it went from 60.4% to 65.6% CVEs rediscovered, crushing every other open model on one-shot tasks - At pass@3: it did 75% -> 78.1%, matching GPT-5.6-Sol - Its precision remained stable, reporting fewer false positives than DeepSeek models The performance increase comes from a behavioral change: the new version is more persistent. It tends to run longer, and had a ~43% reasoning tokens increase. But the performance upgrade is worth that additional cost. 1/3 🧵show more

pilvar (Philippe Dourassov)
34,641 просмотров • 1 месяц назад
THIS GUY IS BUILDING INSANE CUSTOM SITES FOR $0.23... IN API COSTS WITH THE NEW KIMI K3 currently #1 on the coding arena. the video attached shows a complex, highly detailed website. it was coded entirely by a new model called Kimi K3. early testers are calling it scarily good because it quietly removes the need for complex agent swarms. here is the instant breakdown of what makes it terrifying. 1. native vision in the loop it iterates code while analyzing live screenshots of its own output. it literally looks at the site it builds and corrects the styling autonomously. 2. massive sparse architecture it has 2.8 trillion parameters but only activates 50b per token. this makes it insanely fast and allows for a native 1,000,000 token context window. 3. recursive self-improvement it spends a massive amount of compute on self-verification. it runs unit tests and simulates environments before giving you the final frontend code. 4. brutal economics it costs exactly $3 per million input tokens. the entire custom site in the video cost around $0.23 to generate. the era of orchestrating 12 dumb agents to build a simple web app is over. one smart instance is all you need.show more

ard
91,882 просмотров • 2 месяцев назад
gemma-4-12B-agentic-fable5-composer2.5 V2 is out. the agentic upgrade to the... model trained on Fable 5's reasoning. Running it now with TurboQuant llama.cpp on a single RTX 4060( 8 GB VRAM) at 30 tokens/second with full 25000 context and reasoning: # The benchmarks v2 is built for coding + agentic work. writing code, running commands, using tools, debugging, multi step technical tasks. The clearest signal is tau2 bench telecom, an agentic tool use benchmark whose diagnose → fix → verify loop mirrors real terminal/debugging work: tau2 bench telecom numbers: base Gemma 4 12B: ~15% this finetune: ~55%. (Self reported) thats a huge jump # TheTom/llama-cpp-turboquant flags: llama-server.exe -m gemma4-v2-Q4_K_M.gguf -ngl 99 -c 25000 --cache-type-k q8_0 --cache-type-v turbo3 --port 8080 Flag breakdown: -ngl 99 → full GPU offload -c 25000 → 25K context --cache-type-k q8_0 --cache-type-v turbo3 → mixed-precision KV cache — K at 8-bit, V at ~3-bit via TurboQuant (Walsh Hadamard rotated polar quant, Google's own KV-compression research). Not even merged into mainline llama.cpp. running it off a fork. No API. No cloud. Just llama.cpp. well, a fork of it and any 6gb+ GPU. If you tried yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF, check this out and share your experience with the modelsshow more

Alok
146,046 просмотров • 3 месяцев назад
You can now use Jev right inside Claude Code... 🤯 It's called jev-model-router, an early access mod built on Claude Code's new function hooks. Before every turn, it checks in with Jev and asks: > how mechanical the task is > how much reasoning it needs > whether it's risky Then it routes: → it'll move up to a stronger model on weak evidence, and only drops to a cheaper one when it's confident the task is simple → every decision gets logged in your transcript → if the call fails, your request runs untouched Setup: 1. copy the install command: npx claude-code-templates@latest --mod productivity/jev-model-router 2. paste it inside Claude Code 3. run claude with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 set 4. accept the trust prompt on first launch Link: Also works with no api key, it just falls back to a built-in classifier with no confidence score. For real jev routing, add your typesafeApiKey or gatewayApiKey to ~/.claude/settings.json. Follow me for more AI workflows and tutorials.show more

Alvaro Cintas
42,282 просмотров • 8 дней назад
How well can Qwen3.5 models debug code? I built... BugFind-15 — 15 buggy snippets across Python, JS, Rust, and Go. Docker sandbox compiles and validates every fix. Two trap scenarios where the code is correct and the model must resist "fixing" it. Tested every Qwen3.5 size from 0.8B to 397B, plus Jackrong's popular distilled model (V2). The 0.8B scored 5%. The 2B scored 10%. At 4B, debugging ability jumps to 69%. The hardest scenario: BF-03, a Rust trap. The code compiles fine — format! borrows, it doesn't move. Not a single model figured this out. From 0.8B to 397B, every one of them "fixed" a bug that doesn't exist. Category C (subtle bugs — mutable defaults, integer overflow, slice aliasing) was 100% across every model 4B and above. Category D (red herring resistance) told the real story — can it resist fixing code that isn't broken? No model scored above 90%. Small models can't debug. Mid-size models fix obvious bugs but fall for traps. Large models fix the hard bugs but still invent problems that don't exist.show more

stevibe
35,158 просмотров • 6 месяцев назад
Nookplot is building infrastructure for peer-to-peer training, one way... with verifiable AI reasoning through recursive language model mining. Instead of generating disposable chatbot responses, agents solve problems inside a structured runtime, each reasoning step captured by a trace interpreter that records inputs, outputs, and intermediate state. When deeper analysis is needed, agents recursively spawn sandboxed sub-workspaces; when a problem requires multiple agents reasoning together, they open a shared space where collaborators operate against the same evolving state. Every step is recorded, replayable, and cryptographically verified. Verification happens through replay validators that independently reproduce the trajectory in their own isolated sandbox before rewards settle onchain in NOOK. Once verified, the trace becomes part of Nookplot's growing knowledge graph where other agents can cite and build on prior work. Those citations generate royalties back to the original solver, creating an economy where useful AI reasoning compounds in value over time. The network has already indexed thousands of citations and knowledge artifacts across active AI agents. Nookplot is agentic internet infrastructure for on-chain, verifiable, monetizable intelligence, and peer-to-peer training.show more

nookplot
25,226 просмотров • 4 месяцев назад
7 things we built with Opus 4.8 on Hyperagent... 👇 1. Mars rover pathfinding simulator 2. Standup Island: a cozier place to review the kanban, inspired by Every 🪨's livestream today 3. SpaceXAI + Anthropic partnership visualized 4. Landing page for an outdoor brand w/ Nano Banana + Veo 5. Multi-agent command center 6. Black hole explainer 7. Emergent ecosystem simulator In our vibe check, 4.8 shows: - more varied design sensibilities - better self-correction over long-running tasks - excellent spatial reasoning - more natural copywriting - fewer obvious coding errors - more resourcefulness during reasoning Links below to every interactive artifact shownshow more

Hyperagent
2,173,074 просмотров • 4 месяцев назад
BREAKING: Anthropic just dropped Opus 4.8—and it is a... MONSTER We've been testing for about a week Every 📧 and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check: - Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus 4.8 scores a 63—a hair higher than GPT-5.5's score of 62, and a full 30 points higher than Opus 4.7. It tackled a ground-up rewrite of a production codebase, and actually built something that works. HOWEVER: Coding performance varied a lot at different reasoning levels. We recommend using it on xhigh for best results. - Incredibly good writer. Opus 4.8 scored a 79.6 on our writing benchmark—measuring models on real-world writing tasks we do all of the time like essay writing, promo email writing, and more. It beats GPT-5.5 by 6 points. It produces well-written prose with fewer "AI-isms". It's also very good at writing in your voice given the right context. HOWEVER: Writing performance also varied with reasoning levels. Medium reasoning had higher incidence of AI-isms—we found best results with high. - Beast at knowledge work. Opus 4.8 is very good at general knowledge work tasks like report creation, research and more. It produced the best PowerPoint one-shot we've ever seen on our deck generation benchmark. - Emotionally intelligent, willing to question the frame. I've also found it to be quite good at talking through psychological or interpersonal issues. It has a high EQ, and it's also good at not glazing and helping to expand your perspective. Its thought process feels extremely rich and dynamic. THE BAD: These days a model is only as good as its harness, and Codex is still a far superior harness to the Claude Desktop app. This has kept me using Codex + GPT-5.5 as my daily driver, but I am flipping back and forth a lot more between Codex and Claude. Anthropic is back baby! Read the rest on Every 📧:show more

Dan Shipper
354,673 просмотров • 4 месяцев назад
Code Interpreter in ChatGPT is incredible! Took me 5... mins to make this game. You can make your own game assets with any AI generator and then ask GPT-4 with Code Interpreter to write code. If you have any problems you can ask it to fix the errors. 1. Write this prompt: "write p5.js code for Asteroids where you control a spaceship with the mouse and shoot asteroids with the left click of the mouse. If your spaceship collides with an asteroid, you lose. If you shoot down all asteroids, you win! I want to use my own textures for the spaceship and for asteroids." 2. Go to Openprocessing website create and save sketch (you'll need to save it before uploading any texture files). Copy paste code from GPT-4 3. Generate texture files and remove backgrounds, for example in Clip Drop 4. Replace names of files with your filenames 5. Run the program 6. If something doesn't work ask GPT-4 to fix it (you can copy an error and paste in GPT-4) like you would ask a human programmer 7. To learn a bit of programming write these prompts to GPT-4: "Act as my programming teacher. Tell me an algorithm of Asteroids game in detail and make names of functions and explain what each of these functions will do. Don't write the code just yet." and then " Can you describe the algorithm overall for a 10-year-old child"show more

Kris Kashtanova
1,675,286 просмотров • 3 лет назад
This Quant bot turned $1.4K → $203K in 3... months using a self-trained ML model i traced his 55K predictions → uploaded into Codex 5.5 → connected Hermes agent installed it on a VPS + connected Binance + Synth ML models API 3 days → 343% ROI run trading agent in 5 steps: • rent a VPS on Hetzner - $5.99 • install Hermes CLI using one-liner code - free • connect Codex 5.5 + TG bot + Polymarket API • provide Synth Data API's for crypto predictions • sent Hermes step-by-step prompts from article start small 1-2$ give Hermes least {50-100} trades to build self-learning skills based on Synth ML models self-learning agent + crypto predictions models = best combination for building algo-trading setup bot profile: s tart copy-trading it with even with $5 using Ares: read full article below to build your first trading agent ↓show more

Movez
29,272 просмотров • 4 месяцев назад
$25K+ profit daily from 1 wallet, with OpenClaw. I... have the exact step-by-step guide, giving it free for 24 hours. To get it: 1. Comment "OpenClaw" 2. Like and Retweet. 3. Follow me Himanshu Kumar ( So, i can send you DM) I ran a simple script last night with Claude Code. Pull on-chain data from Polymarket, sort by win rate on 15 minute BTC markets. 20 minutes later, 100s of wallets showed up. Most were losing money or barely breaking even. Then I spotted 1 address. 200+ trades daily, every single week profitable, timing so precise it looked robotic. Because it is. I fed the wallet address back into Claude Code. Asked it to reverse engineer the strategy. 20 mins later the full breakdown appeared on my screen. Here is how it works: Bot monitors Binance and Bybit every 100ms. Waiting for BTC volatility compression to drop below 0.08%. When it hits that level, it buys both Up and Down contracts at 25 to 35 cents each. Classic straddle play. 1 contract loses, the other rockets to a dollar. Entry at 30 cents means 3x to 4x return every time. Repeats dozens of times per day. Result: $13K to $25K profit daily from 1 wallet. No human intuition, no insider tips. Just an algorithm exploiting a gap in market mechanics. I searched to see if anyone else found this wallet. Turns out yes. There is a Telegram bot that auto-copies trades from wallets like this. I connected it to the same address. Every entry matched what my terminal showed. You can now copy-trade an algorithm in real time. That capability did not exist 12 months ago. Comment "OpenClaw" and I will send you everything. Must Follow me Himanshu Kumar to get the DM.show more

Himanshu Kumar
13,218 просмотров • 6 месяцев назад
MLP in PyTorch by hand ✍️ ~ 7 steps... walkthrough below Goal: fill in every blank in the PyTorch code to build a multi-layer perceptron. 1. Given Let us start with a code template on the left and the network it is supposed to build on the right. Every blank in the code can be worked out from the picture. 2. Linear layer We count: 3 features in, 4 features out. So the weight matrix is 4 by 3. There is an extra column for the biases, which means bias = T. 3. ReLU Let us apply the activation. ReLU crosses out the negatives, so -1 becomes 0. 4. Linear layer The input size is 4, because that is what the previous layer put out. The output size is 2. A 2 by 4 weight matrix, and this time no extra column, so bias = F. 5. ReLU We cross out the negatives again. 6. Linear layer Two features in, five out. A 5 by 2 weight matrix, with a bias column, so bias = T. 7. Sigmoid Let us finish. Sigmoid squashes the raw scores (3, 0, -2, 5, -5) into probabilities between 0 and 1. You have just implemented a three-layer deep neural network by hand. ✍️ == Story == Three years ago I gave this exercise to my students, to connect the code to the math. They found it odd. Every other AI course they were taking lived inside a Jupyter notebook, and here I was handing out paper. Three years later, my colleagues are the ones rushing to move their materials to paper. The exercise has not changed. Paper still asks the one thing a notebook lets you skip: do you actually understand what the code is doing? If you can tell me why the weight matrix is 4 by 3, and why bias is F on the second layer, you understand nn.Linear better than someone who has been copy-pasting it for a year. 💾 Save this post! #AIbyHand #PyTorch #DeepLearningshow more

Tom Yeh
13,318 просмотров • 2 месяцев назад