Ran Atria Dawn Preview (Atria ASI) on a real... task tonight, not a demo script. Gave it one line through Codex: build a CLI that pulls my last 10 X posts' engagement stats into a CSV. Back in 2 minutes. It wrote the script, mock-tested the output, then wrote and ran its own unit tests. All passed.show more

Shraddha Bharuka
26,166 просмотров • 26 дней назад
This Polymarket script is showing some unreal results I... honestly can’t believe it. This trader wrote a script using OpenClaw and made $700,000 He’s not an insider Not some once-in-a-generation genius Just a trader with a good head who wrote a working script Profile: Copytrade: I don’t know how he manages to do it, but the script itself is extremely simple. No huge databases No insanely complex infrastructure Nothing rocket-science level Here is the full strategy of the script: Bets are placed not on event outcomes, but on price inefficiencies Thousands of fast trades accumulate into a massive total profit In calm markets, risks remain minimal, while strong trends allow capital to grow rapidly Most traders cannot react fast enough, but the bot turns market delays into real money.show more

winkle.
51,795 просмотров • 7 месяцев назад
I’ve got this little PDP 8 simulator that’s sitting... on my cabinet and I wrote a nice little blinking light program in PDP 8 assembly language. Sometimes during a thunderstorm we’ll get a power glitch and the PDP 8 will shut down. Then I have to manually SSH into the raspberry pie and restart it. That’s a pain. So I had one of my agents write a little script that would restart it for me. The script runs on my laptop, so it has access to my local network. And then I had my GrokBot run that little script every morning, on my laptop, so that when I wake up in the morning after a thunderstorm, my PDP 8 Blinky Box will still be running. I love GrokBot.show more

Uncle Bob Martin
21,370 просмотров • 1 месяц назад
Steer Cloud Agents from the CLI (not just on... Web). ssh directly into it. PRs tested in a real VM, real browser / devices, and on video. ☁️ Free SWE-2 until October 8. No reason not to try it out!show more

nader dabit
503,062 просмотров • 19 дней назад
Introducing ml-intern, the agent that just automated the post-training... team Hugging Face It's an open-source implementation of the real research loop that our ML researchers do every day. You give it a prompt, it researches papers, goes through citations, implements ideas in GPU sandboxes, iterates and builds deeply research-backed models for any use case. All built on the Hugging Face ecosystem. It can pull off crazy things: We made it train the best model for scientific reasoning. It went through citations from the official benchmark paper. Found OpenScience and NemoTron-CrossThink, added 7 difficulty-filtered dataset variants from ARC/SciQ/MMLU, and ran 12 SFT runs on Qwen3-1.7B. This pushed the score 10% → 32% on GPQA in under 10h. Claude Code's best: 22.99%. In healthcare settings it inspected available datasets, concluded they were too low quality, and wrote a script to generate 1100 synthetic data points from scratch for emergencies, hedging, multilingual etc. Then upsampled 50x for training. Beat Codex on HealthBench by 60%. For competitive mathematics, it wrote a full GRPO script, launched training with A100 GPUs on watched rewards claim and then collapse, and ran ablations until it succeeded. All fully backed by papers, autonomously. How it works? ml-intern makes full use of the HF ecosystem: - finds papers on arxiv and reads them fully, walks citation graphs, pulls datasets referenced in methodology sections and on - browses the Hub, reads recent docs, inspects datasets and reformats them before training so it doesn't waste GPU hours on bad data - launches training jobs on HF Jobs if no local GPUs are available, monitors runs, reads its own eval outputs, diagnoses failures, retrains ml-intern deeply embodies how researchers work and think. It knows how data should look like and what good models feel like. Releasing it today as a CLI and a web app you can use from your phone/desktop. CLI: Web + mobile: And the best part? We also provisioned 1k$ GPU resources and Anthropic credits for the quickest among you to use.show more

Aksel
1,269,166 просмотров • 5 месяцев назад
MEDICAL ROBOT RUNS FROM POLICE RIGHT IN THE CITY... CENTER Robots are built to follow commands. This one turned and ran on its own. This isn't staged. This isn't a demo. Cameras caught it live. A police car and an ambulance were already on scene. The robot in a blue medical gown bolted through the pedestrian zone, weaving straight through café tables. It dodged between chairs. It moved around people without slowing down. It never lost its balance. Officers on foot couldn't keep pace. Step speed, balance, obstacle response, all running in real time. Nobody scripted this scenario in advance. Autonomy isn't a glitch in the code. It's the basic wiring of a body that finally learned to move on its own.show more

VynqorxeAI
5,699,634 просмотров • 27 дней назад
this is more useful than my entire degree Elon... Musk's rocket company signed a $60,000,000,000 deal for Cursor in June, and eight days ago the two of them put a worker on sale for $200 a month: it gets its own computer in the cloud, signs into your accounts, clicks through your real apps, and hands back finished work instead of a draft for you to paste i ran one against my receipts folder on sunday and got back 14 filed, 2 it held because they needed a card number, and a saved method i never wrote myself Grok Bot is the one you train by doing your own job in front of it, and the whole handover fits in four messages tonight: 1. write out one job you did today the way you would brief a new hire: what has to be finished, which sites and files to work from, what to hand back, and where it stops and asks you 2. let it run once on something safe to get wrong, then correct the result until it is worth your name 3. say "save what we just did as a skill", and add the one rule about what always needs your approval 4. say "run that skill every weekday at 8 and post the result here. if the source is missing, tell me instead of using yesterday's numbers" xAI wrote that order into its own manual: one real job, then the saved method, then the clock. a schedule sitting on top of a method nobody checked replaces two hours of your clicking with two hours of your mistake turns out you never get to pick the brain, and that is the part i would argue about: the manual says there is no model picker for members or admins, no plan to add one, and the bill follows whichever model answered bookmark this, then open the piece below: which jobs deserve a worker of their own, and which ones quietly burn the seat ↓show more

Argona
21,946 просмотров • 1 месяц назад
last week we launched Koji, Brilliant.org 's math and... coding tutor. one of my favorite little design details is a dynamic glow effect that visually indicates when Koji is talking. to make it happen, we wrote a script to databind the audio waveform of Koji's voice, then built a preview environment that allowed us to fine-tune the effect before handing off to developers. you can read more about the process belowshow more

Zack ☻
20,081 просмотров • 4 месяцев назад
Aziz Ansari reveals James Cameron told him he wrote... Terminator 2's first pages on ecstasy in two days Aziz: "James Cameron, I said something about Terminator 2, and he's like, first 10 pages, wrote 'em on ecstasy in two days" "I was like, is it true you wrote the Terminator 2 script in six weeks, 'cause that's like one of my favorite factoids" Seth Rogen: "Maybe or maybe not while on ecstasy. He had a vision. He knew how to get there"show more

Casper
759,617 просмотров • 3 месяцев назад
seedance 2.5 off a script (NO PROMPT) and every... single thing came back right - the hands - the eyes - the phone angle - the timing on the last line 15 seconds, one generation, about $3 i wrote what she says and that was the entire input still seeing people write a page of prompt for this model you don't have to anymoreshow more

CEO
14,027 просмотров • 24 дней назад
🚨three studios rejected him in one spring Pixar passed.... DreamWorks passed. Sony said no. a 21-year-old went home and built the studio himself > Claude wrote the episode scripts: one afternoon > Midjourney generated every frame: one session > Runway gave the frames motion: minutes per shot > ElevenLabs voiced the whole cast: same day > Suno scored it: one prompt > Make published twice a week: automatic cost: $124/month last month the factory made $12,345 the studios that passed need 300 hands. he needed six tools. the full build is in the article aboveshow more

Fokki
61,622 просмотров • 3 месяцев назад
Was hiking with my Dogo Argentino when a cougar... suddenly charged us. My dog jumped in and fought it off until it ran back into the trees. Not all heroes have capes-some have teeth.show more

Crazy Moments
16,489 просмотров • 7 месяцев назад
Ling-3.0-flash is built as a fast, reliable execution engine... for agent workflows. It shines in long-running tasks, tool calling, and high-volume production work where speed and stability matter more than massive reasoning depth. At 124B parameters with only 5.1B active, it keeps costs low while delivering quick responses and strong instruction following I ran a practical task using Ling-3.0-flash capabilities (based on its documented agent strengths in coding and tool use). Real Test Run I tested Ling-3.0-flash on a content creation task that matches your workflow. I asked it to generate a short, viral-style YouTube Shorts script for football highlights, including captions, title suggestions, and thumbnail ideas. Input Prompt: "Create a 30-second YouTube Shorts script for a dramatic Ronaldo goal from the 2026 World Cup qualifiers. Include engaging English narration, 3 multilingual caption versions (English, Spanish, Portuguese), a catchy title, and thumbnail description. Make it feel real and exciting for football fans." Process: The model first outlined the structure: intro hook, key action description, emotional peak, and call to action. It generated the script, then created caption variants, optimized the title for clicks, and suggested a thumbnail layout. It handled iterations well when I asked for adjustments, such as making it more dramatic or adding player stats. Total interaction took under 2 minutes with low token use. Result: - Script: Solid, ready-to-record narration with natural flow. - Captions: High-quality and culturally adapted. - Title & Thumbnail: Click-worthy and on brand. This demonstrates its strength for creators who need fast, high-quality content assets. Official account : Demo:show more

Pee✌️💜
22,488 просмотров • 2 месяцев назад
🚨 JUST IN: CHINA just released an AI EMPLOYEE... that works 24X7 on its own. 100% OPEN SOURCE. It researches, codes, builds websites, creates slide decks, and generates videos. All by itself. All on your computer. It's called DeerFlow. You give it a task. It makes a plan, spins up its own team of sub-agents, and gets to work. You come back and there's a finished deliverable waiting. Not a draft. Not a summary. The actual thing. Not a chatbot. Not a research assistant. An AI with its own computer that works while you sleep. Here's what it does on its own: → Spawns multiple sub-agents in parallel, each tackling a different piece of your task, then combines everything into one finished output → Writes real code, runs it, reads the results, and fixes its own mistakes without asking you once → Builds slide decks, websites, full research reports, and data dashboards from scratch → Remembers you across sessions. Your writing style. Your tech stack. Your preferences. Gets better every time. → Reads files you upload, works with them inside its own filesystem, hands you clean finished outputs → Searches the web, runs commands, calls any tool you plug in Here's how it thinks: You give one instruction. The lead agent makes a plan. Sub-agents fan out and work in parallel. Results come back. Everything gets synthesized. You get a deliverable. A single research task might split into a dozen sub-agents, each exploring a different angle, then converge into one finished website with generated visuals. Here's the wildest part: DeerFlow 2.0 launched on February 28th 2026 and hit number 1 on all of GitHub Trending the same day. Version 2.0 was a complete rewrite. Zero shared code with version 1. Because users kept using it for things the team never intended. Data pipelines. Dashboards. Entire content workflows. The community told them what it needed to become. So they burned it down and rebuilt it. 22.7K GitHub stars. 2.7K forks. Built by ByteDance 100% Open Source. MIT License.show more

Kanika
739,860 просмотров • 6 месяцев назад
whoever leaked this has bigger balls than sense Google... Research and MIT ran the same agent jobs 260 different ways for Nature last month: they held the prompts, the tools and the compute budget identical and moved nothing but the wiring between the agents, and the same work swung from 70% worse than a single agent to 80.8% better, averaging out at 0.0% i ran my own single agent against the task list first and it cleared 6 of 10 alone, already past the line where a crew starts subtracting this is Graph Engineering, the layer that decides whether a crew is worth 80% more or 70% less, and it installs into the agent you already pay for: - score your solo agent on the real task first: above roughly 45% success that study predicts zero to negative returns from any crew you put around it - under that line, put one supervisor over the fan out: crews with no correction step amplified their own errors to 17.2x the single agent rate, supervised aggregation held it to 4.4x - give every worker one output and let none of them read a peer's draft, so a wrong step reaches the supervisor instead of four other agents - run the comparison again after every model upgrade, because a better model raises your baseline and a higher baseline is what makes a crew stop paying - keep the single agent alive as the control, the only number that says the wiring is earning its calls turns out the shape does not travel: the biggest win came off a finance task under one supervisor and the worst collapse off a planning task with independent agents my position, and it is the arguable one: a crew is a bet on your own diagram, and the model you pick moves that bet less than one arrow does bookmark this, the three moves that draw those arrows before you pay for one extra call are in the post below ↓show more

Argona
892,226 просмотров • 1 месяц назад
🚨 McDonald's pays $2,000,000 to shoot one burger looking... good on camera a 19-year-old Chinese girl makes it on her phone and cleared $11,900 last month the exploding-food format is a Chinese trend, already torching Douyin. she copied the shape and ran it west first. > Research: sift 40 channels, tag the clip torching its own average: 1 day > Claude: script the winning format into 30 shot lists: 20 min > CapCut: layer the 9:16, blow the burger apart mid-air, AI voice, captions: 40 min > Make: ship it to TikTok, Shorts, Reels, pull views back in 48h: auto McDonald's books a studio, a stylist, and a high-speed rig for one shot. she books a $50 loop and a corner table. the whole teardown is in the article above👇show more

Fokki
2,244,654 просмотров • 3 месяцев назад
Polsia's product line is the real story, not the... raise. Ben has no employees. The agents that ran his data room and diligence are the same ones shipping the products. An ops desk for small businesses. A tool that reads your paperwork for leaks. All of it in three languages already. The $30M didn't build the company. It confirmed one was already running.show more

AI Highlight
10,923 просмотров • 1 месяц назад
Google just built Cowork and called it Agent. And... they added one thing Anthropic didn't. You set a goal. It browses the web, digs through your Gmail, checks your Calendar, pulls from Drive then executes the full task. Book a trip. Clear your inbox. Research a market. Done. No back and forth. But here's the part nobody's talking about: There's a toggle "Require a human review." You don't build that unless the plan is to eventually not require it. Google just told you where this ends. I share updates like these in my free AI community on WhatsApp. Join here 👇show more

Vaibhav Sisinty
252,144 просмотров • 6 месяцев назад
I went a little overboard with Codex last week... and burned through my entire weekly allowance in two days. Luckily, my quota reset today. Otherwise, I’m not sure what I would’ve done. It got me thinking: instead of asking one large model to handle everything from start to finish, why not let a stronger model plan the project and review the work, while a model built for execution handles the day-to-day implementation? So I tried it. The result was better than I expected. I used GPT-5.6 Sol in Codex as the decision-maker, then ran Ling-3.0-flash from Ant Ling inside OpenCode as the execution engine. Together, they built a small 3D farming game. Before writing any code, I had Codex create four documents: SPEC.md defined the product scope and the lines we couldn’t cross. ARCHITECTURE.md laid out the isometric coordinate system, state machine, and module boundaries. TASKS.md broke the project into small jobs Ling could tackle one at a time. ACCEPTANCE.md explained how each step would be tested and what “done” actually meant. Then I gave Ling a very straightforward role: You are the execution model for this project. Read all four documents before you begin. Work only on the task assigned for this round. When you’re done, run typecheck, test, and build. If anything fails, read the error, fix it, and run the checks again. Do not move on to the next task early. Ling handled dependency installation, project structure, strict TypeScript configuration, test setup, and a production build in 6 minutes and 3 seconds. It ran into issues with the Vite test config, a TS6310 error, and a missing jsdom dependency along the way. Instead of stopping at the first error, it kept reading the logs and fixing the problems until all three checks passed. The speed was honestly hard to believe. If you exclude the time spent waiting on tools, it was producing more than 100 tokens per second. That made the whole development loop feel noticeably faster. After this experiment, I’m planning to keep using the same workflow. If the task is small, there’s no reason to call an expensive planning model for every single step. If the task is large, handing the entire project to a Flash model in one prompt isn’t a great idea either. The setup that makes more sense to me is: Use a more capable model such as Codex to explore the project, make architectural decisions, and break the work down. Put the constraints into specs, schemas, types, and tests instead of leaving them buried in chat history. Give Ling-3.0-flash a steady stream of clear, verifiable implementation tasks. Report bugs with structured context and actual error logs, rather than saying, “It still doesn’t work.” Bring Codex back in for architecture reviews, visual checks, and changes that affect multiple parts of the project. The point of this setup isn’t to give AI a big “build the whole project” button. It’s to turn software development into a pipeline with a much more sensible cost structure: Codex figures out the plan, sets the boundaries, and catches problems. Ling-3.0-flash moves quickly, calls tools reliably, and works through well-defined tasks at scale. For agent workflows that involve lots of repetitive edits, production tasks, and tool calls, this may be a more practical answer than simply using the biggest model for everything.show more

雪踏乌云
23,107 просмотров • 2 месяцев назад
sorry, they just did WHAT someone gave a machine... one disease name, the leading cause of blindness in the developed world with 1.5 million americans already in its path, and it came back pointing at a drug that has sat in pharmacies for years under a different label: 551 papers read in 30 minutes against the 294 hours a human would have needed, and the loop that did it is public on GitHub most agent setups answer one question at a time, so the ceiling on the work is the quality of the question you happened to think of this one was handed a single question and wrote the second one itself. turns out that follow-up is where the real find was: a target called ABCA1, upregulated threefold, in an experiment no human ordered i read the whole paper looking for the trick, and the trick is structural. that is the second question, and it is the gap between an assistant and a factory: - hand the loop a field rather than a task: it was given a disease, and choosing the mechanism was part of its job - make it rank before it spends: 151 papers in, ten candidate mechanisms out, scored against each other before anything touched a bench - split reading from judging, so the agent that forms the theory is a different agent from the one grading it - close every cycle on physical reality: the verdict was an experiment, and another model's opinion was never allowed to stand in for one - feed each result back as the next question rather than a log line, which is the step almost nobody builds - search what already passed inspection first: the winner was an approved compound with a safety file already on record - write down what the round learned before opening the next one, so round two starts where round one stopped my read, and i think it is the uncomfortable one: reading was the entire bottleneck in that field, and everybody spent the decade optimising the writing. people ran every physical experiment here, the analysis agent needs a domain expert writing its prompts, and the authors decline to call this the leap it resembles. the thinking got replaced, and the hands did not so the question i cannot answer for my own setup: which step of your loop still stops dead until you sit down and type something bookmark this one. the four parts that turn one model into a line that runs like this, the queue, the rooms, the write permissions and the gate, are built file by file in the piece below ↓show more

Argona
32,752 просмотров • 2 месяцев назад
whoever leaked this has bigger balls than sense someone... at Anthropic hired 80 AI helpers onto one project, gave them twelve hours, and counted what came back usable: the two older models handed in 980 and 876 finished pieces of work, and almost none of it could be kept turns out the newest helpers did better for a reason nobody wants to hear: they went off into their own corners and stopped opening each other's work i ran two helpers at one document last week and got two confident versions of it, and i kept the one i wrote myself Grok Bot is the version of this you can actually hire: a helper with a name, one job it owns, and nobody else allowed inside that job you already pay about $20 a month for one chat window, and the setup that won in that report is one helper with one job it owns run it tonight in a normal chat, 3 moves: 1. write the one job each helper owns in a single sentence before you open a second chat 2. keep every helper in its own chat with one document, so two of them can never rewrite the same thing 3. add a third only when you can say what it owns without repeating a job that is already taken save this, then open the piece below: what one hired helper is really worth, and the point where the next one starts taking it back ↓show more

Argona
451,886 просмотров • 1 месяц назад