Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

fine-tuning open-source models without this step is just burning GPU credits. inspect your data first. catch duplicates, empty outputs, bad rows. an AI agent can build the checker for you in minutes. i tried it in Qoder. i gave its desktop agent one task: build a local checker for...

141,237 Aufrufe • vor 9 Tagen •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

New open-source agent harness just landed! I got early access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.

elvis

11,303 Aufrufe • vor 1 Monat

Most people think their AI is not smart enough. It is smart enough. It just knows nothing about you. Every chat starts from zero. You paste the doc. You explain the project. You explain the same project again tomorrow. I finally fixed that part: 👇 I have been running Littlebird for a while now. The idea is simple. It is a desktop app for Mac and Windows that reads the text on your active window and sits in on your calls. So it already knows what you have been working on before you ask it anything. You do not brief it. You just ask. What that looks like in a normal day: 1/ Chat that already has the context No pasting. I ask what changed in a brief last week and it answers from the actual document that was open on my screen. 2/ Meeting Notes that write themselves It transcribes the call, then hands me the decisions and the action items. I get to stay in the conversation instead of typing through it. 3/ Routines that run on a schedule A morning briefing. A weekly summary of what I actually shipped. It shows up on its own. 4/ Hummingbird for the small stuff It appears right where you are working, so a quick question does not cost you a window switch and ten minutes of drift. The point is not the notes. Plenty of apps take notes. The point is that I stopped re-explaining my own work to a machine fifteen times a day. Free plan if you want to test the idea before deciding. Link is in the first comment, along with a discount for new users.

Mushfiq Sajib

66,683 Aufrufe • vor 24 Tagen

Met my girlfriend's parents for the first time. Her dad asked what I do for work. I said I build trading systems. He said like Wall Street? I said no. 6 AI agents. They work while I sleep. He laughed. So robots are making you money? I did not argue. I opened my laptop. Showed him the terminal. 6 agents running. 47 mispriced markets caught in the first week alone. His face changed. That is not gambling. That is automation? Exactly. Then I showed him how it works. Built the whole thing in 6 hours. Agent 1: Monitoring Runs 24/7. Watches Polymarket for mispriced markets. Spots an anomaly. Writes to memory and pings me on Telegram instantly. Agent 2: Research Parses news, X, macro data via browser tool on a cron schedule. Every morning I have a full digest on all open positions before I check my phone. Agent 3: Trading Reads the research agent memory. Sees the market has not reacted yet. Acts. Execution tool in gateway mode with a whitelist. No full access on a live server. Agent 4: Watchdog Heartbeat every 5 minutes. Monitoring running. No errors. Positions up to date. Something breaks. Immediate Telegram message. All of this. One Gateway. One config file. Isolation via per-agent scope. The token trick: stopped dumping everything into one file. Critical rules in bootstrap. Markets, patterns, past trades in memory. Semantic search pulls it when needed. Token spend dropped 3x. From $0.40 per request to $0.13. First week running: → 47 mispriced markets caught before Polymarket adjusted → Average entry edge 8 to 12 cents per position → Watchdog fired 3 times and caught a broken RPC before it cost me anything The whole system is plain text files. Open an editor. Change one line. Agent behaves differently. No deploy. No build. Her dad went quiet. Then he asked can you teach this? Her mom asked for the setup guide. I built the entire framework. Six agents. Full deployment. Memory architecture. Telegram alerts. You only need Claude + device + 1 hour per day. Giving this free for 24 hours. To get it: 1. Comment the word "Claude" 2. Like and retweet this 3. Follow me Himanshu Kumar so I can DM you Save this post. Deploy the 6-agent system this week. Start with $200. Scale on evidence.

Himanshu Kumar

47,590 Aufrufe • vor 3 Monaten

this is more useful than my entire degree Elon Musk's rocket company signed a $60,000,000,000 deal for Cursor in June, and eight days ago the two of them put a worker on sale for $200 a month: it gets its own computer in the cloud, signs into your accounts, clicks through your real apps, and hands back finished work instead of a draft for you to paste i ran one against my receipts folder on sunday and got back 14 filed, 2 it held because they needed a card number, and a saved method i never wrote myself Grok Bot is the one you train by doing your own job in front of it, and the whole handover fits in four messages tonight: 1. write out one job you did today the way you would brief a new hire: what has to be finished, which sites and files to work from, what to hand back, and where it stops and asks you 2. let it run once on something safe to get wrong, then correct the result until it is worth your name 3. say "save what we just did as a skill", and add the one rule about what always needs your approval 4. say "run that skill every weekday at 8 and post the result here. if the source is missing, tell me instead of using yesterday's numbers" xAI wrote that order into its own manual: one real job, then the saved method, then the clock. a schedule sitting on top of a method nobody checked replaces two hours of your clicking with two hours of your mistake turns out you never get to pick the brain, and that is the part i would argue about: the manual says there is no model picker for members or admins, no plan to add one, and the bill follows whichever model answered bookmark this, then open the piece below: which jobs deserve a worker of their own, and which ones quietly burn the seat ↓

Argona

21,946 Aufrufe • vor 1 Monat

whoever leaked this has bigger balls than sense Google Research and MIT ran the same agent jobs 260 different ways for Nature last month: they held the prompts, the tools and the compute budget identical and moved nothing but the wiring between the agents, and the same work swung from 70% worse than a single agent to 80.8% better, averaging out at 0.0% i ran my own single agent against the task list first and it cleared 6 of 10 alone, already past the line where a crew starts subtracting this is Graph Engineering, the layer that decides whether a crew is worth 80% more or 70% less, and it installs into the agent you already pay for: - score your solo agent on the real task first: above roughly 45% success that study predicts zero to negative returns from any crew you put around it - under that line, put one supervisor over the fan out: crews with no correction step amplified their own errors to 17.2x the single agent rate, supervised aggregation held it to 4.4x - give every worker one output and let none of them read a peer's draft, so a wrong step reaches the supervisor instead of four other agents - run the comparison again after every model upgrade, because a better model raises your baseline and a higher baseline is what makes a crew stop paying - keep the single agent alive as the control, the only number that says the wiring is earning its calls turns out the shape does not travel: the biggest win came off a finance task under one supervisor and the worst collapse off a planning task with independent agents my position, and it is the arguable one: a crew is a bet on your own diagram, and the model you pick moves that bet less than one arrow does bookmark this, the three moves that draw those arrows before you pay for one extra call are in the post below ↓

Argona

892,088 Aufrufe • vor 1 Monat

jev + opus 5.5 for design is insane... a launch video used to mean an agency, weeks of back and forth and a fat invoice. now Opus 5.5 makes the video, the ad cuts and the landing page from one brief. the video below: made in 3 minutes. at 0:09 it drops into edit mode - dot grid, selection handles, pixel sizes on every card - so you see exactly how it was built. but most people get the same mid result: centered text, gradient background, everything fading in. That's Opus's default look when it has nothing to copy. here's the pipeline that makes it look like an agency did it: 1. Build the brand brain first → logo, colors, fonts, screenshots of your real product, your best ad copy → have Opus turn it into a DESIGN.md. The video and the site both follow it 2. Steal a reference, don't describe one → pick 1-2 launch videos with the pacing you want → "make it like this" beats "make it modern and punchy" every time 3. Give it a renderer → install HyperFrames or Remotion: every scene is code, rendered straight to mp4 → a change = one line and a re-render, not starting over 4. Use real components → buttons, cards and UI from 21st_dev instead of empty placeholder boxes → fake-looking UI is the fastest way to look cheap 5. Storyboard before motion → 3 storyboard options, you pick one, then one still per scene → fixing a still takes seconds, fixing a render takes a re-render 6. Run it as a 4-person team + jev → breakdown: pull pacing, type and transitions from the reference → storyboard: the 3 options and the stills → build: every scene in code, music, draft render → review: watch it second by second, list the problems, touch nothing → jev: scores every review note against DESIGN.md (fix now / later / skip) in milliseconds, so build only touches what matters 7. Cut the ads from the same project → same scenes, different hooks: pain first, result first, offer first → 9:16 for reels, 1:1 for feed, 16:9 for YouTube. No reshoots 8. Ship the landing page from the same DESIGN.md → the hero is a looping cut of the launch video → before launch: mobile, every button state, zero placeholder text 9. Give notes like a director → "slow this zoom by half", "hard cut here", "push in on the button" → "make it better" gets random changes paste this into Claude Code 👇 "Read my brand folder and write DESIGN.md. Break down the reference video I give you: pacing, type, transitions. Show me 3 storyboards, then one still per scene for the one I pick. Build every scene in HyperFrames, add music, render a draft. Review it second by second without editing and list the problems. Pass every note to Jev and have it score each one against DESIGN.md: fix now / later / skip. Fix only "fix now". After I approve: cut 3 ad variants with different hooks in 9:16, 1:1 and 16:9, then build the landing page from the same DESIGN.md with the video loop as the hero." everyone has the same model. bookmark this and start using the edge.

Archive

39,253 Aufrufe • vor 4 Tagen