Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing CoreWeave ARIA, the first AI research agent that runs autoresearch in your W&B dashboard. It reads your runs, finds what's working, and launches the next experiment itself. See it on Andrej Karpathy's nanochat, proposing configs and launching real training runs. Watch👇 Chapters 0:00 The setup, nanochat runs on...

88,176 görüntüleme • 2 ay önce •via X (Twitter)

7 Yorum

Weights & Biases profil fotoğrafı
Weights & Biases2 ay önce

You can ask ARIA something real, like whether dropout after the attention layers helps the overfitting you're seeing. Go get coffee ☕️ Come back to find it's formed a hypothesis, written the config, launched on W&B Launch, scored against your baseline, and proposed next steps.

Weights & Biases profil fotoğrafı
Weights & Biases2 ay önce

ARIA knows your project before you type a word. Open it from a failed run, ask what happened, and it's already reading the CUDA OOM in your logs, no pasting context. Every answer comes back as a live W&B dashboard you can pin and share, not a wall of text.

Weights & Biases profil fotoğrafı
Weights & Biases2 ay önce

Here are some use cases to get you going: - Why did my run fail?! 🫠 - Build a view of my runs and outliers - Why did these OOM? Check the logs and metrics - What hyperparams to try next? Launch it - Find a teammate's run and import it - How do I overlay metrics on a chart?

Weights & Biases profil fotoğrafı
Weights & Biases2 ay önce

You're already sitting on the experiments. ARIA turns them into the next model. It works on mobile, too. 😉 Public preview is live now. Open any W&B project, hit the agent icon, and ask ARIA something real. More info below!

akkiisfrommars profil fotoğrafı
akkiisfrommars2 ay önce

@karpathy This is legit really cool! Can’t wait to try it out.

Samian profil fotoğrafı
Samian2 ay önce

@karpathy ok autoresearch on nanochat is a real flex. my instinct is always to ship faster and skip this layer, then regret it when my eval pipeline turns into garbage

Jason傑森 🇭🇰 | 🛠️ profil fotoğrafı
Jason傑森 🇭🇰 | 🛠️2 ay önce

@karpathy 自主自动调咯先让它玩玩。

Benzer Videolar

Today we are re-introducing Warp, the AI-native employee management platform for payroll, compliance, benefits, HR, and IT that runs itself. We’ve raised $85M to rebuild the last great enterprise software category. We shipped more in the last 3 months than our first two years combined, so we made a video to tell our story and showcase the platform. Warp brings payroll, HR, compliance, IT, and benefits into one platform, from your first hire to your ten-thousandth. Payroll runs on its own, accurate and on time across all 50 states, and in 150-plus countries in local currency. Hire in a new state and Warp registers you with every agency automatically. When a tax notice shows up, Warp reads it, drafts the response, and files it across more than 10,000 jurisdictions. 80% resolve without anyone on your team touching them. The rest go to our in-house tax experts, not back to you. Onboarding runs itself. The moment an offer is signed, agreements, direct deposit, tax forms, benefits, and every account, app, and device get provisioned automatically. The day someone leaves, all of it deactivates in one action. Benefits work the same way. Warp is your licensed broker across every carrier for health, dental, vision, and 401k. You pick the plan. Enrollment, renewals, and every deduction into payroll happen on their own. Revenue has grown 100x in 2 years. 1,200+ companies run payroll, HR, and compliance on Warp. We're on track to process $1.7B in payroll this year. This is just the beginning of what AI-native employee management looks like. We can't wait to build the rest of it with you. Run your company at warp speed:

Ayush S

31,475 görüntüleme • 2 ay önce

HOW TO USE AI LOOPS TO RUN YOUR BUSINESS 24/7 A lot has been written about loop engineering for building products. Almost nothing about using loops to run the business itself. That's the bigger idea. A loop is when you give an agent a goal, a way to check its own work, and permission to keep trying until it hits that goal. Build. Verify. Repeat. Stop when the condition is met. Here's what it looks like in practice: 1/SEO loop You're position 30 for a term you want. The loop runs once a month, makes changes, checks where you rank, and keeps pushing until you're on page one. This is running in production right now on Inbox Zero. 2/Ads loop You're spending $100 a day and losing money. The loop tests creative, checks profitability, kills what fails, and keeps going until the account is in the black. 3/Eval loop Your AI feature is only 88% accurate. The loop keeps adjusting the prompt and swapping the model until it passes 90%. 4/LLM visibility loop People search in ChatGPT now, not just Google. Same loop, new scoreboard. Are we the answer or not? The whole thing hinges on one thing: a metric that comes back black and white. Where do I rank? Did it hit profitability? Did the evals pass? Give an agent that scoreboard and it runs for months. Loops used to run for 30 minutes. These run for a year. Take a step, sleep, wake up next month, take another one. You're basically hiring an agency that never sleeps, gets paid in tokens instead of invoices, and undoes its own mistakes when the number goes down. Full episode on The Startup Ideas Podcast (SIP) 🧃 watch

GREG ISENBERG

83,349 görüntüleme • 2 ay önce