正在加载视频...

视频加载失败

Introducing PRAXIST Beta, your autonomous research team🚀 Define the objective, constraints, and what success looks like. PRAXIST discovers the path. From there, PRAXIST takes on the experimental research loop. Multiple Research Peers explore competing approaches in parallel, share useful findings, and build on accumulated evidence across experiments and generations....

283,095 次观看 • 14 天前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

auto-research is starting to gain traction as a very viable paradigm for creating useful research discovery. now, that paradigm is still in its infancy and the infrastructure to hold all that trail of context as the agents blaze through experiments isn't well defined (to say the least). on that topic, I had the chance to chat with my boys francesco and giulio from paradigma about what underlying infra is needed to make this paradigm work. the paradigma's paradigm, which involves copious amount of DAGs, make this auto-research paradigm a paradigmatic case of essential infrastructure. here's the full video in full: - 0:00 - what is missing from auto-research? - 2:02 - giulio and francesco ai journey - 8:10 - research infra is the bottleneck? - 10:18 - paradigma vision of autonomous research - 13:17 - “important discovery per joules” - 17:15 - why is DAG the unit of research for auto-research? - 20:40 - is paradigma trying to replace the research publication? - 24:50 - how does knowledge is shared between experiments in the DAG? - 27:34 - what is even auto-research lol? - 33:53 - the value of the human mind in this auto-research future. - 37:00 - how do you reconcile hallucination in this auto-research paradigm? - 41:33 - the adoption of auto-research across varied fields? - 47:30 - ✨ introduction to the auto-research infrastructure. ✨ - 56:55 - where is the code? - 59:10 - full IDE next? - 1:03:20 - the place of the human in this DAG / code quality? manual node? token spent? - 1:16:02 - who’s the user for auto-research? - 1:18:13 - how to validate bad DAG? - 1:20:18 - ✨ auto-research agent results ✨ - 1:22:53 - ✨ how a big research DAG looks like? ✨ - 1:25:10 - how to get the canonical DAG for the final result? - 1:27:50 - the auto-research DAG being the new pre-print? - 1:30:05 - what’s next for paradigma and the auto-research infra? - 1:35:00 - what are they excited about research wise? enjoyyyyy my guys 🌹

Yacine Mahdid

12,091 次观看 • 3 个月前

The Agentic Literature Review is Now a Reality. 🚀 I watched a student from King’s College London dismantle a task that used to take weeks. His mission? Deconstruct 47 complex academic papers for his dissertation. The old way: ❌ Endless skimming & highlighting ❌ Messy, unsearchable notes ❌ Citation chaos across 5 different tools The new way? He used a single platform and finished the core analysis in an afternoon. This isn't just another "AI summarizer." ResearchCollab is an AI co-pilot for research. I tested it against the standard "academic grind." The difference was staggering: 🔴 Traditional Workflow: Scattered PDFs, chaotic notes, mental burnout. 🟢 ResearchCollab Workflow: • AI instantly surfaces key insights from 250M+ papers without manual prompting. • Auto-organizes and relates everything personalized to the user. • Generates perfect citations (APA, MLA) in one click. • Brainstorms new research directions you might have missed. 👇 See how it works (No Credit Card Needed): Start your free trial → Why this is a silent revolution for knowledge workers: 1️⃣ Students are cutting literature review time by up to 80%. 2️⃣ Research teams are collaborating in real-time, killing version control nightmares. 3️⃣ The "blank page syndrome" is solved. AI helps you generate outlines and spark unique ideas instantly. The most compelling part? It’s not just about speed. It’s about clarity. ✔️ Finds connections between papers you'd never see. ✔️ Keeps your entire research universe in one searchable place. ✔️ Works 24/7 for less than the cost of your monthly coffee budget. This is the "Copilot" moment for academia and R&D. The barrier to high-quality, organized research has just collapsed. 👉 Support ResearchCollab Product Hunt launch : PS: I've compiled a short guide on "The 5-Day Research Sprint" methodology that this enables. Like 👍 and Comment "Research" and I'll DM you the link.

Anuj

10,780 次观看 • 10 个月前

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

132,623 次观看 • 29 天前

Hedgeye Launches New Mobile App, Giving Investors Instant Access to Market Research 🚨 FOR IMMEDIATE RELEASE Stamford, CT – [March 4, 2025] – Hedgeye Risk Management is excited to announce the launch of the Hedgeye App, a powerful new tool that puts Hedgeye’s market-leading investment research directly in the hands of subscribers. Available now for free download on both the Apple App Store and Google Play, the app delivers a seamless, on-the-go investing experience. With a clean, intuitive interface, the Hedgeye App makes it easier than ever to find the content you care about from Hedgeye’s vast daily research production. The app will help investors to stay ahead of important market trends, manage risk, and make smarter, data-driven decisions—all from their mobile devices. Hedgeye’s "Go Anywhere" full-cycle investing strategy and approach is now truly everywhere—giving users the ability to: ✅Receive real-time notifications on the latest actionable research and investment ideas ✅Watch or listen to The Macro Show, The Call, and other HedgeyeTV programs anytime, anywhere ✅Bookmark and save favorite research reports for easy reference ✅Access complimentary content like "Real Conversations" and additional research insights to explore Hedgeye’s research before subscribing While subscribers unlock full access to Hedgeye’s deep research and market insights, non-subscribers can still benefit from the app by accessing select free content, including market commentary, interviews with top investors, and a preview of Hedgeye’s research. It’s a valuable way to experience Hedgeye’s trusted investing process before committing to a full membership. "Investors need reliable, real-time insights to navigate today’s fast-moving markets," explains Keith McCullough, Founder and CEO of Hedgeye. "The Hedgeye App gives our subscribers instant access to my team’s research and risk management tools they need to protect and grow their wealth—anytime, anywhere." Download the Hedgeye App today on the App Store or Google Play, and take Hedgeye’s market signals, macro insights, ETF and stock ideas, and portfolio coaching wherever you go. Apple ➡️ Google ➡️ ABOUT HEDGEYE Founded in 2008, Hedgeye Risk Management is an independent investment research firm trusted by institutional and retail investors worldwide. With a team of over 40 seasoned analysts covering more than 1,000 stocks and ETFs across 20+ sectors, Hedgeye delivers in-depth, data-driven insights to help investors navigate and profit in any market condition. The firm is renowned for its risk management approach, combining macroeconomic research with bottom-up stock analysis to provide clear, actionable investing signals.

Hedgeye

37,572 次观看 • 1 年前

Is AxonDAO's Psylosis, the First of Its Kind? Psylosis is one of the pioneering DeSci initiatives in the US integrating AI, blockchain, and IRB-approved ethical protocols. This ensures the process and data security meet the highest standards for safety, privacy, and data integrity. Did you know that from 1955 to 2021, there have only been 19 studies on microdosing psychedelics, with an average of 22 participants per study. There have been no (0) known clinical trials combining microdosing psilocybin with voice AI analysis. Microdosing Research: Most studies focus on full (Macro) doses due to regulatory preferences for larger, measurable impacts. Voice AI Technology: AI-driven voice analysis forl health is new and experimental in psychedelics research. Pioneering Efforts: Psylosis bridges traditional methods with innovative AI and DeSci approaches. Biometric data provided by your voice for Psylosis, into CureOS is an objective, quantifiable measurement of psilocybin's effects, complementing subjective reports. This multimodal approach improves understanding of its impact on the brain and body, advancing safe and effective therapies. Psilocybin, the active ingredient in magic mushrooms, has been increasingly studied. By 2021, 13 phase II trials worldwide examined its use (in full doses) for treating depression. These studies reflect rising interest in its therapeutic potential. However, this field is still emerging. Ongoing trials are needed to better understand psilocybin's safety and efficacy across medical conditions, and proper legal jurisdictions. To sign up for the Psylosis Project, updates, and its commitment to ethical research practices, visit our site axondao(dot)io

AxonDAO

20,304 次观看 • 1 年前

More footage of the first documented cougar family in Minnesota in the past century. Volume up for the full experience. More to come soon! Our goal is to learn as much as we can about these cougars in the coming months. But we could really use some help covering costs associated with this research. For instance, we collected 9 scats at this kill and they are on their way to a lab for genetic analysis to try to get individual genetics and determine what western population the mom and dad originated from. Genetic samples cost ~$55-70 per sample, depending on the type and quality of the sample. Your support helps us cover costs like this, and gives us the ability and resources to study these individuals, and any others out there we might learn of. By donating at the link below, you directly support this research. Plus, the support helps us have the capacity to send in any samples we collect in the coming months.Once we have results, we will share with everyone! Notably, we also analyze the genetic samples from every adult wolf we collar, pup we tag, or dead wolf we come across. That work has been supported ENTIRELY by folks donating to our project, and the results have provided a wealth of information on wolf pack and population dynamics. And this work will only continue if generous folks continue to support our work. E.g., a $70 donation ensures we can get the genetics of a wolf. So please donate to our annual fundraiser to support our research, help us cover these costs, and keep this research going! Donate here:

Voyageurs Wolf Project

51,635 次观看 • 3 个月前

New short course: Collaborative Writing and Coding with OpenAI Canvas! Explore new ways to write and code with OpenAI Canvas, a user-friendly interface that allows you to brainstorm, draft, and refine text and code in collaboration with ChatGPT. In the short course, created with OpenAI, and taught by , a research lead at OpenAI, you’ll learn to use Canvas to enhance your workflows. Canvas lets you go beyond simple chat interactions. It provides a side-by-side workspace where you and ChatGPT can edit and refine text or code collaboratively. This makes brainstorming, drafting, and iterating as you write feel more natural and effective. As the first major update to ChatGPT’s visual interface since its launch in 2022, Canvas gives a new, innovative approach to collaboration with AI. For instance, after writing the first version of your code, Canvas can review it and give suggestions for improvement. It can also help with debugging by adding logging, identifying problems to fix, and writing comments. In addition, you'll also learn what it takes to train the model for an interface like Canvas. In this video-only short course, you’ll: - Learn how to ask for in-line feedback and control the iteration of your work by directly editing selected areas of your text or code from the model’s output. - Learn how to access quick automation tools in a shortcut menu that allows you to modify your writing tone and length, enhance your code, and restore previous versions of your work. - Learn how to use Canvas as a research assistant tool with an example of asking the model to reason through the screenshot of a plot to write a research report, in which you can ask questions within the created report. - Ask the model to write Python code to replicate the graph seen on a screenshot image. - Go behind the scenes of how you can create a video game, such as Space Battleship, from scratch, edit it, and display it in one self-contained HTML file. - Get a real-world application example of creating a SQL database from the image of its architecture. - Understand the model training and design processes that power Canvas! Please sign up here:

Andrew Ng

128,180 次观看 • 1 年前

LLM Wikis are being slept on. I argue that creating knowledge bases with LLMs or coding agents is one of the most valuable applications of AI today. It's about being intentional in building and scaling your intelligence stack. To showcase this, I wanted to share an LLM Wiki I have built over the last couple of months. It's called PaperWiki, and I use it across all my research workflows, along with my research agents. In fact, I also use it to curate papers I share with my communities, newsletter, and on X. The PaperWiki is updated regularly with automations, so I basically have agents on a loop maintaining it. All the entries are ingested from different sources and stored in a vault (Obsidian) and further indexed using qmd. And then further presented via an HTML artifact. So all of it is easily accessible to all my agents and easily searchable through full-text search and rich semantic search. The structure of the wiki has proven significantly useful to start interesting and exciting cutting-edge research projects with my research agents (from building tiny and more efficient gpt/difussion llms to building out SoTA harnesses and memory systems). It turns out that agents love markdown files and can more easily navigate the papers given the rich metadata structure of the wiki. I am just getting started on this, but it's clear to me that we should all be experimenting with LLM Wikis. Here's why: Building LLM knowledge bases gets you into the habit of leveraging AI outputs in all kinds of creative ways. It's the good kind of tokenmaxxing we should all be pushing for. LLM Wikis can be maintained automatically in a loop. I use an automation that updates the wiki every day based on papers I curate. The curation is another automation I run in a loop (with a bit of human in the loop), so I get to build on all my previous knowledge and expertise, and all of it compounds the deeper the integration/layers. One interesting result of this process is that I feel like I can better spot high-quality papers and remove noise more easily. Social media could never solve that. And most paper aggregators use metrics I simply don't trust. I like that agents can help with the noise vs. signal problem. This is important for research. Lots of people consider agents to produce mostly slop. But it doesn't have to be that way. Careful curations, prompts, automations, verifiers, and human-in-the-loop can produce some astonishing results. And you really don't need frontier models for this. I use a combination of frontier models (opus-4.8) and open-weight models (deepseek-v4-flash) to maintain this. An exciting future work (we are working on this DAIR.AI) is to tune specialized models on top of this to allow LLMs to quickly understand cutting-edge research ideas and can better conceptualize research strategies that further accelerate scientific research agents. I plan to open-source a bunch of this work, including the artifact, but this is currently work in progress, and I was excited to share some thoughts as I continue working on it. Sharing more as I go. Stay tuned!

elvis

55,747 次观看 • 2 个月前

I just vibe-coded a Meta ad research app in Claude Code 🤯 One keyword → winning ads analyzed, creative briefs written, trends mapped, and 10 ad variations generated for your brand. All inside Claude Code. Perfect for DTC brands and agencies still running competitor research by hand inside the Meta Ad Library. This app eliminates the entire loop: → Search and filter winning Facebook ads by niche, country, language, and performance tier → Watch video ads inline without ever leaving the app → Pull the top performers and generate a full creative brief tailored to your brand → Run a trend radar across 100+ ads to see which formats, CTAs, and landing pages are winning your niche → Pick any winning video and Gemini watches it, reverse-engineers the creative DNA, and writes 10 new variations for your product No scrolling the Ad Library for hours. No manual note-taking. No rewriting briefs from scratch. What you get: → Creative briefs built from real winning ads in your niche → Trend analysis: format distribution, top CTAs, landing page intel, and AI insights → 10 brand-specific ad variations from any winning video, with hooks, scripts, and Nano Banana prompts → A full app you host in Replit and customize for your team Built 100% in Claude Code with the gethookd.ai API + Gemini. I put together a free playbook with every prompt I used to build this from scratch, so you can build your own. Want it for free? > Like this post > Comment "META" And I'll send it over (must be following so I can DM)

Mike Futia

13,038 次观看 • 2 个月前

The hardest problems in AI aren't research problems anymore. They're deployment problems. It’s how we actually deliver real value, today, to build the future people want. That’s why, after 20 years in AI, my next step was inevitable: make robots do useful work for and alongside people, right now. Today, I am delighted to announce the launch of Walden Robotics to tackle just that. We started this year and are coming out of stealth today with a $300M seed round backed by some of the most serious companies and investors in the world. They have seen firsthand our general-purpose robots being useful in production on day one, and getting better every day after. You can see a glimpse of what we've been building in the video below. Physical AI has gone through a rapid phase transition, in part thanks to pioneering research from my friends and co-founders Russ Tedrake , Ben Burchfiel , Siyuan Feng, Rareș Ambruș , and many others at Walden. But from our long experience working together with co-founders Kerri Fetzer-Borelli and Dave Johnson, we learned how hard it is to deploy cutting-edge AI in a real, live, incredibly sophisticated production environment with an intricate ballet of automation and human ingenuity. That’s why we deliberately created Walden Robotics as a full-stack, human-centric, customer-focused robotics company from the start: we seeded the company with a world-class team across hardware, software, AI, deployment, operations, product, and business talent, so we could continuously optimize our whole system end-to-end, deeply and purposefully, from real-world experience with real customers. The efficacy of this strategy speaks for itself: since February, our general-purpose robots have been doing useful work in production at a Toyota plant in North America, moving from first pilot to real work in under two months. Not a lab. Not a demo. Not a future promise. Real work on a real line, today, at one of the best large-scale manufacturers in the world, with general-purpose robots that get better every day. And this is just the beginning. Two ways to find us: If you run a manufacturing or logistics business and want robots that are widely useful now, not someday, let's talk. We own “ for a reason! And if you want to build them: we're hiring across the company, from software, to hardware, AI, ops, product, business, and more. In particular, as the Chief Strategy Officer at Walden, I am recruiting for three incredibly impactful founding roles to fuel our agent-native go-to-market engine. Check out Let’s build together!

Adrien Gaidon

64,339 次观看 • 1 个月前

Everything you love about generative models — now powered by real physics! Announcing the Genesis project — after a 24-month large-scale research collaboration involving over 20 research labs — a generative physics engine able to generate 4D dynamical worlds powered by a physics simulation platform designed for general-purpose robotics and physical AI applications. Genesis's physics engine is developed in pure Python, while being 10-80x faster than existing GPU-accelerated stacks like Isaac Gym and MJX. It delivers a simulation speed ~430,000 faster than in real-time, and takes only 26 seconds to train a robotic locomotion policy transferrable to the real world on a single RTX4090 (see tutorial: The Genesis physics engine and simulation platform is fully open source at We'll gradually roll out access to our generative framework in the near future. Genesis implements a unified simulation framework all from scratch, integrating a wide spectrum of state-of-the-art physics solvers, allowing simulation of the whole physical world in a virtual realm with the highest realism. We aim to build a universal data engine that leverages an upper-level generative framework to autonomously create physical worlds, together with various modes of data, including environments, camera motions, robotic task proposals, reward functions, robot policies, character motions, fully interactive 3D scenes, open-world articulated assets, and more, aiming towards fully automated data generation for robotics, physical AI and other applications. Open Source Code: Project webpage: Documentation: 1/n

Zhou Xian

3,820,909 次观看 • 1 年前

There is a beautiful story that just happened in AI so let me share it for a lighter tone weekend post among all the doom stories in our AI field this week. It’s a story of people on three continents building and sharing in the open a new small efficient and state-of-the-art AI model. It started a couple of months ago when a new team in the AI scene released their first model from their headquarters in Paris (France): Mistral 7B. Impressive model, small and very strong performances in the benchmarks, better than all previous models of this size. And open source! So you could build on top of it. Lewis in Bern (Switzerland) and Ed (in Lyon, in the South of France) both from the H4 team, a team of researchers in model fine-tuning and alignment were talking about it over a coffee, in one of these gatherings that often happen at Hugging Face to break the distance between people (literal distance as HF is a remote company). What about fine-tuning it using this new DPO method that a research team from Stanford in California just posted on Arxiv, says one? Hey, that’s a great idea, replies the other. We've just build a great code base (with Nathan, Nazneen, Costa, Younes and all the H4 team and TRL community) let's use it! The next day they start diving in the datasets openly shared on the HF hub and stumble upon two interesting large and good quality fine-tuning datasets recently open-sourced by OpenBMB, a Chinese team from Tsinghua: UltraFeedback and UltraChat. A few rounds of training experiments confirm the intuition, the resulting model is super strong, by far the strongest they have ever seen in their benchmarks from Berkeley and Stanford (LMSYS and Alpaca). Join Clementine, the big boss of the open evaluation leaderboard. Her deep dive into the model capabilities confirms the results: impressive performance. But the H4 team also hosts a famous faculty member, Pr. Sasha Rush, Associate Professor at Cornell University in his daytime, hacker at HF in his nighttime. Joining the conversation, he proposes to quickly draft a research paper to organize and share all the details with the community. A few days later, the model, called Zephyr (a wind like Mistral), paper, and all details are shared with the world. Quickly other companies, everywhere in the world starts to use it. LlamaIndex, a famous data framework and community, shares how the model blew their expectations on real-life use-case benchmarks, while researchers and practitioners discuss the paper and work on the Hugging Face hub. All this happened in just a few weeks catalyzed by open access to knowledge, models, research, and datasets released all over the world (Europe, California, China) and by the idea that people can build upon one another work in AI to bring real-world value with efficient and open models. Stories like this are numerous everywhere around us and make me really proud of the AI community and see how we can build amazingly useful things together. [the video is just me reading this Friday post hahah]

Thomas Wolf

169,276 次观看 • 2 年前

In my past research experience, finding or developing an appropriate simulation environment, dataset, and benchmark has always been a challenge. Missing features, limited support, or unexpected bugs often occupied my days and nights. Moreover, current simulation platforms are relatively fragmented—making it challenging to replicate the success of the RT-X dataset in unifying community efforts. Introducing RoboVerse, we provide a unified platform, dataset, and benchmark for scalable and generalizable robot learning. We hope to build a shared foundation to combine the community efforts. RoboVerse includes: MetaSim: We carefully designed a configuration system and a universal interface to align current robotic simulators. With MetaSim, you can use any simulator with the same code—bringing together the community’s diverse efforts under one framework! RoboVerse Dataset and Benchmark: We unify popular simulation environments and benchmarks into a single cohesive system and introduce the RoboVerse dataset—a large-scale, high-quality synthetic dataset. Additionally, we propose a standardized benchmark across both imitation learning and reinforcement learning. A cool feature enabled by our unified framework: Hybrid Simulation! You can now integrate physics engines and renderers from different simulators—e.g., using MuJoCo precise physics with Isaac photorealistic rendering. This not only elevates simulation fidelity but also significantly enhances real-world transfer performance across complex robotic applications. Hopefully, our team’s efforts could serve the robotic community to thrive vibrantly in the years to come. RoboVerse is open-sourced🥳!!! Project Page: Documentation: Github Repo: Paper:

Haoran Geng

84,318 次观看 • 1 年前

I just vibe-coded a Meta ad research app in Claude Code 🤯 One keyword search → winning ads analyzed, creative briefs generated, trends mapped, and 10 ad variations written for your brand. All inside Claude Code. Perfect for DTC brands and agencies who are still doing competitor ad research manually inside the Meta Ad Library. If you're clicking through ads one by one, watching videos, taking notes in a Google Doc, trying to reverse-engineer what's working, and then rewriting briefs from scratch every time... This app eliminates the entire loop: → Search and filter winning Facebook ads by niche, country, language, and performance tier → Watch video ads inline without leaving the app → Fetch top-performing ads and generate a full creative brief tailored to your brand → Run a trend radar across 100+ ads to see which formats, CTAs, and landing pages are dominating your niche → Pick any winning video ad and Gemini watches it, reverse-engineers the creative DNA, and generates 10 new ad variations for your product No scrolling the Ad Library for hours. No manual note-taking. No rewriting briefs from scratch. What you get: → Creative briefs generated from real winning ads in your niche → Trend analysis with format distribution, top CTAs, landing page intel, and AI insights → 10 brand-specific ad variations from any winning video — with hooks, scripts, and Nano Banana prompts → A full app you host in Replit and customize for your team Built 100% in Claude Code with the gethookd.ai API + Gemini. I put together a free playbook with every prompt I used to build this app from scratch, so you can build it yourself. Want it for free? > Like this post > Comment "META" And I'll send it over (must be following so I can DM)

Mike Futia

53,602 次观看 • 5 个月前