Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

🔍 Introducing BackSearch. LLMs are increasingly asked to predict the future, but a good backtest requires a snapshot of the internet at a point in time. BackSearch allows LLMs to search the web as it was on a particular date. It’s great for: 🔮 Forecasting and prediction markets. 📈...

419,098 Aufrufe • vor 1 Monat •via X (Twitter)

35 Kommentare

Profilbild von will brown
will brownvor 1 Monat

oh hell yeah this is gonna print

Profilbild von General Reasoning
General Reasoningvor 1 Monat

Details on how to try BackSearch out here:

Profilbild von Daniel Lougen
Daniel Lougenvor 1 Monat

Wayback machine for AI, fuck thats smart. I wonder what kind of training data we might get out of this

Profilbild von Sam Z Liu
Sam Z Liuvor 1 Monat

but upgraded?

Profilbild von Cruncher Jean
Cruncher Jeanvor 1 Monat

You will still inject lookahead bias from the model if it's not completely retrain from scratch. Good for basic RL env but bad for rigorous backtesting. Check ChronosLLM:

Profilbild von Sriraam
Sriraamvor 1 Monat

We were just talking about this yesterday lol 🔥🔥

Profilbild von Josh Harris
Josh Harrisvor 1 Monat

This is exactly what we need when we were building our frontier financial judgement benchmark - will try to integrate for v2 :)

Profilbild von Vivek Iyer
Vivek Iyervor 1 Monat

best data product released in a minute damn

Profilbild von jakedineenasu
jakedineenasuvor 1 Monat

This is what we target in our paper "Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters", although we look at frozen slices of Reddit as a proxy for general markets against replayed Polymarket events.

Profilbild von Muratcan Koylan
Muratcan Koylanvor 1 Monat

The timing is perfect; I really needed this.

Profilbild von Albert Catalán Tatjer
Albert Catalán Tatjervor 1 Monat

I was just looking for something like this

Profilbild von M
Mvor 1 Monat

watch out ppl are gonna do so much prime-rl on this

Profilbild von λux
λuxvor 1 Monat

this is super cool!! 🙌🏻

Profilbild von Anthony Tamasi
Anthony Tamasivor 1 Monat

This is a great idea, but models are pretrained on info up to some cutoff date. Wouldn’t there still be data leakage from future events even without tool use?

Profilbild von prady
pradyvor 1 Monat

Woah!

Profilbild von Alex Peng
Alex Pengvor 1 Monat

This is so cool

Profilbild von Robin Salimans
Robin Salimansvor 1 Monat

this looks super useful

Profilbild von Anirudh Ravichandran
Anirudh Ravichandranvor 1 Monat

excellent, I can totally see the hill climbing with reproducible reward on search environments going brrr @willcb

Profilbild von Mani 🪽
Mani 🪽vor 1 Monat

this is super cool !!!!

Profilbild von Vantix AI Agency
Vantix AI Agencyvor 1 Monat

Searching the web from any past date changes how models get tested

Profilbild von Shreyas Pimpalgaonkar
Shreyas Pimpalgaonkarvor 1 Monat

Wow, this is actually so useful! Congratulations

Profilbild von Nayeem
Nayeemvor 1 Monat

Incredible

Profilbild von Thomas Mattimore 🇺🇸
Thomas Mattimore 🇺🇸vor 1 Monat

Will be a cool way to track various predictions made over time

Profilbild von Restitutor
Restitutorvor 1 Monat

Awesome, I can think of other uses for this

Profilbild von Oscar Hong
Oscar Hongvor 1 Monat

this is genius. always wondered why there wasn’t a benchmark for predictions. does this only work for events that happened after a model’s training date?

Profilbild von Matan Halevy
Matan Halevyvor 1 Monat

unreal, we’re working on a benchmark this would be useful for. Excited to give it a try

Profilbild von Luigi Pagani
Luigi Paganivor 1 Monat

Wow, so cool!

Profilbild von Jia Ming (إحسان)
Jia Ming (إحسان)vor 1 Monat

Google should be doing this

Profilbild von callum
callumvor 1 Monat

super cool! and very useful

Profilbild von आशीष खरे
आशीष खरेvor 1 Monat

This is new. I like the general idea. Great one!

Profilbild von Yusuf
Yusufvor 1 Monat

Need this

Profilbild von Ha Hoang
Ha Hoangvor 1 Monat

is this way back machine but for AI agent?

Profilbild von ethan
ethanvor 1 Monat

great idea

Profilbild von 2cents
2centsvor 1 Monat

@teortaxesTex This is very useful utility

Profilbild von BowTiedStingray
BowTiedStingrayvor 1 Monat

Beautiful - thank you for this

Ähnliche Videos

🚀Introducing VisualWebBench: A Comprehensive Benchmark for Multimodal Web Page Understanding and Grounding. 🤔What's this all about? Why this benchmark? > Back in Nov 2023, when we released MMMU ( a comprehensive multimodal understanding benchmark, we received feedback that it included very few UI screenshots. Considering the growing importance of UI understanding, especially with the rise of powerful agents like Devin ( which is built on the strong vision capability of #GPT4, we recognized the need for a benchmark focused on UI screenshot understanding.📸👀 > Multimodal #LLMs have significantly boosted web agents' performance on benchmarks like Mind2Web and WebArena. For instance, the SeeAct agent ( showcases the power of integrating vision into web agents. However, these benchmarks primarily evaluate the end-to-end task execution ability of web agents rather than their understanding of web pages. 🌉 Bridging the Gap with VisualWebBench > To provide a comprehensive evaluation of multimodal LLMs' web page understanding capabilities, we introduce VisualWebBench. Our benchmark spans 139 websites 🌐 across 12 domains 🏷️ and 87 sub-domains 🔍, ensuring a diverse and representative dataset. It assesses MLLMs at three levels: website-level, element-level, and action-level 📊, and encompasses seven tasks designed to evaluate understanding, OCR, grounding, and reasoning abilities 🧠💡. 😮 Surprising Findings > 🎉 Open-source models are catching up: Even though closed-source MLLMs are still leading the leaderboard, we are happy to see open-source models like LLaVA 1.6 34B achieve comparable performance to Gemini Pro. > 🧠 Grounding ability, crucial for developing MLLM-based web applications, is a weakness for most MLLMs. > 🖼️ Importance of Image Resolution: The limited image resolution handling capabilities of most open-source MLLMs restrict their utility in web scenarios, where rich text and elements are prevalent. > 🧱 Relatively strong correlation with general understanding benchmarks like MMMU but weak correlation with web agent benchmarks like Mind2Web. Web agent benchmarks primarily evaluate the end-to-end task execution ability of web agents, which involves a series of actions to accomplish a goal. In contrast, VisualWebBench emphasizes evaluating the foundational skills of MLLMs such as understanding and grounding web page elements. 💡Fun Fact > Claude Sonnet is better than Opus on our benchmark :) 🎓 Conclusion > VisualWebBench serves as a valuable resource for the community, driving research and development in the field of multimodal web page understanding and grounding. As MLLMs continue to evolve and improve, we look forward to seeing new applications and breakthroughs. We believe that our benchmark will contribute to the development of more powerful MLLMs in the web domain, ultimately leading to a more intuitive and efficient user experience on the web. Kudos to the student leads Junpeng Liu Yifan Song and the team Bill Yuchen Lin, Wai Lam, Graham Neubig, Yuanzhi Li! 👏 Check out more details in the Junpeng's thread👇

Xiang Yue

56,696 Aufrufe • vor 2 Jahren