Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Until now, adding web search to open-source models meant hand-wiring orchestration, managing separate keys, and paying latency taxes on every round trip. Today, we're solving that. We're excited to introduce Baseten Hosted Tools and Baseten Grounded Inference to bring real-time web search server-side to open models running on Baseten...

29,259 Aufrufe • vor 6 Tagen •via X (Twitter)

11 Kommentare

Profilbild von Andrey Styskin
Andrey Styskinvor 6 Tagen

Baseten x Keenable ❤️

Profilbild von Baseten
Basetenvor 6 Tagen

🙌

Profilbild von Ishan Goswami
Ishan Goswamivor 6 Tagen

Exa 🤝 Baseten

Profilbild von kes
kesvor 6 Tagen

Great working with you guys on this!

Profilbild von Baseten
Basetenvor 6 Tagen

💚

Profilbild von Exa Developers
Exa Developersvor 6 Tagen

💚

Profilbild von Teo Gonzalez
Teo Gonzalezvor 6 Tagen

Love to see @baseten 🤝 @ExaAILabs

Profilbild von Baseten
Basetenvor 6 Tagen

@ExaAILabs 💚

Profilbild von Manish Tyagi
Manish Tyagivor 6 Tagen

Excited to see this unfold!

Profilbild von Turk 🇺🇸
Turk 🇺🇸vor 6 Tagen

@saranormous Seems biased. Mentions Dreamforce happening in SF and says nothing about Grok galaxy which has more overall attendees than any baseball game happening this week in SF

Profilbild von RemoteBrowser
RemoteBrowservor 6 Tagen

The latency tax framing is the real pain point. The round trips add up fast when the model decides to search mid-generation. Does Hosted Tools keep the search call inside the same inference loop, or is it still a separate hop?

Ähnliche Videos

🚨 THIS IS ACTUALLY INSANE Your AI agent can have access to the web. But if it can't reliably read what’s actually on the page, that access is almost useless. We looked at Firecrawl as the web layer for AI agents and the numbers are hard to ignore. The setup is simple: Give it a URL, search query, or website. Firecrawl handles the ugly part — crawling, scraping, rendering, extracting, and turning web content into something an AI model can actually use. The headline numbers: → 173,000+ GitHub stars → Search, scrape and interact with the web at scale → Supports web pages, PDFs, DOCX and other content → Structured data extraction for AI workflows → MCP support for connecting it directly to AI agents The workflow looks like this: Search → Scrape → Crawl → Extract → Feed the agent Three things stand out: 1. Scraping becomes an infrastructure layer Instead of maintaining your own pile of HTTP clients, parsers, browser automation and retry logic, you can treat web access as an API. 2. Agents get more than raw HTML The goal isn't just downloading a webpage. It's turning messy web content into clean context that an LLM can reason over. 3. The same layer works across different agent workflows Research agents. RAG pipelines. AI search. Competitive intelligence. Web-data extraction. The interesting shift: AI agents don't just need better models. They need better access to the information those models are supposed to reason about. Firecrawl is building that layer. Save this repo.

Vikas gupta

18,053 Aufrufe • vor 18 Tagen

What happens when there are more agents than eyeballs browsing the internet? How will the economics of the internet be reinvented? Parag Agrawal thinks Shapley values may hold the answer. Parag built Twitter over a decade, eventually becoming CEO and selling the company to Elon. He's spent the last three years building Parallel Web Systems, a search engine built for agents instead of humans. His core argument: 1) human click data is a bug. Agents aren't just a new technology, they're a distinct customer – and the feedback loop that made Google great is the wrong signal for the thing actually doing the work; 2) the ad-funded web assumed scarce human attention. If agents show up instead of eyeballs, the business model underneath the internet has to be rebuilt, or good content stops being published. The conversation covers: — why he shipped a search agent before a search engine, and how that let him grow the index incrementally instead of buying a full web crawl up front — the billion-to-billion matching problem: pull the right 1,000 tokens out of a trillion web pages, and your agent uses under half the tokens — going from a 3-second compute budget to 200 milliseconds — why he doesn't think Parallel is a neo-lab: "our output is a complement to a model" — the Google Cloud deal — on GCP, your grounding options are now Google Search or Parallel Search — reinventing the economics of the internet, Shapley values as a payment rail for publishers, and why the company was originally incorporated as Shapley Inc — why routing 2-10% of inference spend to web data would dwarf every content business outside the walled gardens — the web going from pull to push: "call me if this happens" 00:00 Introduction 03:25 What Is Web Search 05:17 Why Start a New Index 07:52 Search Agents First 10:17 Not a Neolab 13:14 Agents vs Google Search 19:38 Inside the Search Stack 28:59 Search Multipliers With Agents 30:21 Meeting Prep Agent Workflows 31:46 Quality Cost Latency And Turbo 32:42 Are Agents Overtaking Humans 34:28 Ads Model Meets Agent Web 37:20 New Incentives For Content 40:48 Shapley Values Attribution 47:46 Parallel Web And Future Vision Hosted with my very unwilling co-host Andrew Reed and Sequoia Capital

Sonya Huang 🐥

127,791 Aufrufe • vor 28 Tagen