Loading video...
Video Failed to Load
Until now, adding web search to open-source models meant hand-wiring orchestration, managing separate keys, and paying latency taxes on every round trip. Today, we're solving that. We're excited to introduce Baseten Hosted Tools and Baseten Grounded Inference to bring real-time web search server-side to open models running on Baseten... show more
29,259 views • 6 days ago •via X (Twitter)
11 Comments

Andrey Styskin6 days ago
Baseten x Keenable ❤️

Baseten6 days ago
🙌

Ishan Goswami6 days ago
Exa 🤝 Baseten

kes6 days ago
Great working with you guys on this!

Baseten6 days ago
💚

Exa Developers6 days ago
💚

Teo Gonzalez6 days ago
Love to see @baseten 🤝 @ExaAILabs

Baseten6 days ago
@ExaAILabs 💚

Manish Tyagi6 days ago
Excited to see this unfold!

Turk 🇺🇸6 days ago
@saranormous Seems biased. Mentions Dreamforce happening in SF and says nothing about Grok galaxy which has more overall attendees than any baseball game happening this week in SF

RemoteBrowser6 days ago
The latency tax framing is the real pain point. The round trips add up fast when the model decides to search mid-generation. Does Hosted Tools keep the search call inside the same inference loop, or is it still a separate hop?
