Загрузка видео...
Не удалось загрузить видео
How do we run evals with LLMs? Interview with Tanishq Singh. This is a real interview question from a big tech company, asked to a candidate in their technical interview round. The video explains the answer in roughly 11 minutes. 00:00 Question - Reliability with LLMs 01:00 Observability &... show more
35,790 просмотров • 1 месяц назад •via X (Twitter)
Комментарии: 10

Good short one

A clean eval setup needs two loops: offline tests for repeatable comparisons, and production monitoring for failures the test set did not anticipate. Reliability comes from connecting both, not treating a benchmark score as the finish line.

Starting by, I dont know what ur company does & then explaining like an expert, how the bot answers reliably creates trust issues. Evals: Mentioning websearch for flight info & tying it to payment system is literally impossible... post serach only link click or summary feasible

The industry is starting to realize that LLM applications without structured evaluations are just expensive guessing games. The real engineering challenge is treating evals like traditional CI-CD unit tests. If you are not automating your relevancy and groundedness checks directly into your deployment pipeline, you are just waiting for a production incident.

Really good question to ask candidates. I'd push further though — the thing that actually separates people is whether they mention the eval set silently going stale as the model updates. Most eval pipelines rot within a quarter and nobody notices until production drifts.

Wonderful. Thanks you posted it.

Fantastic

evals on llms are all about handling uncertainty. nobody talks about how much data you actually need to converge on a reliable result, or the tradeoff between precision and speed

People just love making a simple concept sound overly complex. The tech industry in general.

Ai reliability has to include validation of accuracy,hallucination rate, latency, reliability…. Preceded by prompt, regression, bias testing Observability should include cost, latency, context passed, failure, retries, reasoning & human intervention at each layer Didn’t hear
