正在加载视频...

视频加载失败

Finta uses AI across its accounting platform for startups, with one rule: accuracy over automation. Every day, Finta uses Respan to test prompts across models, manage golden datasets, and run evals before shipping. See how Andy and his team build reliable AI accounting ↓

54,459 次观看 • 10 天前 •via X (Twitter)

13 条评论

Respan 的头像
Respan10 天前

@andywang Thanks to Andy and the Finta team for doing this!

Barun 的头像
Barun10 天前

@finta @andywang Going from not knowing what happened to running tests daily is a big shift in how a team operates.

Frank Chen 的头像
Frank Chen10 天前

@finta @andywang Watching someone walk through their own workflow tells you more than any feature list!

Michael Mao 的头像
Michael Mao10 天前

@finta @andywang Prompt versions, evaluators, and thresholds all reading from the same production data is what makes this loop possible.

Hendrix Liu 的头像
Hendrix Liu10 天前

@finta @andywang "it was a struggle to know what was happening" is close to word for word what we hear on most first calls.

Ruiqing Yu | Building Product @ Respan(KeywordsAI) 的头像
Ruiqing Yu | Building Product @ Respan(KeywordsAI)10 天前

@finta @andywang Adding new evaluators as you find new failure modes is how this should work. The test set grows with the product.

Michael Wong 的头像
Michael Wong10 天前

@Andydy42 @finta @andywang killer product!

Dylan 的头像
Dylan10 天前

@finta @andywang Hearing how a team actually uses the product day to day is different from describing features. Good one to watch.

Sanjana 的头像
Sanjana10 天前

@finta @andywang The best sign a product works is when someone uses it for hours a day without being asked to.

E_genius 的头像
E_genius10 天前

@finta @andywang The evals-before-shipping bit is the actual product moat here. Reliable AI accounting won’t come from the flashiest demo, it’ll come from fewer silent mistakes.

Alex Liu 的头像
Alex Liu10 天前

@finta @andywang Running tests against a threshold before shipping is the part most teams skip. Good to see it as a daily habit here.

Chetan 的头像
Chetan10 天前

@finta @andywang Prompt versions and eval scores tied together is what makes a threshold mean anything. Otherwise you're just tracking diffs.

E_genius 的头像
E_genius10 天前

@finta @andywang The golden dataset + eval loop is the part most AI demos skip. Shipping fast is cute, shipping without regression roulette is the actual product.

相关视频