Model benchmarks have a blind spot: they measure general knowledge, not whether a language model can actually do the job you hired it for. Vals AI is attacking that gap, and it just raised $40.0M in a Series A round led by Andreessen Horowitz.
The startup’s platform evaluates models on the specific, industry-level tasks where they’ll be deployed-think regulatory filings for finance or clinical notes for healthcare-rather than generic trivia. That focus on real-world utility is what caught the attention of the storied venture firm, which is backing the company’s push to make model selection less of a guessing game.














.png&w=3840&q=75)

