- Vals AI raised $40 million in a Series A round at a $400 million valuation.
- Revenue has grown eightfold since 2025, with both the customer base and the team expanding rapidly.
- Founders Rayan Krishnan and Langston Nashold are building AI’s independent scorekeeper.
Every frontier model now claims to be the smartest. Vals AI is betting that nobody can prove it, and that someone independent needs to.
The San Francisco startup has closed a $40 million Series A at a $400 million valuation, led by Andreessen Horowitz, which has previously backed OpenAI, Anduril, Databricks, and Stripe. Existing backers 8VC, Pear VC, and Bloomberg Beta returned, joined by new investors HRT Ventures and Next Ladder Ventures.
Why benchmarks stopped working
Academic benchmarks worked for years, giving the industry a common way to compare models. That has broken down as frontier systems increasingly ace the same tests: public datasets get saturated, leak into training data, or become the exact thing a model is optimized against.
A model can look brilliant on a leaderboard and still fall apart on the messy, multi-step work it’s actually deployed to do. The stakes are higher now that models are moving from answering questions to doing the work itself, often as agents running unsupervised for hours or days at a time.
Rayan Krishnan and Langston Nashold, Vals’ CEO and CTO, studied computer science together at Stanford and were already working on real-world measurement problems before starting the company. Their team has also worked at Palantir, Microsoft, NVIDIA, Meta, and Hudson River Trading.
Grading models on the job
The startup works by pairing domain experts across law, finance, healthcare, and coding with automated grading systems that score model output to professional standards, rather than testing whether a model can pass a bar exam or ace a textbook problem. Its test sets stay private and are run in limited numbers to prevent the contamination and gaming that have undermined public leaderboards.
Vals says it can turn around benchmark results within hours of getting access to a new model, and it retires tests once they stop separating strong models from weak ones — in May, it swapped a corporate-finance benchmark called CorpFin for a new Excel-modeling test for exactly that reason.
Its evaluations have been cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI, and enterprises use its scores to decide which models go into production.
Tech Funding News recently covered LMArena’s $150 million raise at a $1.7 billion valuation, which relies on user-driven comparisons, and Cambridge-based Trismik’s £2.2 million raise to apply psychometrics to the same problem. Datacurve’s $15 million Series A for private coding datasets is another piece of the same puzzle.
Vals’ vision is that a static benchmark won’t cut it, and it wants to keep rebuilding the test as fast as the models improve.
An independent scorekeeper for a trillion-dollar industry
Announcing the round, a16z general partner Jennifer Li said every market eventually needs an independent scorekeeper, comparing Vals to Moody’s for credit markets or auditors for public companies: once sellers know more than buyers and have every incentive to look good, an outside referee makes the market work.
Vals will use the new money to expand that infrastructure. Alongside the funding, it launched Vals Smith, which lets customers build coding benchmarks from their own GitHub repositories; a set of frontier-risk benchmarks covering cybersecurity, mental health, and AI safety; and Vals Index 2.0, which extends its measurement across the broader economy.
The company says revenue is up eightfold in all of 2025, its customer base has doubled, and its team has tripled in six months. It previously raised $5 million in seed funding from 8VC, Bloomberg Beta, and Pear VC.