Vals AI, a startup dedicated to becoming the independent evaluator for artificial intelligence, has successfully closed a $40 million Series A funding round, achieving a $400 million valuation. The investment was led by Andreessen Horowitz (a16z), a prominent venture capital firm with a history of backing major technology companies. This new capital injection will fuel Vals AI's mission to provide reliable and objective performance metrics in a rapidly evolving AI landscape.
Addressing the Benchmark Breakdown
The AI industry has long relied on academic benchmarks to compare models, but this system is now facing a crisis of credibility. As frontier AI systems become more advanced, they increasingly master the same public tests, rendering them less effective as differentiators. This issue is compounded by training data contamination, where test datasets inadvertently leak into the models' training corpora, skewing results.
A model that excels on a standardized leaderboard may still struggle with the complex, multi-step workflows required in real-world applications. The stakes are escalating as AI transitions from simple query-answering to performing autonomous tasks as unsupervised agents. This gap between benchmark performance and practical capability highlights the urgent need for a new evaluation paradigm that reflects genuine utility.
A New Standard for AI Evaluation
Founded by Stanford computer science alumni Rayan Krishnan and Langston Nashold, Vals AI is pioneering a novel approach to AI assessment. The company develops private, proprietary test sets that measure a model's ability to perform professional work in fields like law, finance, and software engineering. This method prevents the gaming and contamination that have undermined public leaderboards, ensuring the integrity of the evaluation process.
Vals AI combines the knowledge of domain experts with sophisticated automated grading systems to score model outputs against professional standards. The company maintains the relevance of its benchmarks by retiring tests once they no longer distinguish between strong and weak models. This dynamic approach has earned the trust of major AI labs, including OpenAI, Google, and Meta, which have cited Vals' evaluations in their official model releases.
Strategic Investment and Market Validation
The Series A round saw participation from existing investors 8VC, Pear VC, and Bloomberg Beta, alongside new backers HRT Ventures and Next Ladder Ventures. Andreessen Horowitz general partner Jennifer Li compared Vals AI's role to that of Moody's in credit markets, emphasizing that every mature market requires an independent referee. This investment validates the growing demand for a trusted third-party scorekeeper in the trillion-dollar AI industry.
The company has demonstrated remarkable growth, with revenue increasing eightfold since 2025 and its customer base doubling in the last six months. During the same period, the team has tripled in size, drawing talent from leading tech firms like Palantir, Microsoft, and NVIDIA. This rapid expansion underscores the market's confidence in Vals AI's solution and its critical role in the technology ecosystem.
Future Growth and Product Expansion
Vals AI will use the new funding to expand its evaluation infrastructure and enhance its product offerings. The company announced three major releases in conjunction with the funding news to broaden its impact. These launches are designed to empower developers, address emerging AI risks, and provide deeper insights into the economic impact of artificial intelligence across various sectors.
The new products include Vals Smith, a tool allowing customers to create custom coding benchmarks from their own GitHub repositories. The company also released Frontier Risk Benchmarks to assess AI safety in areas like cybersecurity and mental health. Finally, Vals Index 2.0 was launched, offering an expanded measurement of AI capabilities across the broader economy.
This substantial funding round positions Vals AI to solidify its role as the definitive authority on AI performance measurement. As enterprises and governments increase their reliance on AI, the need for independent, rigorous, and reliable evaluation has never been more critical. By building the essential measurement infrastructure for this transformative technology, Vals AI is helping to ensure that progress is both demonstrable and trustworthy.