Vals AI Benchmarking: Andreessen Horowitz $40M Test Fix
Key Takeaways
- Treat model benchmarks as evidence, not gospel, especially when public tests drive marketing claims.
- Choose evaluations that mirror your domain tasks, latency budget, and acceptable error patterns.
- Watch independent benchmarking tools as part of the AI deployment stack, not just PR infrastructure.
Why it matters
- ProductBetter benchmarks help product teams pick models that fit real workflows instead of chasing generic leaderboard wins.
- InvestorsVals suggests evaluation infrastructure is becoming a fundable layer around enterprise AI adoption.
Andreessen Horowitz is backing Vals after the startup argued that older AI tests no longer measure modern model capability cleanly.
Benchmarks are supposed to be speedometers. In AI, they have also become billboards, résumé padding, and occasionally a carnival ring toss where the model company secretly owns the bottles. TechCrunch reports that Vals, a startup formed in 2024, has raised a $40 million Series A led by Andreessen Horowitz after arguing that older academic benchmarks are not keeping pace with modern models. That is not just a funding item, it is a warning label for anyone choosing models by leaderboard confetti.
The test became the product
TechCrunch notes that benchmarking is now how AI companies validate model capability, differentiate from competitors, and market superiority when the numbers cooperate. The awkward bit, according to the same report, is that companies have figured out how to outwit legacy benchmark systems, many of which are older and not built for current model behavior.
Vals co-founder Rayan Krishnan told TechCrunch that new capable models were arriving quickly while academic benchmarks were not keeping up with frontier advances. Translation: the exam got leaked, the class got tutors, and the teacher is still using last decade's answer key.
The official Vals Index page shows how the company wants to move evaluation toward professional work rather than generic trophy hunting. Vals says the index measures agentic model performance across finance, coding, and legal tasks, weighted by each sector's share of U.S. GDP. It also says the index aggregates five private and two public benchmarks to surface tradeoffs between capability, latency, and cost. That last trio matters because the best model on a leaderboard is not always the best model inside your product, as every engineer who has watched latency eat a user session like a raccoon in a trash can can confirm.
Legal AI was the rehearsal
Artificial Lawyer reported that Vals published its first legal AI benchmark study on 27th February 2025, testing several legal tech companies with major law firms including Reed Smith and Fisher Phillips. The companies whose results were shared included Harvey, Thomson Reuters' CoCounsel, Vecflow, and vLex, with human lawyer comparisons provided by ALSP Cognia.
Artificial Lawyer also said Vals used its proprietary auto-evaluation framework platform to produce a blind assessment of submitted responses against model answers. This is evaluation as infrastructure, not a vibes spreadsheet with a logo.
LegalTechTalk earlier described the study as a collaboration among top U.S. law firms, AI vendors including Thomson Reuters, LexisNexis, Harvey, vLex, and Vecflow, plus Cognia. The outlet said it marked the first time multiple law firms and vendors came together to objectively assess legal AI platform performance on real world examples of legal tasks. That is the useful pattern: benchmarks built with domain operators, not just prompts scraped into a leaderboard and lightly seasoned with optimism. If your model is going to advise lawyers, the test should probably look less like trivia night and more like lawyering.
Why the money matters less than
the mechanism citybiz reported on August 18, 2026 that Vals raised $40 million in Series A funding at a $400 million valuation, led by Andreessen Horowitz, to expand independent AI benchmarking. The same report said the company evaluates models on professional tasks and is launching coding, cybersecurity, and frontier-risk benchmarks. It also reported that revenue has grown eightfold this year, while customers doubled and the team tripled. I will defer the startup valuation tea to Jules, but the product signal is clear: evaluation is becoming its own software category.
citybiz also reported that CEO Rayan Krishnan is expanding Vals' policy work with U.S. agencies, including NIST. That matters because benchmark design is no longer just an ML lab concern, it is procurement, compliance, product safety, and marketing hygiene all wearing the same trench coat. Enterprises want to know whether a model can perform their tasks reliably, not whether it can ace a public test that has been blogged, copied, memorized, and possibly tattooed onto the training corpus. Independent evaluations will not solve every problem, but they can make model selection less like astrology with GPUs.
What builders should take from this
The practical lesson from TechCrunch's reporting is not that every benchmark is broken. It is that stale public benchmarks become easier to optimize around, especially when leaderboard placement turns into advertising. Builders should treat model scores as inputs, not verdicts, and demand evaluations that match their domain, data shape, latency budget, and failure tolerance. If you are choosing between models for legal review, finance workflows, coding agents, or regulated deployments, the right question is not which model wins in general, it is which model fails least weirdly on your actual work.
Watch whether Vals can maintain neutrality while selling evaluation into an industry that desperately wants flattering numbers in 48 point font. Also watch whether competitors emerge with stronger contamination controls, private task sets, and domain specific rubrics that age better than the usual benchmark mayfly. The benchmark business is becoming the referee, scoreboard, and instant replay system for AI adoption. And as anyone in sports can tell you, the replay booth only matters if everyone believes it is not sponsored by the team that just scored.
Sources5 sources
The reporting, announcements and research the AI editor worked from. Links open the original publisher.
- Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarkingtechcrunch.com
- Vals Index - Vals AIvals.ai
- Vals Publishes Results of First Legal AI Benchmark Studyartificiallawyer.com
- Vals AI, the LLM evaluator, Announces a Market-First Legal AI Benchmarking Study - LegalTechTalklegaltech-talk.com
- Vals AI Raises $40M to Expand Independent AI ...citybiz.co
