A critical assessment of the $40M Series A that Andreessen Horowitz led into a confidential, pay-to-be-tested AI benchmarking startup — and the issuer-pays conflict sitting at the center of its “independent” pitch.
The labs Vals grades are also the customers who pay Vals, then reprint its scores in their own model cards. Vals has already disclosed a “customer relationship” with benchmark participants. That is closer to issuer-pays auditing than to a neutral rating agency — the exact structure that discredited Moody’s and S&P in 2008.
Vals keeps its test materials confidential to prevent contamination. But a benchmark can be auditable or uncontaminated — rarely both. Secrecy makes any given Vals score un-verifiable by the reader, so trust moves from “is the test well-built?” to “is this company trustworthy?”
Benchmarks are consumables — every model release erodes them, and independent analysis finds private test sets don’t durably stop saturation. Free public leaderboards like LMArena (~$1.7B) anchor one flank; enterprises’ own private evals anchor the other. Reputation is the only moat, and it is fragile.
Key Finding: The need is real and the team is technically strong — AI buyers genuinely lack a trustworthy way to tell which model does the work. But Vals is priced richly on unverified traction (the $400M valuation and “8× revenue” are not independently confirmed), and its central claim of independence is in structural tension with how it makes money. Watch for a public benchmarking dispute or a lab defection as the falsifying event.
a16z pitches Vals as the “trust layer” — the Moody’s, the S&P, the UL of AI. But follow the money and credibility, and the same loop that broke those analogies in 2008 appears.
The frontier lab is the paying customer — it commissions the evaluation of its own model.
Test materials stay confidential to prevent contamination — and to keep the score un-auditable by outsiders.
Vals produces a number. The buyer cannot inspect the test that produced it.
The lab reprints Vals’s number in its own model card — lending Vals credibility, and gaining leverage over it.
Enterprises rely on the “independent” score — a two-way dependency between grader and graded.
a16z’s own analogy cuts against the deal. “Credit markets developed Moody’s and S&P… product manufacturers rely on UL” — but Moody’s and S&P are the canonical case of issuer-pays ratings failing precisely when the ratings mattered most. FourWeekMBA: this “puts it closer to an audit firm than to a rating agency… Whether that distinction survives as the company grows and its relationships with the labs deepen is something to watch.”
Vals’s confidentiality is its anti-contamination moat. It is also what makes its scores impossible for a downstream reader to verify. Trust shifts from “is this test well-designed?” to “do I trust this 25-person startup?” — and independent analysis (Pebblous) found private test sets did not systematically prevent saturation anyway; age and size predicted it better than secrecy. The moat Vals sells does not durably solve the problem it sells against.
The eval/benchmark space is crowded and, above Vals, better capitalized. Vals’s niche — confidential, third-party, professional-task benchmarking — is the narrowest and the most conflict-laden.
Free public crowd-vote leaderboard — the de-facto standard Vals wants to displace. Raised a $150M Series A at ~$1.7B valuation. Anchors the free flank.
Private/expert evals plus SWE-Bench Pro from a multi-billion-dollar data giant. Far deeper pockets and existing lab relationships.
Eval + observability infra for teams building on LLMs. $80M raise at ~$800M valuation (Feb 2026).
Arize: LLM observability + eval, ~$70M Series C. Patronus: automated model/agent eval, $50M Series B. Both attack the enterprise-eval wedge.
Independent nonprofit benchmarks (e.g. FrontierMath). Donor-funded, no issuer-pays conflict — a credibility contrast Vals must answer.
The recurring 2026 consensus: “no general benchmark replaces testing a model on your own data.” The strongest competitor is the buyer’s internal eval — free and trusted.
The squeeze: free public leaderboards on one side, in-house private evals on the other, and better-funded observability players in the middle. Vals is betting that a paid, secret, third-party grade is more trustworthy than any of them — the exact claim its business model undermines.
Seven structural risks the $40M does not resolve.
Labs pay for the test, then market the score. Vals has already disclosed a customer relationship with benchmark participants. No public firewall (pay-to-test-but-not-for-outcome governance) has been described. This is the Moody’s-2008 failure mode, not the aspiration a16z cites.
Secret tests mean no third party can audit whether a Vals score is fair. Trust is purely reputational — and one credible, public dispute could be existential for a self-styled “gold standard.”
Revenue, credibility (via model-card citations), and even an investor (HRT, where a co-founder worked) overlap with the graded ecosystem. A “neutral” grader whose customers are the graded parties is capturable.
Every model release erodes prior tests; private test sets don’t durably prevent saturation. A treadmill cost structure that pressures margins and undercuts a ~$400M price on a company that has tripled its team but disclosed no revenue base.
“8× revenue” comes from a company X post with no baseline (8× of a tiny number is trivial); the $400M valuation is unconfirmed; customer-doubling is self-reported. The headline story underpinning the raise is not independently sourced.
LMArena (free, larger, better-funded) anchors the public leaderboard; enterprises increasingly run their own evals. Vals is squeezed from both sides with reputation as its only moat.
Benchmarking biosecurity, cybersecurity, “law of armed conflict,” and a federal-agency program raises liability, data-handling, and politicization risks a ~25-person startup is thinly resourced to govern — and Vals publishes no visible privacy or data-handling policy for the confidential customer prompts and eval data it processes.
Vals is selling trust in a market that badly needs it — using a business model that quietly erodes the very thing it sells. The gap is real and the founders are sharp. But before any buyer treats a Vals score as gospel, the diligence question is simple: who paid for the test, and can anyone else check it? On today’s evidence, the answer is the lab — and no.
Based entirely on publicly available information, including the TechCrunch announcement of September 19, 2026. Company-supplied figures (valuation, revenue growth, customer counts) are labeled unverified and are not presented as fact.