Vals AI: Who Grades the Graders?

A critical assessment of the $40M Series A that Andreessen Horowitz led into a confidential, pay-to-be-tested AI benchmarking startup — and the issuer-pays conflict sitting at the center of its “independent” pitch.

ProofStory Research September 19, 2026

$40M Series A Led by Andreessen Horowitz — September 19, 2026

Vals AI runs paid, largely confidential benchmarks that test frontier models on professional work — law, finance, coding, and “frontier-risk” domains. It wants to be the “trust layer” between the labs and the buyers who rely on them.

$40M
Series A · a16z
~$400M
Reported Valuation (Est.)
8→25
Employees In 2026
Self-Reported Revenue Growth

Three Core Questions

01

“Is This Real Independence?”

The labs Vals grades are also the customers who pay Vals, then reprint its scores in their own model cards. Vals has already disclosed a “customer relationship” with benchmark participants. That is closer to issuer-pays auditing than to a neutral rating agency — the exact structure that discredited Moody’s and S&P in 2008.

02

“Can a Secret Test Be Trusted?”

Vals keeps its test materials confidential to prevent contamination. But a benchmark can be auditable or uncontaminated — rarely both. Secrecy makes any given Vals score un-verifiable by the reader, so trust moves from “is the test well-built?” to “is this company trustworthy?”

03

“Where Is the Moat?”

Benchmarks are consumables — every model release erodes them, and independent analysis finds private test sets don’t durably stop saturation. Free public leaderboards like LMArena (~$1.7B) anchor one flank; enterprises’ own private evals anchor the other. Reputation is the only moat, and it is fragile.

Key Finding: The need is real and the team is technically strong — AI buyers genuinely lack a trustworthy way to tell which model does the work. But Vals is priced richly on unverified traction (the $400M valuation and “8× revenue” are not independently confirmed), and its central claim of independence is in structural tension with how it makes money. Watch for a public benchmarking dispute or a lab defection as the falsifying event.

The Numbers

Founded
2023–2024, San Francisco (sources conflict on year)
Founders
Rayan Krishnan (CEO, 25; ex-Palantir intern, Stanford AI) & Langston Nashold (CTO; left Stanford AI master’s together)
Funding
$40M Series A led by Andreessen Horowitz; raised ~Aug 2026, announced Sept 19, 2026
Investors
a16z (lead); 8VC & Bloomberg Beta (returning seed); HRT Ventures & Next Ladder Ventures (new); Pear VC cited by secondary press
Product
The Vals Index — confidential, in-house benchmarks (Finance Agent, Terminal-Bench, Vibe Code Bench, Legal Research/HLAB), weighted by each sector’s share of U.S. GDP
Business Model
Labs pay to be tested — “like a student might pay the College Board to take the SAT” (Krishnan). Scores then cited in labs’ own marketing
Valuation
~$400M per Dealroom / secondary press — not disclosed by the company; treat as reported, not confirmed
Use of Funds
Expand benchmark coverage, hire 10–15 staff, scale a new federal-agency evaluation program

The Issuer-Pays Loop

a16z pitches Vals as the “trust layer” — the Moody’s, the S&P, the UL of AI. But follow the money and credibility, and the same loop that broke those analogies in 2008 appears.

How the Money and Credibility Circulate

01

Lab Pays Vals

The frontier lab is the paying customer — it commissions the evaluation of its own model.

02

Vals Tests in Secret

Test materials stay confidential to prevent contamination — and to keep the score un-auditable by outsiders.

03

Score Is Issued

Vals produces a number. The buyer cannot inspect the test that produced it.

04

Lab Cites the Score

The lab reprints Vals’s number in its own model card — lending Vals credibility, and gaining leverage over it.

05

Buyer Trusts

Enterprises rely on the “independent” score — a two-way dependency between grader and graded.

a16z’s own analogy cuts against the deal. “Credit markets developed Moody’s and S&P… product manufacturers rely on UL” — but Moody’s and S&P are the canonical case of issuer-pays ratings failing precisely when the ratings mattered most. FourWeekMBA: this “puts it closer to an audit firm than to a rating agency… Whether that distinction survives as the company grows and its relationships with the labs deepen is something to watch.”

Auditable or Uncontaminated — Rarely Both

Vals’s confidentiality is its anti-contamination moat. It is also what makes its scores impossible for a downstream reader to verify. Trust shifts from “is this test well-designed?” to “do I trust this 25-person startup?” — and independent analysis (Pebblous) found private test sets did not systematically prevent saturation anyway; age and size predicted it better than secrecy. The moat Vals sells does not durably solve the problem it sells against.

Better-Funded on Both Flanks

The eval/benchmark space is crowded and, above Vals, better capitalized. Vals’s niche — confidential, third-party, professional-task benchmarking — is the narrowest and the most conflict-laden.

01

LMArena

Free public crowd-vote leaderboard — the de-facto standard Vals wants to displace. Raised a $150M Series A at ~$1.7B valuation. Anchors the free flank.

02

Scale AI — SEAL

Private/expert evals plus SWE-Bench Pro from a multi-billion-dollar data giant. Far deeper pockets and existing lab relationships.

03

Braintrust

Eval + observability infra for teams building on LLMs. $80M raise at ~$800M valuation (Feb 2026).

04

Arize & Patronus

Arize: LLM observability + eval, ~$70M Series C. Patronus: automated model/agent eval, $50M Series B. Both attack the enterprise-eval wedge.

05

Epoch AI

Independent nonprofit benchmarks (e.g. FrontierMath). Donor-funded, no issuer-pays conflict — a credibility contrast Vals must answer.

06

The Enterprise’s Own Evals

The recurring 2026 consensus: “no general benchmark replaces testing a model on your own data.” The strongest competitor is the buyer’s internal eval — free and trusted.

The squeeze: free public leaderboards on one side, in-house private evals on the other, and better-funded observability players in the middle. Vals is betting that a paid, secret, third-party grade is more trustworthy than any of them — the exact claim its business model undermines.

Weaknesses & Threat Vectors

Seven structural risks the $40M does not resolve.

High

Issuer-Pays Conflict of Interest

Labs pay for the test, then market the score. Vals has already disclosed a customer relationship with benchmark participants. No public firewall (pay-to-test-but-not-for-outcome governance) has been described. This is the Moody’s-2008 failure mode, not the aspiration a16z cites.

High

Confidentiality Destroys Verifiability

Secret tests mean no third party can audit whether a Vals score is fair. Trust is purely reputational — and one credible, public dispute could be existential for a self-styled “gold standard.”

High

Dependency on the Labs It Grades

Revenue, credibility (via model-card citations), and even an investor (HRT, where a co-founder worked) overlap with the graded ecosystem. A “neutral” grader whose customers are the graded parties is capturable.

Medium

Benchmarks Are Consumables

Every model release erodes prior tests; private test sets don’t durably prevent saturation. A treadmill cost structure that pressures margins and undercuts a ~$400M price on a company that has tripled its team but disclosed no revenue base.

Medium

Unverified Traction

“8× revenue” comes from a company X post with no baseline (8× of a tiny number is trivial); the $400M valuation is unconfirmed; customer-doubling is self-reported. The headline story underpinning the raise is not independently sourced.

Medium

Weak Defensibility vs. Free & Internal

LMArena (free, larger, better-funded) anchors the public leaderboard; enterprises increasingly run their own evals. Vals is squeezed from both sides with reputation as its only moat.

Medium

Frontier-Risk Scope Creep & a Data-Policy Gap

Benchmarking biosecurity, cybersecurity, “law of armed conflict,” and a federal-agency program raises liability, data-handling, and politicization risks a ~25-person startup is thinly resourced to govern — and Vals publishes no visible privacy or data-handling policy for the confidential customer prompts and eval data it processes.

Assessment Matrix

Market Opportunity
High
AI eval/trust infrastructure is real and growing; buyers genuinely lack a trustworthy way to compare models on real work
Product Defensibility
Low-Medium
Credible in-house benchmarks, but they are consumables — easily replicated and not durably protected by secrecy
Model Integrity
Low
Issuer-pays economics + disclosed customer relationships + an ecosystem investor is the weakest link in an independence pitch
Competitive Moat
Low-Medium
Reputation is the only moat and it is fragile; better-funded rivals sit on the free-public and enterprise-private flanks
Team
Medium-High
Young but technically strong (Stanford AI, Palantir/Microsoft); already shipping benchmarks cited by top labs; unproven at governance/scale
Traction Quality
Unverified
8→25 headcount is confirmed; the 8× revenue and $400M valuation are single-source and not independently verified
Investor Signal
High
a16z lead (Jennifer Li, Yoko Li, Raghu Raghuram, Shangda Xu); Bloomberg Beta and 8VC returning
Overall
Medium
A real need and a capable team, priced richly on unverified growth, built on a model whose central “independence” claim fights how it earns

Vals is selling trust in a market that badly needs it — using a business model that quietly erodes the very thing it sells. The gap is real and the founders are sharp. But before any buyer treats a Vals score as gospel, the diligence question is simple: who paid for the test, and can anyone else check it? On today’s evidence, the answer is the lab — and no.

Research Sources

Based entirely on publicly available information, including the TechCrunch announcement of September 19, 2026. Company-supplied figures (valuation, revenue growth, customer counts) are labeled unverified and are not presented as fact.

  1. TechCrunch — “Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking” (September 19, 2026) — primary announcement: round, headcount, founder, business-model quote.
  2. Andreessen Horowitz — “Investing in Vals” — lead investor thesis and the Moody’s / S&P / UL / auditor analogy; deal partners named.
  3. Vals AI website (vals.ai) — positioning, “independent evaluation, unbiased benchmarks,” the Vals Index, providers evaluated; no visible privacy/data-handling policy.
  4. Vals Legal AI Report & github.com/vals-ai/legal-research-bench — an open-sourced benchmark and Vals’s own disclosure of a customer relationship with benchmark participants.
  5. FourWeekMBA — critical analysis: auditable-vs-uncontaminated, issuer-pays vs. rating agency, HRT investor conflict, “no valuation disclosed.”
  6. Pebblous (eval-as-infrastructure) — saturation data, benchmarks-as-consumables, lab-dependency, “$400M = infrastructure pricing.”
  7. Dealroom, CryptoBriefing, TechFundingNews, The AI Insider — secondary reports of the ~$400M valuation and expanded investor list (HRT Ventures, Next Ladder, Pear VC).
  8. Bloomberg (April 2024) — early profile corroborating founding-era timeline (sources conflict on 2023 vs. 2024).
  9. Competitor funding: SiliconANGLE (LMArena $150M / ~$1.7B; Patronus $50M), Braintrust ($80M / ~$800M), Arize ($70M Series C), Scale AI / SEAL, Epoch AI (nonprofit funding).
  10. Artificial Lawyer / LawSites — independence and conflict critique in legal-AI benchmarking.