V
VCBench
LeaderboardHow it worksModelsPostsResearch using VCBench
Repositories
Think-Reason-Learn
Papers
VCBenchGPTreePolicy InductionRandom Rule ForestReasoned Rule Mining
V
VCBench
← All models

Claude Opus 5.5 on VCBench

AnthropicReasoning
Board as of 2026-10-07

Claude Opus 5.5 scores F0.5 24.7 on VCBench, the venture capital benchmark from the University of Oxford and Vela Research, with 77.7% precision and 6.7% recall on the 4,500-founder private test set. That is rank 17 of 32, 2.3× the F0.5 of tier-1 VCs and 2.9× Y Combinator, at $5.33 per 1,000 founders scored.

Claude Opus 5.5 on the VCBench leaderboard today
Rank
#17 of 32
F0.5
24.7
Precision
77.7%
Recall
6.7%
Cost / 1k founders
$5.33

Scored on the 4,500-founder private test set, mean over three folds. That is 2.3× the F0.5 of tier-1 VCs. See the full leaderboard.

Compared with the reference rows

EntryPrecisionRecallF0.5Cost / 1k
Claude Opus 5.577.7%6.7%24.7$5.33
Think-Reason-Learn Ensemble40.6%30.1%37.9$0.87
Tier-1 VCs23.0%5.2%10.7n/a
Y Combinator14.0%6.9%8.6n/a
Random Classifier9.0%9.0%9.0$0

Human rows are normalized to the dataset's 9% success rate, so every entry is compared on the same base rate. Costs use list prices on 2026-09-24.

How it was scored

An LLM reads each founder profile at scoring time and returns a prediction. Claude Opus 5.5 used about 452 input and 176 output tokens per founder on claude-opus-5-5, measured from the API usage fields. Cost is tokens times the provider's list price on 2026-09-24 ($4 input and $20 output per million tokens), so it recomputes when prices change. Build costs such as question or policy generation are excluded.

Read more

  • VCBench Leaderboard Update: Accuracy vs Cost, Plus GPT-6 and Claude Opus 5.5: VCBench now plots F0.5 against the cost of scoring 1,000 founders, with GPT-6, Claude Opus 5.5, new submissions and the Think-Reason-Learn Ensemble.

Frequently asked questions

What does Claude Opus 5.5 score on VCBench?
Claude Opus 5.5 scores F0.5 24.7 on VCBench, with 77.7% precision and 6.7% recall, rank 17 of 32 as of 2026-10-07.
Does Claude Opus 5.5 beat human investors at predicting founder success?
Yes. Its F0.5 is 2.3 times that of tier-1 VCs (10.7) and 2.9 times Y Combinator (8.6), after both are normalized to the dataset's 9% base rate.
How much does Claude Opus 5.5 cost per 1,000 founders?
About $5.33 per 1,000 founders at list prices on 2026-09-24. An LLM reads each founder profile at scoring time and returns a prediction. Claude Opus 5.5 used about 452 input and 176 output tokens per founder on claude-opus-5-5, measured from the API usage fields. Cost is tokens times the provider's list price on 2026-09-24 ($4 input and $20 output per million tokens), so it recomputes when prices change. Build costs such as question or policy generation are excluded.

About VCBench

VCBench is the first benchmark for venture capital. It tests how well AI models, AI-native venture capital methods and human investors predict which startup founders will succeed, on 9,000 anonymized founder profiles. It was built by the University of Oxford and Vela Research, the research arm of Vela Partners, an AI-native quant venture capital firm in San Francisco. Many methods on the leaderboard are open source in Think-Reason-Learn.