← All posts

Introducing VCBench, the First Benchmark for Venture Capital

VCBench is the first benchmark for venture capital. It tests how well AI predicts founder success, in a domain where signals are sparse, outcomes take years, and even top investors are usually wrong. It was built by the University of Oxford and Vela Research, the research arm of Vela Partners, an AI-native quant venture capital firm. VCBench gives venture capital predictions a shared, reproducible yardstick for the first time.

Why venture capital

Benchmarks such as SWE-bench and ARC-AGI showed how a shared dataset speeds up progress in AI. Nothing comparable existed for decisions under uncertainty in economic settings. Venture capital is a demanding test. The information is incomplete, the stakes are high, and at the earliest stage only about 1.9% of founders reach a major outcome. Y Combinator picks winners at 3.2% precision and tier-1 venture firms at 5.6%.

The dataset

VCBench pairs each of 9,000 founders with the most recent company they founded, mostly US companies founded between 2010 and 2018. A founder counts as successful if that company was acquired or went public above a $500M valuation, or raised more than $500M. An unsuccessful founder raised between $100K and $4M at inception and saw no exit, IPO or major follow-on round within eight years. 810 founders (9%) are successful.

Profiles come from licensed and public sources, including LinkedIn for education and work history and Crunchbase for funding and outcomes. Only information available before the company was founded is used. Each profile is available as prose for language models and as structured JSON for machine-learning models.

Making it hard to cheat

Large language models have read much of the internet, so a model might simply recognize a famous founder. VCBench removes names, companies, locations and dates, groups exits, clusters industries into 61 categories, replaces universities with QS rankings and turns job dates into durations. Profiles then go through repeated adversarial tests, and any founder a model identifies more than once is removed.

In testing on 300 successful founders, re-identification fell from 17.2% to 1.3% for an offline model and from 77.0% to 15.1% for a model with web search.

How models are scored

The data is split into six folds that each keep the 9% success rate. Half is public for anyone to build with. The other half is private, and leaderboard scores are computed on it. The headline metric is F0.5, which counts precision twice as heavily as recall, because false positives cost an investor more than missed opportunities.

Results at launch

These are the results published with the paper, averaged over six folds.

ModelPrecisionRecallF0.5
GPT-4o29.116.225.1
DeepSeek-R137.68.422.1
GPT-4o-mini29.510.121.2
o342.47.020.9
Gemini-2.5-Pro17.259.020.1
Claude-3.5-Haiku16.948.619.4
GPT-553.74.316.2
Gemini-2.5-Flash12.669.115.1
DeepSeek-V359.13.011.8

GPT-4o reached the highest F0.5, and its 29% precision was 3.2× the precision baseline, ahead of tier-1 VCs at 2.9×. DeepSeek-V3 reached over six times the baseline precision but rarely backed anyone. The authors caution against extrapolating directly from the 9% dataset to the real-world 1.9% success rate.

What came next

VCBench was designed as a living benchmark. Since launch, more than 60 research teams from universities, funds and independent labs have requested the dataset, and new entries have joined the leaderboard, including reasoning methods such as Policy Induction, Random Rule Forest and Reasoned Rule Mining.

Think-Reason-Learn Ensemble on the VCBench leaderboard today
Rank
#1 of 31
F0.5
37.9
Precision
40.6%
Recall
30.1%
Cost / 1k founders
$0.87

Scored on the 4,500-founder private test set, mean over three folds. That is 3.5× the F0.5 of tier-1 VCs. It is the current leader, as of September 2026. See the full leaderboard.

See the live leaderboard, or read how the board now weighs accuracy against cost.

Read the paper

Frequently asked questions

What is VCBench?
VCBench is the first benchmark for venture capital. It was built by the University of Oxford and Vela Research and holds 9,000 anonymized founder profiles, 9% of them successful.
How does VCBench define a successful founder?
A founder is successful if their most recent company was acquired or went public above a $500M valuation, or raised more than $500M.
Which model led VCBench at launch?
At launch in September 2025, GPT-4o had the highest F0.5 at 25.1, with 29.1% precision. The current leader on vcbench.com is the Think-Reason-Learn Ensemble.