NewAccuracy vs cost, plus GPT-6 and Claude Opus 5.5

Can AI predict which founders will succeed?

VCBench is an academic benchmark from the University of Oxford and Vela Research. It scores AI models and top investors on the same real founder outcomes.

Submit a model
F0.5 on 4,500 held-out founders
Best humansTier-1 VCs
10.7
Best ML modelPvalyou
31.5
Best frontier LLMGPT-6 Sol
33.7
Best on VCBenchThink-Reason-Learn
37.9

The best system scores 3.5× tier-1 VCs. Human scores are normalized to the dataset's 9% success rate.

Leaderboard

Every entry is scored on the same 4,500 held-out founders and ranked by F0.5.

Who gets the best score for the money?

Up and to the left is better. Highlighted entries are the best value, because nothing cheaper scores higher.
↖ Better$0$0.10$0.30$1$3$10$30010203040Cost per 1,000 founders scored (USD, log)Score (F₀.₅)Best score 37.9Tier-1 VCs 10.7

Rankings

Sorted by F₀.₅. The mark in each bar is Tier-1 VCs at 10.7.
#
Model
1
Think-Reason-Learn EnsembleVela + Oxford · Hybrid · P 40.6 · R 30.1
37.9
40.6
30.1
$0.87
2
GPT-6 SolOpenAI · Reasoning · P 45.1 · R 17.0
33.7
45.1
17.0
$1.45
3
Policy InductionVela + Oxford · Reasoning · P 34.9 · R 27.2
33.0
34.9
27.2
$0.30
4
GemVC-v0Madhusudhana Naidu · Reasoning · P 39.4 · R 20.3
32.9
39.4
20.3
$2.87
5
Reasoned Rule MiningVela + Oxford · Reasoning · P 34.8 · R 24.9
32.2
34.8
24.9
$0.31
6
Pvalyou Founder Model (ML)Itay Attar · Pvalyou · Classical ML · P 38.2 · R 18.8
31.5
38.2
18.8
$0
7
Random Rule ForestVela + Oxford · Reasoning · P 37.9 · R 18.5
31.3
37.9
18.5
$0.13
8
Verifiable-RLVela + Oxford · Classical ML · P 29.1 · R 30.6
29.4
29.1
30.6
$0
9
Structured-Rule-StumpYagiz Ihlamur · Classical ML · P 32.8 · R 18.0
28.1
32.8
18.0
$0
10
verifiable-reasoningVela + Oxford · Classical ML · P 30.6 · R 21.0
27.7
30.6
21.0
$0
11
Claude Sonnet 5Anthropic · Reasoning · P 69.3 · R 8.1
27.7
69.3
8.1
$2.20
12
large-founder-model-v0Vela + Oxford · Classical ML · P 31.7 · R 17.5
27.2
31.7
17.5
$0
13
GPT-6 LunaOpenAI · Reasoning · P 33.3 · R 14.1
26.1
33.3
14.1
$0.11
14
GPT-4oOpenAI · Reasoning · P 30.0 · R 16.3
25.7
30.0
16.3
$1.55
15
FinGPT-VC2Columbia · Reasoning · P 24.4 · R 27.2
24.9
24.4
27.2
n/a
16
Claude Opus 5.5Anthropic · Reasoning · P 77.7 · R 6.7
24.7
77.7
6.7
$5.33
17
GPT-6 AstraOpenAI · Reasoning · P 56.8 · R 7.4
24.3
56.8
7.4
$7.49
18
GPT-4o-miniOpenAI · Reasoning · P 31.5 · R 11.1
23.0
31.5
11.1
$0.09
19
Claude Haiku 4.5Anthropic · Reasoning · P 21.2 · R 30.1
22.6
21.2
30.1
$0.97
20
FinGPT-VC1Columbia · Reasoning · P 21.8 · R 24.2
22.2
21.8
24.2
n/a
21
o3OpenAI · Reasoning · P 43.2 · R 7.4
21.5
43.2
7.4
$9.25
22
GPTreeVela + Oxford · Reasoning · P 19.4 · R 27.2
20.6
19.4
27.2
$0.12
23
Gemini 2.5 ProGoogle · Reasoning · P 17.1 · R 58.0
19.9
17.1
58.0
$11.51
24
DeepSeek-ReasonerDeepSeek · Reasoning · P 31.8 · R 6.9
18.4
31.8
6.9
$1.36
25
Claude 3.5 HaikuAnthropic · Reasoning · P 15.8 · R 46.4
18.2
15.8
46.4
$0.52
26
GPT-5OpenAI · Reasoning · P 59.1 · R 4.2
16.2
59.1
4.2
$11.23
27
Gemini 2.5 FlashGoogle · Reasoning · P 12.5 · R 68.4
14.9
12.5
68.4
$2.87
28
DeepSeek-ChatDeepSeek · Reasoning · P 80.6 · R 3.0
12.1
80.6
3.0
$0.17
29
Tier-1 VCsHumans · P 23.0 · R 5.2
10.7
23.0
5.2
n/a
30
Random ClassifierBaseline · P 9.0 · R 9.0
9.0
9.0
9.0
$0
31
Y CombinatorHumans · P 14.0 · R 6.9
8.6
14.0
6.9
n/a

How VCBench works

A model reads an anonymized founder profile (education, career, prior companies) and predicts whether the company will become a major success. Venture capital is a hard test for AI, because the information is incomplete, the outcomes take years, and even expert investors are usually wrong.

Success
The company exits or IPOs above a $500M valuation, or raises more than $500M.
Data
9,000 founder profiles from LinkedIn and Crunchbase, 9% of them successful. Profiles are standardized, enriched and anonymized, which cut re-identification by over 90% in adversarial tests.
Scoring
A held-back private test set of 4,500 founders, averaged over three sequential folds. Entries are ranked by F0.5, which weights precision over recall, because a bad bet costs an investor more than a missed one.

The method and dataset are described in the VCBench paper. To submit a model, report an error or join the benchmark committee, email benchmark@vela.partners.