V
VCBench
LeaderboardHow it worksModelsPostsResearch using VCBench
Repositories
Think-Reason-Learn
Papers
VCBenchGPTreePolicy InductionRandom Rule ForestReasoned Rule Mining
V
VCBench
← All models

Gemini 2.5 Flash on VCBench

GoogleReasoning
Board as of 2026-10-07

Gemini 2.5 Flash scores F0.5 14.9 on VCBench, the venture capital benchmark from the University of Oxford and Vela Research, with 12.5% precision and 68.4% recall on the 4,500-founder private test set. That is rank 28 of 32, 1.4× the F0.5 of tier-1 VCs and 1.7× Y Combinator, at $2.87 per 1,000 founders scored.

Gemini 2.5 Flash on the VCBench leaderboard today
Rank
#28 of 32
F0.5
14.9
Precision
12.5%
Recall
68.4%
Cost / 1k founders
$2.87

Scored on the 4,500-founder private test set, mean over three folds. That is 1.4× the F0.5 of tier-1 VCs. See the full leaderboard.

Compared with the reference rows

EntryPrecisionRecallF0.5Cost / 1k
Gemini 2.5 Flash12.5%68.4%14.9$2.87
Think-Reason-Learn Ensemble40.6%30.1%37.9$0.87
Tier-1 VCs23.0%5.2%10.7n/a
Y Combinator14.0%6.9%8.6n/a
Random Classifier9.0%9.0%9.0$0

Human rows are normalized to the dataset's 9% success rate, so every entry is compared on the same base rate. Costs use list prices on 2026-09-24.

How it was scored

An LLM reads each founder profile at scoring time and returns a prediction. Gemini 2.5 Flash used about 284 input and 1,113 output tokens per founder on gemini-2.5-flash, estimated from the prompt and the saved responses; reasoning models are assumed to think about 1,000 tokens per founder. Cost is tokens times the provider's list price on 2026-09-24 ($0.3 input and $2.5 output per million tokens), so it recomputes when prices change. Build costs such as question or policy generation are excluded.

Read more

  • VCBench Leaderboard Update: Accuracy vs Cost, Plus GPT-6 and Claude Opus 5.5: VCBench now plots F0.5 against the cost of scoring 1,000 founders, with GPT-6, Claude Opus 5.5, new submissions and the Think-Reason-Learn Ensemble.

Frequently asked questions

What does Gemini 2.5 Flash score on VCBench?
Gemini 2.5 Flash scores F0.5 14.9 on VCBench, with 12.5% precision and 68.4% recall, rank 28 of 32 as of 2026-10-07.
Does Gemini 2.5 Flash beat human investors at predicting founder success?
Yes. Its F0.5 is 1.4 times that of tier-1 VCs (10.7) and 1.7 times Y Combinator (8.6), after both are normalized to the dataset's 9% base rate.
How much does Gemini 2.5 Flash cost per 1,000 founders?
About $2.87 per 1,000 founders at list prices on 2026-09-24. An LLM reads each founder profile at scoring time and returns a prediction. Gemini 2.5 Flash used about 284 input and 1,113 output tokens per founder on gemini-2.5-flash, estimated from the prompt and the saved responses; reasoning models are assumed to think about 1,000 tokens per founder. Cost is tokens times the provider's list price on 2026-09-24 ($0.3 input and $2.5 output per million tokens), so it recomputes when prices change. Build costs such as question or policy generation are excluded.

About VCBench

VCBench is the first benchmark for venture capital. It tests how well AI models, AI-native venture capital methods and human investors predict which startup founders will succeed, on 9,000 anonymized founder profiles. It was built by the University of Oxford and Vela Research, the research arm of Vela Partners, an AI-native quant venture capital firm in San Francisco. Many methods on the leaderboard are open source in Think-Reason-Learn.