V
VCBench
LeaderboardHow it worksModelsPostsResearch using VCBench
Repositories
Think-Reason-Learn
Papers
VCBenchGPTreePolicy InductionRandom Rule ForestReasoned Rule Mining
V
VCBench
← All models

Policy Induction on VCBench

Vela + OxfordReasoning
Board as of 2026-10-07

Policy Induction scores F0.5 33.0 on VCBench, the venture capital benchmark from the University of Oxford and Vela Research, with 34.9% precision and 27.2% recall on the 4,500-founder private test set. That is rank 4 of 32, 3.1× the F0.5 of tier-1 VCs and 3.8× Y Combinator, at $0.30 per 1,000 founders scored.

Policy Induction on the VCBench leaderboard today
Rank
#4 of 32
F0.5
33.0
Precision
34.9%
Recall
27.2%
Cost / 1k founders
$0.30

Scored on the 4,500-founder private test set, mean over three folds. That is 3.1× the F0.5 of tier-1 VCs. See the full leaderboard.

Compared with the reference rows

EntryPrecisionRecallF0.5Cost / 1k
Policy Induction34.9%27.2%33.0$0.30
Think-Reason-Learn Ensemble40.6%30.1%37.9$0.87
Tier-1 VCs23.0%5.2%10.7n/a
Y Combinator14.0%6.9%8.6n/a
Random Classifier9.0%9.0%9.0$0

Human rows are normalized to the dataset's 9% success rate, so every entry is compared on the same base rate. Costs use list prices on 2026-09-24.

How it was scored

An LLM reads each founder profile at scoring time and returns a prediction. Policy Induction used about 7,255 input and 1,227 output tokens per founder on Typesafe System One (jev-1.13.0), measured from the API usage fields. Cost is tokens times the provider's list price on 2026-09-24 ($0.042 input and $0 output per million tokens), so it recomputes when prices change. Build costs such as question or policy generation are excluded.

Read more

  • Policy Induction: Teaching an LLM an Investment Policy in Plain English: Policy Induction learns a readable investment policy by in-context learning and uses it to predict startup success, with 20x random precision in the paper.

Frequently asked questions

What does Policy Induction score on VCBench?
Policy Induction scores F0.5 33.0 on VCBench, with 34.9% precision and 27.2% recall, rank 4 of 32 as of 2026-10-07.
Does Policy Induction beat human investors at predicting founder success?
Yes. Its F0.5 is 3.1 times that of tier-1 VCs (10.7) and 3.8 times Y Combinator (8.6), after both are normalized to the dataset's 9% base rate.
How much does Policy Induction cost per 1,000 founders?
About $0.30 per 1,000 founders at list prices on 2026-09-24. An LLM reads each founder profile at scoring time and returns a prediction. Policy Induction used about 7,255 input and 1,227 output tokens per founder on Typesafe System One (jev-1.13.0), measured from the API usage fields. Cost is tokens times the provider's list price on 2026-09-24 ($0.042 input and $0 output per million tokens), so it recomputes when prices change. Build costs such as question or policy generation are excluded.

About VCBench

VCBench is the first benchmark for venture capital. It tests how well AI models, AI-native venture capital methods and human investors predict which startup founders will succeed, on 9,000 anonymized founder profiles. It was built by the University of Oxford and Vela Research, the research arm of Vela Partners, an AI-native quant venture capital firm in San Francisco. Many methods on the leaderboard are open source in Think-Reason-Learn.