V
VCBench
LeaderboardHow it worksModelsPostsResearch using VCBench
Repositories
Think-Reason-Learn
Papers
VCBenchGPTreePolicy InductionRandom Rule ForestReasoned Rule Mining
V
VCBench
← All models

Think-Reason-Learn Ensemble on VCBench

Vela + OxfordHybrid
Board as of 2026-10-07

Think-Reason-Learn Ensemble scores F0.5 37.9 on VCBench, the venture capital benchmark from the University of Oxford and Vela Research, with 40.6% precision and 30.1% recall on the 4,500-founder private test set. That is rank 1 of 32, 3.5× the F0.5 of tier-1 VCs and 4.4× Y Combinator, at $0.87 per 1,000 founders scored.

Think-Reason-Learn Ensemble on the VCBench leaderboard today
Rank
#1 of 32
F0.5
37.9
Precision
40.6%
Recall
30.1%
Cost / 1k founders
$0.87

Scored on the 4,500-founder private test set, mean over three folds. That is 3.5× the F0.5 of tier-1 VCs. See the full leaderboard.

Compared with the reference rows

EntryPrecisionRecallF0.5Cost / 1k
Think-Reason-Learn Ensemble40.6%30.1%37.9$0.87
Tier-1 VCs23.0%5.2%10.7n/a
Y Combinator14.0%6.9%8.6n/a
Random Classifier9.0%9.0%9.0$0

Human rows are normalized to the dataset's 9% success rate, so every entry is compared on the same base rate. Costs use list prices on 2026-09-24.

How it was scored

A mix of families: LLM-written rules and classical models are combined into one score, so an LLM still runs at scoring time. Think-Reason-Learn Ensemble used about 20,800 input and 7,448 output tokens per founder on Typesafe System One (jev-1.13.0), measured from the API usage fields. Cost is tokens times the provider's list price on 2026-09-24 ($0.042 input and $0 output per million tokens), so it recomputes when prices change. Build costs such as question or policy generation are excluded.

Read more

  • VCBench Leaderboard Update: Accuracy vs Cost, Plus GPT-6 and Claude Opus 5.5: VCBench now plots F0.5 against the cost of scoring 1,000 founders, with GPT-6, Claude Opus 5.5, new submissions and the Think-Reason-Learn Ensemble.

Frequently asked questions

What does Think-Reason-Learn Ensemble score on VCBench?
Think-Reason-Learn Ensemble scores F0.5 37.9 on VCBench, with 40.6% precision and 30.1% recall, rank 1 of 32 as of 2026-10-07.
Does Think-Reason-Learn Ensemble beat human investors at predicting founder success?
Yes. Its F0.5 is 3.5 times that of tier-1 VCs (10.7) and 4.4 times Y Combinator (8.6), after both are normalized to the dataset's 9% base rate.
How much does Think-Reason-Learn Ensemble cost per 1,000 founders?
About $0.87 per 1,000 founders at list prices on 2026-09-24. A mix of families: LLM-written rules and classical models are combined into one score, so an LLM still runs at scoring time. Think-Reason-Learn Ensemble used about 20,800 input and 7,448 output tokens per founder on Typesafe System One (jev-1.13.0), measured from the API usage fields. Cost is tokens times the provider's list price on 2026-09-24 ($0.042 input and $0 output per million tokens), so it recomputes when prices change. Build costs such as question or policy generation are excluded.

About VCBench

VCBench is the first benchmark for venture capital. It tests how well AI models, AI-native venture capital methods and human investors predict which startup founders will succeed, on 9,000 anonymized founder profiles. It was built by the University of Oxford and Vela Research, the research arm of Vela Partners, an AI-native quant venture capital firm in San Francisco. Many methods on the leaderboard are open source in Think-Reason-Learn.