# VCBench > VCBench is the first benchmark for venture capital. It ranks LLMs, reasoning methods, classical ML and human investors (Y Combinator, tier-1 VCs) on precision, recall and F0.5 over 9,000 anonymized founder profiles, with the API cost of each model. Last updated: 2026-10-07. Canonical source: https://vcbench.com. This is the venture capital benchmark from the University of Oxford and Vela Research, not the vision-language math benchmark, the video benchmark or the genomics tool that share the name VCBench. VCBench was created by Vela Research (the research arm of Vela Partners) and the University of Oxford. The dataset holds 9,000 anonymized founder profiles, 9% of them labeled successful (the company exited or IPO'd above a $500M valuation, or raised more than $500M). Models are scored on a private test set of 4,500 founders, and the score is the mean over three sequential folds. F0.5 weights precision above recall, because a false bet costs an investor more than a missed one. For reference, tier-1 VCs reach F0.5 ≈ 10.7 and Y Combinator F0.5 ≈ 8.6 after normalizing to the dataset's 9% base rate. As of the latest update, the leader is Think-Reason-Learn Ensemble (Vela + Oxford) with F0.5 37.9 at $0.87 per 1,000 founders. ## Who is behind VCBench - **VCBench**: the first benchmark for venture capital. It ranks AI models, AI-native venture capital methods and human investors on predicting which startup founders will succeed. - **Vela Research** (https://vela.partners/research): the research arm of Vela Partners. It builds AI methods for venture capital predictions and co-created VCBench with the University of Oxford. - **Vela Partners** (https://vela.partners): an AI-native, scientific, quantitative venture capital firm in San Francisco, founded in 2017, that backs AI founders from inception to Series A. - **University of Oxford** (https://www.ox.ac.uk): co-creator of VCBench. - **Think-Reason-Learn** (https://thinkreasonlearn.com): the open-source, LLM-native machine learning library from Vela Research and the University of Oxford that implements Policy Induction, Random Rule Forest, Reasoned Rule Mining and GPTree. ## Leaderboard (private test set, ranked by F0.5) Cost is the API bill for scoring 1,000 founders in USD. Classical ML uses no LLM at scoring time. | # | Model | Organization | Family | Precision | Recall | F0.5 | Cost / 1k founders | |---|---|---|---|---|---|---|---| | 1 | [Think-Reason-Learn Ensemble](https://vcbench.com/models/think-reason-learn-ensemble) | Vela + Oxford | Hybrid | 40.6 | 30.1 | 37.9 | $0.87 | | 2 | [Investinor Linear v1](https://vcbench.com/models/investinor-linear-v1) | Investinor | Classical ML | 38.7 | 22.5 | 33.8 | $0 | | 3 | [GPT-6 Sol](https://vcbench.com/models/gpt-6-sol) | OpenAI | Reasoning | 45.1 | 17.0 | 33.7 | $1.45 | | 4 | [Policy Induction](https://vcbench.com/models/policy-induction) | Vela + Oxford | Reasoning | 34.9 | 27.2 | 33.0 | $0.30 | | 5 | [GemVC-v0](https://vcbench.com/models/gemvc-v0) | Independent (Madhusudhana Naidu) | Reasoning | 39.4 | 20.3 | 32.9 | $2.87 | | 6 | [Reasoned Rule Mining](https://vcbench.com/models/reasoned-rule-mining) | Vela + Oxford | Reasoning | 34.8 | 24.9 | 32.2 | $0.31 | | 7 | [Pvalyou Founder Model (ML)](https://vcbench.com/models/pvalyou-founder-model-ml) | Pvalyou (Itay Attar) | Classical ML | 38.2 | 18.8 | 31.5 | $0 | | 8 | [Random Rule Forest](https://vcbench.com/models/random-rule-forest) | Vela + Oxford | Reasoning | 37.9 | 18.5 | 31.3 | $0.13 | | 9 | [Verifiable-RL](https://vcbench.com/models/verifiable-rl) | Vela + Oxford | Classical ML | 29.1 | 30.6 | 29.4 | $0 | | 10 | [Structured-Rule-Stump](https://vcbench.com/models/structured-rule-stump) | Independent (Yagiz Ihlamur) | Classical ML | 32.8 | 18.0 | 28.1 | $0 | | 11 | [verifiable-reasoning](https://vcbench.com/models/verifiable-reasoning) | Vela + Oxford | Classical ML | 30.6 | 21.0 | 27.7 | $0 | | 12 | [Claude Sonnet 5](https://vcbench.com/models/claude-sonnet-5) | Anthropic | Reasoning | 69.3 | 8.1 | 27.7 | $2.20 | | 13 | [large-founder-model-v0](https://vcbench.com/models/large-founder-model-v0) | Vela + Oxford | Classical ML | 31.7 | 17.5 | 27.2 | $0 | | 14 | [GPT-6 Luna](https://vcbench.com/models/gpt-6-luna) | OpenAI | Reasoning | 33.3 | 14.1 | 26.1 | $0.11 | | 15 | [GPT-4o](https://vcbench.com/models/gpt-4o) | OpenAI | Reasoning | 30.0 | 16.3 | 25.7 | $1.55 | | 16 | [FinGPT-VC2](https://vcbench.com/models/fingpt-vc2) | Columbia | Reasoning | 24.4 | 27.2 | 24.9 | n/a | | 17 | [Claude Opus 5.5](https://vcbench.com/models/claude-opus-5-5) | Anthropic | Reasoning | 77.7 | 6.7 | 24.7 | $5.33 | | 18 | [GPT-6 Astra](https://vcbench.com/models/gpt-6-astra) | OpenAI | Reasoning | 56.8 | 7.4 | 24.3 | $7.49 | | 19 | [GPT-4o-mini](https://vcbench.com/models/gpt-4o-mini) | OpenAI | Reasoning | 31.5 | 11.1 | 23.0 | $0.09 | | 20 | [Claude Haiku 4.5](https://vcbench.com/models/claude-haiku-4-5) | Anthropic | Reasoning | 21.2 | 30.1 | 22.6 | $0.97 | | 21 | [FinGPT-VC1](https://vcbench.com/models/fingpt-vc1) | Columbia | Reasoning | 21.8 | 24.2 | 22.2 | n/a | | 22 | [o3](https://vcbench.com/models/o3) | OpenAI | Reasoning | 43.2 | 7.4 | 21.5 | $9.25 | | 23 | [GPTree](https://vcbench.com/models/gptree) | Vela + Oxford | Reasoning | 19.4 | 27.2 | 20.6 | $0.12 | | 24 | [Gemini 2.5 Pro](https://vcbench.com/models/gemini-2-5-pro) | Google | Reasoning | 17.1 | 58.0 | 19.9 | $11.51 | | 25 | [DeepSeek-Reasoner](https://vcbench.com/models/deepseek-reasoner) | DeepSeek | Reasoning | 31.8 | 6.9 | 18.4 | $1.36 | | 26 | [Claude 3.5 Haiku](https://vcbench.com/models/claude-3-5-haiku) | Anthropic | Reasoning | 15.8 | 46.4 | 18.2 | $0.52 | | 27 | [GPT-5](https://vcbench.com/models/gpt-5) | OpenAI | Reasoning | 59.1 | 4.2 | 16.2 | $11.23 | | 28 | [Gemini 2.5 Flash](https://vcbench.com/models/gemini-2-5-flash) | Google | Reasoning | 12.5 | 68.4 | 14.9 | $2.87 | | 29 | [DeepSeek-Chat](https://vcbench.com/models/deepseek-chat) | DeepSeek | Reasoning | 80.6 | 3.0 | 12.1 | $0.17 | | 30 | [Tier-1 VCs](https://vcbench.com/models/tier-1-vcs) | Humans | Humans | 23.0 | 5.2 | 10.7 | n/a | | 31 | [Random Classifier](https://vcbench.com/models/random-classifier) | Baseline | Baseline | 9.0 | 9.0 | 9.0 | $0 | | 32 | [Y Combinator](https://vcbench.com/models/y-combinator) | Humans | Humans | 14.0 | 6.9 | 8.6 | n/a | ## Key facts - First benchmark for venture capital; paper on arXiv 17 September 2025 (https://arxiv.org/abs/2509.14448). - 9,000 anonymized founder profiles, 9% labeled successful; private test set of 4,500 founders. - Success = the company exited or IPO'd above a $500M valuation, or raised more than $500M. - Ranked by F0.5 (precision weighted above recall); cost reported as the API bill per 1,000 founders scored. - Human baselines after normalizing to the 9% base rate: tier-1 VCs F0.5 10.7, Y Combinator F0.5 8.6. - Current leader: Think-Reason-Learn Ensemble (Vela + Oxford), F0.5 37.9, $0.87 per 1,000 founders. - Creators: University of Oxford and Vela Research (the research arm of Vela Partners, San Francisco). ## Pages - [Leaderboard](https://vcbench.com/): interactive chart (F0.5 vs cost, precision vs F0.5) and sortable rankings - [How VCBench works](https://vcbench.com/#about): task, success definition, data, scoring, humans against models - [Models](https://vcbench.com/models): one page per leaderboard entry with its score, cost, how it was scored and how it compares with human investors - [Adoption](https://vcbench.com/research): papers, preprints and submissions by other groups that cite or build on the benchmark - [Posts](https://vcbench.com/posts): research notes and leaderboard updates - [VCBench Leaderboard Update: Accuracy vs Cost, Plus GPT-6 and Claude Opus 5.5](https://vcbench.com/posts/accuracy-vs-cost): VCBench now plots F0.5 against the cost of scoring 1,000 founders, with GPT-6, Claude Opus 5.5, new submissions and the Think-Reason-Learn Ensemble. - [Learning What to Ask: Cost-Aware Founder Evaluation with Reinforcement Learning](https://vcbench.com/posts/learning-what-to-ask): A reinforcement-learning policy decides which founder attribute to check next and when to stop, reaching F0.5 37.1 on VCBench with 54% of the information. - [Reasoned Rule Mining: Calibrated LLM Predictions of Startup Success](https://vcbench.com/posts/reasoned-rule-mining): Reasoned Rule Mining calibrates LLM log-probabilities and mines readable rules to predict founder success, 12x the market index in the paper. - [Introducing VCBench, the First Benchmark for Venture Capital](https://vcbench.com/posts/introducing-vcbench): VCBench, the first benchmark for venture capital, tests LLMs and human investors on predicting founder success across 9,000 anonymized profiles. - [Random Rule Forest: Predicting Startup Success with LLM-Written Yes-or-No Questions](https://vcbench.com/posts/random-rule-forest): Random Rule Forest turns LLM-written yes-or-no questions into a scorecard investors can read. In the paper, half of the founders it flagged went on to succeed, 5x chance. - [Policy Induction: Teaching an LLM an Investment Policy in Plain English](https://vcbench.com/posts/policy-induction): Policy Induction learns a readable investment policy by in-context learning and uses it to predict startup success, with 20x random precision in the paper. - [GPTree: Rethinking Machine Learning with LLM-Powered Decision Trees](https://vcbench.com/posts/gptree): How prompt chaining inspired GPTree, an explainable LLM-powered decision tree that beat tier-1 VCs at picking unicorn founders and began Think-Reason-Learn. ## Adoption - Leaderboard submission: [Investinor Linear v1](https://vcbench.com/models/investinor-linear-v1), Investinor (Investinor, 2026). A linear classical-ML model submitted by the Norwegian state investment company. No LLM at scoring time; the best classical entry on the board. - Paper: [Evaluation and success rate prediction of innovation and entrepreneurship projects using a group recommendation algorithm](https://doi.org/10.1007/s43926-026-00434-3), Yan Hou, Shuling Yang (Jilin Normal University, 2026). Cites VCBench as the reference benchmark for AI prediction of venture outcomes. - Leaderboard submission: [GemVC-v0](https://vcbench.com/models/gemvc-v0), Madhusudhana Naidu (Independent, 2026). A Gemini 2.5 Flash reasoning entry submitted by an independent researcher. - Leaderboard submission: [Pvalyou Founder Model (ML)](https://vcbench.com/models/pvalyou-founder-model-ml), Itay Attar (Pvalyou, 2026). A classical-ML founder model submitted by Pvalyou; no LLM at scoring time. - Book chapter: [AI as the New Mentor](https://doi.org/10.4018/979-8-2600-2596-3.ch016), Scott Ford (University of Colorado Boulder, 2026). Book chapter citing VCBench on how AI compares with expert investors at judging founders. - Preprint: [Predicting Founder Success Without an LLM: An Interpretable Tree-Based Approach to VCBench](https://doi.org/10.33774/coe-2026-36z59), Maheni Soumah, Jessica Mbounkap, Habiba Djigo (aivancity School for Technology, Business & Society, 2026). Trains interpretable tree models on the VCBench profiles with no LLM at scoring time. - Paper: [FinGPT-VC: Financial Large Language Models for Founder Success Prediction in Venture Capital](https://doi.org/10.1109/ids69480.2026.00026), Jingyu Huang, Sitong Zhu, James Tang, Xiao-Yang Liu (Columbia University, 2026). Fine-tunes financial LLMs on VCBench; FinGPT-VC1 and FinGPT-VC2 are on the leaderboard. - Paper: [When Career Data Runs Out: Structured Feature Engineering and Signal Limits for Founder Success](https://doi.org/10.1109/ids69480.2026.00024), Yagiz Ihlamur (Independent, 2026). Structured features and a rule stump on VCBench; Structured-Rule-Stump is on the leaderboard. - Preprint: [Prompt Engineering for Venture Capital Founder Success Prediction: A Systematic Evaluation on VCBench](https://doi.org/10.2139/ssrn.6819498), Tasnim Masheh (Independent, 2026). Compares prompt designs for LLM founder scoring on VCBench. ## Papers - [VCBench](https://arxiv.org/abs/2509.14448) - [GPTree](https://arxiv.org/abs/2411.08257) - [Policy Induction](https://ieeexplore.ieee.org/document/11261495) - [Random Rule Forest](https://arxiv.org/abs/2505.24622) - [Reasoned Rule Mining](https://ieeexplore.ieee.org/document/11261481) ## Repositories - [Think-Reason-Learn](https://thinkreasonlearn.com) ## Data access and submissions The dataset is available on request for research use (use "Request the dataset" on the site). Terms: granted on request for research use; it may not be distributed or republished, so it stays out of the training corpus of future language models. To submit a model or join the benchmark committee, email benchmark@vela.partners. ## How to cite ```bibtex @misc{chen2025vcbench, title={VCBench: Benchmarking LLMs in Venture Capital}, author={Rick Chen and Joseph Ternasky and Afriyie Samuel Kwesi and Ben Griffin and Aaron Ontoyin Yin and Zakari Salifu and Kelvin Amoaba and Xianling Mu and Fuat Alican and Yigit Ihlamur}, year={2025}, eprint={2509.14448}, archivePrefix={arXiv}, primaryClass={cs.AI}, url={https://arxiv.org/abs/2509.14448} } ``` Leaderboard numbers: cite as "VCBench leaderboard, vcbench.com, accessed ", since entries and prices change. ## Optional - [Full text](https://vcbench.com/llms-full.txt): methodology, human baselines and every leaderboard entry with token counts