Notes from the benchmark.
Leaderboard updates, method write-ups and what the results say about AI in venture capital.
VCBench Leaderboard Update: Accuracy vs Cost, Plus GPT-6 and Claude Opus 5.5
VCBench now plots F0.5 against the cost of scoring 1,000 founders, with GPT-6, Claude Opus 5.5, new submissions and the Think-Reason-Learn Ensemble.
Read post →Learning What to Ask: Cost-Aware Founder Evaluation with Reinforcement Learning
A reinforcement-learning policy decides which founder attribute to check next and when to stop, reaching F0.5 37.1 on VCBench with 54% of the information.
Read post →Reasoned Rule Mining: Calibrated LLM Predictions of Startup Success
Reasoned Rule Mining calibrates LLM log-probabilities and mines readable rules to predict founder success, 12x the market index in the paper.
Read post →Introducing VCBench, the First Benchmark for Venture Capital
VCBench, the first benchmark for venture capital, tests LLMs and human investors on predicting founder success across 9,000 anonymized profiles.
Read post →Random Rule Forest: Predicting Startup Success with LLM-Written Yes-or-No Questions
Random Rule Forest turns LLM-generated yes-or-no questions into a transparent scorecard for founder success, 5.6x random precision in the paper.
Read post →Policy Induction: Teaching an LLM an Investment Policy in Plain English
Policy Induction learns a readable investment policy by in-context learning and uses it to predict startup success, with 20x random precision in the paper.
Read post →GPTree: Rethinking Machine Learning with LLM-Powered Decision Trees
How prompt chaining inspired GPTree, an explainable LLM-powered decision tree that beat tier-1 VCs at picking unicorn founders and began Think-Reason-Learn.
Read post →