Notes from the benchmark.

Leaderboard updates, method write-ups and what the results say about AI in venture capital.

VCBench Leaderboard Update: Accuracy vs Cost, Plus GPT-6 and Claude Opus 5.5

VCBench now plots F0.5 against the cost of scoring 1,000 founders, with GPT-6, Claude Opus 5.5, new submissions and the Think-Reason-Learn Ensemble.

Read post →

Learning What to Ask: Cost-Aware Founder Evaluation with Reinforcement Learning

A reinforcement-learning policy decides which founder attribute to check next and when to stop, reaching F0.5 37.1 on VCBench with 54% of the information.

Read post →

Reasoned Rule Mining: Calibrated LLM Predictions of Startup Success

Reasoned Rule Mining calibrates LLM log-probabilities and mines readable rules to predict founder success, 12x the market index in the paper.

Read post →

Introducing VCBench, the First Benchmark for Venture Capital

VCBench, the first benchmark for venture capital, tests LLMs and human investors on predicting founder success across 9,000 anonymized profiles.

Read post →

Random Rule Forest: Predicting Startup Success with LLM-Written Yes-or-No Questions

Random Rule Forest turns LLM-generated yes-or-no questions into a transparent scorecard for founder success, 5.6x random precision in the paper.

Read post →

Policy Induction: Teaching an LLM an Investment Policy in Plain English

Policy Induction learns a readable investment policy by in-context learning and uses it to predict startup success, with 20x random precision in the paper.

Read post →

GPTree: Rethinking Machine Learning with LLM-Powered Decision Trees

How prompt chaining inspired GPTree, an explainable LLM-powered decision tree that beat tier-1 VCs at picking unicorn founders and began Think-Reason-Learn.

Read post →