← All posts

VCBench Leaderboard Update: Accuracy vs Cost, Plus GPT-6 and Claude Opus 5.5

The leaderboard now plots each entry's F0.5 against the cost of scoring 1,000 founders, with new results for six OpenAI and Anthropic models.

See the new leaderboard

Today we're updating the VCBench leaderboard. Alongside accuracy, every entry is now plotted against what it costs to score 1,000 founders. We've also added results for six new models from OpenAI and Anthropic, including GPT-6 and Claude Opus 5.5, new submitted scores, and a Think-Reason-Learn Ensemble.

VCBench is the first benchmark for venture capital, built in collaboration between the University of Oxford and Vela Research, the research arm of Vela Partners.1 It has 9,000 anonymised founder profiles drawn from LinkedIn and Crunchbase. A founder counts as a success if their company later reached an IPO or acquisition above $500M, or raised more than $500M. In the real market about 1.9% of founders get there, and even tier-1 venture firms pick winners at only 2.9 times that rate. Entries are scored on 4,500 held-back profiles using F0.5, which weights precision above recall. More than 60 research teams, from universities, funds and independent labs, have requested the dataset so far.

Among the new entries, GPT-6 Sol scores 33.7, the best yet for a directly prompted model and well above the previous best of 25.7 from GPT-4o. Itay Attar of Pvalyou submitted the new best machine learning score, an F0.5 of 31.5. As of 27 September 2026, the Think-Reason-Learn Ensemble leads VCBench with an F0.5 of 37.9 on the 4,500-founder private test set.

On VCBench, Claude Opus 5.5 and GPT-6 Astra score lower on F0.5 (24.7 and 24.3) because they rarely back anyone, but when they do back a founder they are right 78% and 57% of the time, far ahead of the entries that lead on F0.5. The rankings on the site list precision next to F0.5 for exactly this reason, and we rank on F0.5 as a balance between the two, although another investor might weight them differently. While precision matters most near the final investment decision, a model that backs almost no one leaves too few founders to consider. Most other methods on the board produce a score, so the cut-off can be moved to optimise whichever metric an investor chooses. A directly prompted model only gives a yes or a no, which is much harder to tune.

With the leading approaches this close, what a prediction costs matters as well as how good it is. ARC-AGI, the reasoning benchmark run by the ARC Prize Foundation, already plots every entry against its cost per task.2 In venture the reason is practical. VCBench uses the information an investor sees at the top of the funnel, a founder's career and education, so a model like this works as a first filter across thousands of founders before anyone takes a meeting. At that scale, a method that costs ten dollars per thousand founders is a different product from one that costs thirty cents.

What changed on vcbench.com

We have updated vcbench.com with a “Score vs cost” view that compares accuracy with cost. Each entry is plotted by its F0.5 against what it costs to score 1,000 founders.8 The original view, which ranks on accuracy alone and ignores cost, is still available under “Score only”.

37.9 F0.5
Think-Reason-Learn Ensemble, top of the board, at $0.87 per 1,000 founders
33.7 F0.5
GPT-6 Sol, the best directly prompted model, at $1.45 per 1,000 founders
8.6× cheaper
The ensemble against GPT-6 Astra at $7.49 per 1,000 founders, while scoring 13.6 points higher

Itay Attar's model is a tabular method, and it beats the previous best traditional ML score of 29.4. Because it needs no language model to score a founder, it costs nothing to run through an API.

The entry at the top is itself a mix of the two. The Think-Reason-Learn Ensemble draws on Think-Reason-Learn, an open-source library developed by Vela Partners and the University of Oxford, and combines Verifiable-RL, a classical model that learns from the profiles without calling a language model at scoring time, with four reasoning methods that do.3 On this data the best single entries are reasoning methods, but the combination does better than any of them.

EntryFamilyF0.5Cost per 1,000 founders
Think-Reason-Learn EnsembleHybrid37.9$0.87
GPT-6 SolReasoning33.7$1.45
Policy InductionReasoning33.0$0.30
Reasoned Rule MiningReasoning32.2$0.31
Random Rule ForestReasoning31.3$0.13
Claude Sonnet 5Reasoning27.7$2.20
GPT-6 LunaReasoning26.1$0.11
GPT-4oReasoning25.7$1.55
Claude Opus 5.5Reasoning24.7$5.33
GPT-6 AstraReasoning24.3$7.49
Gemini 2.5 ProReasoning19.9$11.51
Full VCBench leaderboard, as of 27 September 2026 (31 entries)
#EntryOrganisationFamilyPrecisionRecallF0.5Cost / 1k
1Think-Reason-Learn EnsembleVela + OxfordHybrid40.630.137.9$0.87
2GPT-6 SolOpenAIReasoning45.117.033.7$1.45
3Policy InductionVela + OxfordReasoning34.927.233.0$0.30
4GemVC-v0IndependentReasoning39.420.332.9$2.87
5Reasoned Rule MiningVela + OxfordReasoning34.824.932.2$0.31
6Pvalyou Founder Model (ML)PvalyouClassical ML38.218.831.5$0
7Random Rule ForestVela + OxfordReasoning37.918.531.3$0.13
8Verifiable-RLVela + OxfordClassical ML29.130.629.4$0
9Structured-Rule-StumpIndependentClassical ML32.818.028.1$0
10verifiable-reasoningVela + OxfordClassical ML30.621.027.7$0
11Claude Sonnet 5AnthropicReasoning69.38.127.7$2.20
12large-founder-model-v0Vela + OxfordClassical ML31.717.527.2$0
13GPT-6 LunaOpenAIReasoning33.314.126.1$0.11
14GPT-4oOpenAIReasoning30.016.325.7$1.55
15FinGPT-VC2ColumbiaReasoning24.427.224.9n/a
16Claude Opus 5.5AnthropicReasoning77.76.724.7$5.33
17GPT-6 AstraOpenAIReasoning56.87.424.3$7.49
18GPT-4o-miniOpenAIReasoning31.511.123.0$0.09
19Claude Haiku 4.5AnthropicReasoning21.230.122.6$0.97
20FinGPT-VC1ColumbiaReasoning21.824.222.2n/a
21o3OpenAIReasoning43.27.421.5$9.25
22GPTreeVela + OxfordReasoning19.427.220.6$0.12
23Gemini 2.5 ProGoogleReasoning17.158.019.9$11.51
24DeepSeek-ReasonerDeepSeekReasoning31.86.918.4$1.36
25Claude 3.5 HaikuAnthropicReasoning15.846.418.2$0.52
26GPT-5OpenAIReasoning59.14.216.2$11.23
27Gemini 2.5 FlashGoogleReasoning12.568.414.9$2.87
28DeepSeek-ChatDeepSeekReasoning80.63.012.1$0.17
29Tier-1 VCsHumansHumans23.05.210.7n/a
30Random ClassifierBaselineBaseline9.09.09.0$0
31Y CombinatorHumansHumans14.06.98.6n/a

Why the top of the board is also the cheap end

Who gets the best score for the money?

Up and to the left is better. Highlighted entries are the best value, because nothing cheaper scores higher.
↖ Better$0$0.10$0.30$1$3$10$30010203040Cost per 1,000 founders scored (USD, log)Score (F₀.₅)Best score 37.9Tier-1 VCs 10.7
Every VCBench entry with a known cost, plotted by F0.5 on the private test set against the cost of scoring 1,000 founders. Classical ML entries have no API cost and sit in the $0 column on the left. The highlighted entries are the best value, because nothing cheaper scores higher. They are the best classical ML entry ($0, 31.5), Policy Induction ($0.30, 33.0) and the Think-Reason-Learn Ensemble ($0.87, 37.9). The best directly prompted model, GPT-6 Sol, scores 33.7 at $1.45. The dotted line marks Tier-1 VCs at 10.7. Hover or tap a point for its name and numbers.

The directly prompted models sit to the right because the largest models charge far more per token. At the low reasoning setting we used, GPT-6 Astra barely thinks at all, yet it still costs $7.49 per 1,000 founders because each token is priced at $10 in and $50 out per million. The less obvious part is why the Think-Reason-Learn methods, which do far more reasoning overall than a single prompt, cost well under a dollar per 1,000 founders.

These methods do the slow thinking once, up front. A frontier model studies the training data and writes the questions, and scoring a founder is then a long list of quick ones.

Daniel Kahneman's distinction between slow, deliberate reasoning (System 2) and fast, intuitive judgement (System 1) describes the difference well.4 Deciding what separates founders who went on to $500M companies from those who did not is System 2 work. Checking one founder against a clear question is closer to System 1. Think-Reason-Learn methods turn the first into many of the second. Random Rule Forest, for example, has frontier models write a bank of yes-or-no questions from the training data, keeps the ones that separate successes from failures, and lets a transparent model weigh the answers.5 Here is a typical question from one of its VCBench question banks.

“Has the founder worked as a professor or researcher in a related field?”

This is not a lookup in the data. The model has to judge whether the founder's research is related to the company they started. A structured feature cannot capture that, because it takes an understanding of both the research and the business. But it is still a quick, bounded judgement with a yes or no answer, which is the kind of question a fast model answers well. Policy Induction, Reasoned Rule Mining and GPTree work the same way.6 In each case the expensive reasoning happens once, and scoring a new founder comes down to answering a few dozen to a few hundred small questions.

That is why a System One model fits so well. Jev, from TypeSafe AI,7 does not generate text at all. It takes a state and a question and returns a probability, priced at $42 per billion input tokens. Random Rule Forest sends it about 3,200 tokens per founder, which costs $0.13 per 1,000 founders. Policy Induction and Reasoned Rule Mining ask more questions, about 7,300 tokens per founder, which comes to around $0.30 each. With GPTree and Verifiable-RL, which has no API cost, these methods make up the Think-Reason-Learn Ensemble. Together they cost $0.87 per 1,000 founders and score 4.2 points higher than GPT-6 Sol, the second-best entry on the board, which costs $1.45.

Submit with a cost

Anyone can submit to VCBench, and you can request the data at vcbench.com. With your predictions, tell us what it cost to score 1,000 founders, or the model you used and how many tokens it used per founder, and we will work out the cost. Your entry then appears on the leaderboard, and if nothing cheaper scores higher, it is highlighted as the best value. A method that matches the leaders at a fraction of the cost is as useful to us as one that beats them. Itay Attar's model at Pvalyou is a good example, scoring 31.5 with no API cost at all.

Frequently asked questions

Why does the board show cost rather than tokens?
Tokens are not comparable across providers, since tokenizers differ and a token of Jev costs 240 times less than a token of GPT-5 output, whereas dollars are what a fund pays. The tokens are stored underneath so the dollars can be recomputed when prices change.
What is a System One model?
A model that returns a decision, a probability or a choice, instead of generating text. Jev is the first we have used in production scoring.
Is $0 really zero for traditional ML?
The API bill is zero. The compute is an embedding pass or a tree lookup per founder, which is not metered for any entry on the board, including the linear layer behind the Think-Reason-Learn methods.

References and notes

  1. Chen, Ternasky, Kwesi, Griffin, Yin, Salifu, Amoaba, Mu, Alican and Ihlamur. VCBench: Benchmarking LLMs in Venture Capital. arXiv:2509.14448. The live board is at vcbench.com.
  2. ARC Prize Foundation. ARC-AGI leaderboard, which plots each entry's score against its cost per task.
  3. Think-Reason-Learn, the open-source library from Vela Partners and the University of Oxford behind these methods, at thinkreasonlearn.com (code at github.com/Vela-Research/think-reason-learn).
  4. Daniel Kahneman. Thinking, Fast and Slow. Farrar, Straus and Giroux, 2011.
  5. Random Rule Forest (RRF): Interpretable Ensembles of LLM-Generated Questions for Predicting Startup Success. arXiv:2505.24622.
  6. Policy Induction: Predicting Startup Success via Explainable Memory-Augmented In-Context Learning; Reasoned Rule Mining: Precision-Optimised Weighted LogProb Classification; GPTree: Towards Explainable Decision-Making via LLM-powered Decision Trees.
  7. TypeSafe AI, makers of Jev.
  8. Costs use providers' list prices in September 2026. Costs for the Think-Reason-Learn methods and for the models we ran this month are measured from usage; for older directly prompted models they are estimated from token counts.