← All posts

Random Rule Forest: Predicting Startup Success with LLM-Written Yes-or-No Questions

Random Rule Forest asks a language model to write simple yes-or-no questions about founders, keeps the questions that separate successes from failures, and turns the best of them into a transparent scorecard an investor can read and edit.

The idea

Only about 1.9% of founders reach a $500M outcome, yet one of them can return an entire fund. A model that helps find them has to be accurate, and it has to be explainable to the people who act on it. Embeddings and word-count features are hard for an investor to read. A list of questions is not.

Random Rule Forest (RRF) uses a language model as a generator of readable features, and keeps the final decision rule deliberately simple. It is one of the methods Vela Research built for AI-native venture capital at Vela Partners, where venture capital predictions have to be both accurate and easy for investors to check.

How Random Rule Forest works

  1. Generate questions. The model is shown 20 founders, half successful, and asked for 10 yes-or-no questions that might tell them apart. Repeated across batches, this produces 250 candidate questions.
  2. Filter and rank. The model answers every question for every training founder. Near-duplicate questions are collapsed, leaving about 65, and each is ranked by its own F0.5 on the training data.
  3. Vote. The top questions form a scorecard. A founder is predicted to succeed if enough of them come back yes. In practice the model settles on 6 to 11 questions with a threshold of 4 to 8 yes answers.

No single attribute is required. Accumulating enough independent positive signals is what counts, which is close to how many investors already think.

Here are some questions from the paper.

  • “Has the founder been involved in any venture capital or private equity firms?”
  • “Is the founder's university ranked among the top 50 globally?”
  • “Has the founder previously led a startup to a high-value exit?” (expert-written)

What the paper found

The current version of the paper evaluates RRF on 9,192 anonymized founder profiles of US companies founded between 2010 and 2016, with a realistic 1.9% success rate and 100 outer folds of nested cross-validation.

  • RRF reached 10.9% precision, 8.2% recall and an F0.5 of 0.102, 5.6× better than random chance.
  • Adding eight expert-written questions to the candidate pool raised precision to 16.4%, 8.5× random chance, with F0.5 0.125.
  • On a second task, predicting the outcome of Phase I clinical trials, RRF achieved the best PR-AUC (0.638) and ROC-AUC (0.596) of all published methods.

The authors note that the founder comparisons with other methods are not statistically separable, and list the limits plainly. Question generation is stochastic, questions can reward pedigree, and data shift over time matters.

The paper was first posted in May 2025 by Ben Griffin (University of Oxford), Joseph Ternasky, Fuat Alican and Yigit Ihlamur (Vela Research). Later versions add Aaron Ontoyin Yin, Diego Vidaurre and Ugur Koyluoglu, extend it beyond startups and retitle it for predicting success from unstructured data. The numbers above are from the current version.

Where Random Rule Forest stands on VCBench today

Random Rule Forest on the VCBench leaderboard today
Rank
#7 of 31
F0.5
31.3
Precision
37.9%
Recall
18.5%
Cost / 1k founders
$0.13

Scored on the 4,500-founder private test set, mean over three folds. That is 2.9× the F0.5 of tier-1 VCs. Today's figures come from a rerun of the method on VCBench, with costs measured from usage. See the full leaderboard.

Random Rule Forest is part of the open-source Think-Reason-Learn library, with code, data and prompts at github.com/Vela-Research/random-rule-forest. See also Policy Induction and Reasoned Rule Mining.

Read the paper

Frequently asked questions

What is Random Rule Forest?
Random Rule Forest (RRF) uses a language model to write yes-or-no questions about founders, keeps the questions that best separate successes from failures, and predicts success when enough of the top questions are answered yes.
How accurate is Random Rule Forest?
In the current version of the paper, RRF reaches 10.9% precision against a 1.9% base rate, 5.6 times random chance, and 16.4% precision when eight expert-written questions are added.
How does Random Rule Forest score on VCBench?
On the current VCBench leaderboard, Random Rule Forest scores an F0.5 of 31.3 with 37.9% precision and 18.5% recall, at about $0.13 per 1,000 founders scored.