← All posts

Reasoned Rule Mining: Calibrated LLM Predictions of Startup Success

Reasoned Rule Mining treats a language model's confidence as evidence to be calibrated, not a probability to be trusted. It mines readable rules about founders, converts the model's scores into calibrated probabilities, and picks the decision threshold that maximizes precision.

The problem with raw model confidence

When a language model says it is 90% sure, that number is not a true probability. It shifts with prompt wording, sampling and the class balance the model assumes. In a task where only 2% of founders succeed, poorly calibrated confidence ruins precision. Reasoned Rule Mining (RRM) treats this as a calibration problem rather than a prompt-engineering one. Calibrated probabilities are what AI-native venture capital needs, because venture capital predictions only help if a 20% score really means a one-in-five chance.

How Reasoned Rule Mining works

The published abstract describes a framework that produces evidence from each stage as scores, calibrates those scores into probabilities with logistic calibration, adds the calibrated evidence on the log-odds scale, and selects the operating point that maximizes F0.5.

  1. Mine rules. The model analyzes training founders step by step, and each rationale becomes a short if-then rule about the founder. Rules the model was unsure of, measured by perplexity, are dropped, and the rest are merged into a readable policy.
  2. Screen, then re-evaluate. A cheap rules-first pass scores every founder. Only those rated highly with strong confidence go to a stricter second stage, which cut second-stage calls from 8,000 to 698.
  3. Calibrate and combine. The model's log-probabilities from each stage are calibrated and added on the log-odds scale, which is the Bayesian way to combine independent evidence. When the stages are correlated, a small learned combiner sets the weights instead.
  4. Choose the threshold. Founders are ranked by calibrated probability, and the cut-off is set where F0.5 is highest.

Each prediction comes out as a calibrated probability and a yes-or-no call, and the rule policy behind it stays readable and editable by an expert.

What the paper found

The paper evaluates RRM on a reviewed set of 8,000 founders with 160 labeled successes, a 2% success rate, drawn from a pool of 9,892. At the selected operating point, RRM reached 24.5% precision and 15% recall, 12.25× the market index.

The paper was written by Jack Preuveneers (Department of Engineering Science, University of Oxford) and Yigit Ihlamur (Vela Research, the research arm of Vela Partners), and published at the 2025 IEEE International Conference on Cyber Security and Cloud Computing (CSCloud) in New York.

Where Reasoned Rule Mining stands on VCBench today

Reasoned Rule Mining on the VCBench leaderboard today
Rank
#5 of 31
F0.5
32.2
Precision
34.8%
Recall
24.9%
Cost / 1k founders
$0.31

Scored on the 4,500-founder private test set, mean over three folds. That is 3.0× the F0.5 of tier-1 VCs. Today's figures come from a rerun of the method on VCBench, with costs measured from usage. See the full leaderboard.

Reasoned Rule Mining is part of the open-source Think-Reason-Learn library, with Policy Induction and Random Rule Forest.

Read the paper

Frequently asked questions

What is Reasoned Rule Mining?
Reasoned Rule Mining (RRM) is a two-stage framework that mines readable rules with a language model, calibrates the model's scores into probabilities, combines them on the log-odds scale and picks the threshold that maximizes F0.5.
How accurate is Reasoned Rule Mining?
In the paper, on 8,000 founders with a 2% success rate, RRM reached 24.5% precision and 15% recall, 12.25 times the market index.
How does Reasoned Rule Mining score on VCBench?
On the current VCBench leaderboard, Reasoned Rule Mining scores an F0.5 of 32.2 with 34.8% precision and 24.9% recall, at about $0.31 per 1,000 founders scored.