← All posts

Learning What to Ask: Cost-Aware Founder Evaluation with Reinforcement Learning

Investors never see a founder all at once. They learn in steps, through interviews, reference calls and background checks, and every step costs time. Learning What to Ask turns founder evaluation into that kind of sequential decision. A reinforcement-learning policy chooses which attribute to look at next and when it has seen enough to decide.

Why founder evaluation should be sequential

Most models for venture capital predictions treat a founder as a single row of data where every field is known up front. Real diligence does not work that way. Each extra question costs time and money, and a firm that screens thousands of founders a year cannot run full diligence on all of them. The useful question is not only whether a founder will succeed, but what to look at next and when to stop looking.

Learning What to Ask, by Yuhang Ye (University of Oxford) with Fuat Alican, Ben Griffin, Aaron Ontoyin Yin and Yigit Ihlamur, answers that question for AI-native venture capital. It is part of Vela Research's work on machine learning augmented by language models, at Vela Partners.

How it works

The paper frames evaluation as a finite sequence of steps. At each step the agent either reveals one more attribute of the founder, such as education, role or industry, or stops and commits to a prediction from a small neural classifier. Five pieces make this work.

  1. A reward that reflects real costs. A false positive risks capital and a false negative loses an opportunity, so each outcome is rewarded differently. Every query carries a small cost, and repeated queries cost more, which models the price of diligence.
  2. Looking ahead. The only reward comes at the end, so the policy estimates the value of each possible next question by simulating how the rest of the evaluation would play out. This reveals questions that only pay off in combination with later ones.
  3. Two kinds of decision. Choosing the next question is compared across options, while stopping is judged on an absolute scale of confidence. Separating the two is what makes the policy risk-aware.
  4. Stable training. Standard policy-gradient training collapsed to predicting failure for everyone. Instead, the improved policy from the look-ahead step is distilled into the network, which trains reliably.
  5. A language model as advisor. When the policy's top two choices are close, GPT-5.2 is asked for a recommendation and the policy leans toward it. The model never sets rewards, stops the evaluation or overrides the policy.

What the paper found

The method was evaluated on VCBench, the first benchmark for venture capital, with 9,000 anonymized founders and a 9% success rate. Averaged over ten random test splits, it reached 43.9% precision, 23.0% recall and an F0.5 of 37.1, a 4.88× lift over the base rate.

Method (as reported in the paper)PrecisionF0.5
Random classifier9.0%9.0
Tier-1 VCs23.0%10.7
Reasoned Rule Mining87.5%21.0
GPT-4o30.0%25.7
Policy Induction41.0%34.0
Learning What to Ask43.9%37.1

The ablations show why each design choice matters.

VariantF0.5Information used
One-shot neural net, all features34.8100%
Policy gradient (PPO)0% precisionn/a
One-step lookahead22.821.7%
No LLM supervisor36.566.6%
Full model37.154.0%

The full model beats a neural network that sees every feature, while looking at roughly half the information. The language-model advisor matters most for efficiency. Without it, the policy asks for 66.6% of the information instead of 54.0%.

Diligence that follows the evidence

The learned policy spends effort where it pays. Founders with clear negative signals are rejected after two or three questions. Borderline cases are investigated further, and strong candidates lead the policy to explore execution history and industry fit. That gives a firm a principled way to set diligence budgets, instead of collecting the same data points for everyone.

Every decision is auditable. Each prediction comes with its full trajectory, so a partner can see which attributes were requested, in what order, and where the policy stopped. The decision threshold is exposed as a setting, so a firm can tune it to its own tolerance for risk.

The authors note the limits. Results vary across random splits, from 26.3 to 40.8 F0.5, because VCBench has few positive examples. The look-ahead simulations are expensive to train, though not needed at inference, and the method assumes a fixed set of attributes that can be queried.

Read more

Frequently asked questions

What is Learning What to Ask?
Learning What to Ask is a reinforcement-learning method that evaluates founders step by step. At each step it decides which founder attribute to look at next, or stops and makes a prediction, weighing the cost of each extra question.
How accurate is Learning What to Ask on VCBench?
In the paper, averaged over ten random test splits of VCBench, it reached 43.9% precision, 23.0% recall and an F0.5 of 37.1, while using 54% of the available founder information on average.
What role does the language model play?
GPT-5.2 acts as an advisor when the policy is unsure between its top two choices, nudging the policy toward a recommended question. It never sets rewards, stops the evaluation or overrides the policy.
Where was Learning What to Ask published?
The paper by Yuhang Ye, Fuat Alican, Ben Griffin, Aaron Ontoyin Yin and Yigit Ihlamur was released in March 2026 and accepted at an ECML PKDD 2026 workshop.