GPTree: Rethinking Machine Learning with LLM-Powered Decision Trees
In April 2024, almost everyone building with large language models was chaining prompts together by hand. One prompt asked a question, the next acted on the answer, and the logic of the whole system lived in that chain. Watching it, we saw a decision tree hiding in plain sight. GPTree came from asking what would happen if the model built the chain itself, and the chain stayed readable.
Where the idea came from
Prompt chaining worked, but it was fragile. Every chain was designed by a person, tuned by trial and error, and hard to explain to anyone else. At the same time, the most trusted model in finance was still the decision tree. Investors like trees because they can follow every branch and see why a company was scored the way it was. Trees fail, though, the moment the data is messy text, which is exactly what a founder's story is.
A prompt chain is a sequence of questions. A decision tree is a sequence of questions. The difference is that a tree learns which questions to ask from data, while a prompt chain waits for a person to write them. GPTree joins the two. The language model proposes the questions, the data decides which ones split successful founders from unsuccessful ones, and the result is a tree anyone can read.
Rethinking machine learning from scratch
Classical machine learning starts after the interesting work is done. Someone decides which features matter, turns them into numbers, and the model learns weights on top. With language models, the model can take on that first step. It can read a founder's career in words and ask the questions a good investor would ask.
That changes what explainability means. In a GPTree model, there is no separate explanation bolted on after the fact. The questions are the model, so reading the tree tells you exactly how it decides. When an expert disagrees, they can change the model in plain language. This was the idea that set the direction for everything Vela Research built next.
How GPTree works
- Describe the task. One short prompt sets the context, such as asking the model to act as a venture capital analyst looking for patterns in successful founders.
- Generate insights. The model summarizes what successful founders have in common, in batches, and a second pass merges those summaries into one list of insights.
- Propose questions. For each part of a founder's profile, the model writes candidate yes-or-no questions. A question can be answered by the model itself, by generated code, or by grouping categories.
- Choose the best split. Each candidate question splits the founders in two. GPTree keeps the question that separates successes from failures most cleanly, then repeats the process on each branch until the tree is complete.
- Refine with an expert. An investor can collapse a node, rebuild a branch, or add advice in plain words, such as asking the model to consider whether a founder worked at a large technology company or studied at a top-ranked university.
A new founder is then routed down the tree by the answers to its questions, and the leaf they land in gives the prediction.
What the paper found
The paper, by Sichao Xiong (University of Oxford and Vela Research), Yigit Ihlamur, Fuat Alican and Aaron Ontoyin Yin (Vela Research), was posted to arXiv on 13 November 2024. It evaluated GPTree on 9,892 founders of companies founded between 2010 and 2016, where success means an IPO or acquisition above $500M, or more than $500M raised.
- On the dataset's 9.9% base rate, GPTree reached 37.3% precision, and 40.8% with expert refinement. Prompting GPT-4o directly reached 15.7%.
- Scaled to the real-world base rate, GPTree with expert refinement reached 7.8% precision in picking unicorn founders at inception. Tier-1 VCs reach 5.6%, Y Combinator 3.2% and random picks 1.9%.
- Training a tree on 6,000 founders with GPT-4o mini took about ten hours and cost about $30.
The paper was candid about the limits. Generated code can be unreliable, similar questions phrased differently can get different answers, and the model can hallucinate when a profile is thin. Each of those limits became the starting point for the work that followed.
What came next
GPTree was the first step. Random Rule Forest took the questions out of a single tree and into an ensemble that votes. Policy Induction moved the reasoning into editable policies written in plain text. Reasoned Rule Mining added calibration, so a model's confidence means what it says. All of them now live in Think-Reason-Learn, the open-source library from Vela Research and the University of Oxford, and the ensemble of these methods leads VCBench today.
- Rank
- #22 of 31
- F0.5
- 20.6
- Precision
- 19.4%
- Recall
- 27.2%
- Cost / 1k founders
- $0.12
Scored on the 4,500-founder private test set, mean over three folds. That is 1.9× the F0.5 of tier-1 VCs. The methods it inspired now hold four of the top seven places, led by the Think-Reason-Learn Ensemble. See the full leaderboard.
Where this goes
We believe the next generation of machine learning will be built on reasoning rather than on hand-made features. Models will ask their own questions, show their work, and take correction from experts in plain language. Venture capital is a demanding place to prove that, because the data is sparse, the outcomes take years, and every decision has to be defended. That is why Vela Partners runs as an AI-native venture capital firm, and why we built VCBench, the first benchmark for venture capital, in collaboration with the University of Oxford, to measure progress in public.
The same approach applies wherever decisions need to be both accurate and explainable, from healthcare diagnosis to hiring and lending. As foundation models improve, models like GPTree improve with them, without losing the ability to show exactly how they decide.
Read the paper
- GPTree: Towards Explainable Decision-Making via LLM-powered Decision Trees (arXiv, November 2024)
- GPTree on Vela Research
Frequently asked questions
- What is GPTree?
- GPTree is an explainable decision tree whose nodes are questions written and answered by a large language model. It needs no feature engineering or hand-built prompt chains, only a short description of the task.
- Where did the idea for GPTree come from?
- The idea began at Vela in April 2024, when builders were chaining prompts together by hand. A prompt chain is a sequence of questions, like a decision tree, so GPTree lets the model propose the questions and the data choose which ones to keep.
- How well did GPTree predict founder success?
- In the paper, GPTree with expert refinement picked unicorn founders at inception with 7.8% precision, ahead of tier-1 VCs at 5.6%, Y Combinator at 3.2% and random selection at 1.9%.
- How does GPTree score on VCBench?
- On the current VCBench leaderboard, GPTree scores an F0.5 of 20.6 with 19.4% precision and 27.2% recall, at about $0.12 per 1,000 founders scored. The methods it inspired, led by the Think-Reason-Learn Ensemble, now hold four of the top seven places.