ArticlesApr 28, 2026 · 9 min read

Admissions Intelligence: Modelling Yield Without Modelling People

There's a version of admissions analytics that predicts whether an offer converts, and a version that predicts what kind of applicant is worth an offer in the first place. We will only build the first one.

educationadmissionsethics

The first question an admissions office asks us is almost always framed the same way: "can you help us predict who's going to enroll?" It's a reasonable question and a genuinely hard modelling problem, and we can usually improve on whatever spreadsheet-and-gut-feel process a mid-sized institution is currently running to plan its incoming class. But it's worth being precise about what that question is and isn't, because the same underlying capability — a model that scores an applicant — can be pointed at two different targets that look similar from the outside and are not similar at all in what they do to the people involved.

One target is: given this admitted student, and everything we know about how offers convert, what's the probability they enroll if we admit them? That's a yield model. It's predicting institutional behavior in aggregate — how a cohort with certain observable characteristics responds to an offer, financial aid package, and enrollment deadline — in order to help an institution plan class size, aid budget, and outreach timing. The second target is: given this applicant, what kind of student are they likely to be — how capable, how much of a flight risk, how much financial aid can we get away with offering — and that's not a yield model anymore. That's scoring a person on dimensions the institution has decided matter, using a model that will inevitably encode whatever biases sit in the historical admissions data it was trained on, and then using that score to make a decision about that person rather than a decision about the class as a whole.

Where we draw the line, mechanically

The distinction isn't philosophical — it shows up as a concrete difference in what the model is allowed to take as input and what it's used to decide, and we hold to three rules on every admissions engagement:

We model the offer, not the applicant's merit. A yield model's job is to predict whether a specific offer — at a specific price point, with a specific aid package, under specific deadline pressure — converts to an enrolled student. It should never be asked to answer "should this person have gotten an offer," because that's an admission decision, and admission decisions in every jurisdiction we work in carry legal and ethical obligations around fairness that a black-box model cannot discharge on an institution's behalf. The model runs after the admit decision is made by a human committee, not before it, and it never sees the raw application file — only the offer terms and the applicant's response to logistics like aid, deadline, and communications.

We refuse race, gender, and proxies for either as model inputs — including proxies an institution didn't ask for. This sounds obvious until you look at what's actually predictive in enrollment data: zip code, high school name, and first-generation status are all strong yield predictors, and all three are tight enough proxies for race and class that including them reproduces exactly the discriminatory pattern a fairness policy is supposed to prevent, just one layer removed from a protected category. We screen inputs for proxy risk before a model ever gets built, not after a bias audit flags it, because the audit-after approach means the model has already been used on a live admissions cycle by the time anyone checks.

Yield scores inform aid strategy and outreach timing, not admission itself. A yield model tells an enrollment office "this segment of admits typically needs to hear from financial aid within five days or conversion drops sharply" — that's an operational insight about process, and acting on it makes the institution's own behavior better, not the applicant's treatment worse. It never tells an admissions committee "admit fewer students like this one," because that use crosses from predicting an outcome into shaping who gets the opportunity to produce one.

The case where an institution asked us to cross the line

We've had exactly one engagement where an admissions office asked, directly, for a model that scored applicants on "likelihood of academic success" to use during the admit decision, not after it — reasoning, not unsympathetically, that they were drowning in applications and wanted a way to prioritize file review. We declined to build the input-side model and instead built something narrower: a triage tool that flagged incomplete files and missing required documents for faster processing, which solved the actual bottleneck (review-queue throughput) without touching the judgment the institution was trying to outsource. It took longer to agree on scope than it would have taken to just build the scoring model they initially asked for, and it's the engagement we'd point to first when explaining why this line matters in practice rather than in a policy document — the institution got real capacity relief, and no applicant was ever scored on anything but whether their file was complete.

Why this discipline is also, separately, good modelling practice

It would be convenient to say this restraint costs accuracy and we hold to it anyway on principle, but that's not quite honest — a yield model trained cleanly on offer terms, aid, and response timing tends to outperform a kitchen-sink model stuffed with applicant demographic proxies, because the proxy variables are noisy correlates of yield behavior rather than causal drivers of it, and noisy inputs degrade a model's stability across admissions cycles even when they look predictive in a single year's backtest. The institution that measured an 18 percent improvement in yield against modelled offers didn't get there by scoring applicants more aggressively. It got there by finally being able to answer, with evidence instead of instinct, which offers were underpriced on aid and which outreach sequences were arriving too late to matter — questions about the institution's own process, which is the only thing a yield model should ever have been asked to explain.