Responsible Prediction in Education: What We Refuse to Build
Most vendor conversations about responsible AI in education stay abstract on purpose. Ours doesn't — here are the specific things institutions have asked us to build that we've turned down, and why.
"Responsible AI" in education usually gets discussed at the level of a values statement — fairness, transparency, human oversight, the vocabulary every vendor now puts on a slide. It's not wrong, it's just not specific enough to do any work, because it doesn't tell you what to say no to when an institution asks for something that's technically buildable and operationally tempting. So instead of a values statement, here's a list of things we've actually been asked to build in education engagements, and turned down, with the mechanism for why.
An individual-level dropout risk score used to deny services
We build early-warning signals — that's core to what we do — but there's a specific, common request we don't fulfill: a single risk score per student, used to gate access to something the student needs, like priority registration, a scholarship renewal, or eligibility for a support program. The mechanism problem is that a risk score built for detection (flag this student for outreach) is being asked to do allocation (decide who gets a resource), and those are different jobs with different error costs. A false positive in detection costs an unnecessary check-in email. A false positive in allocation costs a student a scholarship they needed and would have kept if given the chance. We'll build the first. We won't build the second without an explicit, separate fairness and appeals process that the institution — not us — owns and stands behind, because gating access based on a model's prediction of future behavior is a decision with legal and human consequences that should never live inside a vendor's black box.
A model trained to predict which students will file complaints or transfer
An institution asked us, once, for a model to flag students "likely to become a retention risk due to dissatisfaction" — trained partly on advising-note sentiment and support-ticket history — so that staff could "get ahead of it." We turned this down because the actual use case that emerged in the follow-up conversation was preemptive management of students who might complain, not support for students who were struggling, and those produce identical-looking training data (a student who's unhappy) while serving opposite purposes. We proposed measuring advising-note sentiment as an input to the existing week-six early-warning signal instead — folded into a model whose output is a support outreach, not a flag visible to the people that student might later have a grievance against.
Faculty-facing scores that rank instructors by predicted student outcomes
Outcome models are useful for identifying which interventions work. They are a bad tool for ranking faculty, because student outcomes are confounded by course difficulty, prerequisite preparation, and section composition in ways no model can fully strip out — and a ranking presented with false precision becomes a personnel-evaluation tool wearing a data-science costume. We've declined two requests to build "faculty effectiveness scores" derived from section-level outcome data, and instead scoped both engagements down to course-level outcome tracking with no instructor attribution, which answered the institution's real underlying question — where should curriculum redesign effort go — without creating a metric that would get misused in tenure and merit conversations it was never validated for.
Predictive admissions scoring on the input side of the decision
We've written about this one at length elsewhere: yield modelling on offers, never scoring applicants on likelihood of success as an input to the admit decision itself. It belongs on this list because it's the request we get most often, usually reframed as "just help us prioritize file review," and it's the clearest example of why the refusal has to be specific to the use, not the technique — the modelling techniques involved are the same regardless of which side of the admit decision they're applied to. What changes is who bears the cost of the model being wrong.
The common thread
Every item on this list is a case where the underlying technique is unremarkable — classification, scoring, risk modelling, the same tools we use for the early-warning and outcome work that is the core of our education practice — and the refusal isn't about the math. It's about where the prediction lands and who absorbs the error when it's wrong. A model that helps an advisor reach a struggling student five weeks earlier fails safely: the cost of a false positive is a wasted check-in. A model that gates a scholarship, ranks a faculty member, or shapes an admit decision fails expensively, onto a person who had no say in being scored. We build the first kind. We say no, in writing, to the second — and we'd rather lose the engagement than build something that looks like data science and functions like an unaccountable decision.

