Churn models, attrition models and early-alert systems for students share a shape: a score for each person, a threshold, a list. Building the score is the part data teams enjoy and the part that gets demonstrated. It is also, in our experience, the part that matters least. The organizations that get value from these models made three decisions before training anything.
Decide what happens to a flag
A flag is a request for someone's time. Who receives it, how many they can handle in a week, and what they are expected to do must be settled first, because those answers determine the threshold. A team that can reach forty people a week needs a list of forty, ranked, not a list of four hundred coloured red. The model's job is to rank; the program's job is to decide capacity.
Explain every score
An advisor, a manager or an account owner will not act on a number they cannot explain to the person in front of them. Each score should come with the two or three signals that drove it, in plain language: missed assignments and no platform activity for a week; a drop in hours and a pay band at the bottom of the range; a fall in order frequency and an open complaint. Explanations also keep the model honest, because a signal that makes no sense to the people using it is usually a data problem.
Measure the follow-up, not the flag
The question that justifies the program is not how many people were flagged but whether the people who were reached stayed, returned or renewed at a higher rate than comparable people who were not. That means recording who was contacted, when, and what was done, and comparing outcomes by group a term, a quarter or a year later. Without that record the model can be accurate and the program can still be worthless.
Design privacy in
Risk scores about people are sensitive by definition. Who may see a score, how long it is kept, whether the person it concerns can see it, and which signals are off limits are decisions for the institution, not the data team. Making them early, and visibly, is what allows the program to exist at all in a school board, a hospital or a bank.
Keep the model boring
Most of these problems are solved well by transparent models trained on a few years of history and tested on data the model has never seen. Accuracy by horizon, checked every term or quarter, tells you when the model needs retraining. Complexity should be added only when a simpler model is demonstrably missing people who were later lost.
The test
If the people who receive the alerts can say what they did with last month's list and what happened to those people, the program is working. If the answer is that the list was reviewed, it is not yet a program; it is a report.