Decision support, not a verdict. Risk flags indicate that a student may benefit from timely support and should be reviewed by an educator or advisor. All data shown is synthetic.
Eight short steps that walk through the whole system using one real student. Every number is explained where it appears — nothing here assumes you know what a model is.
Why does this system need to exist at all?
Online courses lose a lot of students — often a third or more never finish. The frustrating part is that the warning signs are usually already in the system: someone stopped logging in, stopped handing work in, went quiet in the discussion forum.
Nobody is watching those signals week by week, so the problem gets noticed at the exam — far too late to help.
So the problem isn't missing data. It's that nobody is looking at it in time. This system does the looking.
What it is, in one sentence
What goes in? Where does the information come from?
Every week, for every student, the system records 25 numbers describing what they did. They fall into four groups:
Activity
How much they did
e.g. logins, minutes on the site, pages viewed
Assessment
How their work is going
e.g. quiz scores, work handed in, late, missed
Interaction
Whether they're connected to others
e.g. forum posts, replies, asking peers for help
Timing
When they work, not how much
e.g. study times all over the place, weekend cramming
Four of their signals, week by week, in real units — this is exactly what the model is handed.
| Signal | wk 1 | wk 2 | wk 3 | wk 4 | wk 5 | wk 6 | wk 7 | wk 8 |
|---|---|---|---|---|---|---|---|---|
| Logins | 7 | 3 | 5 | 5 | 5 | 5 | 1 | 5 |
| Assignments submitted | 1 | 0 | 2 | 1 | 0 | 2 | 0 | 2 |
| Discussion posts | 3 | 2 | 1 | 0 | 2 | 0 | 0 | 0 |
| Average quiz score | 52.91 | 71.11 | 74.71 | 74.71 | 63.69 | 63.69 | 63.69 | 63.69 |
Faded columns are weeks after the prediction point — the model is not allowed to see them. Red numbers are well below the cohort average for that week.
In plain English
When predicting week 5, the system may only use weeks 1 to 5. Never week 6.
How it's worked out
Enforced in four places, including inside the model itself, where later weeks are mathematically unreachable rather than merely withheld.
What it tells you
The week-by-week results are real, not the result of accidental cheating.
Why it's here at all
It's the easiest way to get beautiful results that mean nothing, and it's invisible unless you test for it.
It is entirely invented. 500 fake students are simulated over 8 weeks — some steady, some drifting away slowly, some vanishing suddenly, some struggling early then recovering.
Crucially, the simulation decides who dropped out based on how they behaved, plus a large amount of deliberate randomness. That randomness means two students with identical behaviour can still end differently — like real life. Without it the problem would be trivially easy and every result would be a lie.
You pick a student and a week. What comes back, and how do you read it?
Your input
Student student_high_risk_001 · Week 5
Meaning: “It is currently week 5. Based only on what has happened so far, how worried should I be about this student?”
Risk score
39%
Out of 100 students who behaved like this one, about 39 did not finish the course.
In plain English
Out of 100 students who behaved like this one, roughly this many did not finish the course.
How it's worked out
The model reads the student's weeks so far and produces a number, which is then corrected (“calibrated”) so the percentage means what it says.
What it tells you
A description of a pattern, not a judgement about a person. A high score means this student's behaviour resembles students who struggled.
Why it's here at all
It gives a teacher a place to start when they have 500 students and time to contact ten.
Risk band
Above the flag line — this student would appear on a teacher's list.
In plain English
A traffic light — low, medium or high — so nobody has to read percentages.
How it's worked out
“Medium” starts at the exact point where the system would flag someone. “High” is a good margin past that point.
What it tells you
Low = looks like the rest of the cohort. Medium = the system would flag them. High = flagged with room to spare.
Why it's here at all
Bands track the flag point automatically, so they can't drift out of step with the actual decision.
The most important thing on this page
How to read this: Each dot is a separate prediction made using only the weeks up to that point. The shaded band is how much the model wavers. The dashed line is the flag point at 26%.
In plain English
The score above which a student gets flagged for attention.
How it's worked out
Chosen on a separate group of students the model never trained on, tuned to favour catching people, and capped at 40% of the cohort.
What it tells you
Everyone above the line appears on the teacher's list.
Why it's here at all
0.5 is an arbitrary default that suits almost no real problem.
A number alone is useless. What is the reasoning behind it?
The system's own explanation
At week 5, this student shows some early warning signs (risk score 39%). This is an opportunity for a light-touch check-in. The model draws on the whole history so far fairly evenly rather than singling out any particular week. The signals contributing most are that assignments submitted is below the cohort average (1.0 vs 1.3 over weeks 3-5); late submissions is above the cohort average (1.0 vs 0.2 over weeks 3-5); irregularity of study times is above the cohort average (0.57 vs 0.47 over weeks 3-5). On the positive side, course material views is above the cohort average. Model confidence is 84% with moderate uncertainty - repeated stochastic passes agree on this estimate.
How to read this: Bars to the right pushed the risk up; bars to the left pulled it down. For example, "Assignments submitted" moved this student's risk by 11.6%.
This is a sensitivity measure — how much the model's output depends on a signal. It is not a claim that the signal caused anything.
In plain English
One thing the model watched — logins, assignments handed in, forum posts — and how much it moved this student's score.
How it's worked out
Take that one signal, replace it with the cohort average, run the model again, and measure how far the score moves. A big drop means the signal was doing a lot of work.
What it tells you
“Assignments submitted, +8%” means: if this student had been average on assignments, their risk would be 8 points lower.
Why it's here at all
It turns a score into something a teacher can check for themselves and disagree with.
How to read this: Taller bars are weeks that mattered more to this particular judgement.
In plain English
How much the model leaned on each individual week when judging this student.
How it's worked out
Read directly out of the model's attention layer — this is the model reporting on itself, not a guess about it.
What it tells you
A tall bar on week 4 means the model's opinion rests heavily on what happened in week 4.
Why it's here at all
It points a teacher at when things changed, not just that they did.
How to read this: the model decides for itself whether to judge a student on recent swings or a long slow trend. For this student it leaned most on the 1-week window.
The point of the whole system. A score with no suggested action is just anxiety.
For each of the 25 signals: take that one signal, replace it with the cohort average, and run the model again. If the risk drops a lot, that signal was doing a lot of work.
This asks the model rather than telling a story about the data — which is why it can be trusted as a description of the model's reasoning, even though it says nothing about cause and effect in the real world.
Every model is sometimes wrong. Does this one admit it?
Confidence
84%
The model was run 30 times with small random changes. The runs disagreed noticeably, so treat this score with care.
In plain English
How much the model agrees with itself.
How it's worked out
The model is run 30 times with small random changes switched on. If all 30 runs land close together, confidence is high; if they scatter, it's low.
What it tells you
High confidence means the score is stable. Low confidence means the model is genuinely unsure about this particular student.
Why it's here at all
A score with no confidence attached invites more trust than it deserves.
What counts as good
Higher is steadier — but low confidence is honesty, not failure.
Needs human review
Not flagged
The repeated runs agreed and the score isn't sitting on the borderline.
In plain English
The system saying: don't act on me alone for this student.
How it's worked out
Triggered when the 30 runs disagree too much, or when the score sits right on the borderline where a nudge either way would flip the decision.
What it tells you
A prompt to look at the student's situation yourself, not a conclusion.
Why it's here at all
A system that never admits doubt gets trusted in exactly the cases where it shouldn't be.
Why a system that says “I don't know” is better
Calibration error
0.038
When the system says "60%", the real figure is off by about 4 percentage points on average. Small enough that the percentages can be read literally.
In plain English
Whether “60%” actually means 60%.
How it's worked out
Group all predictions by their stated percentage, then check what fraction of each group really dropped out, and measure the gap.
What it tells you
0.04 means the stated percentages are off by about 4 points on average — close enough to read literally.
Why it's here at all
Without it, the model's numbers rank students correctly but the percentages are meaningless.
What counts as good
Below 0.05 is good.
Does it treat every group of students the same way?
A system like this can quietly go wrong by flagging one group far more often than another. So fairness is measured and published, in three different ways, including where the answer is unflattering.
Counterfactual check
0.000
We secretly rewrote every student's gender, age group and subject to every possible value and re-ran the model. Not a single score changed — those details never reach the model.
In plain English
If we secretly changed a student's gender, age group or subject and re-ran the model, would their score change?
How it's worked out
Rewrite every sensitive attribute of every student to every possible value, re-run, and record the largest change.
What it tells you
Ours is exactly zero — those attributes are never fed to the model, and this proves it rather than assuming it.
Why it's here at all
It converts “we excluded that field” from a claim into a measurement.
What counts as good
Exactly 0.
Consistency
88.2%
Students who behaved almost identically get almost the same score, even when their demographics differ.
In plain English
Do two students who behaved almost identically get similar scores?
How it's worked out
For each student, find the five students whose behaviour was most similar — deliberately ignoring their demographics — and check how close their scores are.
What it tells you
0.88 means similar students usually get similar scores.
Why it's here at all
This is fairness at the level of the individual, not the group.
What counts as good
Closer to 1 is fairer.
Flag-rate gap
0.200
The biggest gap between two groups in how often they get flagged is 20 percentage points, in "academic background". Whether that is unfair depends on whether the groups genuinely differ — which is why the real dropout rate is shown alongside it on the Fairness page.
In plain English
Does the system flag one group far more often than another?
How it's worked out
Work out the share flagged within each group, then take the biggest gap between two groups.
What it tells you
0.13 means one group is flagged 13 percentage points more often than another.
Why it's here at all
A gap isn't automatically unfair — groups can genuinely differ — which is why the real dropout rate is shown next to it.
What counts as good
Closer to 0 is more even.
Something we built that didn't work
How well does it actually work, and compared to what?
ROC-AUC
0.845
Given one student who dropped out and one who didn't, it correctly identifies the riskier of the two about 84% of the time.
In plain English
Show it one student who dropped out and one who didn't. How often does it correctly rank the dropout as the riskier of the two?
How it's worked out
Compare every possible pairing of a dropout and a non-dropout, and count how often the ranking is right.
What it tells you
0.84 means it gets that pairing right 84% of the time.
Why it's here at all
It measures whether the ranking is useful, without depending on where you draw the flag line.
What counts as good
0.5 is a coin flip. 1.0 is perfect. Above 0.8 is genuinely useful.
Recall
0.816
Of the students who really did drop out, it caught about 82% of them.
In plain English
Of the students who really dropped out, how many the system managed to catch.
How it's worked out
Correct flags divided by the number of students who actually dropped out.
What it tells you
0.82 means it caught roughly 8 in every 10 students who needed help, and missed 2.
Why it's here at all
This is the number that matters most here — a missed student is a real harm.
Precision
0.602
Of the students it flagged, about 60% really did drop out. The rest got an unnecessary check-in.
In plain English
Of the students it flagged, how many really did drop out.
How it's worked out
Correct flags divided by total flags.
What it tells you
0.60 means 6 in every 10 flagged students really were at risk — the other 4 got an unnecessary check-in.
Why it's here at all
It tells you how much of a teacher's time the system wastes.
Why precision is lower than recall — on purpose
In plain English
We chose on purpose to catch more struggling students, accepting more false alarms.
How it's worked out
The flag line is set to favour recall — and capped so the system can never flag more than 40% of a cohort.
What it tells you
More unnecessary check-ins, fewer missed students.
Why it's here at all
Missing someone who needed help is worse than sending a friendly email that turned out to be unnecessary. The cap exists because a list of 300 names is not a signal.
The honest range
0.75 – 0.92
We only tested on 82 students, so the true score is somewhere in this range rather than exactly 0.845. The width of that range is part of the result.
In plain English
The honest range the true value probably sits in.
How it's worked out
Re-run the scoring thousands of times on random re-samples of the test students, and take the middle 95% of the answers.
What it tells you
“0.84 (0.75–0.92)” means the real number is probably somewhere in that range — we tested on only 82 students.
Why it's here at all
A single number implies a precision we don't have. The width of the range is part of the result.
This is the whole point of an early-warning system, so it's worth checking.
At week 2
0.79
Only 2 weeks of behaviour to go on
By week 8
0.88
The full picture
It is already useful at week 2 and improves as the course runs — which is exactly the shape you want. A system that only works at the end is useless, because the student has already gone.
How do the pieces connect?
YOU pick a student and a week
|
v
The website asks the API for that student's risk
|
v
The API loads only weeks 1..t <-- later weeks unreachable
|
v
Signals are put on a common scale
|
v
The model reads the weeks:
catch sudden changes -> track the longer trend
-> decide which weeks matter -> produce a raw number
|
v
Run it 30 times with small random changes <-- confidence comes from here
|
v
Average them, then correct the scale so 60% really means 60%
|
v
Ask "why?": remove each signal in turn, see what moves
|
v
Write it in English, pick suggested actions
|
v
YOU SEE a score, a reason, a confidence level and a next stepWhat it does
Scores every student every week, and explains each score.
Why it exists
The warning signs are already in the data; nobody is watching them in time.
Why it's useful
It turns 500 students into a short, explained list a teacher can act on today.
Now go and use it
You know what every number means. The rest of the site is the same information, laid out for daily use rather than for learning.
In plain English
Whether this student was judged on recent week-to-week swings or on a long, slow trend.
How it's worked out
The model learns to make this choice itself, separately for each student and each week, rather than being told a fixed window.
What it tells you
A high '1-week' share means recent changes dominated. A high '7-week' share means the slow trend mattered more.
Why it's here at all
Students don't all move at the same pace, so a single fixed window fits nobody well.