It's the most familiar scene in intelligence projects: the metrics look excellent at acceptance, three months later the business starts complaining that it "doesn't work well any more", six months later the system is shelved, and the conclusion is "AI isn't reliable". But look back carefully and the problem usually isn't the algorithm — it's that the environment changed and the model didn't.
What drift actually looks like in engineering
The term "data drift" sounds academic, but how it behaves on site is very plain.
- A supplier changed: same model of equipment, but the incoming material reflects light differently, and the vision model's segmentation boundary starts to shift
- The season changed: temperature and humidity shift, the sensor baseline moves as a whole, and the anomaly threshold stops working
- The process changed: takt time sped up, and samples that used to be captured at rest now carry motion blur
- The operator changed: handling habits differ, and the distribution of part orientation no longer matches the training set
- A sensor aged: drift accumulated slowly, crossed the decision threshold over several months, and on no single day did anything look "suddenly broken"
None of these changes is a fault, so no monitoring alert fires. But the model's input distribution has already moved away from the training distribution, and decay begins there.
Why these problems only show up after delivery
Because under project-based delivery, validation is concentrated into two weeks to a month. That window is too short for drift to become visible. And acceptance criteria usually look only at current accuracy, never at the stability of that accuracy.
| How it is accepted | What you see | What you miss |
|---|---|---|
| Offline test-set accuracy | The model's static performance on historical data | Performance once the data distribution shifts |
| Short on-site trial run | Usability in the current environment | Slow variables such as seasonality and batch changes |
| Feature checklist sign-off | Whether the features were built | Whether the effect can be sustained over time |
Turning drift into a routine you can manage
Our approach is to institutionalise three things in stage two (the closed loop), rather than hoping someone will fix it when it breaks.
- Monitor the input, not just the output: track the distribution of incoming data continuously against the training baseline, and raise a warning when it passes a threshold
- Keep a human feedback channel: let the people on site mark a wrong judgement at the lowest possible cost — those labels are the most valuable retraining data you will ever get
- Make retraining a routine: agree a fixed calibration cycle and trigger conditions (drift warnings, accuracy dropping below threshold), and make sure model updates can be staged and rolled back
Treat a model as a one-off deliverable and it will fail within months. Treat it as an asset that needs maintaining and it gets better with time.
Three questions for anyone evaluating a proposal
Before signing an intelligence project contract, it's worth asking these three clearly:
- After go-live, how would we notice that model performance is declining? And who notices?
- What does retraining require? Where does the data come from? How often?
- How are model updates deployed? If something goes wrong, can we roll back to the previous version quickly?
If the answers are vague, then the accuracy figure quoted at acceptance is probably the best this system will ever be.