In conversations about intelligence retrofits, the most common opening line is "we want to bring in an AI system — which model should we use?" But push the project forward and you find the model is rarely what decides the outcome. Something earlier does: whether the data on site can flow out steadily, under one consistent definition.
Onboarding is underrated because it doesn't look like "intelligence"
What does onboarding work actually produce? A data path that works, a clear protocol mapping table, and an edge program that doesn't lose readings when a device drops offline. None of it looks impressive in a demo, and it's hard to fit on a slide. But skip it, and every later stage comes back as rework — in one form or another.
- Model training lacks samples — because collection was intermittent and sample rates were uneven
- Performance is unstable after go-live — because site data drifted away from the training set and nobody noticed
- The system raises false alarms constantly — because a device going offline was read as an abnormal state, with no link-health check
- Acceptance turns into an argument — because "is the data accurate" was never defined in the first place
The three things onboarding must actually solve
We break onboarding into three goals that all have to be met. Miss one and the step isn't finished.
| Goal | What it means | How it is accepted |
|---|---|---|
| Data can be extracted | Device protocols are fully parsed and every key field can be read | During continuous collection, the missing-field rate stays below the agreed threshold |
| Data flows reliably | Network drops and device restarts recover automatically, with no loss of critical data | Disconnect the network deliberately: data buffers at the edge and backfills once the link returns |
| Definitions are clear | Every field's meaning, unit, precision and sampling period is explicitly defined | A data dictionary is issued, and the business and implementation sides read the same field the same way |
Why this step deserves to be delivered as a stage of its own
Because its value stands on its own. Even if a customer later decides to pause the model work, what was built during onboarding keeps producing value: automatic collection replaces manual logging, device status goes from invisible to visible, and anomalies move from "found afterwards" to "known at the time".
That is also why we split the roadmap into stages. Each stage should be the smallest unit that can deliver value independently — not a half-finished thing that only means something at the finish line.
If stopping halfway leaves you with nothing of value, the stage was defined wrongly from the start.
A simple test for whether onboarding is ready
Before moving on to model work, answer one question: if you exported the site data as a spreadsheet right now, could the business side read every column and understand what it is, and why it holds that value?
If the answer is "I'd have to ask the colleague who owns it", then onboarding isn't yet solid enough to build capability on top of. Clean up the definitions and the data path first, and every step afterwards goes faster.