Why Predictive Models Need to Understand How Sensors Fail

Published on 2026-10-05

A predictive model cannot be more trustworthy than its interpretation of the measurements it receives. In physical systems, an apparent change in a time series may represent deterioration in the monitored asset, a genuine change in operating conditions, or a change in the measurement chain itself. These cases can produce similar data, but demand very different actions.

A rising temperature trend may indicate increasing friction in a bearing. It may equally be caused by sensor drift, altered sensor mounting, contamination, a changed sampling interval, or replacement of the sensing element with a device having a different response characteristic. A model trained to identify early equipment degradation may assign high confidence to the wrong explanation unless the sensing system is treated as part of the modelled system.

This is a fundamental limitation in predictive analysis based on physical-world data. The model observes measurements, not the underlying state directly.


Sensor failure is often gradual and plausible

Complete sensor failure is comparatively easy to detect. Values may become absent, fixed, out of physical range, or inconsistent with neighbouring instruments. The more difficult faults are those that remain plausible.

Drift is a progressive change in sensor response over time. It can arise from ageing, chemical exposure, contamination, component degradation or changes in installation conditions. A temperature probe may develop an offset; a pressure transmitter may slowly lose zero stability; an optical sensor may respond differently as its window becomes fouled. Because the resulting values can remain within expected operating ranges, simple range checks will not identify the problem.

Bias has similar consequences but need not evolve gradually. An installation error, incorrect scaling configuration, damaged cable, compensation error or calibration adjustment can introduce a persistent offset. If a model learns from biased historical data, it may encode that bias as normal operation. If the bias appears after deployment, it may generate false warnings or distorted remaining-useful-life estimates.

Noise presents a different failure mode. Increased random variation may be caused by electrical interference, loose connections, unstable illumination, vibration, failing electronics or an unsuitable filtering configuration. Smoothing can reduce visible variation, but indiscriminate filtering may also remove short-lived events that are operationally significant. It can also create artificial delay, causing the apparent event time to diverge from when the physical event occurred.

The distinction matters because a model that treats measurement degradation as asset degradation can trigger unnecessary intervention. One that filters or suppresses the wrong signal can miss a developing fault.


Dropout changes the meaning of a time series

Missing data is not merely an inconvenience for a predictive pipeline. The pattern of absence may itself carry information about the measurement system, communications infrastructure or operating environment.

A sensor can drop out intermittently because of power instability, radio interference, connector faults, network congestion or firmware defects. Data may be absent during precisely the conditions that matter most, such as high temperature, high vibration or heavy process load. In this case, missingness is not random. Replacing gaps with interpolated or forward-filled values can create a deceptively smooth record and conceal a condition associated with both asset stress and telemetry failure.

A model should therefore retain explicit information about data availability: observation time, expected sampling interval, late-arrival status, quality flags and the method used to handle missing values. These fields may be model inputs, controls for downstream use, or both. The correct choice depends on the intended decision and the evidence available from validation.

Time also requires careful treatment. A value with an inaccurate timestamp can be more damaging than an absent value when several sensor streams are correlated. Misalignment can make one variable appear predictive of another simply because the recorded order of events is wrong. This is especially important where data is buffered at the edge, transmitted intermittently, or aggregated from separate devices.


Replacement creates a change point, not a continuation

Sensor replacement is commonly treated as routine maintenance. For predictive systems, it is a potentially material change in the data-generating process.

Even nominally identical instruments can differ because of calibration state, manufacturing tolerance, mounting geometry, firmware version, signal conditioning or local environment. A new sensor may be more accurate than its predecessor, but its readings may not be directly comparable with the earlier series. The discontinuity can look like an abrupt process change and may affect models that rely on long-term trends, baselines or residuals.

Systems should record sensor identity, installation and removal times, calibration status, configuration version and relevant maintenance events. This enables a model or monitoring service to segment the series, apply an appropriate baseline, or suspend certain inferences until sufficient post-change data has been gathered. It also permits later investigation of whether an apparent operational event coincided with a change to the measurement chain.

This is not an argument for discarding historical data after every replacement. It is an argument for preserving the context needed to establish whether the data remains comparable.


Measurement-aware models need engineering boundaries

There is no universal algorithm that can reliably separate sensor fault from process change using a single signal. A model needs additional evidence and explicit boundaries on what it can infer.

Redundant sensing can help, provided the sensors do not share the same failure mechanism. A second temperature measurement mounted beside the first may confirm a discrepancy, but two sensors supplied by the same power rail and installed in the same contaminated environment are not independent evidence. Cross-variable checks can also be valuable. A pressure change may be more credible when accompanied by corresponding changes in flow, actuator position and energy consumption, subject to a known physical relationship.

Physics-informed constraints, sensor diagnostics and model residual monitoring provide further controls. If a value violates a feasible rate of change, energy balance or known operating envelope, the system should distinguish a suspect measurement from a confirmed process event. That distinction should be visible to the operator rather than hidden behind a single prediction score.

For higher-consequence use, the decision logic should define degraded operation. It should specify what occurs when critical inputs are stale, missing, out of calibration, inconsistent, or known to have been replaced. Continuing to issue apparently precise predictions from an impaired measurement chain is often less safe than reducing confidence, withholding an automated recommendation, or reverting to a defined conservative mode.


Train on the system that will actually exist

Predictive models are often evaluated against curated historical data in which sensor faults have been removed or corrected. This can produce attractive accuracy figures while leaving the deployed system unprepared for ordinary field conditions.

Validation should include representative measurement faults: realistic drift rates, intermittent dropout, timing errors, noise changes, configuration changes and sensor replacement events. It should test not only prediction accuracy but also fault detection, confidence behaviour, alarm burden, recovery after repair and the ability to reconstruct why a result was produced. Synthetic fault injection is useful where labelled incidents are scarce, but its assumptions should be documented and tested against observed device behaviour.

A dependable predictive system does not assume that every numerical input is a faithful statement about the physical world. It establishes what evidence supports each measurement, detects when that evidence weakens, and constrains its own outputs accordingly. In systems where predictions inform maintenance, safety or regulated decisions, understanding sensor failure is not a peripheral data-quality task. It is part of the model's engineering basis.

Copyright © 2026 Obsidian Reach Ltd.

UK Registered Company No. 16394927

3rd Floor, 86-90 Paul Street, London,
United Kingdom EC2A 4NE

020 3051 5216