The Difficult Engineering Problem of Updating Devices in the Field

Published on 2026-09-28

A field device is often most difficult to update precisely when an update is most necessary. It may be on an exposed asset, behind an intermittent cellular connection, powered by a constrained battery system, or installed in a controlled environment where site access requires planning and authorisation. A failed update can therefore become an operational outage, a safety concern, or an expensive recovery visit.

The engineering task is not to deliver a new firmware file. It is to move a device from one known, functioning state to another, while preserving a defensible path back to service if any part of that transition fails.


An update is a controlled state transition

A robust update design treats firmware deployment as a state machine, not a single action. The device must distinguish, at minimum, between an active image, a candidate image being received, a candidate image ready for activation, an image undergoing first boot, and a confirmed working image.

This distinction matters during failure. A connection may drop halfway through transfer. Power may be lost after an image has been written but before metadata has been committed. The new firmware may boot, initialise a peripheral incorrectly, and then fail before it can report status. If these states are not explicit and durably recorded, recovery behaviour becomes ambiguous.

A common architecture uses two independently bootable application slots. The running image remains intact while the new image is downloaded into an inactive slot. Only after complete receipt and verification does the bootloader mark that slot as a candidate. On the next reboot, the bootloader starts it under a trial condition. The new application must explicitly confirm successful operation within a defined period. If it does not, the bootloader restores the previously confirmed image.

This consumes non-volatile memory, and may require a more capable bootloader, but it avoids overwriting the only known-good software image. For a device that cannot be readily reached, that margin is normally worth the resource cost.


Authenticity begins before transfer

Transport encryption is useful, but it is not a sufficient firmware security control. A device must independently establish that the image was authorised for that device and has not been altered. Otherwise, an endpoint compromise, misconfigured server, or intercepted distribution route may be enough to install unauthorised code.

The normal mechanism is cryptographic signing. The bootloader, or another trusted component below the application, contains a trust anchor such as a public verification key or certificate. It verifies a signature over the firmware image or a signed manifest before permitting installation. The trust anchor must itself be protected from unauthorised modification, commonly through hardware security features, secure boot configuration, or carefully controlled write protection.

The signed metadata needs more than an image hash. It should identify the target hardware, image version, compatible bootloader version, size, and applicable dependencies. Without a hardware compatibility check, a valid image intended for one variant can be installed on another. Without version controls, an attacker may be able to replay an old but correctly signed image containing a known vulnerability.

Key management is part of the system design. Signing keys require controlled access, separation from build infrastructure where appropriate, revocation and replacement procedures, and an approach to trust-anchor rotation. A signature check is only as credible as the process governing the authority that creates signatures.


Interrupted transfer must be ordinary, not exceptional

Field networks fail routinely. Devices should expect incomplete downloads, duplicate blocks, delayed acknowledgements and long periods without connectivity. A reliable update protocol should support resumable transfer, recording which verified blocks are present in persistent storage rather than assuming a continuous session.

Writing directly to executable flash during download is risky. Staging data in a separate area allows the system to calculate and verify the complete image digest before activation. It also allows received data to be discarded safely if the manifest does not validate.

Storage integrity requires attention at the implementation level. Metadata that selects the active slot or records a successful boot must survive reset at any instruction boundary. Techniques include redundant records, sequence counters, checksums, append-only metadata pages and atomic flag changes supported by the underlying flash technology. The correct choice depends on the memory's erase and write behaviour. An apparently small update marker can be the single point of failure that determines whether a device remains recoverable.


Rollback needs a definition of success

Automatic rollback is valuable, but the condition for confirming a new image must be carefully chosen. A device that confirms immediately after reaching its main loop can preserve defective firmware that later fails under actual load. A device that waits for every external dependency may roll back unnecessarily when a network or sensor is temporarily unavailable.

The confirmation criteria should reflect the device's intended function. They may include completion of critical self-tests, successful configuration migration, availability of required measurement channels, stable operation for a defined interval, or completion of a secure check-in. The criteria need to be bounded. An update should not leave the device indefinitely in a trial state because an unrelated remote service is unavailable.

Some changes cannot be safely reversed by simply restoring old code. Firmware may alter persistent data structures, calibration storage, radio settings or attached-module firmware. Compatibility planning is therefore essential. Where possible, migrations should be forward-compatible, staged over separate releases, and reversible. Where not possible, the update package and recovery plan must explicitly account for that constraint.


Recovery is an operational capability

A dual-image design does not remove the need for recovery. Bootloaders can contain defects, flash can wear out, and an erroneous update policy can distribute a valid but unsuitable image to a fleet. Devices should retain a minimal recovery path that is isolated from normal application behaviour. Depending on the product, this may be a protected serial interface, a local maintenance mode, a hardware boot selector, or a network recovery protocol with tightly constrained authority.

Remote recovery must not become an unauthenticated back door. It needs the same or stronger authentication and authorisation controls as the normal update route, and should expose only the functions needed to restore a trusted image.

Operational evidence matters as well. The service should record the device identity, current and target versions, package identifier, signature-verification result, transfer outcome, activation time, trial result and rollback reason. This supports fleet management, incident investigation and controlled release decisions. It also makes it possible to distinguish a device that never received an update from one that received it, rejected it correctly, and remained safely on its prior version.

Relevant guidance exists for this class of problem. IETF SUIT defines an architecture and manifest model for firmware updates in constrained devices, while NIST provides broader guidance on platform firmware resilience. Neither document eliminates product-specific engineering decisions about hardware, threat model and failure recovery.

A dependable field-update mechanism is therefore built into the device architecture from the beginning. Once a product has only one writable firmware image, no durable update state and no trusted recovery path, adding remote updates later is not a deployment feature. It is a redesign of the system's security and failure behaviour.


References

Copyright © 2026 Obsidian Reach Ltd.

UK Registered Company No. 16394927

3rd Floor, 86-90 Paul Street, London,
United Kingdom EC2A 4NE

020 3051 5216