Writing · Connected products and strategy
The OTA problem: firmware strategy for connected products
Over-the-air update is not just a deployment mechanism. It is product strategy, compliance and risk.
Over-the-air (OTA) firmware update tends to be treated by engineering teams as a deployment mechanism. It is that. It is also a product strategy decision, a compliance management problem and a risk management challenge, and the organisations I have worked with consistently underestimate it until the first deployment cycle makes the gap visible.
The technical infrastructure for delivering updates over the air is the easier part. The harder questions tend not to be answered until they are urgent. Which firmware version is running on which device in which market? Who is responsible for compliance with regulations that changed last month in a jurisdiction that was not on the original roadmap? What does the recovery procedure look like when a bad update reaches ten thousand deployed units before anyone notices?
Why it matters commercially
The commercial case for investing properly in OTA capability is straightforward once the alternative comes into view. A connected device that can be updated in the field has a longer useful life than one that cannot. Features can be added after sale, security vulnerabilities can be patched without a field service visit, and reliability improvements can be deployed to the existing installed base, not only to new units coming off the production line. In a subscription or service revenue model, the ability to keep improving the product for customers who have already bought it is not a nice-to-have. It is the mechanism that makes the business model work.
The inverse is also worth stating clearly, because I have seen it play out more than once. A product that cannot be updated remotely accumulates technical debt in the field. The firmware running on deployed units diverges from current development over time, until supporting the existing base and developing the next version are effectively separate engineering efforts. Field service visits to update firmware are expensive and slow. In some deployment contexts (urban shared-mobility fleets, remote industrial monitoring, agricultural sensors spread across hundreds of hectares) they are operationally impractical at any meaningful scale. The CTO who treats OTA as an engineering convenience rather than a strategic product capability tends to arrive here earlier than expected.
The regionalisation problem
This is where connected products in emerging or fragmented regulatory markets meet a challenge that the software world has largely solved and the hardware world has not. When a product crosses a jurisdictional boundary, the firmware may need to be different. The OTA system then becomes the mechanism by which the product stays compliant, not simply the mechanism by which it stays current.
E-scooters are a clear example, and one I have direct experience with. When I wrote this in 2025, speed limits for shared mobility devices varied by city, by state and by country. Maximum motor power was regulated differently in the EU, the UK, the US and Australia. Geofencing was mandated in some markets and absent in others. Radio frequency approvals meant the wireless stack might need to operate within different parameters depending on where the device was deployed. In the US and Australia specifically, the regulatory picture was fragmented at state or territory level. The specific rules have moved on since then, and will keep moving. That is the point.
In markets where the regulatory framework is still forming, which describes most of the connected mobility space and a good deal of the wider IoT market, the rules change mid-lifecycle. A device that was compliant at launch has to be updated to stay compliant as regulations mature, sometimes on timescales that do not align with normal firmware release cycles. A monolithic firmware binary that cannot be configured for different operational parameters without a full rebuild does not solve this problem. It is a source of escalating operational complexity that becomes harder to manage with every new market the product enters.
The codebase architecture decision
The strategic choice that needs to be made early, ideally before the firmware architecture is committed and certainly before the first multi-market deployment, is how regional variation will be managed in the codebase. There are two broad approaches, each with real trade-offs, and choosing between them after the architecture is established is considerably more expensive than choosing before.
A single codebase with a configuration layer. The firmware contains all capabilities, and a configuration layer determines which are active in a given deployment. Speed limits, power caps, geofencing rules and radio parameters are configuration values rather than compiled-in constants. The OTA system can update configuration independently of firmware, so a compliance change in one market does not require a full firmware release and does not touch devices in other markets.
Separate firmware branches. Maintaining separate branches for different markets or regulatory environments is understandable when the differences between markets are large and structural, but it compounds quickly. Four markets with different regulatory requirements, across three hardware generations, produces twelve firmware variants to maintain, test and deploy. The risk of a variant-specific defect going undetected is proportional to that complexity.
Neither approach is universally correct. The right answer depends on how different the markets actually are, how often the regulations change and how large the installed base in each market is likely to become. What is generally true is that teams who think this decision through before the architecture is committed have considerably more options than those who discover the question when the first international deployment is already being planned.
The update risk
Every OTA deployment carries the risk that the update contains a defect not caught in testing. In software, a bad deployment can typically be rolled back quickly, and the impact is measured in degraded user experience. In connected hardware the situation is different in ways that matter. A firmware update that makes a device behave incorrectly in the field may not be detected until the problem has spread through a significant portion of the installed base. In safety-relevant products (anything with a motor, a battery under active management, or a sensor acting on the physical world) a bad update is not an inconvenience. It is a liability.
Staged rollouts are the standard engineering response, and they are necessary. In my experience they are rarely sufficient on their own, because the observation phase depends on three things: knowing what to look for, having telemetry that surfaces the relevant signals quickly enough to act on, and having a decision process fast enough to halt the rollout before the damage is large. All three need to be in place before the first staged rollout, not designed in response to the first incident that makes their absence visible.
The rollback strategy is equally important and equally often underprepared. Can a device be restored to its previous firmware state over the air? What is the procedure if a device is in a state where it cannot receive an update at all? What does the field intervention path look like when remote recovery has failed? These questions carry real operational cost, and they are much better answered in the architecture phase than in the middle of an incident.
Where this sits in the CTO's work
OTA tends to be underinvested because, from the engineering side, it looks like a deployment concern. Meanwhile the questions that make it genuinely difficult (compliance tracking, regional configuration management, update risk, rollback planning) look like engineering concerns from the product and commercial side. In practice they fall between the two, and neither answers them fully.
I have seen this lead to situations that were genuinely difficult to recover from: a product in market with no clear record of which firmware version is running on which cohort, a regulatory change in a key market that cannot be deployed quickly because the architecture was not built to support it, an update-induced fault discovered after it has already reached most of the installed base.
These failure modes do not appear because the engineering team was careless or the commercial team was inattentive. They appear because the questions that produce them are not naturally owned by either function, and because early in a product programme, when they are cheapest to address, the pressure to ship makes them easy to defer. The CTO is the person best placed to own them, and the right time to own them is before the first device leaves the factory.
Four things worth taking seriously
For boards: treat OTA as a strategic product capability, not an engineering convenience. In a service or subscription model it is what lets the product keep improving after sale.
For engineering leaders: decide how regional variation will be handled in the codebase before the firmware architecture is committed, and certainly before the first multi-market deployment.
For product and compliance teams: know which firmware version is running on which cohort in which market. Without that record, a regulatory change becomes an emergency.
For anyone approving a release process: put the telemetry, the halt decision and the rollback path in place before the first staged rollout, not after the first incident.
I would be interested to hear who owns these questions in your organisation, and whether that ownership was decided deliberately or arrived with the first incident.
© 2025 Catherine Ives-Yim. All rights reserved.