Writing · Connected products and strategy

What makes an embedded system reliable

Embedded reliability is decided early, in hardware, architecture and process, before the field finds out.

Embedded systems are the processing intelligence built into physical devices: the firmware in a connected scooter, the real-time controller in an automotive braking system, the motor control loop in an industrial robot. They are invisible when they work and noticed only when they do not. That is one reason the discipline that produces them is underappreciated relative to its importance.

I have worked in embedded and firmware development for most of my career, most recently on e-scooter platforms where the firmware sits at the intersection of safety-relevant behaviour, BLE connectivity and real-time motor control. The pattern holds everywhere: reliability engineering only becomes visible to everyone else when something has gone wrong in the field.


Types of system and their timing demands

The useful distinctions between embedded systems are about timing and connectivity, not application domain.

Standalone systems perform specific tasks independently, without external dependencies: a digital watch, a basic appliance controller.

Real-time systems divide into hard and soft variants. Hard real-time systems have absolute timing constraints, where a missed deadline is a failure, not an inconvenience: pacemakers, automotive braking systems, motor control loops. A 10-millisecond response that takes 15 milliseconds is not a slow response. It is a wrong response. Soft real-time systems tolerate occasional timing lapses without catastrophic consequences: video streaming buffers, user interfaces, non-safety-critical monitoring.

Networked systems are connected to other devices or infrastructure: home automation, industrial IoT sensors, vehicle telematics.

Mobile systems balance processing capability against power and size: smartphones, wearables, portable medical devices.

Most modern connected products span more than one of these categories at once, which is where the interesting engineering problems live. Knowing where a system sits in this taxonomy before architecture decisions are made is not background knowledge. The wrong classification produces the wrong design.


Reliability starts below the firmware

Discussions of embedded reliability tend to focus on the firmware layer: test coverage, safe language subsets, defensive coding. That focus is necessary but not sufficient. In a connected electro-mechanical product, the firmware runs on hardware that will experience voltage transients, thermal cycles, mechanical stress and the accumulated effects of years in the field. The reliability of the finished product is the product of the reliability of every layer. The hardware layer is where the most consequential decisions are made earliest, and where they are hardest to reverse once the design is fixed.

De-rating

The most consistent hardware reliability practice is also the one that most surprises engineers from a software background: specifying components to operate well below their rated maximum. A capacitor rated at 50V in a 24V circuit is not over-specified. It is de-rated, and that margin is not waste.

Component failure rates are not linear with operating stress. They rise sharply as conditions approach rated limits, and thermal and electrical stress are the dominant drivers of long-term degradation. Standard de-rating practice targets operating voltage at around 80% of rated value and junction temperatures well below the component maximum, typically aiming for 85°C on a component rated at 125°C or higher.

The bathtub curve describes component failure rates over time: early failures declining to a long steady-state period, then rising as wear-out begins. De-rating shifts it substantially. The useful-life phase extends and wear-out onset is deferred. In a product designed for five to ten years in the field, consistent de-rating is frequently the difference between a product that reaches end of life gracefully and one that accumulates progressive field failures throughout its service life.

Self-monitoring

A system that can detect its own fault conditions before they become customer-visible failures is qualitatively different from one that cannot.

Watchdog timers are the most familiar mechanism: a hardware circuit that resets the processor if the firmware fails to signal it within a defined window. It catches the class of fault where the software has entered an unexpected state and cannot recover itself. The same principle extends across the system. Voltage rail supervisors hold the processor in reset if any supply rail falls outside its acceptable range, preventing code from executing on unstable power and producing undefined behaviour. Thermal monitoring can reduce load or start a controlled shutdown before junction temperatures reach the point where degradation accelerates. CRC verification on stored calibration and configuration data catches memory corruption that would otherwise produce subtly wrong behaviour with no detectable error event.

I have seen the absence of these mechanisms produce field failures that took weeks to diagnose. The failure was intermittent, the system had no record of its own state when it occurred, and the only evidence was the customer's report that the product sometimes behaved strangely. Designing self-monitoring in at the architecture stage is not expensive. Reconstructing the failure mode without it is.

Redundancy

For most commercial connected products, full hardware redundancy is neither practical nor cost-justified. The component cost, design complexity and physical space make it a solution for environments where the criticality of the function warrants it: safety-critical infrastructure, high-availability industrial systems, medical devices, aerospace. There, hardware redundancy is the architectural answer to the requirement that a single component failure must not produce a system failure.

Active redundancy keeps both paths operating and switches on failure detection. Standby redundancy keeps one path dormant until needed, which reduces power consumption but adds switching complexity and the risk that the backup does not activate cleanly when it matters.

A flight monitoring system I designed shows the approach in a context where the cost was justified. Two independent units, physically separated, each took live data simultaneously and continuously cross-checked each other. If one lost power or its connection dropped, the surviving unit alerted the system administrator and absorbed the full monitoring load without interruption. When the failed unit recovered, it resynchronised against the data it had missed from its peer and rejoined the active pair.

The design rested on two principles: the system administrator should be the first to know, not the last, and recovery should be automatic and self-healing rather than needing intervention. In practice the system has not missed a data point since deployment, despite occasional connection drops and power failures on both units.

Failure mode analysis

None of these practices is applied reliably without a systematic method for identifying what can fail. Failure Mode and Effects Analysis (FMEA), and its extension FMECA, which adds criticality scoring, is the discipline that forces that enumeration. It is a structured walkthrough of the design at component level, asking for each item what the possible failure modes are, what their effects on the system would be, how severe and detectable they are, and what controls exist.

The value is not mainly in the document it produces. It is in the thinking it requires: the conversations between hardware engineers, firmware engineers and the people who will service the product in the field, which surface failure modes none of them would have found working separately. Applied early, FMEA changes what gets built. Applied late, or not at all, it becomes the forensic tool used to explain why the product failed in ways that were, in retrospect, entirely predictable.


Language choices matter more than they seem

The language question comes up early in any embedded project and tends to be underestimated. It shapes not only how code is written but how easy it will be to audit, test and reason about for correctness over the life of the product.

C. My default for most embedded work has been C, and broadly it still is, though my position is less settled than it used to be. The depth of tooling, the library ecosystem and the density of accumulated expertise make C hard to argue against for most production embedded work, particularly where the team already has strong C skills and the schedule has no room for a substantial learning curve.

Ada. I have more respect for Ada than most of the engineers I talk to. It is regularly dismissed as a relic of defence and aerospace procurement, which is not entirely unfair as history. But its emphasis on type safety, and its habit of catching at compile time the errors C will happily let through into production, has always struck me as a serious engineering position rather than a historical accident. I would not necessarily reach for it first on a commercial consumer product. For safety-critical work where the cost of a field failure is genuinely high, I think it deserves more serious consideration than it usually gets in the conversations I have been part of.

Rust. This is the one I am watching with increasing interest. Memory safety guarantees without a garbage collector are not trivial in resource-constrained embedded work. The embedded Rust community has matured quickly enough to produce real production experience, not just enthusiastic conference talks, which suggests it is past the point of being an interesting experiment. I have not yet run a full production embedded project in Rust, but I am closer to recommending it for new projects where long-term maintenance matters than I was two years ago.

Python. Python's role at the embedded end of the spectrum has expanded in ways that would have seemed implausible when I started in this field. I use it regularly for tooling, test scaffolding and host-side simulation of embedded behaviour, and MicroPython has carved out a real space on microcontrollers for prototyping. Its runtime efficiency is still a constraint in many embedded contexts, but in the higher tiers of an application stack, where development speed matters more than every last cycle, it earns its place.


Operating systems, or the deliberate lack of one

In my experience, whether to use an operating system at all is decided too quickly. Teams either adopt an RTOS reflexively because that is what they used before, or run bare-metal because the application seems simple enough that an OS feels like overhead. The timing requirements are almost always the right starting point. The essential question is whether deterministic behaviour under load is required, or whether a richer software environment is needed and slightly less predictable latency is acceptable.

RTOS. For the safety-relevant real-time control I work with on e-scooter platforms, an RTOS is generally the right answer. FreeRTOS in particular is well understood, well documented and well certified against standards, and it has become something of a default for good reason. Its determinism is not hypothetical. It is what makes it possible to write code that provably behaves within its timing budget under load, which matters when the system interacts with physical safety.

Embedded Linux. This is more interesting than it used to be for complex products, particularly where the application needs substantial networking or a rich driver ecosystem. The standard Linux scheduler is not a real-time scheduler, but the PREEMPT_RT patches have improved latency enough that it is a reasonable choice where hard real-time guarantees are not the absolute priority. The trade-off is complexity: Linux brings a great deal of capability and a great deal of surface area to manage.

Bare-metal. This remains the right answer for genuinely latency-critical, resource-constrained applications where an OS would be overhead without benefit. I reach for it less often now than earlier in my career, but for simple controllers or specialised signal-processing tasks there is still an elegance to having complete control over every cycle.


The standards landscape

Taken seriously, the standards that govern safety-critical embedded development are an enormous amount of work. They also represent the accumulated knowledge of a great many engineers who have dealt with failures in the field. Even on projects where strict compliance is not formally required, I have found them worth reading carefully rather than treating as a compliance checkbox somebody else handles.

ISO 26262 is the one I encounter most directly, given the automotive and micro-mobility context of my current work. It is increasingly referenced in e-mobility contexts that technically sit outside its original automotive scope, which any firmware developer working on something that moves people should know. IEC 61508, the broader functional safety standard from which 26262 derives, is worth understanding for the model it gives of why these standards are structured as they are, not just what they require.

MISRA I have mixed feelings about in practice. The intent is sound: the guidelines define a constrained subset of C that avoids constructs associated with unreliable or undefined behaviour, and compliance catches real things. But enforcing MISRA strictly on a mature codebase is expensive, and the constraints occasionally push towards code that is technically compliant but harder to read than the version it replaced. The worst case is a team forced to refactor working, well-understood code into a form that satisfies the analyser but is less legible to the engineers who maintain it. Compliance has been achieved. The practical reliability of the development process has not obviously improved.

I apply MISRA with judgement rather than mechanically, and I take it more seriously as a baseline where I am working in a supply chain with contractual compliance requirements. The principle that matters extends well beyond any individual standard: process traceability and risk management matter as much as the technical output itself.


Where reliability is actually won or lost

The technical choices above matter, but in my experience the practices around them are where projects diverge most sharply. They are also where projects that started with good intentions get into trouble at the end of a development cycle, when the schedule is tight and the temptation to cut corners is strongest.

Testing. Rigorous testing sounds obvious, but the gap between embedded projects that take it seriously and those that nominally do is large. Hardware-in-the-loop testing, host-side unit testing via cross-compilation, and proper integration testing across the software and hardware boundary are all achievable now in ways they were not a decade ago. The investment in test infrastructure pays back disproportionately in the later stages, when changes need to be made quickly and confidently. Continuous integration (automated builds, automated test runs, clear gates before code progresses towards release) tends to make embedded teams faster in the medium term, even though it needs genuine upfront investment and some cultural change in teams that have not done it before.

Code review and static analysis. This is the pairing I find most reliably catches what testing misses. Tools such as PC-lint, Polyspace and Coverity find categories of defect (unreachable code, potential buffer overruns, subtle undefined behaviour) that a test suite will only find when exactly the right code path is exercised at exactly the right moment.

Documentation and traceability. This is the last piece, and the one I see deprioritised most often under schedule pressure. That is particularly unfortunate in embedded work, where the maintenance phase can last a decade or more and the engineers who wrote the original code are rarely still available to explain why decisions were made as they were.


What this means for technical leadership

A CTO in a connected-product business does not need to be a fluent embedded engineer. They do need to understand the cost structure of embedded work well enough to decide where to invest and what to protect. Embedded teams that are under-resourced on testing tend to produce products that are impressive in the lab and unreliable in the field. The field failures arrive after the product has shipped, when the cost of fixing them has multiplied significantly.

Test infrastructure is a leadership decision as much as an engineering one. Hardware-in-the-loop rigs, simulation environments and automated regression suites need upfront investment that competes with delivery pressure. The CTO who understands why that investment exists can protect it when schedules tighten. The CTO who cannot make that case will find the test infrastructure is the first thing cut and the last thing blamed when the product fails.

Language and toolchain choices made early in a product's life persist for years. A team locked into a toolchain poorly suited to the product's safety requirements will pay for it for the whole life of the product. Understanding the trade-offs well enough to ask the right questions in the architecture conversation (C versus Ada versus Rust, bare-metal versus RTOS versus embedded Linux) is not a luxury for a CTO in a connected-product business. It is table stakes.


Four things worth taking seriously

For boards: field failures arrive after shipping, when they cost most to fix. Ask whether test infrastructure is funded and protected, not just whether the product works in the lab.

For engineering leaders: treat de-rating, self-monitoring and failure mode analysis as architecture decisions made early, not features added once the design is fixed.

For founders: language, toolchain and operating system choices will outlive most of your roadmap. Make them against the product's timing and safety requirements, not the team's habits.

For anyone writing embedded code: read the relevant safety standards even where compliance is not required. They record failures other engineers have already paid for.


Which reliability practice has saved you most in the field, and which one did you only adopt after a failure made the case for it?

© 2024 Catherine Ives-Yim. All rights reserved.

Catherine Ives-Yim

Catherine Ives-Yim

Chartered Engineer and independent technical adviser, with a lifetime at the bleeding edge of embedded systems, connected products, data platforms and AI-assisted engineering, who has advised clients across the UK, Europe, the Middle East, the Far East, North America and Africa. Based in Leeds.