Writing · Connected products and strategy

The metrics that hide the failure

The interesting failures sit between domains. Measure the boundaries, not just the parts.

The product was a low-cost intra-product communications system: the internal bus that lets the modules of a device talk to each other. It needed to be reliable. The software engineers built in error correction and packet retry, which is standard practice, and in this case they did it well. Packets that arrived damaged or incomplete were caught, corrected and retransmitted. The system looked robust because it was robust. Every test passed. Everyone moved on.

The hardware had a physical layer problem. The timing of the protocol was marginal in ways that showed up as elevated raw error rates at the electrical level. Not dramatically elevated, but enough to be visible on the bench if anyone had looked. Nobody looked, because the software was quietly fixing everything. The error correction did not report what it was correcting. It just corrected it. That silence was the problem.


What the field revealed

In the field, across a range of operating temperatures the bench had not replicated, the hardware error rate climbed. For a while the software kept up. Then it could not. The product started producing random, intermittent failures, the hardest kind to diagnose, because the symptoms appeared in the software domain and the root cause was in the hardware domain. Both teams had done their jobs. The hardware engineers had met their spec. The software engineers had built reliable error handling. Nobody had been watching the boundary between them.

That kind of failure has a name in complex systems: inter-domain masking. The strong domain hides the weak one until the weak one breaks. By then the product has shipped, the warranty clock is running, and the team is diagnosing in the field rather than in the lab. The metric that would have caught it, raw physical-layer error rate during bench testing, was never logged because there was no obvious reason to log it. The software had made it look unnecessary.


Failures live at the boundaries

I have built products across defence, healthcare and commercial settings, most of them combining electronics, software and mechanical systems. The consistent pattern is that the interesting failures sit at the boundaries between domains, not inside them. A product that scores excellently on software metrics while its thermal management is marginal is not an excellent product. It is a product with a problem that has not declared itself yet.

What follows is the metric set I use to avoid that. It is a diagnostic, not a checklist. Pick what fits the product class. Track across both individual products and time. The trend lines are almost always more useful than the snapshot values.


Electronic systems

Power consumption. How much the device draws under real operating conditions, not lab conditions. It drives battery life, thermals and efficiency. The gap between bench measurement and field reality is where most surprises live.

Heat dissipation. How well the product manages its own heat. Bad thermals are not just a performance problem. They shorten the operating life of every other component nearby. A product that runs hot shortens its own lifespan, quietly and consistently.

Signal integrity. The quality of the signal across the electronic path. Marginal signal integrity is particularly insidious because it is invisible in normal conditions and catastrophic in edge ones. The failure tends to appear in the field, under specific combinations of temperature and load that were never quite reproduced in testing.

Component lifespan. Expected operating life before failure, under realistic conditions. It drives warranty cost, field-service cost and the credibility of any reliability claim made to a customer.


Mechanical components

Wear and tear. The rate of degradation under repeated use. Cheap to measure on a bench. Expensive to learn about in the field, after the customer has already experienced it.

Environmental tolerance. Temperature, humidity, UV, salt, dust: whatever the product will actually see. Not what the spec says it should tolerate, but what the operating environment actually delivers. The two are frequently not the same.

Vibration resistance. How much vibration the product can sustain without losing function or structural integrity. It is critical for anything that moves or travels. The failure modes here often appear at specific resonant frequencies that only surface under real operating conditions.

Material fatigue. How materials behave under prolonged cyclic stress. This is rarely a dramatic failure. It is a slow degradation that only becomes visible when it crosses a threshold, which is exactly why it needs measuring before that point.


Software

Performance. Response time, throughput and resource utilisation under realistic load. The question is not whether the software performs in the demo environment. It is whether it performs on the hardware it has, at the scale it will actually reach.

Bug rate and defect density. The frequency and severity of defects over time. A declining rate signals that the codebase is stabilising. A flat or rising rate under continued development signals that it is not.

Security posture. Open vulnerabilities, patching cadence, CVE backlog. For any connected product this is not a separate concern from product health. It is product health. A single unpatched vulnerability in a shipped device is a liability that sits on the balance sheet whether or not it appears in any report.

User satisfaction. Gathered through structured feedback and usability testing, not inferred from low support ticket volumes. A product can hit every other metric on this list and still be exhausting to use. That is a product health problem.


The integrated view

This is the section that matters most and gets measured least.

The individual domain metrics are necessary but not sufficient. A product that scores well on software, adequately on electronics and poorly on mechanical is not a product with one weak area. It is a product whose weakest area sets its real-world reliability score. The customer does not experience the software in isolation from the chassis. They experience the product.

System reliability (MTBF and MTTR). The overall dependability story. These numbers should be derived from the domain metrics, not measured independently of them. If MTBF is better than the component lifespan data would predict, the product is not healthy. There is a measurement gap.

Interoperability. How the product works with the systems around it. Almost nothing ships standalone any more. A product that functions perfectly in isolation but behaves unpredictably in the customer's actual environment is not a healthy product.

Compliance and standards. Regulatory and industry standards across all three domains. The cross-domain cases, where a mechanical decision affects an electrical certification or a software update touches a safety classification, are where compliance gaps tend to hide.

Lifecycle and sustainability. Energy efficiency, repairability, end-of-life recyclability. This has moved from moral obligation to market signal and, in a growing number of categories, to regulatory requirement. It belongs in the health scorecard.


The board-level translation

A board understands financial health metrics intuitively: margin trend, cash runway, customer acquisition cost. Directors use these to make decisions about risk and investment without needing to understand the underlying detail.

Technical health metrics serve the same function. A CTO who can present a clean trend-line view across these four areas, and explain in plain commercial terms what a deteriorating signal integrity score or a rising defect density means, gives the board something it almost never has: a way to assess technical risk before it becomes a financial event.

The alternative is what happens in most organisations. Technical debt and product risk stay invisible to the board until something fails publicly, a recall is required, or a key customer churns because the product stopped being reliable. By that point the metric that would have caught it early is obvious in retrospect. It always is.

Run these quarterly. Use the trend lines. Build the vocabulary to translate them upward. The goal is not a spreadsheet. It is a shared language for technical risk that the whole leadership team can use before the problem arrives.


Four things worth taking seriously

For boards: ask for technical health as quarterly trend lines, presented alongside the financial metrics, so that technical risk is visible before it becomes a financial event.

For engineering leaders: log what your error handling and other compensating mechanisms are correcting. A strong domain that silently absorbs a weak one is hiding a failure, not preventing it.

For hardware and software teams: someone has to own the boundary between you. Measure under the temperatures, loads and environments the product will actually see, not only those the bench reproduces.

For anyone reading a reliability figure: check it against the domain data underneath. If system reliability looks better than component lifespan would predict, the gap is in the measurement.


I would be interested to hear which metric in your own products you suspect is being quietly compensated for by another domain.

© 2024 Catherine Ives-Yim. All rights reserved.

Catherine Ives-Yim

Catherine Ives-Yim

Chartered Engineer and independent technical adviser, with a lifetime at the bleeding edge of embedded systems, connected products, data platforms and AI-assisted engineering, who has advised clients across the UK, Europe, the Middle East, the Far East, North America and Africa. Based in Leeds.