Writing · Archive
From Silicon to Society
I have spent most of my career at the intersection of embedded systems and organisational leadership, which means I have spent a good deal of time thinking about both kinds of systems simultaneously. The parallels are more instructive than they might appear.
A multi-processor embedded system running a real-time operating system solves a fundamental coordination problem: multiple concurrent processes need to share resources, respond to events, and produce coherent output without interfering with each other in ways that break the overall behaviour. The operating system provides scheduling, priority management, and the synchronisation primitives that prevent race conditions and priority inversions from turning well-written individual processes into an unpredictable system.
Organisations face the same problem and, more often than not, solve it less well.
Race conditions and priority inversions
In embedded systems, a race condition occurs when two processes access a shared resource without proper synchronisation, and the outcome depends on the order of execution rather than the logic of either process individually. Both processes may be doing exactly what they were designed to do. The problem is at the interaction point.
In organisations, this happens constantly. Two teams making decisions that are individually sensible but mutually contradictory, each unaware of what the other is doing until the results collide downstream. The sales team commits to a delivery timeline without checking with engineering. The product team deprioritises a feature without telling the customer who was relying on it. No individual acted wrongly. The system produced the wrong output because there was no synchronisation mechanism at the interaction point.
A priority inversion in embedded systems is when a high-priority task is blocked waiting for a resource held by a low-priority one, effectively causing the urgent work to be delayed by work that matters less. This is extremely common in organisations too, most visibly when senior leadership cannot make a decision because it is waiting on analysis that is stuck behind lower-priority work in someone’s queue. The symptom is that urgent things move slowly, which looks like a management problem but is actually a scheduling problem.
What the analogy is useful for
Framing organisational problems in systems terms is not just a parlour trick for engineers. It changes what to look for when things go wrong. Instead of asking which person made the wrong decision, the question becomes where the synchronisation is missing. Instead of blaming a team for moving too slowly, the question becomes what resource contention or priority inversion is blocking them.
It also suggests specific interventions. A race condition needs a synchronisation primitive: a meeting, a shared document, a decision authority that both parties consult before acting. A priority inversion needs a scheduling change: explicitly freeing the high-priority task from its dependency rather than just adding urgency to the blocked work.
The embedded systems approach to reliability, designing for failure modes, building in redundancy for critical functions, testing interactions rather than just individual components, applies directly to the organisational design problem. Most team failures are integration failures. Most integration failures are design failures. Understanding the system before something goes wrong is more effective than diagnosing it after.
Fault tolerance and redundancy
In embedded systems, fault tolerance is designed in from the start for critical functions. A system that cannot tolerate the failure of a single component is not a resilient system; it is a system waiting for the right failure mode. The design principles are well-established: redundancy for critical functions, graceful degradation when components fail, watchdog timers to detect and recover from stalls, and clear separation between the failure domain of one component and the operation of everything else.
Organisations that have thought about this tend to be more robust under pressure. Single points of failure, one person who holds a critical relationship, one process that nobody else understands, one system that everything depends on but nobody has tested under load, are structural risks that show up as crises rather than as design choices. The crisis always looks sudden. The underlying fragility was always there.
The redundancy question for an organisation is not “do we have a backup?” It is “have we tested the backup?” A redundant system that has never been exercised is not actually redundant. It is a component that has been installed and assumed to work. The organisations that discover this at the worst moment are the ones that treated resilience as a documentation exercise rather than an operational practice.
What this means for a technical leader
The value of systems thinking for a CTO is not that it produces better org charts. It is that it provides a diagnostic vocabulary for problems that would otherwise be attributed to people rather than structure.
When a team is consistently slow, the natural response is to question the team. The systems question is different: what are they waiting for, and why is that dependency not being managed? When two parts of the organisation keep producing conflicting outputs, the natural response is to call a meeting. The systems question asks why there is no synchronisation mechanism between them, and what it would take to build one.
The embedded engineer’s habit of asking “what does this system do when this component fails?” is one of the most useful habits to carry into leadership. The answer to that question, applied to an organisation, reveals more about its actual resilience than any number of risk registers.
© 2024 Catherine Ives-Yim. All rights reserved.