Multi-failure resilience: reconfiguration and recovery
Reason through interacting failures, common causes and cascades, then rebuild safe service step by step.
Mastery objectives
- distinguish independent failure, common cause and cascade
- reason from life-critical functions rather than isolated equipment
- define a minimum safe configuration and recovery plan
- test resilience with multi-failure scenarios and limited resources
1. The second failure changes the problem
A design may tolerate one failure and become fragile after a second. Two redundant pumps do not protect the function if they share power, cooling or software. Resilience therefore begins with common-dependency mapping and the ability to keep a minimum function alive when the nominal configuration is unavailable.
2. Common causes
Common causes include fire, flooding, contamination, software defects, bad part batches, surges and human error. They can disable several redundant channels at once. Physical separation, technology diversity, version diversity and distributed stocks reduce risk, but each has a cost that should be tied to function criticality.
3. Cascades across systems
A power failure can stop pumps and cooling; rising temperature can then damage batteries or electronics; communications loss makes diagnosis harder. A cascade is a causal chain, not a list of independent failures. Procedures should detect early signals and shed selected loads before propagation removes recovery options.
4. Minimum safe configuration
During a crisis the objective is not immediate restoration of the full settlement. First establish a stable state: refuge volume, controlled atmosphere, minimum water, critical power, communications and medical capability. Industrial and scientific functions remain off until reserves and power margin support their return.
5. Reconfiguration
Reconfiguration uses planned alternate paths: switching a bus, connecting a spare pump, isolating a compartment, moving crew or transferring a load to another controller. These actions should be tested before the emergency. An untested emergency connection may prove inaccessible or incompatible when it is most needed.
6. Diagnosis under uncertainty
After multiple failures, some sensors may also be wrong. Diagnosis cross-checks independent measurements, local observation and system behavior. Teams avoid changing many variables simultaneously because that destroys causal information. When time allows, restore observability first and change one thing at a time.
7. Recovery resources
Repair consumes power, people, tools, spares, time and sometimes a safe atmosphere. Resilience planning reserves these resources. Using every available battery watt to maintain a nonessential load can eliminate the ability to restart a critical pump hours later.
8. Progressive return to normal
After the fault is controlled, loads return in stages. Each stage has a stability criterion. Voltage, temperature, pressure, flow and consumption are monitored before the next function is added. Restarting too quickly can recreate the cascade or drain reserves just after the first recovery.
Deepening: beyond N+1
N+1 protects against loss of one element, not necessarily a zone failure or shared-resource failure. Exercises should include combinations such as one generator lost, cooling limited and one technician unavailable. The goal is to reveal dependencies that simple redundancy counts miss.
Deepening: configuration debt after crisis
Emergency workarounds leave an unusual configuration: temporary cables, disabled protections, borrowed spares. That debt must be inventoried and removed. Otherwise the settlement may look nominal while silently operating with fewer barriers than before.
9. Worked example: recovery power
The refuge requires 22 kW and post-failure generation is limited to 35 kW, leaving 13 kW for recovery. A restart pump needs 9 kW for three hours, leaving only 4 kW for other temporary loads. Starting an 8 kW workshop simultaneously would exceed available power. Recovery therefore needs an explicit power sequence.
10. Exercise
A settlement loses one of two main power converters and then detects a water-loop leak. Define the first three actions, the minimum safe configuration and the order in which noncritical services should return.
11. Reasoned solution
First stop the cascade: stabilize power, isolate the leak, preserve water and refuge capability, then verify measurements. Industrial loads remain off. Once reserves and loop status are known, services return by criticality with margin checks. A common-cause diagnosis should precede restart of the second channel if the same cause could damage it.
12. Mini-project
Build a six-hour multi-failure exercise involving power, ECLSS, mobility and communications. Define injected events, available information, expected decisions, success criteria and the data required for debriefing.
