MARS BIBLE — RISK & RESILIENCE DOSSIER
Common-cause failure: why two backup systems can fail together
Redundancy protects against independent failures; it is much weaker when backups share the same vulnerability.
A Mars settlement could multiply duplicate equipment and remain fragile. If two pumps share a design defect, two computers share the same software or three power lines pass through one room, one cause can defeat all redundancy.
Redundancy does not mean independence
A common-cause failure produces several malfunctions from one event or process. NASA notes that a common cause can eliminate an entire redundant set and that additional identical units then provide diminishing reliability gain.
The first question for any N+1 architecture is therefore: what is genuinely different? Common power, software, supplier, location, cooling and maintenance are hidden dependencies.
Teaching example: three identical pumps
Assume three pumps, each able to circulate water. If random failures are independent, three units provide strong tolerance. But if one unsuitable lubricant, contaminant or design defect attacks all three, the third pump is not a third barrier; it repeats the same vulnerability.
Diversity can therefore matter as much as quantity: different technology, supplier, physical separation or a simpler repairable fallback.
Physical, digital and human common causes
Flooding, fire, surge or impact can damage several nearby units. A software defect can act simultaneously on computers running the same code. An incorrect maintenance procedure can insert the same defect into multiple units.
Crew can also become a common cause through fatigue, training gaps or ambiguous documentation that makes multiple operators repeat the same incorrect action.
Diversity + separation + repairability
Risk reduction uses justified component diversity, physical separation, avoidance of unnecessary common buses, protection from external events and the ability to fall back to a simpler mode.
NASA has also studied diverse redundancy for life support: a different backup can survive the cause that defeats the primary equipment. On Mars this should be coupled to local repair.
Common-cause audit
Draw each vital function as a dependency graph and mark shared nodes: power, software, room, cooling, network, reference sensor, maintenance and consumables. Then remove each shared node in turn.
- do not count shared-path redundancy twice;
- physically separate critical functions;
- provide a diverse backup path;
- test common software and maintenance errors;
- retain manual or simplified fallback when it adds real independence.
Small reliability calculation: independence versus common cause
The next engineering step is to identify which part of failure probability is independent and which part is shared. That is why physical separation and technological diversity can matter more than a third identical unit.
The same calculation should always be repeated with an adverse assumption, followed by the question: what real measurement could confirm or reject that assumption? A calculation teaches as much through its limits as through its numerical result.
The dossier keeps measured data, published values, design assumptions and teaching scenarios visibly separate. Mixing those statuses would create false precision.
From calculation to action
Measurement itself needs resilience: backup sensing, independent confirmation, calibration range and a defined response when data are missing. An alarm with no strategy for sensor failure can increase risk.
What must be tested before depending on it
Run the scenario using real hardware or a representative twin, then repeat it with one additional failure. Measure diagnosis time, human errors, consumable use and ability to return to nominal conditions.
Results then update inventory, procedures and design. Safety becomes a learning loop rather than a document frozen before departure.
Questions never to skip
- What event actually starts the failure chain?
- Which functions are lost immediately, then after 10 minutes, 1 hour and 24 hours?
- Which redundant units still share power, software, location or maintenance?
- What degraded mode remains genuinely habitable?
- What must be repairable locally without waiting for Earth?
This dossier in the settlement
Scientific and technical sources
The sources below support the physical phenomena and safety building blocks; settlement architecture remains an explicitly identified prospective synthesis.
- NASA NTRS — Common Cause Failures Dominate and Defeat Redundancy
- NASA NTRS — Diverse Redundant Systems for Reliable Space Life Support