MARS BIBLE — RISK & RESILIENCE
Fatigue, human error, and bad procedures: the risk that crosses every system
A Mars settlement can have sound hardware and still fail if fatigue, cognitive overload, or a poorly designed procedure pushes the crew toward the wrong decision.
Human risk is inseparable from technical risk: sleep debt, interruptions, alarm floods, ambiguous procedures, and time pressure can turn a recoverable anomaly into a cascade. This chapter examines schedules, interfaces, checklists, training, and stop-work authority as engineered barriers.
1 — “Human error” is not a final cause
Writing “operator mistake” ends analysis too early. Why did the interface, procedure, schedule, or organization allow the error to reach the system?
Human factors treats people as system components with predictable capabilities and limits.
2 — Fatigue and sleep debt
Fatigue reduces vigilance, working memory, and the ability to detect subtle anomalies.
A settlement cannot depend on permanent heroism to compensate for chronic understaffing.
3 — Procedures that are too long or ambiguous
A checklist should highlight decisions and traps, not bury users in unreadable pages.
Critical steps benefit from measurable criteria and independent confirmation.
4 — Interfaces that invite the wrong action
Identical adjacent controls, ambiguous units, or alarm floods can create predictable errors.
Human-machine interface is a safety barrier just like a valve or fuse.
5 — Reporting culture
A crew that hides near-misses loses its best preventive data.
The system should allow error reporting without making every report automatically punitive.
6 — Train degraded scenarios
Operators need practice with incomplete data, communication delay, and unavailable equipment.
Training also reveals procedures that cannot be performed in real time.
Learning calculation: turn a reserve into decision time
LEARNING CALCULATION — ASSUMPTIONS ARE EXPLICIT
LEARNING ASSUMPTION: a procedure has 24 steps and a critical response must finish in 6 min.
Average time if steps were equal: 360 s ÷ 24 = 15 s/step. If several steps need 60 s, the procedure is probably unrealistic for that emergency.
This does not measure cognitive load; it reveals a timing inconsistency to test in simulation.
Decision questions specific to this risk
- What fatigue level is predictable during the operation?
- Can the procedure really be completed in available time?
- Does the interface make two dangerous actions easy to confuse?
- Do priority alarms remain visible during alarm floods?
- Are near-misses recorded and turned into improvements?
Design for imperfect humans
On a long mission it is unrealistic to assume that every operator will always be rested, focused and perfectly informed. A good interface makes dangerous actions harder to perform accidentally, distinguishes normal from degraded states and avoids requiring human memory to hold dozens of parameters at once.
This does not remove crew responsibility. Safety comes from a combination of competence, procedure, ergonomics, automation and cross-checking. If one ordinary mistake can destroy a life-critical system, the architecture itself deserves review.
Measure workload before the crisis
Operational logs can reveal accumulating overtime, growing alarm counts or repeatedly deferred maintenance. These weak signals should be treated as risk indicators, not as proof of dedication. An overloaded organization gradually becomes less able to detect its own mistakes.
Mars planning therefore needs human margin just as it needs energy or water margin. Sleep opportunity, duty rotation and the availability of a relief operator are safety resources.
Main primary sources
Connect to other dossiers
Fatigue turns small ambiguities into failures
A procedure that looks perfectly clear at a desk can become dangerous at 3 a.m. after a long EVA and an unexpected alarm. Fatigue reduces attention, working memory and the ability to detect inconsistency. Prevention therefore acts on work organization as well as individual skill.
Interfaces and procedures should remain understandable under load. Consistent codes, explicit confirmations, visible critical steps and limits on irreversible actions reduce the chance that one simple error becomes catastrophic.
A checklist is not an excuse to stop understanding
Checklists are powerful against omission, but dangerous when followed without understanding system state. A degraded situation can leave the assumed scenario and make a normally safe step inappropriate.
Cross-training and exercises should therefore teach the intent behind the procedure: what each step protects, which signs show that an assumption is no longer valid, and when to stop and reassess.
Experience must rewrite procedures
After an incident, the objective is not merely to identify “human error.” Teams should ask why the error was possible: confusing interface, alarm overload, missing information, excessive workload or contradictory procedures. A human cause may be a symptom of system design.
A Mars settlement needs a fast local learning loop. Each near miss becomes data that can change an interface, work order, redundancy or cross-check rule before a more severe event occurs.