Mission team and operations: decide far from Earth
Operate a distant mission with incomplete information
Starting question — How do people make safe decisions when Earth cannot answer in real time?
Intuition. A Mars crew needs roles, handovers, abort criteria and communication habits that preserve intent across delay, fatigue and changing system state.
- Explain the governing idea before calculating.
- Name the units, evidence source and operational boundary of every key quantity.




Many Earth-orbit operations can rely heavily on a real-time ground team. Mars breaks that assumption: light-time alone introduces minutes of delay and outages are possible. Mission teams need intent-based work, complete information packages, adaptable procedures and explicit local authority. This module treats operations as an engineered subsystem.
Zero-prerequisite concepts
operational role
Definition. An operational role links responsibility, authority, information access and competence to a person or function.
Example. The EVA lead may have authority to abort an excursion when predefined safety criteria are met.
Pitfall. A title without decision authority or required information is not a usable operational role.
For each decision, identify who can decide, what information they see and who must be informed.
Guided exercise — operational role
In a Mars mission scenario, identify one situation in which “operational role” changes an engineering or operational decision. State the evidence you would inspect, the mistake you must avoid, and one independent check you would perform before accepting the decision.
Detailed correction — operational role
Core meaning. An operational role links responsibility, authority, information access and competence to a person or function.
Mission example. The EVA lead may have authority to abort an excursion when predefined safety criteria are met.
Error to reject. A title without decision authority or required information is not a usable operational role.
Independent check. For each decision, identify who can decide, what information they see and who must be informed.
- Quantification
- Use the physical unit that belongs to operational role when it is quantitative; if it is qualitative, do not invent a numerical unit.
- Verification
- Compare the conclusion with the mission example, the stated pitfall and the mental check before using operational role operationally.
briefing
Definition. A briefing aligns the team on objective, current state, hazards, roles, success criteria and abort conditions before action begins.
Example. Before a rover sortie, the team reviews route, weather, suit status, communications windows and turnaround criteria.
Pitfall. A briefing should not become a one-way recital that hides unresolved questions.
At the end, each participant should be able to state the mission intent and the conditions that trigger a stop.
Guided exercise — briefing
In a Mars mission scenario, identify one situation in which “briefing” changes an engineering or operational decision. State the evidence you would inspect, the mistake you must avoid, and one independent check you would perform before accepting the decision.
Detailed correction — briefing
Core meaning. A briefing aligns the team on objective, current state, hazards, roles, success criteria and abort conditions before action begins.
Mission example. Before a rover sortie, the team reviews route, weather, suit status, communications windows and turnaround criteria.
Error to reject. A briefing should not become a one-way recital that hides unresolved questions.
Independent check. At the end, each participant should be able to state the mission intent and the conditions that trigger a stop.
- Quantification
- Use the physical unit that belongs to briefing when it is quantitative; if it is qualitative, do not invent a numerical unit.
- Verification
- Compare the conclusion with the mission example, the stated pitfall and the mental check before using briefing operationally.
handover
Definition. A handover transfers responsibility together with system state, open risks, assumptions and pending actions.
Example. The night shift records an intermittent CO₂ scrubber alarm and the temporary workaround for the next crew.
Pitfall. Passing only a task list can lose the reasoning and uncertainty behind those tasks.
Ask whether the incoming operator can explain what changed, what remains uncertain and what must happen next.
Guided exercise — handover
In a Mars mission scenario, identify one situation in which “handover” changes an engineering or operational decision. State the evidence you would inspect, the mistake you must avoid, and one independent check you would perform before accepting the decision.
Detailed correction — handover
Core meaning. A handover transfers responsibility together with system state, open risks, assumptions and pending actions.
Mission example. The night shift records an intermittent CO₂ scrubber alarm and the temporary workaround for the next crew.
Error to reject. Passing only a task list can lose the reasoning and uncertainty behind those tasks.
Independent check. Ask whether the incoming operator can explain what changed, what remains uncertain and what must happen next.
- Quantification
- Use the physical unit that belongs to handover when it is quantitative; if it is qualitative, do not invent a numerical unit.
- Verification
- Compare the conclusion with the mission example, the stated pitfall and the mental check before using handover operationally.
abort criterion
Definition. An abort criterion is a measurable condition that requires stopping, retreating or switching to a safer plan before recovery margin disappears.
Example. A surface team may turn back when oxygen reserve reaches a specified threshold or communications redundancy is lost.
Pitfall. “Abort if it feels unsafe” is not enough for time-critical operations.
The criterion should be observable, linked to a safe alternative and defined before pressure or fatigue distorts judgement.
Guided exercise — abort criterion
In a Mars mission scenario, identify one situation in which “abort criterion” changes an engineering or operational decision. State the evidence you would inspect, the mistake you must avoid, and one independent check you would perform before accepting the decision.
Detailed correction — abort criterion
Core meaning. An abort criterion is a measurable condition that requires stopping, retreating or switching to a safer plan before recovery margin disappears.
Mission example. A surface team may turn back when oxygen reserve reaches a specified threshold or communications redundancy is lost.
Error to reject. “Abort if it feels unsafe” is not enough for time-critical operations.
Independent check. The criterion should be observable, linked to a safe alternative and defined before pressure or fatigue distorts judgement.
- Quantification
- Use the physical unit that belongs to abort criterion when it is quantitative; if it is qualitative, do not invent a numerical unit.
- Verification
- Compare the conclusion with the mission example, the stated pitfall and the mental check before using abort criterion operationally.
Calculation laboratory — formula, units, inverse check and limits
Quantitative mini-lessons
Minimum round-trip latency
- 1 — Concrete question
- What does “t_roundtrip = 2×t_oneway” compute in “Minimum round-trip latency”?
- 2 — Intuition without symbols
- A response from Earth cannot return until the message travels out and the reply travels back.
- 3 — Quantities
- t_roundtrip: round-trip delay; t_oneway: one-way delay
- 4 — Formula
- t_roundtrip = 2×t_oneway
- 5 — Read aloud
- Read “t_roundtrip = 2×t_oneway” by naming every operation, subscript and grouping explicitly.
- 6 — Symbols and meaning
- t_roundtrip: round-trip delay; t_oneway: one-way delay
- 7 — Pronunciation
- The “Read aloud” line above is the oral reference for “Minimum round-trip latency”. Any subscript, exponent or grouping that changes the meaning of the relation should be spoken explicitly.
- 8 — Units
- times in the same unit
- 9 — Convention
- For “Minimum round-trip latency”, substitute values without changing the reference frame, time basis, system boundary or sign convention halfway through the calculation. Stated units: times in the same unit.
- 10 — Why this operation
- In “Minimum round-trip latency”, multiplication combines the factors that directly build the requested quantity; the factors must describe the same case.
- 11 — Assumptions
- The relation “t_roundtrip = 2×t_oneway” applies here only to the scenario described by the card. Inputs must be mutually consistent and satisfy the physical assumptions associated with “Minimum round-trip latency”.
- 12 — Unit check
- times in the same unit Verify that dimensional reduction reaches the unit of the requested output.
- 13 — Numerical case
- With t_oneway = 12 min, t_roundtrip = 24 min.
- 14 — Why the calculation works
- The numerical case applies “t_roundtrip = 2×t_oneway” directly to the stated values. The calculation is meaningful because the quantities are substituted into the same relation before the result is interpreted for “Minimum round-trip latency”.
- 15 — Independent check
- Quick check: for any non-zero factor, dividing the result by that factor should recover the other expected contribution in “Minimum round-trip latency”.
- 16 — Mental estimate
- Before calculating “Minimum round-trip latency” precisely, round the inputs to one useful digit and predict the sign and order of magnitude. The detailed result should remain consistent with that estimate.
- 17 — Interpretation
- This latency requires complete decision packages rather than interactive conversation.
- 18 — What the result does not prove
- For “Minimum round-trip latency”, the number obtained answers only the model “t_roundtrip = 2×t_oneway” under the stated scenario. It does not by itself validate the input data or the model outside those conditions.
- 19 — Sensitivity
- Vary one input at a time around the nominal case to identify what drives the result of “Minimum round-trip latency” and whether that variation can change the mission decision.
- 20 — Guided and autonomous exercises
Guided exercise. t_oneway = 18 min.
Detailed guided correction — open after trying
t_oneway = 18 min. t_roundtrip = 2×18 = 36 min.
Autonomous exercise. t_oneway = 7 min.
Autonomous correction — open after trying
t_oneway = 7 min. t_roundtrip = 14 min.
- 21 — Mission decision
- Pre-authorise decisions that must remain local during this delay.
Time consumed by repeated clarification cycles
- 1 — Concrete question
- What does “t_clarification = n_cycles×t_roundtrip” compute in “Time consumed by repeated clarification cycles”?
- 2 — Intuition without symbols
- Repeated exchanges consume the full communication latency on every cycle, even before human analysis time.
- 3 — Quantities
- t_clarification: minimum accumulated duration; n_cycles: number of question-response cycles; t_roundtrip: delay per cycle
- 4 — Formula
- t_clarification = n_cycles×t_roundtrip
- 5 — Read aloud
- Read “t_clarification = n_cycles×t_roundtrip” by naming every operation, subscript and grouping explicitly.
- 6 — Symbols and meaning
- t_clarification: minimum accumulated duration; n_cycles: number of question-response cycles; t_roundtrip: delay per cycle
- 7 — Pronunciation
- The “Read aloud” line above is the oral reference for “Time consumed by repeated clarification cycles”. Any subscript, exponent or grouping that changes the meaning of the relation should be spoken explicitly.
- 8 — Units
- t_clarification and t_roundtrip in time; n_cycles dimensionless
- 9 — Convention
- For “Time consumed by repeated clarification cycles”, substitute values without changing the reference frame, time basis, system boundary or sign convention halfway through the calculation. Stated units: t_clarification and t_roundtrip in time; n_cycles dimensionless.
- 10 — Why this operation
- In “Time consumed by repeated clarification cycles”, multiplication combines the factors that directly build the requested quantity; the factors must describe the same case.
- 11 — Assumptions
- The relation “t_clarification = n_cycles×t_roundtrip” applies here only to the scenario described by the card. Inputs must be mutually consistent and satisfy the physical assumptions associated with “Time consumed by repeated clarification cycles”.
- 12 — Unit check
- t_clarification and t_roundtrip in time; n_cycles dimensionless Verify that dimensional reduction reaches the unit of the requested output.
- 13 — Numerical case
- With n_cycles = 4 and t_roundtrip = 28 min, t_clarification = 112 min.
- 14 — Why the calculation works
- The numerical case applies “t_clarification = n_cycles×t_roundtrip” directly to the stated values. The calculation is meaningful because the quantities are substituted into the same relation before the result is interpreted for “Time consumed by repeated clarification cycles”.
- 15 — Independent check
- Quick check: for any non-zero factor, dividing the result by that factor should recover the other expected contribution in “Time consumed by repeated clarification cycles”.
- 16 — Mental estimate
- Before calculating “Time consumed by repeated clarification cycles” precisely, round the inputs to one useful digit and predict the sign and order of magnitude. The detailed result should remain consistent with that estimate.
- 17 — Interpretation
- The time cost justifies sending context, assumptions and options in the first message.
- 18 — What the result does not prove
- For “Time consumed by repeated clarification cycles”, the number obtained answers only the model “t_clarification = n_cycles×t_roundtrip” under the stated scenario. It does not by itself validate the input data or the model outside those conditions.
- 19 — Sensitivity
- Vary one input at a time around the nominal case to identify what drives the result of “Time consumed by repeated clarification cycles” and whether that variation can change the mission decision.
- 20 — Guided and autonomous exercises
Guided exercise. n_cycles = 3 and t_roundtrip = 36 min.
Detailed guided correction — open after trying
n_cycles = 3 and t_roundtrip = 36 min. t_clarification = 3×36 = 108 min.
Autonomous exercise. n_cycles = 5 and t_roundtrip = 20 min.
Autonomous correction — open after trying
n_cycles = 5 and t_roundtrip = 20 min. t_clarification = 100 min.
- 21 — Mission decision
- Limit clarification loops by preparing a self-contained anomaly package.
Workload utilisation
- 1 — Concrete question
- What does “U_work = H_tasks / H_available” compute in “Workload utilisation”?
- 2 — Intuition without symbols
- A schedule should count only hours genuinely usable after sleep, meals, exercise and fixed duties.
- 3 — Quantities
- U_work: human-capacity utilisation; H_tasks: required task hours; H_available: genuinely available hours
- 4 — Formula
- U_work = H_tasks / H_available
- 5 — Read aloud
- Read “U_work = H_tasks / H_available” by naming every operation, subscript and grouping explicitly.
- 6 — Symbols and meaning
- U_work: human-capacity utilisation; H_tasks: required task hours; H_available: genuinely available hours
- 7 — Pronunciation
- The “Read aloud” line above is the oral reference for “Workload utilisation”. Any subscript, exponent or grouping that changes the meaning of the relation should be spoken explicitly.
- 8 — Units
- U_work dimensionless; hours over the same period
- 9 — Convention
- For “Workload utilisation”, substitute values without changing the reference frame, time basis, system boundary or sign convention halfway through the calculation. Stated units: U_work dimensionless; hours over the same period.
- 10 — Why this operation
- In “Workload utilisation”, division relates a quantity to a reference, duration or capacity; the denominator must belong to the same case and remain non-zero.
- 11 — Assumptions
- The relation “U_work = H_tasks / H_available” applies here only to the scenario described by the card. Inputs must be mutually consistent and satisfy the physical assumptions associated with “Workload utilisation”.
- 12 — Unit check
- U_work dimensionless; hours over the same period Verify that dimensional reduction reaches the unit of the requested output.
- 13 — Numerical case
- With H_tasks = 54 h and H_available = 72 h, U_work = 0.75 = 75%.
- 14 — Why the calculation works
- The numerical case applies “U_work = H_tasks / H_available” directly to the stated values. The calculation is meaningful because the quantities are substituted into the same relation before the result is interpreted for “Workload utilisation”.
- 15 — Independent check
- Quick check: multiplying the result by the denominator should reconstruct the numerator of “Workload utilisation” within rounding.
- 16 — Mental estimate
- Before calculating “Workload utilisation” precisely, round the inputs to one useful digit and predict the sign and order of magnitude. The detailed result should remain consistent with that estimate.
- 17 — Interpretation
- A high utilisation rate reduces capacity to absorb anomalies even if the nominal schedule still fits.
- 18 — What the result does not prove
- For “Workload utilisation”, the number obtained answers only the model “U_work = H_tasks / H_available” under the stated scenario. It does not by itself validate the input data or the model outside those conditions.
- 19 — Sensitivity
- Vary one input at a time around the nominal case to identify what drives the result of “Workload utilisation” and whether that variation can change the mission decision.
- 20 — Guided and autonomous exercises
Guided exercise. H_tasks = 60 h and H_available = 72 h.
Detailed guided correction — open after trying
H_tasks = 60 h and H_available = 72 h. U_work = 60/72 ≈ 0.8333 = 83.33%.
Autonomous exercise. H_tasks = 40 h and H_available = 64 h.
Autonomous correction — open after trying
H_tasks = 40 h and H_available = 64 h. U_work = 40/64 = 0.625 = 62.5%.
- 21 — Mission decision
- Protect explicit human reserve for diagnosis, recovery and unplanned work.
Available crew-hour margin
- 1 — Concrete question
- What does “H_margin = H_available − H_nominal” compute in “Available crew-hour margin”?
- 2 — Intuition without symbols
- Human margin is the time left after all nominal work is protected, not raw calendar time.
- 3 — Quantities
- H_margin: crew-hour reserve; H_available: usable capacity; H_nominal: nominal planned workload
- 4 — Formula
- H_margin = H_available − H_nominal
- 5 — Read aloud
- Read “H_margin = H_available − H_nominal” by naming every operation, subscript and grouping explicitly.
- 6 — Symbols and meaning
- H_margin: crew-hour reserve; H_available: usable capacity; H_nominal: nominal planned workload
- 7 — Pronunciation
- The “Read aloud” line above is the oral reference for “Available crew-hour margin”. Any subscript, exponent or grouping that changes the meaning of the relation should be spoken explicitly.
- 8 — Units
- all three quantities in crew-hours over the same period
- 9 — Convention
- For “Available crew-hour margin”, substitute values without changing the reference frame, time basis, system boundary or sign convention halfway through the calculation. Stated units: all three quantities in crew-hours over the same period.
- 10 — Why this operation
- In “Available crew-hour margin”, subtraction measures a margin or difference between comparable quantities expressed in the same frame.
- 11 — Assumptions
- The relation “H_margin = H_available − H_nominal” applies here only to the scenario described by the card. Inputs must be mutually consistent and satisfy the physical assumptions associated with “Available crew-hour margin”.
- 12 — Unit check
- all three quantities in crew-hours over the same period Verify that dimensional reduction reaches the unit of the requested output.
- 13 — Numerical case
- With H_available = 72 h and H_nominal = 60 h, H_margin = 12 h.
- 14 — Why the calculation works
- The numerical case applies “H_margin = H_available − H_nominal” directly to the stated values. The calculation is meaningful because the quantities are substituted into the same relation before the result is interpreted for “Available crew-hour margin”.
- 15 — Independent check
- Quick check: adding the subtracted term back to the result should reconstruct the starting quantity in “Available crew-hour margin”.
- 16 — Mental estimate
- Before calculating “Available crew-hour margin” precisely, round the inputs to one useful digit and predict the sign and order of magnitude. The detailed result should remain consistent with that estimate.
- 17 — Interpretation
- A small margin can become negative as soon as an anomaly consumes several specialist hours.
- 18 — What the result does not prove
- For “Available crew-hour margin”, the number obtained answers only the model “H_margin = H_available − H_nominal” under the stated scenario. It does not by itself validate the input data or the model outside those conditions.
- 19 — Sensitivity
- Vary one input at a time around the nominal case to identify what drives the result of “Available crew-hour margin” and whether that variation can change the mission decision.
- 20 — Guided and autonomous exercises
Guided exercise. H_available = 80 h and H_nominal = 66 h.
Detailed guided correction — open after trying
H_available = 80 h and H_nominal = 66 h. H_margin = 14 h.
Autonomous exercise. H_available = 48 h and H_nominal = 44 h.
Autonomous correction — open after trying
H_available = 48 h and H_nominal = 44 h. H_margin = 4 h.
- 21 — Mission decision
- Do not commit to optional work if it consumes the reserve needed for safe return or the next anomaly.
Handover completeness
- 1 — Concrete question
- What does “C_handover = N_transmitted / N_required” compute in “Handover completeness”?
- 2 — Intuition without symbols
- A useful handover preserves assumptions, open anomalies, stop criteria and next actions required for continuity.
- 3 — Quantities
- C_handover: completeness; N_transmitted: critical items transmitted; N_required: critical items required
- 4 — Formula
- C_handover = N_transmitted / N_required
- 5 — Read aloud
- Read “C_handover = N_transmitted / N_required” by naming every operation, subscript and grouping explicitly.
- 6 — Symbols and meaning
- C_handover: completeness; N_transmitted: critical items transmitted; N_required: critical items required
- 7 — Pronunciation
- The “Read aloud” line above is the oral reference for “Handover completeness”. Any subscript, exponent or grouping that changes the meaning of the relation should be spoken explicitly.
- 8 — Units
- counts dimensionless; C_handover may be expressed as a percentage
- 9 — Convention
- For “Handover completeness”, substitute values without changing the reference frame, time basis, system boundary or sign convention halfway through the calculation. Stated units: counts dimensionless; C_handover may be expressed as a percentage.
- 10 — Why this operation
- In “Handover completeness”, division relates a quantity to a reference, duration or capacity; the denominator must belong to the same case and remain non-zero.
- 11 — Assumptions
- The relation “C_handover = N_transmitted / N_required” applies here only to the scenario described by the card. Inputs must be mutually consistent and satisfy the physical assumptions associated with “Handover completeness”.
- 12 — Unit check
- counts dimensionless; C_handover may be expressed as a percentage Verify that dimensional reduction reaches the unit of the requested output.
- 13 — Numerical case
- If N_transmitted = 18 and N_required = 20, C_handover = 0.90 = 90%.
- 14 — Why the calculation works
- The numerical case applies “C_handover = N_transmitted / N_required” directly to the stated values. The calculation is meaningful because the quantities are substituted into the same relation before the result is interpreted for “Handover completeness”.
- 15 — Independent check
- Quick check: multiplying the result by the denominator should reconstruct the numerator of “Handover completeness” within rounding.
- 16 — Mental estimate
- Before calculating “Handover completeness” precisely, round the inputs to one useful digit and predict the sign and order of magnitude. The detailed result should remain consistent with that estimate.
- 17 — Interpretation
- The percentage is meaningful only if the required-item list itself is qualified and risk-oriented.
- 18 — What the result does not prove
- For “Handover completeness”, the number obtained answers only the model “C_handover = N_transmitted / N_required” under the stated scenario. It does not by itself validate the input data or the model outside those conditions.
- 19 — Sensitivity
- Vary one input at a time around the nominal case to identify what drives the result of “Handover completeness” and whether that variation can change the mission decision.
- 20 — Guided and autonomous exercises
Guided exercise. N_transmitted = 27 and N_required = 30.
Detailed guided correction — open after trying
N_transmitted = 27 and N_required = 30. C_handover = 27/30 = 0.90 = 90%.
Autonomous exercise. N_transmitted = 15 and N_required = 16.
Autonomous correction — open after trying
N_transmitted = 15 and N_required = 16. C_handover = 15/16 = 0.9375 = 93.75%.
- 21 — Mission decision
- Do not transfer responsibility while any mandatory critical item is missing or ambiguous.
1. A role needs both authority and information

“The commander decides” is incomplete. Decides what, within what time, using which data and up to what risk? Responsibilities can be distributed across immediate safety, vehicle systems, medical work, navigation, science and maintenance. A small crew may combine roles, but ownership should remain explicit.
An operational role is more than a title. It combines decision authority, access to information and responsibility for handover. If two people both believe they own an urgent action, they may act in parallel; if each assumes the other owns it, nobody acts. Procedures therefore state who may stabilise the system, who authorises irreversible action and who maintains the timeline. Authority can change with mission phase and communications delay. During a Mars emergency, the local crew necessarily has more autonomy than a crew near Earth. That autonomy must be prepared through boundaries, thresholds and training rather than improvised after the anomaly begins.
2. Latency turns conversation into delayed messaging

With t = 12 min one-way light-time, the minimum round trip is about 2t = 24 min. Three sequential clarifications already consume more than an hour before analysis time. Efficient Mars communication therefore bundles context, evidence, actions taken, constraints and the requested decision.
With many minutes of one-way delay to Mars, conversation becomes asynchronous messaging. A poorly formed question can waste an entire light-time cycle before useful analysis even starts. A message to Earth should therefore contain current state, trend, configuration, actions already taken, hypotheses and the decision being requested. Earth should respond in the same style, stating assumptions that make the advice valid. Communications time becomes an operational resource. The local team continues stabilising and measuring while remote analysis proceeds. This prevents Earth from becoming a slow remote control for a system that evolves faster than the dialogue can complete.
3. Procedure, checklist and mission intent are different tools
A procedure describes sequence; a checklist protects critical points from omission; mission intent explains what outcome must survive when the case no longer matches the script. Autonomous teams need all three. A checklist cannot invent a response to a novel failure, while broad intent can be too vague under stress.
A procedure describes a controlled method; a checklist protects against omission of critical items; mission intent explains the outcome to preserve when reality leaves the written case. Confusing these tools creates either enormous lists or uncontrolled improvisation. An irreversible action may require an observable condition and confirmation. A frequent routine may need only a short checklist. When the crew encounters an unanticipated state, intent—preserve habitable pressure before science productivity, for example—supports judgement. Training therefore teaches not only how to execute steps but how to recognise when the procedure no longer matches the actual system state.
4. Handover carries assumptions and capability debt
Shift change should transfer open anomalies, temporary configurations, consumed margins, deferred work and the rationale behind decisions. “Pump B isolated” is less useful than “Pump B isolated after unstable current; Pump A running; no redundancy until the planned 18:00 test.” The log becomes shared operational memory.
A good handover transfers more than open tasks. It carries assumptions, weak anomalies, consumed margin and deferred decisions. A pump may still operate while current slowly drifts; if that fact disappears at shift change, the next alarm will look sudden. The log should separate measured fact, interpretation and action. A practical structure can include current state, changes since the last watch, active risks, prohibited activities, decisions pending and wake-up thresholds. The goal is not to write a novel at every handover but to preserve enough operational memory that the incoming team does not have to reconstruct context under pressure.
5. Schedule real human capacity, not perfect nominal utilisation
Crew time, sleep, expertise and EVA hours are resources. A six-hour two-person repair may be impossible when another anomaly requires continuous monitoring. Plans need reserve capacity for the unplanned rather than scheduling every person to 100 percent nominal utilisation.
Human capacity can be budgeted like electrical energy. A day with 40 theoretical person-hours should not be scheduled to all 40 if the mission needs anomaly response. Reserving 20 percent leaves 8 person-hours unassigned, although the right fraction depends on phase, fatigue and risk. Presence is not the same as competence: four available people do not replace two specialists if a repair requires particular qualification. Schedules also include sleep, meals, exercise, habitat maintenance, training and documentation. A resilient mission accepts that some capacity remains deliberately unproductive during nominal operations because that reserve is what makes off-nominal work possible.
6. Automation should make safe reactions fast and ambiguous decisions explainable
Automation excels at monitoring and bounded rapid response. When judgment is required, it should expose the evidence, affected functions and confidence behind a recommendation. An alarm that says only “critical fault” adds cognitive load; an alarm that explains trend and safe options supports autonomy.
Automation is valuable when it makes a safe reaction faster and repeatable. It becomes dangerous when it hides uncertainty or performs an irreversible action from a weak diagnosis. A good interface shows what triggered the action, which evidence supports the interpretation and what condition permits return. Humans need a practical way to take over when context exceeds the model, but ‘manual control’ cannot mean hundreds of low-level commands that no tired operator can manage. Degraded modes therefore combine minimum trustworthy automation, clear procedures and a readable state representation. The objective is to reduce cognitive load rather than move opacity from one layer to another.
7. Manage anomalies as stabilise, understand, recover
Stabilise prevents further deterioration. Understand gathers evidence and narrows causes. Recover restores functions in controlled sequence. Mixing the phases can produce premature repair commands before the actual state is understood.
Anomaly management can be organised as stabilise, understand, recover. Stabilise means stop escalation even when the exact cause is unknown. Understand means collect evidence, compare hypotheses and identify dependencies. Recover returns the service to a sustainable state and proves that return. Jumping directly to repair can destroy evidence or create a second fault. The crew also preserves a timeline of commands and observations so Earth or the next watch can reconstruct the event. This sequence is particularly useful under delay because it creates a coherent anomaly package while the spacecraft remains in a controlled state.
8. Failure scenario: technical anomaly after poor sleep
A fatigued team does not have the same decision performance as a rested one. If the anomaly can be stabilised safely, the rational response may be to automate monitoring, defer complex repair, prepare tools and restore crew readiness. Human safety includes fatigue management, not just technical compliance.
Fatigue changes perception, working memory, decision speed and risk tolerance. The system should not assume the crew can always execute a complex procedure at peak cognitive performance. Critical work can be delayed when a stable mode exists; urgent actions should be short, observable and supported by clear interfaces. Cross-training also reduces dependence on the most expert individual if that person is ill or exhausted. In a combined technical failure and poor-sleep scenario, the rational first choice may be to stabilise the system and let a rested team conduct the detailed diagnosis later. Human resilience depends on the ability to buy time.
9. Ask Earth questions that remain useful after delay
Earth is valuable for deep independent analysis. The crew can ask “compare repair A and B for our present configuration” instead of “tell us what to do now.” This lets Earth exploit expertise without pretending to close a real-time loop it physically cannot close.
A question sent to Earth should still be useful when the answer returns. ‘What do we do now?’ is weak if the reply arrives forty minutes later and the state has changed. A stronger request defines branches: if temperature exceeds X, the crew will isolate the loop; otherwise it will hold the current mode until a stated time. Earth can then analyse each future state, challenge thresholds and propose actions that remain applicable. This turns delay into parallel work. Messages need units, timestamps, configuration versions and uncertainty so a recommendation is not accidentally applied to a system state that no longer exists.
10. Lessons learned need controlled change
After an incident, procedures may change. The change should retain version, rationale, evidence, approval and rollback. A successful one-off workaround is not automatically a universal rule. Operational knowledge should evolve with the discipline of a well-managed software configuration.
Lessons learned matter only when they change work in a controlled way. After an anomaly, the team separates immediate correction, probable cause, evidence obtained, procedure change and required test. A new version should identify what changed and why. Silently replacing a checklist destroys the ability to understand earlier decisions; never changing it preserves a known weakness. The process therefore combines memory with evolution. On Mars, the same principle applies to physical configuration: a maintenance procedure must know which sensor version, firmware build or locally manufactured part is actually installed before prescribing an action.
Guided case — an anomaly with 18 minutes of one-way light time
A thermal-loop pump shows increasing current but remains functional. Earth is 18 light-minutes away, so even a simple question requires at least 36 minutes before an answer returns, plus analysis time. The crew first stabilises locally: non-essential loads are reduced, the redundant loop is checked and a switchover threshold is defined. The message to Earth includes current trend, temperature, flow, configuration, recent maintenance and actions already completed.
While waiting, the local plan has branches. If temperature exceeds 65 °C, the pump will be isolated and bypass opened; otherwise the system remains under observation for one hour. Earth can analyse a future that will still be useful when the response arrives. Advice based on the assumption that the crew has done nothing would already be obsolete.
The next day, the team updates handover notes and the controlled procedure. It preserves the difference between observed fact, the hypothesis of bearing degradation and the maintenance action selected. The operational lesson becomes reusable without turning an unproven hypothesis into historical truth.
11. Mini-project: run a degraded Mars day
- Define four crew roles.
- Create a schedule with 20 percent time reserve.
- Add a 10:00 fault and two-hour Earth outage.
- Write a stabilisation checklist.
- Prepare the delayed Earth message.
- Plan evening handover.
- State how the procedure changes after review.
The purpose is to show that the organisation still functions when the nominal plan disappears.
12. Mission lab — build an anomaly package for delayed Earth support
A good delayed message lets an Earth team work without five clarification rounds. It includes time and configuration, initial symptom, relevant evidence, actions already taken, results of those actions, remaining resources, crew hypotheses and the decision being requested. If conditions can change before the reply arrives, the package also states which thresholds will trigger local action.
Imagine a pump whose temperature is rising. Instead of transmitting one number, the package includes several hours of trend, motor current, flow, pressure, ambient temperature, recent maintenance and the state of the redundant loop. Earth can test competing explanations while the crew retains authority for urgent thresholds already defined in local procedure.
This style turns communications into an interface between two asynchronous teams. It also makes decisions auditable because the eventual action can be traced to the evidence available at the time rather than reconstructed from memory.
Earth responses should use the same discipline. They can state assumptions, confidence, configuration for which the advice is valid and conditions that would invalidate it. That prevents a well-intended recommendation from being applied after the local state has changed.
13. Budget workload and preserve response capacity
A day scheduled to 100 percent utilisation has no resilience. If four crew members each have ten useful work hours, theoretical capacity is 40 person-hours. Reserving 20 percent for unplanned work leaves 40 × 0.20 = 8 person-hours of response capacity and 32 person-hours of nominal tasks. Twenty percent is not a universal optimum; the point is that human reserve can be budgeted explicitly rather than hoped for.
Skill reserve matters as well. If only one crew member can service a vital system, illness, fatigue or EVA can create a human single point of failure. Cross-training, clear procedures and diagnostic tools reduce that dependency. Critical maintenance should be mapped to at least the number of qualified people required by the mission’s loss-of-crew and fatigue assumptions.
Finally, protect sleep and recovery as functional resources. A tired crew can be safe if the system allows stabilise-and-wait decisions; it becomes vulnerable when every anomaly demands immediate complex manual work. Operations architecture should therefore include states in which machines hold the system stable long enough for humans to regain decision quality.
Mission reasoning laboratory — connect the calculation to a real decision
Design roles around decisions, not job titles
A distant crew cannot wait for Earth to resolve every ambiguity. Operational roles must therefore be tied to specific decisions: who may stop an EVA, who authorises a battery load shed, who accepts a temporary repair, and who can declare the habitat safe after an alarm. The same person may hold several roles in a small crew, but the decision boundaries still need to be explicit. Ambiguity is especially dangerous during fatigue, when people tend to assume that somebody else has authority or that a remote team will answer in time.
Communicate for delayed understanding
A useful Mars message contains a time tag, current configuration, symptoms, actions already taken, uncertainty, resource margins and a precise question. The sender should anticipate that the receiver will read it many minutes later, perhaps after the situation has changed. Messages should therefore distinguish facts from hypotheses and state what the crew will do if no reply arrives. This style reduces repeated question-and-answer cycles and lets Earth work in parallel instead of merely asking the crew for information that could have been packaged initially.
Protect handovers from memory loss
Shift changes are vulnerable moments because the outgoing operator holds tacit context that is not obvious from telemetry. A good handover records abnormal trends, temporary configurations, disabled alarms, open maintenance, assumptions and the next decision point. The incoming operator should read back the critical state and challenge anything unclear. On a Mars base, handover quality affects more than efficiency: a forgotten isolation valve, software inhibit or reduced oxygen reserve can turn a later routine action into a compound emergency.
Budget human capacity as a finite resource
Crew time is a mission resource like water or energy. A nominal schedule that consumes every hour leaves no margin for unexpected maintenance, medical events or complex decision-making. Workload should include preparation, donning and doffing, cleanup, documentation, travel, communications, verification and recovery after demanding tasks. Fatigue also changes the value of an hour: one hour of highly focused troubleshooting at the wrong circadian phase may carry more risk than several hours of routine work. Protect reserve capacity for anomalies.
Use checklists without replacing expertise
Checklists are strongest at catching omissions in well-understood procedures. They are weaker when the situation does not match the assumed state. Operators must know which steps are hard safety barriers and which can be adapted under authority. Challenge-and-response helps confirm critical configuration, while mission intent helps the crew choose among safe alternatives when a step cannot be completed as written. Training should therefore combine exact procedure execution with scenario practice in which the crew must recognise when the procedure no longer fits.
Stabilise before diagnosing deeply
During an anomaly, the first goal is usually to stop the situation from getting worse. That can mean isolating a leaking line, shedding a suspect load, closing a hatch or placing a rover in a stable parked state. Only after immediate hazards are controlled should the team expand diagnosis. This ordering prevents a cognitively attractive search for root cause from consuming the last margin. The team should continually ask which next action is reversible, which information it preserves and which action could close off the safe-return path.
Turn experience into controlled change
After a difficult operation, the debrief should separate facts, interpretations and proposed changes. A lesson is not complete when somebody says “next time be more careful.” It should identify the condition, mechanism and specific barrier that failed or succeeded. Changes to procedure, software, training or configuration then need review and version control so that one local improvement does not create a new incompatibility. On Mars, where local adaptation will be essential, disciplined learning is the bridge between autonomy and uncontrolled improvisation.
Progressive exercises — solve first, then open the correction
Exercise A — clarification cost
Earth–Mars one-way communications delay is 14 minutes. A ground team and crew use four sequential question-and-answer cycles, where each new question waits for the previous answer. Calculate the physical minimum elapsed time caused by light-time alone, then name two reasons the real operational delay will be longer.
Detailed correction — Exercise A — clarification cost
Light-time. One question-and-answer cycle needs 28 minutes for propagation. Four strictly sequential cycles therefore need at least 112 minutes.
Real operations. Message preparation, review, antenna scheduling, network routing, crew workload and decision time add delay beyond the physical minimum.
Beginner vocabulary checkpoint
- mission intent — Concise statement of the purpose and desired end state that guides decisions when detailed instructions no longer fit the situation.
- authority — Formal permission to make a defined class of operational decisions.
- situational awareness — Shared understanding of what is happening, why it matters and what is likely to happen next.
- crew resource management — Methods for using communication, leadership, workload and cross-checks to reduce human-error risk.
- handover log — Structured record used to transfer system state, decisions, risks and open actions between operators or shifts.
- turnaround point — Latest point at which a team can reverse course and still return with protected resource margins.
- contingency — Preplanned response to a credible off-nominal condition.
- procedure — Controlled sequence of steps used to perform a task reproducibly and safely.
- checklist — Compact verification aid that confirms critical items without replacing understanding of the procedure.
- command authority — Defined right to approve, reject or direct an operational action.
- communications delay — Elapsed propagation time that prevents immediate interactive conversation with a distant support team.
- one-way light time — Minimum signal propagation time from sender to receiver, set by distance and the speed of light.
- decision latency — Delay between recognizing a need for a decision and executing the authorised response.
- fatigue — Reduction in alertness and performance caused by insufficient sleep, circadian disruption or sustained workload.
- workload — Amount and complexity of work assigned to a person or team over a period.
- cross-check — Independent verification by another person, method or data source before accepting a critical result.
- challenge-and-response — Checklist method in which one person calls an item and another confirms the required state or action.
- go/no-go criterion — Predefined condition used to decide whether an activity may start or continue.
- debrief — Structured review after an activity to capture facts, deviations, lessons and actions.
- lesson learned — Validated observation from experience that produces a controlled change in training, design or procedure.
Sources and references
Engineering studio — command under latency and human workload
The scenario combines a slow habitat pressure decrease, an immobilised rover and a delayed communications pass. Two crew members are available and Earth cannot contribute a useful decision for more than 20 minutes. Workload at one station is represented by ρ = λs, where λ is task arrival rate in tasks/h and s is average task duration in hours. With λ = 4 tasks/h and s = 10/60 h, ρ ≈ 0.67: the station looks sustainable on average, yet a burst of anomalies can immediately saturate human capacity.
The task therefore goes beyond an average. The crew prioritises the leak, rover safety and communications, decides what may wait, and records authority-transfer conditions. A good procedure states what triggers volume isolation, when an EVA must be aborted and which data must be preserved for Earth. The expected result is an operational loop that remains robust when remote support is not synchronous.
