MARS BIBLE — RISK & RESILIENCE

Combined power + thermal + ECLSS failure: the cascade that can threaten the entire base

A power failure becomes far more dangerous when it also degrades heat rejection and life support: three systems that looked separate now share one survival clock.

This chapter follows a cascade in which electrical generation falls, thermal pumps stop, and ECLSS margins shrink. The goal is to identify shared dependencies, first loads to shed, irreversible thresholds, and restart paths that must remain independent.

1 — A failure can change category within minutes

Power runs pumps, valves, sensors, and computers. When the main bus fails, some equipment transfers to batteries, but their endurance differs.

The emergency becomes thermal when pumps or fans stop, then atmospheric when CO₂, humidity, or temperature leaves its safe band. Event order determines what must be protected first.

Dependency graph connecting power, thermal control and ECLSS
Three networks that appear separate can share the same failure points.

2 — Build the dependency graph before the emergency

For every vital function, identify power, cooling, data, sensing, fluid network, and operator dependencies. A redundant unit sharing the same converter or cooling loop is not independent.

A dependency graph lets engineers remove a node in simulation. If two backup paths disappear together, the design has a common vulnerability.

3 — Measure time to threshold, not only present status

A battery may show 60%, but the useful question is how many minutes remain at current power. A room may be breathable now but cross its CO₂ limit in two hours.

Operators need countdowns associated with energy, temperature, CO₂, pressure, oxygen, and water thresholds.

Each subsystem has a different clock

An electrical fault can stop a pump immediately while temperature or carbon-dioxide concentration crosses a critical threshold later. That difference in dynamics is valuable because it allows actions to be prioritized. Operators need to know which variables evolve over seconds, minutes or hours.

A crisis display should therefore show not only current state but an estimate of time to limit for essential functions. That estimate should be conservative and updated with real measurements. Time becomes a calculable resource alongside remaining energy.

Restarting in the wrong order can recreate the failure

When power returns, the temptation is to switch everything back on. Yet simultaneous startup loads, thermal demands or current peaks may exceed the still-limited capacity. Safe recovery therefore needs sequencing: measurement and control functions first, life-critical loads next, and secondary loads only when margin is demonstrated.

The same is true for ECLSS. A system that has been stopped for a long time may need checks, purge or stabilization before it can be considered nominal. Restoration is an engineering phase in its own right, not simply the reverse of shutdown.

4 — Shed load across multiple resources

Stopping a machine saves electricity but may increase another demand. Stopping a cooling pump saves power yet accelerates overheating; shutting a greenhouse saves energy but creates delayed food cost.

Good load shedding maximizes total survival time, not merely the number of kilowatts removed.

Load-shedding and restart sequence after a combined failure
Buying time requires an explicit restoration order.

5 — Restart can become a second emergency

When power returns, all loads should not start together. Inrush current, flows, and heat loads can trigger another collapse.

A Mars black start should restore measurement, control, and vital functions first, then add loads progressively while checking margins.

6 — Deliberately test cascades

Validation should go beyond “pump A fails, pump B takes over”. Test combinations: pump A failure + wrong sensor + aged battery + high thermal demand.

Combined scenarios reveal dependencies that unit tests miss.

Calculate time before an electrical reserve is exhausted

LEARNING CALCULATION — ASSUMPTIONS ARE EXPLICIT

LEARNING ASSUMPTION: usable battery energy remaining = 420 kWh; vital loads after shedding = 70 kW.

Ideal endurance = 420 kWh ÷ 70 kW = 6 h. Keeping a 20% operational reserve leaves 420 × 0.80 = 336 kWh planable, so 336 ÷ 70 = 4.8 h.

This ignores efficiency, temperature, aging, and maximum discharge power. It teaches the difference between stored energy (kWh) and load power (kW).

Decision questions specific to this risk

  • Which vital function loses cooling when the main bus fails?
  • Which threshold is reached first: battery, temperature, or CO₂?
  • Which load can be stopped without creating a delayed crisis?
  • Do electrical and thermal backups share a converter or controller?
  • Which black-start order prevents a second collapse?

Main primary sources

Connect to other dossiers

Why these three systems can fail together

Life support uses electricity to circulate air, pump water, control valves and process streams. Electrical and biological equipment produce heat that must be rejected. Thermal control itself depends on pumps, fans and automation. A power failure can therefore degrade life support and cooling simultaneously.

This is why three separately “redundant” systems may still share one electrical bus or one heat-rejection loop. Dependency mapping must identify those intersections before the mission.

Time to threshold differs by resource

After power loss, some temperatures rise slowly while carbon dioxide can reach an operational threshold faster. A battery may last hours while local air circulation has only minutes of margin. Crisis response should therefore be ordered by real time constants.

A survival dashboard should show estimated time to threshold and uncertainty, not only red/green lights. Priority goes to the function with the shortest margin or whose loss would make other repairs impossible.

Restart must be sequenced

Restoring every load at once can collapse the grid again. Systems need staged restart: minimum control, vital circulation, cooling, air treatment and only later less critical functions. This sequence should be tested with representative hardware.

A Martian black start is therefore a mission scenario of its own. Operators need to know which sources can start without the grid, which loads come first and how to confirm stability before reconnecting the rest.