DELTA-SIERRAMARSEXPLORE · UNDERSTAND · SETTLE
Support my work
MODULE 35 · ADVANCED MARS CURRICULUM · UNDERSTAND, CALCULATE, VERIFY.

Maintenance, fleet reliability and repair on Mars

Design a base that keeps operating as components age, fail and must be diagnosed, cannibalized, repaired or replaced without an immediate supply chain.

Before you begin — Prerequisites: modules 00–34 as relevant. Every important symbol is defined at first use.

Mastery objectives

  • connect principles to architecture or operational decisions
  • repeat simple calculations and verify units and assumptions
  • identify degraded modes, interfaces and uncertainty
  • produce a verifiable procedure or plan

1. Reliability and maintainability are not the same thing

Reliability describes the probability that equipment performs its function for a given period. Maintainability describes how easily and quickly it can be diagnosed, opened, repaired, tested and returned to service. A highly reliable but irreparable system can become dangerous on Mars; a somewhat less reliable but modular and repairable system may deliver better operational availability.

2. Design for access, not performance alone

A filter hidden behind three assemblies can turn a ten-minute task into a six-hour intervention. Connectors, access panels, handling masses, torque requirements, tools and work zones must be designed from the start. Common interfaces allow one assembly to replace another without rebuilding the whole system.

3. From symptom to cause: troubleshooting and fault trees

Troubleshooting separates symptom, failure mechanism and root cause. A falling water flow can be caused by a pump, a clogged filter, a partly closed valve, a bad sensor or a leak. Crew need telemetry, measurement points and procedures that prevent healthy hardware from being replaced by guesswork.

4. Repair at the lowest reasonable level

On Earth an entire module may be replaced because logistics make that convenient. On Mars spare mass is limited. Repairing a board, motor, seal or pump can save tens of kilograms of inventory. Going too deep, however, increases training, tools and error risk. The correct repair level is an architecture trade.

5. Redundancy, cannibalization and technical debt

Identical equipment provides redundancy and can enable cannibalization: one unavailable asset becomes a donor for another. The strategy must be controlled because it creates technical debt. Every removed part must be recorded so the base knows which assets are no longer complete.

6. Preventive, predictive and corrective maintenance

Preventive maintenance is scheduled by time or operating hours. Predictive maintenance uses vibration, temperature, current or pressure trends to intervene when degradation begins. Corrective maintenance occurs after failure. Mars operations must blend all three to reduce both spare mass and sudden-failure risk.

7. Common-cause failures and genuine independence

Two identical units do not automatically create robust redundancy. If they share the same power converter, software defect, environmental sensor or maintenance error, a single cause can disable both. Reliability work must therefore distinguish independent failures from common-cause failures. On a Mars base, two circulation pumps mounted together might both be lost if fluid contamination, local overheating or an incorrect maintenance procedure affects the whole subsystem. Useful redundancy may require separated power paths, diverse sensing, isolation capability, degraded modes and procedures that confirm the backup channel is actually healthy before load is transferred to it.

8. Sizing spares from risk rather than fear

Carrying a duplicate of every part sounds reassuring but quickly becomes impossible as mission duration, mass and volume grow. Spare provisioning must combine failure frequency, criticality, replacement time, local repair capability, parts commonality and the consequences of stockout. A light but unique seal may justify several copies, while a heavy and highly reliable structure may be better served by a repair strategy than by full replacement. Maintainers must also consider hidden consumables such as lubricants, filters, connectors, fasteners, adhesives, cleaning materials and test equipment. Maintenance logistics therefore becomes a quantitative extension of reliability engineering rather than a list written after design is complete.

9. Closing the loop: failure, lessons learned and configuration change

A base that repairs equipment without learning will accumulate recurring failures. Every significant incident should create an operational record: initial symptoms, telemetry, hypotheses tested, accepted cause, parts replaced, post-repair condition and any procedural change. If the same failure returns, the team must decide whether the root problem is maintenance, environment or design. Permanent changes then have to pass through configuration management so drawings, software, inventories and procedures continue to describe the hardware that actually exists on Mars. Maintenance therefore becomes a continuous-improvement process, not merely a way to restore service.

10. Worked example: availability of a critical function

Assume equipment operates an average 2,000 h between failures and requires 20 h to repair. Approximate availability is A = MTBF/(MTBF+MTTR) = 2,000/(2,000+20) ≈ 0.9901, or 99.01%. If repair time rises to 200 h because access or parts are poor, availability drops to 90.9%. Maintainability can therefore cost almost ten percentage points without changing intrinsic reliability.

11. Progressive exercise

A base owns three identical pumps. Two are required for nominal operation and one is a spare. Propose a rotation, vibration-monitoring and common-spares strategy. Explain what changes if the main seal has a 30% probability of failing during a 500-day interval.

Mini-project

Build a supportability plan for a Mars workshop: critical functions, repair levels, tools, offline documentation, common parts, cannibalization criteria, return-to-service tests and availability indicators.

Primary sources and bridges