Why this reform exists
ANALYSIS Measure 3.04 should not be read as an abolition slogan. Its purpose is to turn a reform intention into a verifiable decision. The issue is no longer whether staff will test AI, but how to build a common framework that avoids a proliferation of uncontrolled tools. DINUM now provides an interministerial AI stack and usage guidance; agencies should build on that shared infrastructure rather than each buying a separate assistant. That distinction is essential: Bible France asks what should change, why, through which legal route and with what net effect for taxpayers and service users. [1]
What the measure actually changes
ANALYSIS The proposal is: Industrialise useful AI for document search, classification, assisted drafting and case analysis, with human validation, data security and measured gains before scale-up. It belongs to the agencies and operators volume, whose general purpose is not to deny public missions but to test the value of each institutional layer. A useful function can be retained while its organisation changes deeply; a small body can also remain autonomous where that autonomy protects expertise or impartiality that cannot credibly be reproduced elsewhere.
Implementation method and timetable
IMPLEMENTATION Each operator would map tasks as automatable, assistable, sensitive or prohibited. Low-risk cases would be tested with before/after protocols; decisions affecting people would remain under human responsibility. Models, logs, data services and compute costs would be pooled where possible. The timetable must include a baseline, target design, transition phase and a date for steady-state measurement. No gain should be claimed while old and new arrangements run in parallel unless that temporary double cost is explicitly separated.
Costing: never confuse funding with savings
COSTING The AI baseline is current human time plus existing software. The target adds inference, hosting, integration, human review, testing, security and maintenance. Net savings only include expenditure or capacity actually avoided after those costs. Time released but neither redeployed nor converted into additional service remains a productivity indicator, not a cash saving.
Net recurring saving = costs removed − costs recreated − transferred liabilities − recurring residual costsControl, data and indicators
CONTROL The reform requires a specific dashboard: Average task time before/after; human rework rate; critical errors; cost per assisted case; share of usage on approved infrastructure; data incidents; hours actually redeployed to service delivery. These indicators are published before and after transformation. Where the objective is qualitative — faster processing, clearer accountability or better data availability — it is measured directly rather than converted into invented monetary value.
Objections and safeguards
ANALYSIS The central objection is serious: The key risk is not only factual error: data leakage, automation bias, illusory productivity and vendor dependence also matter. Claimed savings are meaningful only when time released is measured and actually redeployed to useful work. The safeguard is to document the counterfactual, preserve legal duties and service continuity, then organise independent reviews after twelve and twenty-four months. The reform is corrected if costs merely move elsewhere or service quality deteriorates.
Public decision and success criteria
ANALYSIS Scale-up requires an owner, test set, proportionate human review and before/after measurement for every use case. The report publishes time actually released, rework, critical errors, data incidents and compute cost. A striking assistant that reduces neither delay nor cost is not counted as a successful transformation.
Measure-specific dossier: what must be demonstrated
Start from tasks rather than from the prestige of AI
Each agency maps repetitive document tasks: search, classification, summarisation, data extraction, assisted drafting and anomaly detection. Every use case is scored for expected value, data sensitivity, human verifiability and cost of error. Low-risk tasks can move quickly; decisions affecting rights or sanctions require stronger governance. DINUM's Albert API provides a shared public-sector foundation. [1]
Separate assistance, automation and decision-making
A drafting assistant is not equivalent to a system that ranks a citizen or triggers enforcement. The use register therefore states who validates the output, which data reach the model, what logs are kept and how an employee can challenge it. The EU AI Act and CNIL guidance require risk-sensitive governance rather than a generic label of 'AI'. [3][4]
Measure inference cost and failure rates
A striking demo is not a business case. Tests record minutes saved, human rework, critical errors, false positives, query volume, inference cost and supervision cost. Published service tariffs make it possible to include usage cost instead of treating AI as free. [5] Savings are only recognised when released time is actually redeployed or expenditure is avoided.
Prevent shadow AI and vendor lock-in
Sensitive data remain on authorised infrastructure; critical prompts and models are versioned; incidents are reported; and a degraded mode keeps the service operating without the model. The aim is not to crown a permanent winning model but to make public-sector AI reversible, auditable and compatible with legal duties.
Trace data, models and decisions
Each AI use case records data sources, sensitivity, model and version, essential parameters, log retention and the accountable business owner. Generated assistance is not presented as an administrative decision. Where a system can affect rights, sanctions or controls, the human-validation path is explicit and incidents can be replayed. A faulty model or knowledge base must be removable without stopping the whole service.
Test failure before counting productivity
Pilots are deliberately exposed to incomplete, contradictory and unusual files. Teams measure hallucinations, omissions, classification errors, potential data leakage and user over-reliance. A documented non-AI fallback remains available. Productivity is accepted only after measuring human rework, serious incidents and cost per case at representative volume.
The evidence file that makes the measure challengeable
The public AI register records process, purpose, data category, provider or model, accountable human owner, testing protocol, production date and significant incidents. Protected prompts or personal data need not be public, but the method must show how errors are detected, corrected and assigned to a responsible authority.
Full-scale test: three uses, three risk levels
The pilot combines internal document search, pre-filling a case and ranking a sensitive application. Search can be heavily assisted; pre-fill requires systematic review; sensitive ranking needs legal analysis, bias testing and appeal safeguards. One dashboard publishes both time gains and failure rates so only flattering demos cannot be selected.
This measure in the system
Measure 3.04 is assessed with neighbouring measures in the volume: pooling, merger or reintegration must never count the same saving twice.
Notes and sources
- DINUM — Albert API, socle d’IA générative pour les services publics — institutional document used for the legal, operational or financial baseline of this measure.
- DINUM — infrastructure sécurisée pour l’IA — institutional document used for the legal, operational or financial baseline of this measure.
- CNIL — Guide pratique sur l’intelligence artificielle — institutional document used for the legal, operational or financial baseline of this measure.
- Règlement européen sur l’intelligence artificielle — texte officiel — institutional document used for the legal, operational or financial baseline of this measure.
- DINUM — tarifs et limites d’Albert API — institutional document used for the legal, operational or financial baseline of this measure.

