Technical & digital health · AI, surveillance & governance
Responsible AI in public health surveillance: what WHO guidance actually requires of a national programme
AI is entering surveillance work faster than governance is. WHO has published the guidance a national programme needs; this article turns it into decisions a surveillance officer can actually take, with the limits stated honestly.
Key takeaways
- WHO's 2021 guidance on the ethics and governance of AI for health set out six guiding principles and remains the reference document.
- WHO's 2024 guidance on large multi-modal models addresses generative AI specifically, including risks that pre-generative guidance did not anticipate.
- WHO's EIOS initiative shows what disciplined open-source epidemic intelligence looks like, and was upgraded with partners in October 2025.
- For surveillance, AI is most defensible in triage and prioritisation of signals — with human verification retained for every action that has consequences.
- Governance is concrete: documented purpose, data provenance, human accountability, bias assessment, and an audit trail for every automated decision.
My exploratory research interest in artificial intelligence for public health is narrow on purpose: signal triage, information-quality support and workforce assistance, not autonomous decision-making. That narrowness is the point. The failure mode in surveillance is not that AI performs badly; it is that a plausible-looking output enters a decision chain nobody designed to question it.
1. The two WHO documents that define the ground rules
WHO published 'Ethics and governance of artificial intelligence for health' in June 2021, the product of eighteen months of deliberation among experts in ethics, digital technology, law and human rights together with ministry representatives. It was WHO's first global report on AI in health and set out six guiding principles for the design and use of AI in health systems. In January 2024 WHO followed with guidance on large multi-modal models — AI systems that accept and generate multiple data types — addressing generative AI risks that the 2021 document could not have anticipated.
2. What disciplined AI-supported surveillance looks like: EIOS
WHO's Epidemic Intelligence from Open Sources (EIOS) initiative is the clearest working example of machine-assisted epidemic intelligence at scale. WHO describes it as the leading initiative for open-source intelligence in public health decision-making, built on three pillars — community, technology, and evidence-based practice — with its technology developed through long-standing collaboration with the European Commission's Joint Research Centre. WHO, with the European Commission and Germany's Federal Ministry of Health, launched an upgraded EIOS system in Berlin in October 2025, and an EIOS strategy covering 2024-2026 sets out its direction.
The design lesson generalises to any national programme. EIOS is not an automated alerting machine; it is a filtering and prioritisation layer that puts candidate signals in front of trained epidemiologists, who verify. That division of labour — machine recall, human precision — is the only architecture I would recommend to a national surveillance unit today.
| Workflow step | Machine role | Human role |
|---|---|---|
| Media and open-source scanning | High-recall retrieval, deduplication, translation, clustering | Decide which clusters are worth verification |
| Routine data quality screening | Outlier and inconsistency flagging | Investigate flags against the source register |
| Signal verification | Assemble supporting context | Verify with the reporting unit; this step is never automated |
| Risk assessment and response decision | None, beyond summarising evidence on request | Full human accountability, documented |
| Public communication | Drafting support only | Approval by the accountable officer, always |
3. The risks that matter in a low-resource surveillance system
- Automation bias: an over-worked officer accepts a plausible machine output because verifying it costs time they do not have. This is a workload problem disguised as a technology problem.
- Data provenance and consent: surveillance data are often personal and always sensitive; a model that ingests them without a documented legal basis creates liability that outlives the project.
- Representational bias: signal detection trained on well-covered regions and languages under-detects exactly where health systems are weakest.
- Fabrication in generative models: confident, well-formatted, wrong summaries are the specific hazard of large multi-modal models in an operational context.
- Dependency and continuity: a tool provided under a project that ends leaves a workflow that cannot be run — plan the exit before adoption.
- Opacity: if you cannot explain to a district officer why a signal was prioritised, you cannot expect them to act on it or contest it.
4. A governance checklist you can adopt this quarter
- 01Written purpose statement per use case: what decision this tool supports, and what decision it must never make alone.
- 02Data inventory: what data enters the tool, under what legal basis, retained for how long, and who can access it.
- 03Named accountable human for every automated output that reaches a decision, with the name published internally.
- 04Bias and coverage assessment before deployment: which populations, districts and languages are under-represented in the inputs, and how that is compensated.
- 05Verification rule: define, in writing, which outputs require human verification before any action — for surveillance signals, the answer should be all of them.
- 06Audit trail: log inputs, outputs, the verifying officer and the final action, so a decision can be reconstructed months later.
- 07Performance review on a schedule, including false-positive burden on staff time, not only detection sensitivity.
- 08Exit plan: how the workflow continues if the tool, licence or project disappears.
In surveillance, the value of a machine is recall. The value of a human is refusal. Design so that both are cheap.
5. Do the data plumbing first
Most AI ambitions in national health programmes fail at the data layer rather than the model layer. WHO's Global Strategy on Digital Health 2020-2025, endorsed by the World Health Assembly in 2020, and its SMART Guidelines work — which turns narrative WHO guidelines into digital adaptation kits and FHIR-based implementation guides — describe the sequence properly. Interoperable, standards-based data (HL7 FHIR for exchange, ICD-11 for classification, SNOMED CT for clinical terminology, an OpenHIE-style interoperability layer for architecture) is what makes any analytic layer, AI or otherwise, worth building.
- Standardise exchange before you standardise analytics: FHIR-aligned interfaces outlive individual tools.
- Adopt ICD-11 and terminology standards for comparability, which WHO frames explicitly as an interoperability requirement.
- Use WHO SMART Guidelines digital adaptation kits so clinical logic is implemented consistently rather than reinvented per vendor.
- Fix routine data quality first — an AI layer over unverified routine data industrialises the error.
References and sources
Every factual claim above is drawn from the documents below. Where a figure could not be confirmed against a primary source, the article says so instead of quoting it.
- 01Ethics and governance of artificial intelligence for healthWorld Health Organization · 2021
- 02Ethics and governance of artificial intelligence for health: guidance on large multi-modal modelsWorld Health Organization · 2024
- 03The Epidemic Intelligence from Open Sources (EIOS) initiativeWorld Health Organization · 2025
- 04WHO upgrades its public health intelligence system to boost global health securityWorld Health Organization · 2025
- 05Global strategy on digital health 2020-2025World Health Organization · 2021
- 06SMART GuidelinesWorld Health Organization · 2025
- 07ICD-11 2022 releaseWorld Health Organization · 2022
Work with me
Need this applied to your programme?
I support governments, development partners and research teams on health systems planning, implementation research, evaluation design and health information systems.
Continue reading
Technical & digital health
Making DHIS2 data good enough to decide with: a data-quality workflow for municipal health teams
12 min readHealth policy
The six global health security agendas that will decide the next decade — and what they demand of a country like Nepal
19 min readResearch methods