AI in Humanitarian Work: What Is Useful, What Is Hype, What Is Harmful
The state of play in 2026
The humanitarian sector spent the early 2020s in a period of cautious experimentation with AI. By 2026 the experimentation phase is over. UNHCR, IOM, WFP, OCHA, the major INGOs, and a long list of national-society actors are running AI in production. Some of it is genuinely transformative. Some of it is sold as transformative and is closer to a fancy spreadsheet. Some of it has caused harm severe enough to warrant suspension. The honest picture is that all three are true at once, and the only way to read the field clearly is to look at it use case by use case.
Use case 1: displacement forecasting
The Danish Refugee Council's Foresight model and the Jetson project run by UNHCR's Innovation Service are the best-documented examples. Both ingest a wide range of structured indicators, including conflict events from ACLED, food security data from IPC, market price volatility, weather anomalies, and historical movement records, and produce probabilistic forecasts of cross-border arrivals weeks to months ahead.
The good news is that on stable corridors with consistent historical data, these forecasts are useful. They give country offices enough lead time to pre-position supplies, negotiate with host-government counterparts, and request emergency funding before the movement is visible at the border. The less good news is that they degrade rapidly when the underlying drivers change abruptly, which is exactly the moment humanitarian responders most want a forecast. The model that worked well predicting Somali movements into Kenya in 2019 was less useful in 2023 when drought, conflict in Ethiopia, and a shifting Kenyan policy environment all moved at once.
The practitioner consensus, which our reading of the literature supports, is that displacement forecasts are valuable as one input to a planning process and dangerous as a single source of truth. The error bars matter as much as the central estimate. For the underlying concepts these forecasts are trying to predict, see our guide to understanding displacement.
Use case 2: satellite imagery analysis
Computer vision applied to satellite imagery is one of the clearest wins. Models trained on labelled examples can identify destroyed buildings, expanding informal settlements, agricultural land taken out of cultivation, and changes in IDP camp footprint at a scale that was simply impossible with human analysts. UNOSAT, the Humanitarian OpenStreetMap Team, and several commercial providers now produce damage assessments within hours of new imagery being available.
The limits are worth naming. Cloud cover, dense tree canopy, and dust storms degrade results. Models trained on imagery from one region transfer imperfectly to another. The difference between a destroyed building and one with a tarpaulin roof is sometimes invisible from above. And the cadence of imagery for any specific place is governed by commercial satellite tasking decisions that humanitarian responders do not control.
The pattern that works is to use AI for the wide-area screening pass, then have human analysts verify the flagged tiles. The pattern that fails is to publish a casualty figure based on an unreviewed model output. We use Esri World Imagery in our country dashboards and treat satellite-derived figures as triangulation against ground-truth registration data, never as a substitute. The deeper question of which AI failures actually fall on displaced people is taken up in our piece on AI bias and displacement data.
Use case 3: language and translation
Translation between widely spoken languages is one of the unambiguous successes. A field officer collecting testimony in Arabic, French, or Spanish can now produce an English-language working draft in seconds rather than waiting for a human translator. This is real time saved on real cases.
The failure modes are predictable. Under-resourced languages are translated poorly or not at all. Languages like Rohingya, with limited written tradition, sit largely outside what current models handle. Dialects are flattened toward the majority form. Specialised vocabulary in protection, health, and legal contexts is mistranslated in ways that can change the meaning of a case note.
The right discipline is to use AI translation as a first pass for triage and to commission human translation for any document that will be acted on or published. The wrong discipline, which we have seen in evaluations of asylum-processing systems in several jurisdictions, is to treat machine translation as definitive evidence in a status determination.
Use case 4: case management and triage
Several large agencies now use AI to triage incoming protection cases by urgency, route them to the right caseworker, and surface relevant precedent from the case archive. The labour-saving is real. A caseworker who would have spent two hours triaging a backlog now spends fifteen minutes verifying a model's ranking and another fifteen minutes on the cases that need a closer look.
The risk concentrates in two places. First, any case the model misclassifies as low priority can sit untouched for weeks while the situation deteriorates. Second, models that learn from past caseworker decisions inherit the patterns of those decisions, including the systematic ones. A model trained on a backlog where unaccompanied minors were historically prioritised will reproduce that prioritisation. A model trained on a backlog where a particular nationality was historically deprioritised will reproduce that too, silently, until someone audits it.
The agencies doing this well publish their triage models' performance disaggregated by case type and by demographic group, run regular shadow comparisons with human triagers, and keep a clear path for a caseworker to override the model's ranking without bureaucratic friction.
Use case 5: chatbots for affected populations
Information-distribution chatbots are now widespread. A displaced person can message a service to find out which clinic is open, which border is passable, what documents they need to register, and what their rights are in their country of asylum. At their best, these tools push accurate, current information directly to people who need it, in their language, on a device they already have.
At their worst, they hallucinate eligibility criteria, give wrong addresses, and provide reassurances about safety that turn out to be wrong. Several published audits have found chatbots in the field producing materially incorrect guidance on legal status and on access to services. The cost of a confidently wrong answer in this context is borne by the person who acted on it.
The right pattern is again retrieval-augmented: the chatbot is grounded in a curated, current knowledge base maintained by the agency, with a clear escalation path to a human for anything outside the knowledge base and an unambiguous disclosure that the user is talking to a machine. The wrong pattern is to point a general-purpose model at the public web and hope.
Use case 6: predictive analytics for funding and operations
Donor agencies are increasingly using AI to predict which projects will deliver against their indicators, which countries are likely to slide into crisis, and where to allocate resources for maximum impact. The technology is the same as in other sectors. The risk is also the same: optimising for measurable outcomes shifts attention away from outcomes that are harder to measure, and the politics of which indicators are chosen is often the politics of who gets funded.
Where AI has already caused measurable harm
Three documented examples are worth naming because they recur in different forms in different jurisdictions. The AI bias and displacement data article unpacks why these failures recur structurally rather than as isolated bugs.
Biometric registration systems that perform worse on darker skin tones, on older faces, and on faces of people who have been physically affected by malnutrition or trauma. Where such systems are the gateway to assistance, their error rate is a denial-of-service against the populations they perform worst on. Multiple UNHCR-implemented biometric systems have been independently audited; the gap in match rates across demographic groups is real and has narrowed only partially.
Risk-scoring at borders. Several European and North American jurisdictions deploy AI systems to rank arriving asylum-seekers by perceived risk. Where the training data reflects historical patterns of which nationalities have been refused, the model reproduces those patterns. Where the model's output drives a decision, the patterns are laundered through what looks like a neutral technical process. The EU AI Act now classifies these systems as high-risk and imposes transparency and contestability requirements. Implementation is uneven.
Automated benefit-eligibility decisions. Several governments have used algorithmic systems to determine eligibility for social assistance, including for refugee and migrant populations. The Dutch *toeslagenaffaire* and the Australian *robodebt* scandal are the most-cited cases, but the pattern recurs. When an algorithmic system makes a benefit-denial decision and the appeal process is slow, the harm is immediate and the remedy is delayed.
Principles that have held up in practice
After several years of watching what works and what does not, a small number of principles have emerged from sector practice that are worth taking seriously.
Use AI for triage, not for terminal decisions. A model that flags cases for human attention is useful. A model that closes cases without human review is dangerous.
Ground generative outputs in primary sources. Free-running language models that answer from training memory will hallucinate. Models constrained to read a verified document set before answering will not, or will at least fail in detectable ways.
Disclose the use of AI. Affected populations have a right to know when they are interacting with a machine. Readers of analysis have a right to know when a summary was drafted by a model.
Audit by sub-group. Aggregate accuracy is a deceptive metric. Performance disaggregated by nationality, gender, age, and language is what reveals whether a system is fair.
Keep a path to redress. Every AI-mediated decision needs a path to a human reviewer that does not require the affected person to understand the technology. If that path does not exist, the system is not ready.
How we apply this on Humanity Centered Data
The platform uses AI in three places. We draft synthesis paragraphs for country briefings using a retrieval-augmented pattern that reads only verified primary sources. We use computer vision to flag changes in satellite imagery for our country pages, with human review before publication. We use translation models internally to draft English-language summaries of reports published in other languages, with a clear disclosure on any page where translation is the original source of the text.
Every AI-assisted page carries an amber disclosure banner, an inline citation for every numerical claim, and a link to the editorial standards page that explains the workflow in detail. The reading list above is the foundation. The principles above are the discipline. The reason both matter is that humanitarian data is read by people who make resource-allocation decisions, and the cost of a confident wrong answer in that context is paid by people who never get to see the model that produced it.
Sources
- Danish Refugee Council, Foresight displacement forecasting model. Documented operational forecasting work cited in use case 1.
- UNHCR Innovation Service, Project Jetson. Reference site for the displacement-forecasting pilot cited above.
- ACLED and IPC, primary structured data feeds used by humanitarian forecasting models.
- UNOSAT and the Humanitarian OpenStreetMap Team, satellite-imagery analysis providers referenced in use case 2.
- Amnesty International, Xenophobic Machines: Discrimination through unregulated use of algorithms in the Dutch childcare benefits scandal, 2021.
- Royal Commission into the Robodebt Scheme, Final Report, 2023.
- European Commission, EU AI Act regulatory framework. Cited for the classification of risk-scoring systems.
- ICRC, Handbook on Data Protection in Humanitarian Action. Reference text for the principles section.
- OCHA Centre for Humanitarian Data, Predictive Analytics: Models, Peer Review, Community. Sector-level documentation of the forecasting use cases discussed.
Keep reading
AI Bias, Displacement Data, and the People Who Get Left Out
A model is a mirror of its training data. In displacement work, the people missing from the data are the people most likely to be missed by the model.
Syria's Displacement Landscape: Internal and Cross-Border Trends
Analysis of Syrian displacement data capturing internal displacement and cross-border refugee movements, essential for humanitarian planning and recovery assessment.
Ukraine Displacement Patterns: Data Signals from an Ongoing Conflict
Analysis of Ukraine's displacement crisis showing how internal and cross-border movement data informs humanitarian coordination and host-country planning.
Sudan's Displacement Crisis: Interpreting Humanitarian Data Trends
Analysis of Sudanese displacement data revealing concentrated outflows, uneven settlement patterns, and the importance of continuous monitoring in volatile environments.
Afghanistan Displacement Dynamics: A Data-Driven Perspective
Analysis of Afghan displacement patterns reflecting decades of instability, with insights for humanitarian planning and long-term support strategies.
What Artificial Intelligence Actually Is, In Plain Language
AI is not a thinking machine. It is a pattern matcher trained on human output, and the quality of what it produces is bounded by the quality of the data you give it.
