Skip to content
Sonivance Limited logoInsights . Evidence . Better Decisions

M&E

Using AI in Monitoring and Evaluation Without Weakening Your Evidence

30 September 2026 · 8 min read

AI is already inside most evaluation work, usually quietly. Here is where it helps, where it puts your evidence at risk, and how a Kenyan team can start without losing human oversight.

Analyst reviewing survey data and quality checks on screen in a Nairobi office

AI is already inside most evaluation work

Ask a programme team whether they use AI and the answer is often no. Then ask how the interview recording became a transcript, how the open-ended answers were grouped, or who drafted the summary that opened last quarter's report, and the picture changes. The tools have arrived in the workflow ahead of the policy that would describe them.

Research on the third sector suggests this unevenness is normal rather than local. A systematic review published in 2025 examined 65 studies of AI adoption in non-governmental organisations and found adoption uneven and skewed towards larger organisations, with use cases falling into broad groups such as engagement, decision-making, prediction, management and optimisation. Smaller teams were not less willing, they were less resourced and less guided.

For an NGO, a research firm or a university unit in Kenya, the practical question is therefore not whether to use AI. It is which tasks to hand over, which to keep, and what has to be true before a model touches anything that ends up in a report a donor reads.

Where it earns its place

The tasks where AI tends to pay for itself share a pattern: they are laborious, they have a clear right answer that a person can check, and they sit before a human decision rather than replacing one.

  • Transcribing and translating interviews, so a team working in three languages can read the same material.
  • First-pass coding of open-ended responses, as a starting structure that a researcher then revises.
  • Drafting consent text, interview guides and survey prompts that an evaluator edits rather than accepts.
  • Flagging implausible values, duplicate submissions and skipped patterns inside a dataset for someone to investigate.
  • Summarising long documents, donor reports and previous evaluations into a briefing a team can argue with.
  • Turning a verified finding into plain language for different audiences, with the underlying analysis untouched.

Where it puts your evidence at risk

The EvalSDGs guidance on AI in evaluation sets out seven risks worth taking seriously: bias and fairness, data misinterpretation, privacy concerns, security vulnerabilities, ethical dilemmas, reliability and robustness, and over-reliance on technology. Those read as abstract categories until you translate them into what could actually happen in a study.

A model summarising open-ended answers may quietly flatten dissent, presenting three thoughtful minority responses as one clean theme. Another may produce confident statements about a community it has never seen, in a county whose realities it does not know. Someone may paste identifiable respondent data into a public tool without a lawful basis or any intention of telling the person who gave it. And a finding generated by a model may reach a report with no trail back to a source, which means nobody can defend it when a reviewer asks.

Two things follow. Personal data collected for research in Kenya sits under the Data Protection Act, so processing it through a third-party service needs a lawful basis and clear communication with respondents; check your own compliance position rather than assuming a popular tool is a safe one. And any output that influences a conclusion needs a person who can trace it, because a model's confidence is not evidence.

What a sensible first year looks like

Organisations that have done this well describe an unglamorous sequence. The GAIN alliance, which has published what it learned running internal experiments across programme teams, started with training for staff, then a short policy of do's and don'ts covering data security, then focused workshops with specific teams on specific operational problems. That order matters. Rules written before people understand the tools go unread, and tools adopted before rules exist get used in ways nobody can explain later.

  • Choose one low-stakes task, such as transcribing interviews that are already recorded with consent.
  • Keep sign-off human and name the person responsible for each automated step.
  • Record what was automated, when and by whom, so a reviewer can reconstruct the process.
  • Tell respondents when a third-party service will handle their recordings or responses.
  • Keep identifiable data out of public tools unless your agreements and approvals clearly allow it.
  • Check a sample of outputs against the source every time, not only the first time.

What not to hand over

It is easier to describe the boundary by listing what stays human. Anything that determines whether a programme worked, what caused a change, or how a community should be described in a public document requires a person who can be asked how they reached the conclusion.

  • Attribution and final judgements about what changed and why.
  • Findings that name communities, individuals or specific sites.
  • Anything a reviewer cannot verify against a source they can read.
  • The decision to publish, and responsibility for what the report claims.

How a research partner fits in

Most organisations do not need an AI strategy, they need a few decisions made clearly. Which tasks may be automated, who checks the output, what data may leave the building, and what a reviewer is entitled to see. Writing those down takes an afternoon and prevents most of the arguments later.

Sonivance Limited works with NGOs, research firms and universities on field data collection, monitoring and evaluation, and analytics. Our position is practical rather than ideological: models may assist transcription, first-pass coding and quality flags, while supervision, back-checks and interpretation stay with named people. If your team is deciding where the line sits for a particular study, we can help you set it before fieldwork starts rather than after a finding has already been drafted.

Frequently asked questions

Can AI replace field enumerators?
No. It can speed up transcription, translation and some checking, but community entry, sensitive interviews, judgement about who is actually at home and the discipline of a back-check all stay human. Removing supervision does not make fieldwork cheaper, it makes the results harder to defend.
Is it safe to put survey data into an online AI tool?
It depends what the data contains and what your own agreements allow. Identifiable responses should not go into a public tool without a lawful basis, respondent information and internal approval. Aggregated or clearly de-identified open-ended text is a different risk level, and your organisation's data protection policy, not the tool's marketing, is the thing to follow.
Does AI actually improve data quality?
It helps with first-pass checks such as duplicates, impossible values and suspicious completion times, which is genuinely useful. It does not replace supervision, back-checks and daily review, and those are what catch fabrication and misunderstanding in the field.
Where should a small NGO start if resources are tight?
Pick one administrative bottleneck rather than a transformation, train the two people most likely to use the tool, write one page of rules covering what may and may not be automated, and review after a month. A small, governed use is worth more than an ambition that never gets written down.

Have a Research Question?

Let's turn it into actionable evidence.

Whether you need a baseline study, field data collection, program evaluation, market research, data analysis or a research partner for a larger assignment, Sonivance is ready to help.