The MEL system, layer by layer
Most MEL systems fail in the same place: they are built from the indicator sheet upwards instead of from the decision downwards. This module sets the stack that the rest of the toolkit fills in.
A monitoring, evaluation and learning system is not a document or a database. It is the set of arrangements by which a programme decides what it needs to know, gathers evidence proportionate to that need, judges what the evidence means, and changes something as a result. Everything in this toolkit sits in one of five layers. If a layer is missing, the layers above it will produce activity without insight.
Five questions that must be answered before any indicator is written
Who is the primary intended user?
Name people, not institutions. A programme board, a ministry directorate, a country office lead. Utilisation-focused evaluation holds that use is determined by who is at the table when questions are set, not by report quality.
What decision is waiting on this?
Budget reallocation, extension, redesign, scale-up, exit. If no decision is pending, the exercise is accountability reporting and should be scoped and priced as such.
What is the causal claim?
The theory of change fixes what the programme says it will cause. Without it, evaluation questions drift into description and findings cannot be judged against intent.
How certain and how agreed is the terrain?
Where cause and effect are known and stakeholders agree, performance monitoring works. Where either is low, complexity-aware approaches are needed alongside it. See Module 06.
What is proportionate?
Match method to stake and budget. A 200,000 USD pilot does not need a randomised trial; a 60 million USD flagship needs more than a quarterly output table.
Who else is trying to cause this change?
Contribution, not attribution, is the honest default in most development settings. Name the other actors early so the design can rule alternative explanations in or out.
The Development CAFE has run monitoring, evaluation and learning platforms in more than forty countries for donors including USAID, FCDO, DFAT, DFC, GIZ, the European Union, the World Bank and the ASEAN Secretariat. This toolkit is the working reference we use internally, published so that partners, commissioners and country teams can use the same vocabulary. Everything here is either drawn from public donor guidance and cited, or is DevCAFE original work and marked as such.
Writing the scope of work
The scope of work is where most evaluations are won or lost. DFAT's own review of its evaluations found that weak plans were weak in predictable ways: thin sampling, no ethics, no triangulation, no clear route from evidence to conclusion.
A scope of work (also called terms of reference) is a commissioning instrument, a contract schedule and a design brief at the same time. Written well, it constrains the bidder toward good practice. Written badly, it invites a methods section that says "mixed methods" and means nothing. DFAT's 2024 to 2025 review of evaluation quality found that poor-quality plans gave little detail on how findings would be used, lacked key design elements, paid limited attention to ethics or flexibility, explained sampling strategies insufficiently, made little use of monitoring or secondary data, left data collection tools underdeveloped, and offered no clear process for judging the strength of evidence. Each of those gaps is preventable at the scope of work stage.
The twelve required sections
- Background and contextWhat the intervention is, its budget, duration, geography, delivery chain and current status. Include what has already been evaluated so the bidder does not re-tread it.
- Purpose, audience and intended useName the decision, the decision-maker and the date the finding is needed. State whether the exercise is formative, summative, developmental or a mid-term learning review.
- Evaluation questionsFive to eight, not twenty. Each question should be answerable, should map to a criterion, and should have a plausible method attached. Sub-questions belong in an annex, not the main list.
- CriteriaState explicitly which criteria apply. The European Commission assesses interventions against relevance, effectiveness, efficiency, sustainability, impact, coherence and EU added value. The OECD DAC set of six is the common floor; add value for money, gender equality and inclusion, or localisation where the commissioner requires them.
- Scope and boundariesTime period, components in and out, geographies, and what the evaluation will not cover. Boundaries prevent scope creep and price disputes.
- Approach and methodsSay what is required (for example, a contribution analysis, disaggregated sampling, or a specific participatory element) and what is left to the bidder. Require a sampling rationale and a triangulation strategy, not just a tool list.
- Ethics, safeguarding and data protectionEthical review route, informed consent, handling of personal and sensitive data, safeguarding reporting line, data retention and destruction, and applicable law. In Indonesia this means the Personal Data Protection Law; in the EU, GDPR.
- Deliverables and quality standardsInception report, tools, draft, final, and a two-page summary. Name the quality standard the deliverables will be assessed against and attach it.
- Timeline and level of effortDays by role, with the inception phase given real weight. Under-resourced inception is the single most common cause of weak evaluation design.
- Team compositionRequired expertise, language, national representation and independence. DFAT's standards emphasise a clear specification of the mix of technical skills and attributes required, including the role of the team leader and local team members, and require consideration of the use of local consultants with the requisite skill sets.
- Management, governance and conflict of interestWho approves what, the reference group, and the conflict of interest declaration each team member must sign.
- Use, dissemination and management responseHow findings will be discussed, who owns the management response, and where the report will be published. Build the response step into the timeline, not after it.
Too many questions, no priority order. A scope of work with eighteen evaluation questions and a forty-day budget guarantees a shallow answer to all of them. Rank the questions, mark three as primary, and say plainly that the remainder are secondary and may be answered at lower confidence.
What each commissioner expects in the methods section
| Commissioner | Anchor guidance | What the methods section must show |
|---|---|---|
| USAID legacy Program Cycle | ADS 201, Evaluation Toolkit | Link between evaluation questions and the results framework, data quality standards, conflict of interest disclosure forms, and a statement on how findings feed the next strategy |
| FCDO | Magenta Book (2026), FCDO evaluation strategy and policy | Theory of change, evaluation type declared as process, impact or value for money, feasibility appraisal, and a route from evidence to counterfactual reasoning |
| DFAT | Design and MEL Standards | Explicit sampling strategy, ethics and flexibility, triangulation, a stated process for judging strength of evidence, and how findings will be used |
| European Commission (INTPA and FPI) | Evaluation Handbook (2024); Results Oriented Monitoring | Coverage of the seven EU criteria, an evaluation matrix mapping questions to judgement criteria and indicators, and alignment with the intervention logic in the Action Document |
| GIZ | Central project evaluations, OECD DAC criteria | Assessment along the results model, contribution analysis of the module objective, and a rating scale applied transparently |
| World Bank | PDO and intermediate results indicator structure | Attribution or contribution logic against the project development objective, and treatment of data limitations |
Theory of change
A theory of change is both a process and a product. The product is a diagram and a narrative. The process is the argument the team has about why any of it should work, and that argument is the part with value.
Isabel Vogel's review for DFID remains the reference point for how theories of change are actually used in development. It found that the element shared across all approaches is the assumptions, with the standard instruction to make them explicit, and that getting depth and critical thinking on assumptions is widely agreed to be the crux of the process. Interviewees reported that many theories of change they had seen were superficial and mechanistic, without exploring the deeper questions. The practical implication is that a theory of change workshop that produces a tidy diagram in two hours has probably failed.
The review also noted a shift toward identifying several relevant pathways to impact for a given initiative rather than a single pathway, acknowledging non-linearity and emergence, and treating documented theories and diagrams as subjective interpretations used as evolving organising frameworks.
Building one in eight moves
- Agree the long-term changeOne sentence, in the language of the people who would experience it. Not "improved enabling environment" but "smallholder cocoa farmers earn a living income from the same land".
- Work backwards, not forwardsAsk what must be true immediately before the long-term change, then what must be true before that. Backward mapping surfaces preconditions that forward planning skips.
- Name the social actorsChange happens through people and institutions behaving differently. For each outcome, name who is doing something differently. Vagueness here is the leading cause of unmeasurable outcomes.
- Surface the assumptions at every linkBetween each pair of boxes, ask what has to hold for the arrow to be real. Write these down as numbered assumptions and rate each for confidence and evidence base.
- Map the other actors and factorsWho else is pushing on the same outcome. This is what allows a later evaluation to reason about contribution rather than assert it.
- Identify multiple pathwaysMost programmes have two or three routes to the same outcome. Draw them. A single linear chain is usually a simplification imposed by the reporting template.
- Write the narrativeThe diagram alone is not a theory of change. The narrative explains the causal reasoning, the evidence behind it, the contested points and what the team is uncertain about.
- Set review pointsDiarise when the theory will be revisited, usually annually and after any major context shift. A theory of change that has never been revised has never been used.
Quality check before you sign it off
- Could a sceptical outsider identify what would have to happen for this theory to be wrong?
- Is every outcome expressed as a change in what somebody does, not as an activity in disguise?
- Are the assumptions specific enough to be monitored, and is at least one indicator attached to the riskiest of them?
- Does the narrative say where the evidence base is strong and where the team is guessing?
- Do the people who deliver the programme recognise it, and can they explain it without the diagram in front of them?
- Has it been tested against at least one alternative explanation for the change you expect?
Theory of change explains why change is expected, including assumptions and other actors. Results framework shows the hierarchy of results the programme is accountable for. Logframe is a contractual matrix of results, indicators, means of verification and risks. The theory of change is the reasoning; the other two are the accountability instruments derived from it. Donors sometimes ask for all three. They are not interchangeable and should not be produced by reformatting one into another.
Results frameworks and logframes
Every donor uses a results hierarchy. They differ in how many levels there are, what each level is called, and who is accountable for it. Translating between them is a core practitioner skill.
The underlying logic is stable across systems: resources produce deliverables, deliverables are taken up and used, use produces change in behaviour, and behaviour change produces change in conditions. What varies is nomenclature and the level at which the implementer is held accountable. Getting this wrong in a proposal is an avoidable scoring loss.
The logframe matrix, and what usually goes wrong in it
| Column | What it should contain | Frequent failure |
|---|---|---|
| Results statement | A single change, stated as an achieved state, at one level only | Two changes joined by "and", so it can never be scored cleanly |
| Indicator | The observable signal that the result has occurred, with unit and disaggregation | An activity count parked at outcome level |
| Baseline | Value and date, with source | "To be determined", left unfilled for the life of the programme |
| Target | Value by date, justified by a stated rationale | Round numbers with no basis, set to look ambitious at approval |
| Means of verification | Named instrument, responsible person, frequency | "Project records", which nobody can audit |
| Assumptions and risks | Conditions outside control that must hold, linked to the theory of change | Generic risks copied between programmes, never reviewed |
DG INTPA reviews all Action Documents for consistency of logframes and intervention logic and for uptake of corporate indicators before quality review meetings. Its sector indicators guidance is designed to help teams build solid logical framework matrices aligned to the SDGs, the NDICI regulation, the Gender Action Plan and the Global Europe Results Framework. If you are writing for an EU-funded action, drawing indicators from that corporate set rather than inventing your own materially improves the design score.
SMART indicators and reference sheets
SMART is a floor, not a standard. It tells you whether an indicator is well formed. It does not tell you whether it is worth measuring.
Specific
One phenomenon, one population, one unit. If two people could count it differently, it is not specific.
Measurable
There is a defined method and a data source that exists or can be created within budget.
Achievable
The target is reachable given the intervention's scale and duration. Also read as "attributable" in some donor systems.
Relevant
It measures the result, not a proxy for effort. It answers a question someone will act on.
Time-bound
It has a reporting frequency and a date by which the target applies.
Disaggregated
Sex, age, disability and location as a minimum, and any dimension the programme claims to affect. A number without disaggregation cannot answer an equity question.
The seven tests USAID applies on top of SMART
Legacy USAID Program Cycle guidance asks whether an indicator is direct, objective, useful for management, practical to collect, attributable to the intervention, timely, and adequate in combination with others. The most useful of these in practice are direct (does it measure the result itself, or a distant relative?) and practical (can the field team actually collect it at the required frequency without distorting delivery?).
Five data quality standards, applied to every indicator
USAID assesses performance data against five standards: validity, reliability, timeliness, precision and integrity. These are the questions a data quality assessment answers, and they are the right questions regardless of who is funding the work.
| Standard | The question | What a failure looks like in the field |
|---|---|---|
| Validity | Does the data clearly represent the intended result? | Attendance sheets used as evidence of skills gained |
| Reliability | Is it collected by consistent, unbiased methods over time? | Enumerators change the question wording between rounds |
| Timeliness | Is it available frequently enough for management decisions? | Annual data used to steer a quarterly adaptation cycle |
| Precision | Is the level of detail sufficient for the decision? | National totals when the decision is district-level resourcing |
| Integrity | Are there safeguards against error and manipulation? | Single-person data entry with no verification and incentives tied to targets |
When a target becomes a performance incentive, the measure degrades. Programmes reach numbers by shifting to easier-to-reach participants, by re-counting the same people across activities, or by redefining what counts as a "trained" person. Guard against it with unique identifiers, spot verification, disaggregation that would reveal displacement, and by pairing every headline count with one quality measure.
Indicator types, and how many you need
Output indicators
Count what the programme delivered. Cheap, timely, and prone to being mistaken for progress. Keep them, cap them.
Outcome indicators
Measure change in what social actors do. Slower, more expensive, and the only level at which programme claims can be tested.
Proxy indicators
Stand in where direct measurement is impossible. Must be justified with the reasoning for why the proxy tracks the real thing.
Sentinel indicators
A signal from the wider system that prompts investigation, not a target to hit. See Module 06.
Qualitative indicators
Rubric-scored judgements against defined descriptors. Rigorous when the rubric is set in advance and applied by more than one rater.
How many
One to three per result statement. A logframe with sixty indicators is a data collection burden that will be met by fabrication or by nobody reading it.
Complexity-aware monitoring
Performance monitoring answers whether the plan is being delivered. It cannot see what nobody planned for. Complexity-aware monitoring is the set of approaches that covers the blind spots.
USAID's discussion note on complexity-aware monitoring is the clearest short statement of the problem. It holds that complexity-aware monitoring is appropriate where cause and effect relationships are poorly understood, which makes it difficult to identify solutions and draft detailed implementation plans in advance, and where expected results may need refinement as strategies unfold. It is a complement to performance monitoring, not a replacement for it.
Three principles
Address the three blind spots
Performance monitoring misses unintended outcomes, alternative causes from other actors and factors, and non-linear pathways of contribution including feedback loops.
Synchronise with the pace of change
Use leading, coincident and lagging indicators so that the monitoring rhythm matches how fast the thing you care about actually moves.
Attend to systems concepts
Interrelationships, perspectives and boundaries. What counts as inside the system, and whose account of it are you using?
Five recommended approaches
| Approach | What it does | Blind spot it covers |
|---|---|---|
| Sentinel indicators | A proxy for the system that signals the need for further investigation, described as a canary in a coal mine | Pace of change; early warning of system shifts |
| Stakeholder feedback | Systematically seeks the perspectives of partners, participants and those excluded from a project | Perspectives; unintended effects on non-participants |
| Process monitoring of impacts | Tracks the predicted and emergent processes that transform outputs into results | Non-linear pathways; the gap between delivery and effect |
| Most significant change | Captures results across diverse stakeholder groups and makes each group's perspective on those results explicit | Unintended outcomes; whose values define significance |
| Outcome harvesting | Works backwards from observed change to describe and verify contribution, capturing unexpected results | All three, particularly alternative causes |
Complexity-aware approaches are often read by commissioners as soft or unaccountable. Frame them as risk management. A programme that only monitors its plan is blind to the two things most likely to end it: an unintended harm it did not look for, and a change that happened for reasons other than its work. Budget them as a named line, typically five to fifteen per cent of the MEL budget, and report them alongside the logframe rather than instead of it.
Outcome harvesting
Start from the change, not from the plan. Outcome harvesting collects evidence of what has been achieved and works backwards to establish whether and how the intervention contributed, in the way that forensics or archaeology reason from what is left behind.
Outcome harvesting was developed by Ricardo Wilson-Grau and colleagues, drawing on outcome mapping and utilisation-focused evaluation, and refined through real-life practice. It enables evaluators, grant makers and managers to identify, formulate, verify and make sense of outcomes.
An outcome is a change in the behaviour of one or more societal actors: their actions, activities, relationships, policies or practices. The harvester identifies demonstrated, verifiable changes in behaviour influenced by an intervention, and how the project or programme plausibly contributed to them. Outcomes can be positive or negative, intended or unintended, and the contribution can be large or small. Both the outcome and the contribution must be specific and measurable enough to be verifiable.
Unlike most approaches, outcome harvesting does not necessarily measure progress against predetermined objectives. Each application is customised to the needs of primary intended users and their principal intended uses, with useful and feasible harvesting questions guiding the collection of information. The initial sources are the individuals or organisations whose actions influenced the outcomes, because they know what was achieved and are motivated to share it.
What happens in each step
- Design the harvestHarvest users and harvesters identify usable questions to guide the harvest and agree what information will be collected as the outcome description, in addition to the change in the social actor and how the intervention contributed. Design decisions continue to be made throughout, as new information about outcomes emerges.Yield: harvesting questions, outcome description format, sources map
- Review documentation and draft outcomesDesk review of reports, memos, media, case files, promotional material and websites to identify and formulate draft outcome statements.Yield: a first list of draft outcome statements
- Engage with informantsHarvesters work directly with change agent informants to review the outcome descriptions extracted from files, identify and formulate additional outcomes, and classify them. Informants often consult others inside or outside their organisation who are well informed about the outcomes.Yield: refined, precise, verifiable outcome statements
- SubstantiateObtain the views of one or more independent people knowledgeable about the outcome, or about a representative sample of outcomes, and about how they were achieved, to enhance the validity and credibility of the findings. Substantiation also enriches understanding and can surface outcomes the primary users had not identified.Yield: independently verified subset, plus new outcomes
- Analyse and interpretOrganise outcome descriptions in a database to make sense of them, then analyse and interpret to provide evidence-based answers to the harvesting questions. Classification by actor, type, geography, significance and contribution strength is what makes patterns visible.Yield: evidence-based answers to the harvesting questions
- Support use of findingsPropose points for discussion to harvest users, grounded in the evidence-based answers. Use is designed in from step one; it is not a dissemination afterthought.Yield: decisions, adaptations, and a documented management response
Anatomy of an outcome statement
Outcome descriptions may be as brief as a single sentence or as long as a page. They may also carry the significance of the outcome, its context, the contribution of others, diverse perspectives, and anything else useful for answering the harvesting questions. Three components are non-negotiable.
The change
Who changed, what they now do differently, when, and where. Specific enough that a third party could check it.
The significance
Why this change matters for the goal. This is where the outcome is connected to the wider theory, and it is the part most often left blank.
The contribution
What the intervention did that plausibly influenced the change, stated modestly and specifically, with other contributors named.
OUTCOME (September 2025, Sulawesi, Indonesia) CHANGE The provincial trade agency published a revised electronic licensing standard operating procedure that removes three duplicate document requirements for small exporters, and began applying it in two districts. SIGNIFICANCE Duplicate documentation was the single largest source of clearance delay identified in the baseline. Removing it shortens the median clearance time and lowers the fixed compliance cost that falls hardest on the smallest exporters. It also sets a precedent other provinces have cited. CONTRIBUTION The programme ran the process mapping workshop that produced the duplicate list, drafted two of the four revision options, and funded the legal review. The reform was championed internally by the agency's deputy head, and the national chamber of commerce lobbied in parallel. The programme's role was technical and enabling, not decisive. SUBSTANTIATED BY Deputy head, provincial trade agency (interview, 12 Oct 2025); independent customs broker association (interview, 19 Oct 2025); published SOP document no. 44/2025.
Nine principles behind the six steps
The steps are rooted in nine principles derived from emerging practice: five relate to process and four to content. Rather than telling you what to do, they guide the decisions you make as the harvest proceeds. Two are worth singling out.
Iterative process
Key design decisions are made throughout the harvest as new information about outcomes emerges. A harvest designed entirely in advance is not a harvest.
Strive for less
Principle four reminds harvesters to emphasise learning over collecting a large number of outcomes: fewer outcomes, better understood, are more useful. Resist the pull toward a long inventory.
When to use it, and when not to
Outcomes were not predictable at design
Advocacy, policy influence, network strengthening, systems change, research uptake, capacity development across many actors.
Many actors are contributing
Attribution is impossible but contribution matters. Harvesting handles multiple contributors honestly.
The question is about coverage or cost
Harvesting will not tell you how many people were reached or what a unit costs. Pair it with performance monitoring.
You need a counterfactual
If the commissioner needs an estimate of what would have happened anyway, harvesting is the wrong instrument. Use a quasi-experimental design or contribution analysis.
Outcomes that are really outputs. "Fifty officials trained" is not a change in behaviour. Contribution claims that inflate. If the substantiator would not recognise the programme's described role, rewrite it. Substantiation skipped for budget reasons. Substantiating a purposive sample of the most significant or most contested outcomes is the minimum credible standard, and it is what separates a harvest from a collection of anecdotes.
Most significant change
A story-based, participatory technique in which stakeholders search for significant outcomes and then deliberate on their value in a systematic and transparent way. It answers a question no indicator can: significant to whom, and why.
Most significant change was developed by Rick Davies and refined with Jess Dart, whose 2005 guide remains the standard reference and sets out the technique in ten steps. Dart and Davies describe it as a form of dynamic values inquiry, in which designated groups of stakeholders continuously search for significant programme outcomes and then deliberate on the value of those outcomes, contributing to both programme improvement and judgement.
Field staff and participants write short stories of the most significant change they have seen in a defined domain over a defined period, and say why they chose it. Those stories move up through levels of the organisation. At each level, a group reads them, selects the one it considers most significant, and records the reasons for its choice. The selected stories and the reasons travel upward together. The reasons are the data.
The ten steps
- Start and raise interestIntroduce the technique to stakeholders and secure buy-in. MSC depends on people willingly writing and reading stories; imposed participation produces thin material.
- Define domains of changeBroad, deliberately fuzzy categories, typically three to five, such as changes in people's lives, changes in participation, or changes in institutions. Domains are not indicators and should not be defined precisely.
- Define the reporting periodMonthly, quarterly or half-yearly. Frequency drives the volume of stories and the workload of selection.
- Collect significant change storiesAsk the question directly: looking back over the last period, what do you think was the most significant change in this domain, and why is it significant to you? Record the story and the reason.
- Select the most significant of the storiesGroups read the stories, discuss them and select one, then write down the criteria they applied. This deliberation is the core of the technique.
- Feed back the results of the selectionTell the people who submitted stories which were selected and why. Skipping this step is the most common implementation failure and it kills participation.
- Verify storiesVisit a sample of story locations to confirm the events described. Verification protects credibility and often deepens the account.
- QuantifyOnce a significant change is identified, ask how widespread it is. Stories can be followed by counts of how many others experienced the same change.
- Conduct secondary analysis and meta-monitoringAnalyse the whole body of stories for themes, and monitor the process itself: who submitted, who was selected, whose voice dominates.
- Revise the systemChange the domains, frequency or selection process in light of what the first cycles reveal. MSC is meant to evolve.
Participants often have reservations about what "significance" means, about the acceptability of subjectivity, and about the appearance of competition between stories. These need to be addressed openly at step one. Implementation is also frequently incomplete: changes get described but their significance is left cursory or absent, stories are collected but never put through a selection process, and feedback to the original storytellers is forgotten. If you cannot commit to steps five and six, do not start.
What MSC does and does not deliver
Unintended and unanticipated outcomes
Because the prompt is open, MSC surfaces changes no indicator was watching for.
Explicit values
The selection reasons make visible what different stakeholder groups actually value, which is often not what the logframe rewards.
Representativeness
MSC deliberately seeks the exceptional. It cannot tell you what is typical. Pair with survey or routine data if you need incidence.
A quick win
Story collection, selection at multiple levels, feedback and verification take real facilitation time across several cycles before patterns appear.
The wider method shelf
Method choice follows from the question, the state of knowledge about cause and effect, and what data can honestly be obtained. Filter the shelf below to shortlist candidates.
Filter by what you need
Randomised controlled trial
Random assignment to treatment and control produces an unbiased estimate of average effect. Strong internal validity, weak external validity, and only feasible where assignment can be controlled ethically.
Quasi-experimental designs
Difference-in-differences, propensity score matching, regression discontinuity and interrupted time series construct a comparison where randomisation is impossible. Each rests on assumptions that must be stated and tested.
Contribution analysis
Sets out the theory of change, gathers evidence for each link, assembles the contribution story, seeks out the main alternative explanations, and revises until the claim is credible. The workhorse method where attribution is impossible but a causal claim is still required.
Process tracing
Tests causal mechanisms against evidence using four diagnostic tests: straw in the wind, hoop, smoking gun, and doubly decisive. Suited to single-case policy influence questions where the mechanism matters more than the magnitude.
Realist evaluation
Asks what works, for whom, in what circumstances and why, by building context-mechanism-outcome configurations. Useful where the same intervention produces different results across sites.
Qualitative comparative analysis
Identifies which combinations of conditions are necessary or sufficient for an outcome across a medium number of cases, typically ten to fifty. Bridges case study depth and cross-case generalisation.
Outcome harvesting
Works backwards from verified changes in actor behaviour to plausible contribution. Best where outcomes could not be specified in advance. Covered in full in Module 07.
Most significant change
Story-based values inquiry that surfaces what different stakeholders consider significant and why. Covered in full in Module 08.
Outcome mapping
Plans and monitors change in the behaviour of boundary partners using progress markers graded as expect, like and love to see. Strong for capacity development and network programmes.
Sentinel indicators
A small set of system-level signals watched for movement rather than managed to target. Triggers investigation when they shift. Cheap, and the earliest warning available.
Developmental evaluation
Embeds an evaluator inside an innovation process to feed real-time evidence into ongoing design decisions. Fits pilots, adaptive programmes and social innovation, not accountability reporting.
Process evaluation
Examines whether the intervention was implemented as intended, at what fidelity and dose, and what implementation factors explain variation in results. Often the highest-value evaluation for a mid-term point.
Cost-benefit and cost-effectiveness analysis
Compares costs against monetised benefits, or against units of outcome. Requires a defensible effect estimate first, so it is downstream of an impact design rather than a substitute for one.
Value for money assessment
Assesses economy, efficiency, effectiveness and equity against explicit criteria and rubrics rather than a single ratio. The dominant approach in UK-funded programmes where monetisation is not credible.
Participatory and community-led approaches
Community scorecards, participatory video, ranking and mapping exercises that place judgement with the people affected. Strengthens both validity and legitimacy where the programme claims to serve those communities.
Before and after with plausibility reasoning
The most common design in practice and the weakest. Acceptable only when paired with explicit reasoning about alternative explanations. Never present it as an impact estimate.
Combine, do not choose alone. Nearly every credible design in this field is a mixed one: a counterfactual or contribution backbone, plus a complexity-aware method to catch what the backbone cannot see.
Data systems and data quality
A MEL information system is a chain of custody for evidence. Every handover is a place where quality is lost and where the rights of the people who supplied the data can be breached.
Running a data quality assessment
USAID practice is a useful default even outside US-funded work: indicators reported upward should be assessed for data quality at some point within a three-year period, and new indicators within six months of establishing baseline data. The assessment looks at more than the numbers. It examines the quality of the indicator definition itself, the collection instruments, the collection methods, the database management and the actual data collected.
- Select indicators and assemble the teamPrioritise indicators that drive decisions or carry reputational risk. Include someone who did not collect the data.
- Review the indicator definitionIs it a direct measure? Is it unambiguous about what should be counted and what should not?
- Trace one number end to endTake a single reported figure and follow it back through the database, the submission form and the original record. Most quality problems surface here.
- Test against the five standardsScore validity, reliability, timeliness, precision and integrity, with evidence for each score rather than an assertion.
- Document limitations and actionsRecord what the data can and cannot support, agree corrective actions with owners and dates, and carry the limitations into every report that uses the figure.
Responsible data, in practice
Much MEL data comes from at-risk or underrepresented populations, and it influences decisions that affect those same populations. A responsible data approach that takes ethics and protection into account is not optional. At the design and planning stage, formulate the plan and budget for the full data lifecycle, including the time, systems, people and money involved, and conduct a privacy or risk-benefit assessment to flag risk areas and develop mitigations before collection starts.
Minimum viable data
Collect only what will be used. Over-collection is the norm in this sector and it raises both cost and risk. Challenge every field on the form against a named use.
Consent that is real
In the respondent's language, explaining what happens to the data, who will see it, how long it is kept, and how to withdraw. Recorded, not assumed.
Separate identifiers
Keep direct identifiers apart from response data with a linking key held by a named custodian. Publish only aggregates that cannot re-identify small groups.
Named data owner
One person accountable for each dataset, with access review at least annually and on every staff change.
Retention and destruction
A schedule, applied. Data held beyond its purpose is liability without value.
Do no harm review
Ask who could be endangered if this dataset leaked, and design backwards from that answer. In conflict-affected or civic-space-restricted settings this review is the design.
Tool landscape
| Layer | Commonly used | Choose on |
|---|---|---|
| Mobile data collection | KoboToolbox, ODK, SurveyCTO, CommCare | Offline reliability, enumerator workflow, quality control features, cost at scale |
| Routine and health systems | DHIS2 | Government alignment; use national systems where they exist rather than building parallel ones |
| Programme and indicator management | DevResults, TolaData, ActivityInfo, structured spreadsheets | Indicator hierarchy support, disaggregation handling, audit trail, donor reporting exports |
| Analysis | R, Python, Stata, NVivo, Dedoose, MAXQDA | Team capability and reproducibility. Prefer scripted analysis so results can be re-run and checked |
| Visualisation | Power BI, Tableau, Metabase, Apache Superset, Looker Studio | Licensing cost, hosting and data residency requirements, and who will maintain it after handover |
Before selecting any platform, ask who will run it when the programme closes and whether they can afford the licence. Systems that die at handover were never MEL systems; they were reporting conveniences for the implementer.
Dashboards that get used
The sector has an epidemic of dashboards built to do everything and therefore tailored to nobody. A dashboard is a decision instrument for a named audience, or it is wallpaper.
The problem is well documented in the MERL Tech community: dashboards fail because design principles from graphic design and user experience are not applied, and because tools are built to do all the things without being tailored enough to be usable by the most important audiences. Prototype before you build. A paper prototype of a dashboard, tested with the person who is supposed to act on it, saves months.
Eight rules we apply
- Name the audience and the decision cycle firstA board that meets quarterly needs a different artefact from a field manager who acts weekly. Build two views, not one compromise.
- Lead with a verdict, not a chartThe top of the screen should say what the situation is in words. Charts justify the verdict; they are not the verdict.
- Show exceptions, hide the compliantIf an indicator is on track, the dashboard should be quiet about it. Attention is the scarce resource.
- Pass the five-second testShow it to someone unfamiliar for five seconds and ask what they took away. If they cannot answer, the hierarchy is wrong.
- Cap the filtersTwo filters, chosen deliberately. Every additional control transfers analytical work to a user who does not want it.
- Put disaggregation one click awayEquity questions are answered by breakdowns. If the breakdown is three menus deep, nobody will look.
- Date and source every tile"As at" date, source system, and the data quality flag if one is open. Trust collapses the first time a number cannot be explained.
- Prototype on paper, then buildTest a sketch with the actual decision-maker. Build only what survived the test.
Change over time: line. Comparison across categories: horizontal bar, sorted by value, not alphabetically. Part of a whole: stacked bar, not a pie, unless there are two or three parts. Distribution: histogram or box plot. Geography: map only when location itself is the pattern, otherwise a sorted bar chart is more readable. Two variables: scatter with the relationship stated in words next to it.
FRAME: responsible AI in MEL
Generative AI is now in the working practice of most evaluation teams, usually undeclared. FRAME is The Development CAFE's approach for making that use deliberate, documented and defensible.
FRAME was developed by The Development CAFE and is set out in full in the forthcoming Routledge volume From Algorithms to Evidence: Using Generative AI in Evaluation Practice. The checklist below emerged directly from questions raised by participants in DevCAFE's AI for MEL e-learning course as they grappled with real deployment decisions. Its structure is modelled on the OECD DAC evaluation criteria: complete it at evaluation inception, review it at key milestones, and adapt it to your organisational context and evaluation type. Items marked not applicable should be documented with a rationale.
Five principles
Accountability architecture
A named human is answerable for every AI-assisted output. Decision authority, validation responsibility and documentation responsibility are assigned before any tool is used.
Stakeholder engagement
The people whose data and words are processed, and the people who will use the findings, know that AI is in the workflow and have a route to object.
Epistemic integrity
Claims must remain traceable to evidence. Fluent text that cannot be sourced is not a finding. Hallucination management is a protocol, not a hope.
Transparency
The report says what AI did, at which stage, and what a human verified. Silence about AI use is itself a disclosure failure.
Proportionality
The intensity of governance matches the stakes. A literature scan and a judgement about a programme's future do not warrant the same controls.
The line that does not move
Evaluative judgement is made by humans, not by AI. AI may support synthesis and flag patterns. It does not decide whether a programme worked.
Seven evaluation phases
| Phase | What FRAME requires |
|---|---|
| Design | AI appropriateness assessment across contextual sensitivity, data characteristics, stakeholder expectations, capacity and risk-benefit; plus tool selection criteria covering capability match, transparency and data governance |
| Structuring and planning | Method specification, quality assurance protocols, hybrid human and machine workflow design, alignment to the theory of change |
| Data collection | Enhanced informed consent, real-time disclosure to respondents, input validation, ground-truthing protocols |
| Analysis | Structured and documented prompting, multi-stage iterative validation, hallucination management, confidence calibration, and separate documentation of AI contribution |
| Judgment | Human evaluative conclusions; AI restricted to judgement support such as synthesis and pattern flagging |
| Reporting | Process disclosure, limitation acknowledgement, verification statement, version control |
| Utilization | Accuracy preservation, audience-appropriate disclosure, feedback integration, query monitoring |
The FRAME checklist
Tick as you go. The counter is for your own tracking and resets when the page reloads, so print or export before you close it.
FRAME approach checklist for responsible AI use in evaluation
AI is not responsible; it has no agency. The practical question is how we become responsible creators and users of it. FRAME is written for the people, not the tools.
Learning, adaptation and use
An evaluation that changes nothing has failed regardless of its methodological quality. Use is designed in from the scope of work, not appended at dissemination.
USAID's collaborating, learning and adapting framework treats strategic collaboration among internal and external stakeholders, continuous learning and adaptive management as the connective tissue between all components of the programme cycle. The 2026 Magenta Book update makes a comparable move for UK government evaluation, shifting emphasis from measuring what happened toward learning how to improve in real time, with new guidance on the responsible and ethical use of AI in social research and evaluation and a new section on place-based evaluation.
Four mechanisms that actually produce change
Pause and reflect, scheduled
A recurring session, in the calendar before the year starts, where the team reviews evidence against the theory of change and decides what to change. If it is not diarised, it does not happen.
Learning agenda
A short list of questions the programme genuinely does not know the answer to, owned by named people, with a route by which each will be answered and a date. Not a research wish list.
Management response and tracker
Each recommendation gets accept, partially accept or reject, with a rationale, an owner and a date. Reviewed at the next governance meeting. This single artefact does more for use than any dissemination plan.
After action review
Immediately after a milestone or a shock: what was supposed to happen, what happened, why the difference, what we do differently. Thirty minutes, no blame, written up.
Designing for use from the start
- Identify the primary intended users by name at inception and involve them in setting the questions.
- Ask each of them, at the start, what finding would change their decision. If nothing would, rescope.
- Deliver an emerging findings session before the draft report, so the first time a commissioner meets an uncomfortable finding is not in writing.
- Write recommendations that name an actor, an action and a timeframe. "Strengthen coordination" is not a recommendation.
- Limit recommendations to those the commissioner has the authority to implement, and mark the rest as being for other actors.
- Publish. Evidence that stays inside the commissioning organisation cannot be built on by anyone else.
Some of the most durable effects of an evaluation come from the process rather than the report: teams that learn to think evaluatively, partners who see their own data for the first time, a theory of change that was finally argued about properly. DFAT's own review recommended broadening the definition of evaluation use to take account of the range of ways programmes use evaluations, including learning, signalling intentions and contributing to the field, and to recognise the use of both the evaluation process and its outputs. Record process use deliberately, or it goes unreported.
Donor requirements compared
The same evaluation, written for six different commissioners, is six different documents. This is the crib sheet.
| Dimension | USAID legacy | FCDO / UK | DFAT | EU (INTPA and FPI) | GIZ |
|---|---|---|---|---|---|
| Core guidance | ADS 201 Program Cycle; Evaluation Toolkit | Magenta Book (2026); Green Book; FCDO evaluation strategy and policy | Design and MEL Standards; Development Evaluation Policy | Evaluation Handbook (2024); Results Oriented Monitoring | Central project evaluation guidance; OECD DAC criteria |
| Evaluation criteria | OECD DAC criteria, applied selectively by question | Process, impact and value for money evaluation types | OECD DAC criteria plus gender equality, disability and social inclusion | Relevance, effectiveness, efficiency, sustainability, impact, coherence, EU added value | OECD DAC criteria, rated on a defined scale |
| Results terminology | Goal, development objective, intermediate result, sub-IR | Impact, outcome, output in a logframe | End of programme outcome, end of investment outcome, intermediate outcome, output | Impact, outcome, output in the Action Document logframe | Results model with module objective and indicators |
| Data quality | Five standards: validity, reliability, timeliness, precision, integrity; periodic DQA | Evidence quality and strength of evidence appraisal per the Magenta Book | Standards address sampling, triangulation and strength of evidence explicitly | ROM assessment against eight monitoring criteria including the four core DAC criteria | Data quality addressed through evaluability and evidence strength ratings |
| Learning system | Collaborating, learning and adapting (CLA) across the programme cycle | 2026 Magenta Book emphasises real-time learning and improvement | Learning is named in the standards themselves (DMEL) | Results agenda, annual consolidation of ROM knowledge | Capitalisation of results and cross-project learning |
| Independence | External evaluators submit conflict of interest disclosure forms | Independent evaluation with published reports; ICAI scrutiny of UK aid | Independent evaluation with mandatory publication and management response | Independent external evaluations managed centrally or by delegations | Central evaluations independent of the operational unit |
| Publication | Reports posted to the Development Experience Clearinghouse | Published on gov.uk; development data published to IATI | Published with management response | Published on the Commission evaluation registers | Published evaluation reports and management responses |
| AI in evaluation | No single consolidated policy in legacy guidance | New Magenta Book guidance on responsible and ethical AI use in social research and evaluation | Addressed through ethics and quality provisions | Addressed through Better Regulation and ethics provisions | Addressed through institutional digital and data policies |
USAID Program Cycle guidance, in particular ADS 201 and the data quality standards, remains one of the most detailed public MEL frameworks in the sector and is widely used as a reference standard well beyond US-funded work. Following the 2025 restructuring of US foreign assistance, some of these resources now sit on archive domains and the operational policy landscape is in transition. Treat the methodological content as durable and verify the current compliance position directly with the contracting officer for any live US-funded award.
Template library
Nine copy-ready structures. Use the copy button, paste into your own document, and delete what does not apply.
TERMS OF REFERENCE
[Evaluation type] of [intervention name]
1. BACKGROUND
1.1 The intervention: objectives, budget, duration, geography, partners
1.2 Current status and what has already been evaluated
1.3 Policy and strategy context
2. PURPOSE, AUDIENCE AND USE
2.1 Purpose: formative / summative / developmental / learning review
2.2 Primary intended users (named roles)
2.3 The decision this evidence will inform, and the date it is needed
2.4 Secondary audiences
3. SCOPE
3.1 Time period covered
3.2 Components, geographies and partners included
3.3 Explicitly out of scope
4. EVALUATION QUESTIONS
4.1 Primary questions (3 maximum, ranked)
4.2 Secondary questions
4.3 Criteria applied to each question
5. APPROACH AND METHODS
5.1 Required design elements
5.2 Expected data sources, including existing monitoring and secondary data
5.3 Sampling strategy and rationale expected from the bidder
5.4 Triangulation strategy and how strength of evidence will be judged
5.5 Participation of programme stakeholders and affected communities
6. ETHICS, SAFEGUARDING AND DATA PROTECTION
6.1 Ethical review route and standard applied
6.2 Informed consent and handling of personal or sensitive data
6.3 Applicable data protection law
6.4 Safeguarding reporting line and incident procedure
6.5 Data retention, storage location and destruction
7. DELIVERABLES
7.1 Inception report including evaluation matrix and tools
7.2 Data collection tools for approval
7.3 Emerging findings session
7.4 Draft report
7.5 Final report and 2-page summary
7.6 Datasets and codebooks
7.7 Quality standard against which deliverables will be assessed
8. TIMELINE AND LEVEL OF EFFORT
Days by role and phase, with inception weighted realistically
9. TEAM
9.1 Required expertise, languages and national representation
9.2 Team leader role and responsibilities
9.3 Independence and conflict of interest requirements
10. MANAGEMENT AND GOVERNANCE
10.1 Contract manager and technical approver
10.2 Reference group composition and role
10.3 Approval points
11. USE AND DISSEMINATION
11.1 Management response process and owner
11.2 Dissemination products and audiences
11.3 Publication commitment
12. BUDGET AND SUBMISSION
12.1 Ceiling or indicative budget
12.2 Proposal requirements and evaluation criteria for bids
12.3 Deadline and submission route
THEORY OF CHANGE NARRATIVE
1. THE PROBLEM
What is wrong, for whom, and why it persists. Evidence base for this
diagnosis, and where the diagnosis is contested.
2. THE LONG-TERM CHANGE
One sentence, in the language of the people who would experience it.
3. THE PATHWAYS
Pathway A: [name]. Sequence of changes, and which actors change.
Pathway B: [name].
Where the pathways interact or depend on each other.
4. WHO CHANGES, AND HOW
For each outcome: the actor, what they currently do, what they would do
differently, and what would make that visible.
5. ASSUMPTIONS
A1 [statement] Confidence: high / medium / low Evidence: [source]
Monitored by: [indicator or method]
A2 ...
Rank by risk. The riskiest assumption gets a monitoring commitment.
6. OTHER ACTORS AND FACTORS
Who else is working on this outcome, and what could move it
independently of this programme.
7. WHAT WOULD SHOW THIS THEORY IS WRONG
The observations that would force a redesign.
8. REVIEW SCHEDULE
When this will be revisited, by whom, and what triggers an early review.
PERFORMANCE INDICATOR REFERENCE SHEET Indicator number and title: Result it measures (level and statement): Indicator type: output / outcome / impact / proxy / sentinel Unit of measure: Precise definition: every term defined, including what is excluded Disaggregation required: sex, age, disability, location, other Rationale: why this indicator, and what decision it informs DATA Data source: Collection method: Collection instrument: Responsible for collection: Frequency: Responsible for analysis and reporting: Estimated cost per round: QUALITY Known limitations: Actions taken to address limitations: Date of most recent data quality assessment: DQA result: validity / reliability / timeliness / precision / integrity VALUES Baseline value and date: Baseline source: Targets by period: Target rationale: Actuals by period: CHANGE LOG Date | Change made | Reason | Approved by
MONITORING, EVALUATION AND LEARNING PLAN 1. Purpose of this plan and how it will be used 2. Theory of change and results framework (summary plus annexed full version) 3. Evaluation questions and learning questions 4. Indicator set 4.1 Indicator table by result level 4.2 Reference sheets (annex) 4.3 Complexity-aware components and why they are included 5. Data collection plan 5.1 Methods, instruments, sampling 5.2 Calendar by quarter 5.3 Roles and responsibilities 6. Data management 6.1 Flow from collection to reporting 6.2 Systems and storage 6.3 Quality assurance and DQA schedule 6.4 Data protection, consent and retention 7. Analysis and reporting 7.1 Analysis approach, including disaggregation and triangulation 7.2 Reporting products, audiences and calendar 7.3 Dashboard specification 8. Evaluation plan 8.1 Planned evaluations, timing, type and indicative budget 9. Learning and adaptation 9.1 Learning agenda 9.2 Pause and reflect calendar 9.3 Management response process 10. Responsible AI use (FRAME position for this programme) 11. Budget and staffing 12. Risks to the MEL system itself, and mitigations 13. Annexes: reference sheets, tools, theory of change, evaluation matrix
EVALUATION MATRIX One row per sub-question. Complete at inception, before any tool is drafted. | Evaluation | Criterion | Judgement | Indicators | Data | Method of | Strength of | | question | | criterion | / evidence | sources | analysis | evidence | | | | | sought | | | expected | |------------|-----------|-----------|------------|----------|-----------|-------------| | Q1.1 | | | | | | | | Q1.2 | | | | | | | Judgement criterion: the standard against which "good" will be assessed. Strength of evidence: state in advance what would count as strong, moderate or weak evidence for this question, and how conflicting evidence will be handled. This is the column most often left blank and the one that most determines whether the findings can be defended.
PART A: HARVEST DESIGN BRIEF Harvest users (named): Intended use of the findings: Harvesting questions (2 to 4, usable and feasible): Definition of outcome for this harvest: Time period and boundaries: Change agents to be engaged: Substantiation approach and sample: Classification scheme for analysis: Expected number of outcomes (strive for fewer, better understood): Timeline by step (design, review, engage, substantiate, analyse, support use): PART B: OUTCOME STATEMENT Outcome ID: Date of the change: Location: CHANGE Who changed, what they now do differently, when and where. Specific and verifiable by a third party. SIGNIFICANCE Why this matters for the goal, and for whom. CONTRIBUTION What the intervention did that plausibly influenced this change. Other contributors named. Stated modestly. CLASSIFICATION Actor type: Outcome type: Geography: Intended / unintended: Significance rating: Contribution strength: SUBSTANTIATION Independent source 1 (name, role, date, method): Independent source 2: Documentary evidence: Substantiator's assessment: confirmed / partly confirmed / disputed Notes on any disagreement:
MOST SIGNIFICANT CHANGE STORY Reporting period: Domain of change: Recorded by: Date: Storyteller role: Consent to share obtained: yes / no Consent to use name: yes / no If no, anonymise before circulating THE STORY Looking back over this period, what do you think was the most significant change in this domain? Tell it as it happened: what the situation was before, what happened, who was involved, and what is different now. WHY IT IS SIGNIFICANT Why did you choose this change rather than another one? Why does it matter to you? VERIFICATION (completed later) Verified by: Date: Method: Result: --- SELECTION RECORD (completed by each selecting group) --- Level: Group members: Date of selection: Stories considered: Story selected: REASONS FOR THE SELECTION (this is the data, write it properly): Feedback sent to storytellers on:
DATA QUALITY ASSESSMENT
Indicator:
Reporting period assessed:
Assessment date: Assessors:
Data source and implementing partner:
REVIEWED
[ ] Indicator definition and reference sheet
[ ] Data collection instrument
[ ] Collection method and enumerator protocol
[ ] Database or system holding the data
[ ] A traced sample of records, from source document to reported figure
FIVE STANDARDS
VALIDITY Does the data clearly represent the intended result?
Finding: Rating: adequate / not
RELIABILITY Is collection consistent and unbiased across time and sites?
Finding: Rating:
TIMELINESS Is it available frequently enough for management decisions?
Finding: Rating:
PRECISION Is the level of detail sufficient for the decisions it informs?
Finding: Rating:
INTEGRITY Are there safeguards against error and manipulation?
Finding: Rating:
TRACE TEST
Reported figure: Figure found at source:
Discrepancy and explanation:
CONCLUSION
Overall assessment:
Limitations to carry into all reporting that uses this indicator:
ACTIONS
Action | Owner | Due date | Status
Next assessment due:
PART A: LEARNING AGENDA ENTRY Learning question: Why we do not already know the answer: Who needs the answer, and for what decision: How it will be answered (method, data, who does it): Owner: Answer needed by: Interim checkpoints: What we will do differently depending on the answer: Status: open / partly answered / answered / dropped (with reason) PART B: MANAGEMENT RESPONSE TRACKER Source: [evaluation or review name, date] | No | Recommendation | Response | Rationale | Action agreed | Owner | Due | Status | | | | accept / | | | | | | | | | partly / | | | | | | | | | reject | | | | | | |----|----------------|-----------|-----------|---------------|-------|-----|--------| | 1 | | | | | | | | | 2 | | | | | | | | Review points: [governance meeting dates at which this tracker is tabled] Recommendations addressed to other actors (not for us to implement): Recommendations rejected, with published rationale:
Sources and further reading
Every method statement in this toolkit traces to one of the sources below, or is marked as DevCAFE original work. Links were checked in August 2026.
Outcome harvesting
- Outcome Harvesting Community. About OH. outcomeharvesting.net/about-oh
- Wilson-Grau, R. and Britt, H. (2012). Outcome Harvesting. Ford Foundation MENA Office. Summarised at betterevaluation.org
- INTRAC (2017). Outcome Harvesting (M&E Universe series). intrac.org
- Outcome Mapping Learning Community. Outcome Harvesting. outcomemapping.ca
- Wilson-Grau, R. (2018). Outcome Harvesting: Principles, Steps, and Evaluation Applications. Information Age Publishing. Review at here2there.ca
Most significant change
- Davies, R. and Dart, J. (2005). The Most Significant Change (MSC) Technique: A Guide to Its Use. mande.co.uk
- Dart, J. and Davies, R. (2003). A dialogical, story-based evaluation tool: the most significant change technique. American Journal of Evaluation, 24(2), 137 to 155.
- Better Evaluation. Most significant change. betterevaluation.org
- Asian Development Bank (2009). The Most Significant Change Technique (Knowledge Solutions 25). adb.org
USAID: program cycle, complexity-aware monitoring and data quality
- USAID. ADS Chapter 201: Program Cycle Operational Policy. Copy hosted at preparecenter.org
- USAID (2021). Discussion Note: Complexity-Aware Monitoring. usaidlearninglab.org and archived at 2017-2020.usaid.gov
- USAID Learning Lab. Complexity-aware monitoring approaches. usaidlearninglab.org/approaches
- Better Evaluation. Discussion note: complexity aware monitoring. betterevaluation.org
- USAID. Data Quality Assessment Plan and DQA checklist (five data quality standards). 2017-2020.usaid.gov
- Edu-Links. Complexity-Aware Monitoring summary of the three principles. edu-links.org
United Kingdom: Magenta Book, Green Book and FCDO
- HM Treasury (2026). The Magenta Book: central government guidance on evaluation, 2026 update, led by the Evaluation Task Force. Announcement: civilservice.blog.gov.uk
- HM Treasury (2020). Magenta Book. assets.publishing.service.gov.uk
- HM Treasury (2026). The Green Book: appraisal and evaluation in central government. gov.uk
- FCDO. Evaluation strategy (2025 edition, superseded by the 2026 to 2030 strategy). gov.uk
- Cabinet Office Evaluation Task Force. Strategy and cross-government evaluation guidance. gov.uk
Australia: DFAT
- DFAT. Design and Monitoring, Evaluation and Learning Standards. dfat.gov.au; full text (PDF)
- DFAT (2025). Review of the quality and use of DFAT evaluations: 2024 to 2025. dfat.gov.au
- DFAT. 2024 Review of the Quality and Use of DFAT Evaluations: management response. dfat.gov.au
- Better Evaluation. DFAT design and monitoring and evaluation standards. betterevaluation.org
European Union
- European Commission, DG INTPA and FPI (2024). Evaluation Handbook. capacity4dev.europa.eu
- European Commission. What is monitoring? Results Oriented Monitoring and the eight monitoring criteria. international-partnerships.ec.europa.eu
- European Commission, DG NEAR. Monitoring and evaluation criteria for enlargement and neighbourhood assistance. neighbourhood-enlargement.ec.europa.eu
- DG INTPA.D4. Support to quality of design: logframes, measurement frameworks and sector indicators guidance. capacity4dev.europa.eu
- European Evaluation Society (2025). Review of the Commission's new Evaluation Handbook. europeanevaluation.org
Theory of change and evaluation approaches
- Vogel, I. (2012). Review of the use of Theory of Change in international development. Report for DFID. Full report (PDF); summary at betterevaluation.org
- GSDRC. Review of the Use of Theory of Change in International Development. gsdrc.org
- GEF STAP. Theory of Change: a short literature review and annotated bibliography. thegef.org
- Better Evaluation. Approaches (comparative library covering contribution analysis, process tracing, realist evaluation, QCA, outcome mapping and developmental evaluation). betterevaluation.org
- OECD DAC Network on Development Evaluation (2019). Better Criteria for Better Evaluation: relevance, coherence, effectiveness, efficiency, impact and sustainability.
Data, dashboards and responsible data
- MERL Tech. Responsible Data Management for M&E: Design and Planning. merltech.org
- MERL Tech and CLEAR-AA. Guidance on Responsible Data Governance for Monitoring and Evaluation. merltech.org
- MERL Tech. We need more design thinking in monitoring, evaluation, research and learning. merltech.org
- ICTworks. Four years of dashboard design progress with MERL Tech data. ictworks.org
- Responsible Data. Responsible Data and MERL. responsibledata.io
DevCAFE original work
- Gandhi, V. J., Bruce, K. and Nielsen, S. B. (forthcoming, 2026). The FRAME approach to responsible AI in evaluation. In From Algorithms to Evidence: Using Generative AI in Evaluation Practice. Routledge.
- The Development CAFE. AI for MEL e-learning course. The FRAME checklist emerged from questions raised by course participants.
- The Development CAFE. EvalEdge Podcast and associated practice notes.
- Plates 01, 07 and 08 in this toolkit are DevCAFE reference architectures, informed by the donor guidance cited above.