Recommended
Blog Post
Are Aid Agencies Paying Attention to What Works in Education?
How good are donor organisations at designing their aid programmes? A well-designed programme—one that focuses on the most relevant issues, learns from past successes and failures, and draws on solid evidence about what works—is likely to have more impact than a poorly designed one. Yet organisations differ substantially in the requirements they place on programme sponsors to follow good design principles. We would love to assess donors on the impact of their aid portfolios. But in the absence of an impact measure that is comparable across all sectors and types of interventions, we are developing a methodology to compare the extent to which projects are set up for success.
This blog exploits advances in artificial intelligence models to systematically review 244 documents—operational manuals, appraisal guidance, and other official documents—from 49 aid providers to assess organisations against five key design criteria. We find that, aside from some notable exceptions, including Japan and the UK, multilateral aid providers have stronger procedures for project design. However, these results are preliminary, and we welcome suggestions for improvements and what we may have missed.
New data challenges, new AI opportunities
In 2027, CGD will publish the sixth edition of its Quality of Official Development Assistance (QuODA) index. QuODA quantitatively assesses 49 of the largest bilateral and multilateral donors across four dimensions of aid quality—prioritisation, ownership, transparency and untying, and evaluation. The index uses a number of key publicly available data sources for its assessment, including the OECD’s Creditor Reporting System database, Global Partnership for Effective Development Cooperation monitoring, and the International Aid Transparency Initiative’s data. Since CGD published the fifth edition of QuODA, the OECD DAC has updated its methodology for peer reviews. These used to be systematic, covering various aspects of each bilateral donor’s evaluation systems, but they are now bespoke to each donor, meaning we no longer have standardised, up-to-date information on how each donor evaluates their programmes.
To fill this gap, our first preference was to examine the degree to which donors actually fund high-impact interventions, as measured by cost-effectiveness information in project documents and whether such appraisals are based on high-quality evidence (building on proposals by other CGD colleagues). However, the paucity of project-level documentation across providers, and lack of consensus on the most high-impact interventions outside a few sectors, mean these approaches were unviable for a full comparison across entire aid portfolios.
Instead, we examine how careful agencies are in programme design: whether they have rules in place (such as requirements to review and reference evidence) that make impact more likely. This is upstream of the impact we ultimately care about, but requirements for sound programme design along the dimensions we explore below are more likely to achieve that impact, and are easier to assess. We used AI models to systematically examine appraisal and evaluation guidance, operational manuals, and programme design documentation—244 in total across the 49 providers—to look at the requirements on ex ante project appraisals, monitoring and evaluation, and results-based management. After reviewing a range of evaluation and project appraisal criteria from authoritative organisations including the OECD-DAC’s Evaluation Criteria, the World Bank IEG’s ICR Review guidance, the MOPAN 4 Framework and Analytical Guidance, 3ie’s Global Evidence Commitment, and J-PAL’s generalizability framework, among others, we chose five criteria to capture the core elements of effective project design and evaluation:
- Evidence in selection: Is effectiveness evidence likely to influence which intervention is chosen? Is there a formal requirement to cite it?
- Lessons from past projects: Do donors examine past evaluations of similar programmes? Is there a formal feedback loop from evaluation to design?
- Cost-benefit modelling: Are costs and benefits modelled appropriately? Is there a formal requirement to do so where possible?
- Demonstrated need: Are programmes designed based on diagnostics of unmet needs in the country or sector?
- Clear success criteria: Is there a formal requirement to set performance metrics to assess the success of the programme? Do these include transparency baselines?
Some organisations highlighted the importance of other factors such as transparency, coherence and coordination with other actors, and country ownership, but we omit those here because they are covered in other dimensions of QuODA.
We scored each donor on these five criteria using a rubric anchored to a defined mechanism and, when possible, accompanied by a worked example from a donor document. Every score is supported by a quoted passage, and when we found no evidence, we scored it as 0 out of 5.
We recognise that not all criteria apply equally to each type of provider included in QuODA. For example, projects in social sectors are inherently less likely to calculate economic internal rates of return or other cost-benefit metrics than infrastructure projects, which might disadvantage social-sector focused donors on that dimension. At the same time, there is a larger body of evidence on “smart buys” in such sectors relative to infrastructure, which could point in the other direction. Similarly, QuODA covers many provider types with very different mandates—the IMF, for example, has an operating model that fundamentally differs from bilateral agencies.
Example high-level scoring rubric for criterion 1 — Evidence in selection
| Score | Definition |
|---|---|
| 5 | Gate — the design (including its evidence base) is assessed at entry, and a weak result can block or send back funding, or an evidence review itself decides what is selected. |
| 4 | Validated — the design's evidence base is independently scored/appraised at entry (rating recorded to inform approval), but no explicit score threshold blocks funding. |
| 3 | Required — formal, standard requirement to set out the effectiveness / technical evidence for the chosen intervention in appraisal (self-assessed). |
| 2 | Considered — evidence considered qualitatively; no formal requirement to cite effectiveness evidence. |
| 1 | Minimal — selection driven by priorities/calls; little or no emphasis on external effectiveness evidence. |
| 0 | No evidence of how effectiveness evidence informs selection found. |
Our initial findings suggest that multilateral entities generally have stronger processes for project design across our five criteria, though some bilaterals, including Japan and the United Kingdom, also perform well, aligning their projects to country needs and defining clear metrics to judge success. Climate and health vertical funds, including Gavi, the Global Fund, the Green Climate Fund, and the Global Environment Facility, tend to have the most stringent requirements on the use of evidence in project selection. At the other end of the spectrum, AI couldn’t find any documentation that Iceland, Spain, or Korea require evidence or cost-benefit analysis in their project appraisals.
We have included our preliminary results below. These are likely to change slightly as we refine our methodology and find new documents. In the meantime, to make these results more robust, we welcome input, particularly from agency officials, on the following questions:
- Does our framework for evaluation and project design make sense? Are there other aspects we should consider including?
- Do these scores look reflective of actual provider performance?
- Have we missed relevant documents? Are there internal documents that may influence the scores? (We include a downloadable list of the documents we have on file on this page.)
Preliminary results
DISCLAIMER & PERMISSIONS
CGD's publications reflect the views of the authors, drawing on prior research and experience in their areas of expertise. CGD is a nonpartisan, independent organization and does not take institutional positions. You may use and disseminate CGD's publications under these conditions.
Thumbnail image by: USAID in Africa/ Flickr