CGD NOTE

AI RCTs Are Booming—To Be Useful, They Must Evolve

This note also appears on VoxDev

AI RCTs are on the rise—a trend that’s set to continue and is likely to intensify. But randomising a rapidly evolving, increasingly widely adopted technology poses a distinct set of challenges. So how can we actually garner durable, generalisable, and policy-relevant insights?

Like it or not—you’re in good company either way—a wave of AI RCTs is on the horizon.

Recently, through their Economic Futures Research Fund, Anthropic announced they are committing $200 million to funding large-scale RCTs for ambitious, creative pilots or programme evaluations[1]. This inflow of funding, including funding calls from J-PAL, with more potentially on the horizon, will compound existing trends in economic research.

Leaving aside debates of whether this is the optimal allocation of resources (for some of our takes, see there is no randomising a technological revolution, or adaptation, not adoption, is king), if we accept the coming wave of trials as inevitable, how can we maximise the durability, generalisability, and policy relevance of these studies? We have a number of concrete proposals, but first we outline the current AI RCT trends in economics.

More RCTs + AI = AI RCT boom

Trial registries are a window into the future. By looking at the descriptions of planned and ongoing studies, we can learn what the profession is investing its time and money in. Doing so reveals that AI looms increasingly large.

The number of trials registered in the American Economic Association’s RCT Registry (the economics profession’s largest registry of randomised controlled trials) and the Registry for International Development Impact Evaluations (RIDIE, focused on development economics) has been increasing steadily[2]. From a little under 1,000 pre-registered studies in 2019, the number had risen to around 1,650 in 2024, and a similar number in 2025. By the half-year mark, around 900 had been registered in 2026, suggesting that the upwards trend continues. This could simply reflect more diligent registration—because more journals now require it for publication, say—but likely also indicates a growing number of trials each year. Figure 1 shows that interventions that use AI in one form or another are on track to make up one in five trials—up from one in thirty in less than four years.[3]

Figure 1. Registered trials since 2019, with AI-related records highlighted

AEA + RIDIE registrations by initial registration year; 2026 is a partial year
Registered trials since 2019, with AI-related records highlighted

Notes: Records not identified as AI-related are not confirmed non-AI; classification uses public registry metadata. RIDIE detail-page records are deduplicated against AEA by exact title.

In LMICs, AI studies are less common but rising just as steeply

Randomised trials are used (or registered) somewhat more in high-income countries—despite those only accounting for 17% of the world population.

Figure 2 shows the share of AI studies in the total number of trials, split by country income group, over time. The pattern and trend are clear: the share of AI studies in LMICs is lower, but the trend line is as steep.

Figure 2. AI-related share of registered trials, by income group

Within-group rates: AI-related HIC trials as a share of all HIC trials, and the equivalent calculation for LMICs; 2026 is partial
AI-related share of registered trials, by income group

Notes: Each rate uses its own income-group denominator. LMIC combines LIC, lower-middle-income, and upper-middle-income economies. Country assignment combines registry metadata with reviewed title/description inference; multi-country records receive weight one, split across classifiable countries. Based on the roughly 66.4% of AI-related records that show the study country.

Where is AI being randomised?

Whether AI has the greatest value in the most capacity-constrained settings or whether it's most useful in places with complementary infrastructure remains unclear. The cost-benefit analyses that emerge from these empirical studies should shed light on where productive applications lie—for now, though, we can only infer where researchers see potential by looking at the areas they study.

Figure 3, based on AEA registry categories and assigning each record to its main topic, shows that studies of behaviour, education and labour are the most common subjects overall, and account for over half of AI-related RCTs. We can also see that education, firms and labour research are more common for AI-related projects than for overall registry records, suggesting a broad focus on productive impacts of AI. When comparing the AI records (red markers) with general development research (black markers), as compiled by Jessica Leight for 2021-2025, the top registry categories of behaviour, education and labour are substantially over-represented in these empirical studies.

Figure 3. Primary topics of registered research and development articles

AEA + RIDIE, 2019-2026 (2026 partial): all registered N=10,214; AI-related N=739; development articles n=1,497
Primary topics of registered research and development articles

 

Notes: Each registry effort receives one reviewed primary topic; all three mutually exclusive distributions sum to 100%. Development-article data: Jessica Leight (2026), VoxDev, Figure 3: https://voxdev.org/topic/methods-measurement/state-play-development-economics-research. AI follow-ups are collapsed to research efforts; exact-title RIDIE duplicates matched to AEA are removed.

Where to from here? We should be ambitious

In a worst-case scenario, we could see a wave of one-and-done research projects in LMICs, that are out of date before the results are public, don’t generalise, and aren’t taken up by governments or scaled. In this scenario, the wave of AI RCTs would not lead to a wave of useful insights, but a wave of funding for well-networked senior researchers.[4]

That need not be the case, but ensuring so requires urgently changing with the times. For those questions that can be answered by RCTs, the current model, and the standard research process around it, isn’t set up for success. There are a number of problems to solve, some old, some new. After discussing these, we outline some concrete fundable proposals to maximise our collective learning from the AI RCT wave.

Challenges

Longstanding issues

Scale: Governments = Scale. Research projects without initial government buy-in or local collaborators are unlikely to be adopted, or scaled, by governments.

Generalisability: That many randomised interventions fail to repeat their performance when taken to a different setting remains a key challenge—and is one possible reason why policymakers don’t give much weight to impact evaluations.

Duplication: How do we avoid the scattered, duplicative pilots of, for example, the mobile health era? ‘Pilotosis’ is real—the situation got so bad in Uganda that, in 2012, the government put a moratorium on mHealth pilots.

General Equilibrium: Much of the coming wave of trials will not just involve testing AI interventions themselves, but new models of redistribution and other policies which may become more important in an age of transformative AI—to figure out their actual effects, economy-wide, we need macro. Much of this work may appear unrelated to AI, but it will be helpful to design these macro experiments with the potential for AI transformation in mind.

Issues particularly exacerbated by AI

Iteration: An RCT, from planning over baseline to mid- and endline can sometimes take two years, or even longer. By the time an intervention has been tested, and the results published, the AI model it was based on may no longer be available at all. What lessons can we learn which outlive the RCT, when the underlying technology is improving so rapidly. Faster trials can help, but we probably need to find new ways of evaluating the type of (increasingly common) intervention that itself evolves during the trial window.

Figure 4. Planned AI trial spans cross more than three model generations

AEA AI-related research efforts, all years; span uses the earliest planned start and latest planned end across linked registrations
Planned AI trial spans cross more than three model generations

Notes: N=741 research efforts with valid non-negative dates; mean span=11.3 months. Model-generation blocks use the observed average interval between flagship lab model-release announcements.

Cost: AI could make it far cheaper and easier to develop digitally-delivered interventions. Unless the cost of impact evaluations falls in tandem, evidence generation becomes a bigger share of budgets. To the extent that MEL funding is commonly capped at some part of the total budget, evaluation cost becomes a bigger hurdle to evidence generation than it already is.

Fidelity: With AI interventions, the technical capacity of the research team (i.e. their ability to adapt, iterate, and improve the AI solution) might be more important than the intervention being tested itself. I.e., the difference in impacts between different AI interventions might be lower than the difference in impacts between teams of varying quality delivering the same AI intervention.

Control group: What does ‘control’ mean in an era of rapidly expanding adoption? What would we have learned from randomising access to the internet in a country where overall adoption was rising rapidly, and the internet itself was rapidly evolving? This complicates our interpretation of results, as it’s possible that many trials won’t capture the impact of AI itself, but of the AI being randomly provided, against the AI being silently adopted in the background. Is the agricultural advisory tool built by a team of economists on top of an existing model really better than the next iteration of a general purpose model that becomes available during the trial window?

Solutions

Some of these issues arise from the technology being studied, some arise from the discipline doing the studying. But given the unique moment we’re in, and the changing funding landscape, we have the opportunity to be ambitious and tackle both. Here are some initial proposals (comments, feedback, or further ideas encouraged).

New, or much improved, centralised research tracker for economics

Building on the registry of pre-registered trials, we need a registry of pre-paper results, where researchers upload mid-line, and end-line, results, so we don’t have to wait for the paper to see impacts.

We also need a pilot registry system. This would build on current registries of trials that are already underway, by also publicising those trials that never got underway because the pilot didn’t work. Learning lessons from the pilots that never got off the ground will be extremely valuable in directing future work, but is currently not tracked systematically, so we are missing a whole world of lessons learned. Building this database, with prizes for registering verified pilots that never went further, would help here.

This tracker could also include new variables that every AI RCT is likely to track, e.g.: models used, tokens spent, cost for specific language tokeniser, choices for fine-tuning (hyperparameters, training dataset size and shape), performance gain from base model, cost of training run in compute, and licensing and compensation for any individuals who helped in training dataset. These variables would be helpful for comparing between different interventions on dimensions like: performance gains from training, impact achieved, and cost effectiveness.

New, or much improved, selection mechanisms

It’s natural that funding does not always flow to what would have been the best projects from a policy perspective. Relationships, networks, academic incentives to publish, and seniority all play a role – we are human, after all.

This is particularly inefficient from a policy perspective, as local researchers, who are best placed to build relationships with the government, find it hard to break into these networks. It also prioritises academic novelty, which does not always overlap with policy usefulness – for example, academia does not incentivise testing existing interventions in new settings. Funders should deliberately support studies based on their policy relevance, even if typical academic review boards would choose ‘novel’ studies more likely to end up in a top journal.

For AI, this unwritten arrangement is particularly inefficient. Those at the top of the discipline, sat on review boards, or looked on particularly favourably by them, typically have strong development expertise but are less likely to be the tech-savvy researchers with enough technical ‘taste’ to understand and implement the most interesting and useful applications of AI.

So we need new and improved selection mechanisms, made up of more inter-disciplinary review boards that contain economists, technologists, and policymakers, to fund projects that are economically sound, technologically useful, and policy relevant. These review boards would benefit from the improved, centralised research tracker we outlined above, in order to easily avoid duplication.

Funders should also incentivise co-creation with governments, and prioritise projects with local researchers. This could involve creating joint research agendas with specific governments, and/or investing in national research organisations. Building local capacity is an added bonus, but these projects will also just be better for the world by making scale more likely.

Unfortunately, we have a long way to go here. Across the AEA Registry, 14.1% of all RCTs involved the government, compared with 10.7% of AI-related efforts. AI-related RCTs were therefore about 3.4 percentage points less likely to involve the government.

Figure 5. Government-related research efforts

All registered versus AI-related AEA research efforts, all years; hatched segments show unclear public evidence.
Government-related research efforts

Notes: "Government-related" includes trials with a disclosed government partner or conducted through an explicitly public-sector institution, service, programme, workplace, platform, or facility. The hatched section represents cases with plausible government involvement that the registry information does not establish conclusively. All RCTs: 222 unclear efforts, 1.8% of the total. Counting them all would raise the estimate from 14.1% to 15.9%. AI-related RCTs: 7 unclear efforts, 0.9%. Counting them all would raise the estimate from 10.7% to 11.6%.

An important open question, which is particularly relevant to how funding should be allocated, is how to design RCTs that are model-agnostic. Even if we speed up RCTs so that the results become available while the model being tested is still relevant, the results could become irrelevant in a matter of weeks or months when a new model emerges (on either the capability or cost frontier). We need a set of best practices for designing RCTs that can keep results generalisable across models, but this question does not receive enough attention. Some of these best practices would rely on capability growth being more-or-less predictable (cost for capability goes down at about this rate; voice tokens become available at this rate). It will be more difficult to design RCTs which are relevant after surprising capability gains (e.g. a pre-voice model RCT not taking voice into consideration, or some other paradigm shift).

Deliberate, cross-silo research teams

Funding should flow to inter-methodological, inter-disciplinary research teams. The most valuable RCTs will be run by teams including randomistas, those with experience feeding micro data into macro models, and those with the technical skills to actually adapt AI to local needs. We need to deliberately pair these different types of researchers.

Option C Thinking

This is the John List thinking of A/B/C trials – where the C arm is a version of the intervention that is not optimal, but what would be feasible at scale. This is worth more consideration – i.e. not only testing the best model, but having a C group in the trial which uses a cheaper model, or one that runs on most types of phones, or where the human in the loop is more indicative of the government official who’d be implementing the programme at scale.

Other ideas to explore

  • Clinical trial networks could raise efficiency and shorten trial times.
  • The use of AI for supplementary intervention aspects like training and outcome measurement could be one avenue.
  • Greater emphasis on alternative approaches like A/B testing.

We are not starting from scratch

Some studies are already making good progress. Fab AI, in partnership with Google DeepMind, and with support from Sierra Leone’s Ministry of Education, reported and tested an AI guided learning intervention. In just eight-weeks of this classroom RCT, they tested over 1,700 junior-secondary math students and found a large (+0.26 SD) gain. Whether this intervention was cheap to run is unclear, but it was both fast and more replicable than most.

For it not to be an exception that proves the rule, we need methodological innovation in both the interventions themselves, and also how we assess them. Interventions that leverage AI have the potential to generate detailed, timely and actionable data, but the evidence agenda needs to evolve if it is to become more rather than less relevant as AI applications proliferate.

  1. ^Not all of this funding will fund ‘AI related’ RCTs, as classified in this blog.
  2. ^These registries are repositories of study designs for concrete research projects, often including pre-analysis plans that lay out specific hypotheses and the methods for testing them.
  3. ^Defined as studies which have one of the following terms in their description: AI or artificial intelligence, Machine learning or deep learning, Neural networks, Natural language processing or NLP, Large language models or LLM and/or GPT or ChatGPT. Or one of the following named tools: Claude, Gemini, Copilot, Khanmigo and DALL-E, Generative AI, chatbots or voicebots, Computer vision, Facial recognition, image recognition or object detection.
  4. ^This is also defensible – established researchers are typically better prepared to lead on projects which require substantial investment and expertise in design and implementation.

CITATION

Ohlenburg, Tim, Oliver Hanney, Sharif Kazemi, Joseph Levine, and Shahrukh Wani. 2026. AI RCTs Are Booming—To Be Useful, They Must Evolve. Center for Global Development.

DISCLAIMER & PERMISSIONS

CGD's publications reflect the views of the authors, drawing on prior research and experience in their areas of expertise. CGD is a nonpartisan, independent organization and does not take institutional positions. You may use and disseminate CGD's publications under these conditions.


Thumbnail image by: putilov_denis/ Adobe Stock