
Site selection has always been a CRO's highest-leverage decision. Pick the right investigators and study startup stays on schedule, screen failure rates hold steady, and sponsors see the progress they committed to. Pick the wrong ones and the trial spends a year chasing enrollment across sites that were never going to hit their numbers. For decades, the industry has tried to solve this with feasibility questionnaires: standardized surveys sent to investigators asking how many eligible patients they can enroll. The answers come back optimistic, the trial launches, and a familiar pattern plays out.
Tufts Center for the Study of Drug Development has tracked that pattern for years. In a global sample of Phase II and Phase III trials, roughly one in ten selected sites enrolled zero participants, and more than a third under-enrolled. Trials tended to meet their overall enrollment goals eventually, but often at the cost of nearly doubling their original timelines. That gap between what feasibility surveys promised and what sites actually delivered is where predictive feasibility modeling has emerged as an operational answer.
Predictive feasibility modeling uses statistical and machine learning techniques to forecast site performance before a trial starts. Instead of asking a principal investigator how many patients they think they can enroll, the model looks at how similar sites behaved in similar trials and builds a data-backed projection. The inputs are records the industry already has: historical enrollment rates, screen failure rates, protocol deviation rates, site startup timelines, investigator experience, patient population data pulled from electronic health records or claims databases, and public sources such as ClinicalTrials.gov.
The outputs are what CRO feasibility teams need in order to defend a site list. A ranked shortlist. Predicted enrollment rate per site. Predicted screen failure rate. Predicted activation time. Confidence bands that show how certain the model is about each projection. None of this eliminates human judgment. It gives operations leads a starting point built on evidence rather than optimism, and it makes the reasoning behind each site selection easier to explain to a sponsor. CROs weighing which performance signals matter most can review which recruitment metrics decide site selection when they build the model's evaluation framework.
The core weakness of survey-based feasibility is that it asks people to predict their own future performance under conditions they cannot fully see. Investigators are answering a questionnaire, often without the finalized protocol in hand, and without any real accountability for how their estimate holds up once the study opens. Enthusiasm gets baked in. So does a natural tendency to overestimate how many patients in the practice actually meet a complex set of inclusion and exclusion criteria.
Feasibility teams know this. Most CROs discount investigator estimates by a rule of thumb before they trust them. But a rule of thumb is not a model. It cannot tell you which of two similarly enthusiastic sites is actually more likely to deliver, and it cannot factor in things like competitive trial density in a region, seasonal patient flow, or a site's track record on protocol deviations. Predictive feasibility modeling brings those signals into the same view and produces an estimate that is transparent, testable, and updatable as new enrollment data comes in.
Model quality depends on the data behind it, and CROs have more available signal than they often realize. Internal databases from past studies hold years of site performance history that can be normalized and used as training data. Site-level metrics such as principal investigator experience, past enrollment ratios, screen failure rates, protocol deviation history, and activation timelines give the model a picture of how each site behaves under pressure.
Real-world data sources such as electronic health records, claims databases, and patient registries fill in the patient-level picture: prevalence in a catchment area, comorbidity patterns, demographic distribution, and standard of care pathways. Country and regional data add regulatory approval timelines and the competitive trial landscape, which matters because a strong site in a saturated region may still under-deliver simply because eligible patients are being pulled into competing studies. Feasibility teams that build this profile before the model runs tend to get outputs they trust enough to act on.
Getting from a predicted enrollment number to a working study is a separate operational challenge that begins the moment sites are activated. CROs planning a defensible startup sequence often work in parallel from a CRO's guide to enrolling the first patient faster so the model's site ranking translates into actual first-patient-in dates.
Predictive feasibility is not a crystal ball, and CROs that treat it that way run into the same disappointment they had with surveys, only with more spreadsheets. Novel indications with little historical data are the hardest cases. A first-in-class oncology protocol has no exact precedent, so the model borrows from adjacent indications and produces wider confidence bands. Rare disease trials face a related problem: small sample sizes make conventional models unstable, and epidemiologic overlays or Bayesian methods have to do a lot of the work that historical data usually does.
Data silos across CRO and sponsor systems fragment the signal, and site heterogeneity means academic centers, community practices, and private clinics need different comparison groups. A more serious limitation is ethical. Because historical enrollment has skewed toward large, well-resourced centers, a model that optimizes purely for predicted enrollment can quietly deprioritize sites that serve underrepresented populations. FDA's Diversity Action Plan expectations under the Food and Drug Omnibus Reform Act have made that risk unacceptable, and predictive systems now need explicit fairness constraints to avoid entrenching the very patterns the industry is trying to correct.
A prediction is not a patient. Even the best-tuned model produces a projection at the site level, not a referral in the coordinator's inbox. The path from a forecasted enrollment number to an enrolled participant runs through outreach, matching, pre-screening, and referral, and that is where most trials still lose time. This is why CROs are increasingly pairing predictive feasibility with real-time recruitment signal from AI-assisted matching platforms that surface potentially eligible participants against the specific protocol. The two tools are complementary. The model tells the CRO where enrollment is likely to happen. The matching engine tells the CRO whether it actually is. CROs evaluating this pairing should read what CROs should look for in an AI-assisted patient matching model before committing to a partner.
Volume of referrals alone does not close the gap. A site can receive hundreds of unqualified leads and still miss enrollment, which is why referral quality tends to predict enrollment better than referral volume. Pre-screening throughput is the operational metric that keeps a predictive model honest. When the model says a site will enroll ten patients a month and the referral pipeline is producing three qualified pre-screens a week, feasibility teams have early evidence to reallocate. That kind of feedback loop is what turns a static forecast into a working recruitment plan.
ICH E6(R3), the updated good clinical practice guideline finalized in early 2025 and adopted by FDA later that year, formalizes the shift toward risk-based quality management and quality by design. In practice, that means every meaningful operational decision, including which sites are selected, should have an explicit rationale tied to risk and quality. A defensible predictive model gives CROs a documented basis for site selection that fits the E6(R3) mindset far better than a stack of survey responses ever could. Teams that want a plain-English summary can look at what changes for CRO quality management at scale under ICH E6(R3).
FDA's 2025 draft guidance on the use of artificial intelligence to support regulatory decision-making adds another layer. It introduces a risk-based credibility framework for AI models tied to a defined context of use, which is worth reading for any CRO planning to use predictive outputs as part of regulatory submissions. The direction of travel is clear. Regulators expect the industry to move from opinion to evidence in the operational choices that shape trial quality, and site selection sits near the top of that list.
DecenTrialz is a clinical trial recruitment and pre-screening platform that pairs AI-assisted participant matching with registered nurse-led pre-screening. That combination is where predictive feasibility becomes actionable. The AI matching layer surfaces potentially eligible participants against protocol-specific criteria in real time, generating the referral signal a CRO needs to validate or correct a model's site-level projection. The nurse-led pre-screening layer applies clinical judgment to that signal so the participants who reach the site are more likely to convert to enrolled patients rather than screen failures.
The research site team owns final eligibility determination, informed consent, study walk-through, and enrollment. DecenTrialz does not replace that work, and it is not a substitute for the site's clinical judgment or the sponsor's oversight. What it does is close the last-mile gap between a feasibility forecast and an enrolled participant, and it does that in a way that gives CROs real-time visibility into whether the model's assumptions are holding in practice. Sponsors and CROs interested in how the workflow fits together can learn more at decentrialz.com.
More accurate than survey-based baselines, in most published comparisons, but not a guarantee. Predictive models produce probability-weighted rankings and confidence bands, and their accuracy varies by therapeutic area and by the quality of the training data. The models are designed to be recalibrated as real enrollment data comes in, so the value builds over time.
At a minimum, historical enrollment, screen failure, and startup metrics from past studies, plus access to public data such as ClinicalTrials.gov. Real-world data from electronic health records or claims sources strengthens the model considerably, either through internal access or an established data partner. Even modest internal datasets can be augmented with public and real-world sources.
No. It gives the team a defensible starting point. Predictive outputs still need human validation through relationships with investigators, protocol-specific judgment, and, where useful, site visits. Regulators expect documented human decision-making, so the model is decision support rather than a decision maker.
These are the hardest cases. Small sample sizes call for Bayesian methods and epidemiologic overlays, and patient-level identification from EHR or referral networks becomes essential to compensate for thin historical data. In these settings, pairing the model with real-time patient identification and pre-screening tends to produce better outcomes than the model alone.
Public data, real-world data partnerships, and pre-screening referral signal can substitute for a large proprietary database. A smaller CRO with disciplined data collection from each new study can build a workable model within a few therapeutic areas, especially if it pairs the model with a real-time recruitment layer that provides quick feedback.
Predictive feasibility modeling is not a replacement for the operational judgment CROs have always brought to site selection. It is a way to bring more evidence into that judgment, and to give sponsors a defensible reason to trust the site list. The teams that get the most out of it are the ones that treat the model as one input in a larger recruitment system, one that includes real-time patient matching, nurse-led pre-screening, and continuous feedback from the actual referral pipeline. That is the system predictive models were always meant to support, and it is the reason the industry is finally getting closer to trials that start on time and enroll to plan.
Was this article helpful?


Predictive feasibility modeling has moved from a data science experiment to an increasingl...

Direct-to-patient shipment moves the study drug from the research site to the participant'...

External control arms have moved from a niche design tool to a recurring conversation betw...
Get updates on verified clinical trials, emerging treatments, and research breakthroughs directly in your inbox. No spam, just science that matters.