A payer-sponsored artificial intelligence (AI) referral management pilot should be designed as a multi-organization operations experiment, not a software installation. Choose two to four provider groups, fix a measurable baseline, define data-sharing boundaries and patient-choice safeguards up front, run 12 weeks with weekly operating reviews, and pre-commit the thresholds that decide scale or no-scale.

Quick answer

  • Sponsor at the plan, execute inside provider workflow.
  • Pick cohorts by referral volume, specialty mix, and integration readiness — not enthusiasm alone.
  • Baseline before go-live; a pilot without a pre-period is a demo.
  • Use matched comparison groups where randomization is impractical.
  • Publish scale thresholds in week zero so results are judged, not negotiated.

This is deliberately narrower than a general implementation roadmap: the hard parts of a payer pilot are cross-organizational — data boundaries, provider enablement, and attribution across entities you do not employ.

Why payer pilots fail

  • No baseline. Completion and leakage were never measured the same way before go-live.
  • Volunteer bias. The most engaged group is the least representative one.
  • Ambiguous ownership. Plan network teams, provider operations, and IT each assume another party owns exceptions.
  • Data friction. Coverage and network files arrive late, so matching runs on stale data and providers lose trust in week two.
  • Moving goalposts. Success is defined after results appear.
  • Governance as an afterthought. Compliance review lands in week ten and stalls scale.

12-week pilot plan

PhaseWeeksObjectivesExit criteria
Design and governance0–2Charter, cohorts, evaluation design, data-sharing boundaries, AI governance sign-offSigned charter and approved data flows
Data and integration2–4Coverage/network files, provider directory, EHR interfaces, test transactionsEnd-to-end test referral completes in a test environment
Baseline lock4–5Pre-period metrics computed identically for pilot and comparison groupsBaseline report accepted by all parties
Enablement5–6Role-based training, exception playbooks, escalation paths, patient-choice scriptingEvery role has a named owner and a runbook
Live operations wave 16–9Two specialties per group, weekly operating reviewWeekly KPI pack produced without manual heroics
Expansion wave 29–11Add specialties or sites, tune matching and outreachStable exception volume and provider satisfaction
Evaluation and decision11–12Final measurement, scorecard, scale/no-scale decisionDocumented decision with conditions

Sponsor and RACI

ActivityPlan sponsorPlan network opsProvider operationsPlan IT / dataCompliance / privacyVendor
Charter and scale thresholdsARCCCI
Cohort selectionARCCIC
Data-sharing boundariesCCCRAC
Coverage and network data deliveryICIA/RCC
EHR integrationIIACIR
Provider enablementICA/RIIR
Exception operationsICA/RIIC
AI governance and monitoringCCCCAR
Measurement and reportingARCRIC

A = accountable, R = responsible, C = consulted, I = informed. One accountable name per row, or the row will not happen.

Minimum viable data set

Data elementSourcePurpose
Member eligibility, plan, productPlanNetwork verification at point of referral
Network participation by provider, location, productPlanIn-network determination
Provider directory with specialty, access, languagePlan and providerMatching inputs
Referral orders with specialty and clinical contextProvider EHRWorkflow trigger
Authorization requirements and statusPlan and utilization managementDependency management
Appointment and visit eventsProvider and platformScheduling and completion signals
Specialist result documentsProviderLoop closure
Claims or encounter extracts (pre-period and pilot)PlanCompletion confirmation and leakage

Define refresh cadence and file-level acceptance tests. Stale network files are the single most common reason providers stop trusting recommendations.

Data-sharing boundaries and safeguards

  • Share the minimum necessary to operate the workflow; document the purpose for each element.
  • Keep contracted rate detail out of clinician-facing screens; express cost as relative tiers where appropriate.
  • Segregate data by provider organization so one group's performance is not exposed to another without agreement.
  • Present recommendations as ranked options with visible rationale and one-click override.
  • Record patient preference as a matching input and log every override reason.
  • Confirm business associate agreements, access controls, audit logging, and retention before the first live referral.
  • Align model documentation, monitoring, and incident response to a recognized framework such as the NIST AI Risk Management Framework. See the companion AI governance checklist and our integration and security practices. This is operational guidance, not legal advice.

Pre-pilot baselines

Compute for the 6–12 months before go-live, using one written definition per metric:

  • In-network referral completion rate.
  • Referral-to-scheduled rate and referral-to-completed rate.
  • Time to appointment by specialty.
  • Out-of-network completed care rate and associated cost.
  • Authorization turnaround and first-pass approval rate.
  • Result return and acknowledgment rate.
  • Staff touches per referral, sampled by time study.
  • No-show rate at destination.

Recompute the same metrics for the comparison group over the same period. If a metric cannot be computed pre-period, it cannot be a scale threshold.

Evaluation design

Preferred: randomize at the site or clinic level within participating groups, so patient mix and referral patterns are balanced.

Practical alternative: matched comparison groups selected on pre-period referral volume, specialty mix, payer mix, geography, and baseline completion rate. Report both absolute change and difference-in-differences against the comparison group.

Guardrails: pre-register the primary endpoint (in-network referral completion rate is the usual choice), name secondary endpoints, define the analysis population, and specify how exclusions are handled — clinically justified out-of-network referrals, patient-choice overrides, and cases with incomplete claims runout.

Claims lag: allow at least 60–90 days of runout before final completion figures. Report interim results as provisional and label them as such.

Exception workflows

Every pilot generates exceptions; the question is whether they are visible. Define at minimum:

ExceptionTriggerOwnerTarget action
No in-network match foundMatching returns no qualifying destinationPlan network opsAccess remediation or documented exception
Authorization pending beyond thresholdDays since submission exceeds standardProvider operationsEscalate to utilization management
Patient unreachableOutreach attempts exhaustedProvider operationsAlternate channel or care team follow-up
Scheduling beyond access standardNo appointment within target windowPlan network opsWiden radius, alternate destination, or eConsult
Result not returnedDays post-visit exceeds thresholdProvider operationsDirect retrieval from destination
Patient declines recommendationOverride recordedProvider operationsHonor choice, log reason, monitor pattern

Weekly and final KPIs

Weekly operating pack: referrals routed, share with in-network destination selected, referral-to-scheduled rate, median days to appointment, authorization pending count and aging, exception queue volume and aging, override count and reasons, provider-reported friction items.

Final evaluation: in-network referral completion rate, leakage rate, time to appointment, authorization turnaround and first-pass rate, result return and acknowledgment rate, staff touches per referral, no-show rate, and access equity segmentation by language, geography, and coverage type.

Illustrative scale thresholds

The figures below are illustrative examples only — they are not benchmarks, not commitments, and not observed results. Set your own thresholds from your baseline before go-live.

MetricIllustrative threshold
In-network referral completion rate+5 to +10 percentage points versus baseline and comparison group
Referral-to-scheduled rate+10 percentage points
Median time to appointment−20%
Authorization turnaround−25%
Result return rate≥ 85% within 30 days of visit
Staff touches per referral−30%
Provider satisfactionMajority would continue using the workflow
GovernanceZero unresolved high-severity findings

Scale decision scorecard

DimensionWeightEvidence
Completion and leakage impact30%Primary endpoint versus comparison group
Access improvement15%Time to appointment, scheduling within standard
Provider adoption and satisfaction15%Usage rates, override patterns, interviews
Operational efficiency15%Staff touches, exception aging
Integration durability10%Interface stability, data freshness, defect rate
Governance and safety10%Monitoring results, audit completeness, incidents
Economic case5%Modeled cost avoidance against total cost of ownership

Score, then apply one of three decisions: scale, extend with named conditions, or stop. Write the decision down with its evidence.

Where ReferralPoint fits

ReferralPoint's public documentation describes a network-aware orchestration layer: IdealMATCH™ for insurance- and network-aware specialist selection across clinical need, network status, cost, quality, access, geography, language and social needs, and patient preference; IntelligentDATA™ for the underlying provider and network data; NetworkMANAGEMENT™ for network performance and tiering; Auto PriorAUTH™ for authorization automation; Auto ReferralCOORDINATOR™ for patient outreach, scheduling, and coordination; and Auto 360° VISIBILITY™ for end-to-end status and loop closure. That combination maps to the pilot stages above, which is why we scope pilots by workflow stage rather than by module count. Clinician and patient choice remain intact. See solutions for payers, our product family, the comparison hub, and case studies. Integration methods vary by EHR and version — verify during procurement.

Sources and methodology

Capability statements reflect ReferralPoint's official product, solution, and integration pages as of publication. Program design draws on primary sources: CMS materials on value-based care, the CMS Closing the Referral Loop electronic clinical quality measure, CMS-0057-F interoperability and prior authorization provisions, HL7 Da Vinci Prior Authorization Support implementation guidance, and the NIST AI Risk Management Framework for AI governance structure. All thresholds are labeled illustrative. No customer outcomes, pricing, or competitor performance claims are asserted.

Frequently asked questions

Q: How should a payer pilot AI-driven referral management? A: Sponsor it centrally, run it inside provider workflow, and treat it as an experiment. Select two to four representative provider groups, lock a pre-period baseline computed identically for pilot and comparison groups, define data-sharing boundaries and patient-choice safeguards before go-live, operate 12 weeks with weekly reviews, and evaluate against thresholds published in week zero.

Q: How long should the pilot run? A: Twelve weeks of live operations is usually the minimum that produces stable signal, with two weeks of design and two to three weeks of data and integration work before it. Add 60 to 90 days of claims runout before final completion and leakage figures. Shorter pilots measure novelty; longer ones without decision gates tend to drift.

Q: How many provider groups should participate? A: Two to four. One group cannot separate platform effect from local operating quirks, and more than four multiplies integration and enablement work before you have learned anything. Choose on referral volume, specialty mix, integration readiness, and willingness to share baseline data — not on enthusiasm alone, which biases results.

Q: What is the single most important metric? A: In-network referral completion rate, because it only improves when every upstream step actually happened. Report it against a matched comparison group rather than against the baseline alone, and pair it with time to appointment, authorization turnaround, and staff touches so you can attribute the change to access, administration, or automation.

Q: How do we protect patient choice in a payer-sponsored pilot? A: Present ranked recommendations with visible rationale, keep override available at one click, capture the override reason, and treat patient preference as a matching input rather than an exception. Monitor override rates and patterns weekly. Include patient-choice safeguards in the pilot charter and review them with compliance before the first live referral.

Q: What if we cannot randomize? A: Use matched comparison groups selected on pre-period volume, specialty mix, payer mix, geography, and baseline completion. Report difference-in-differences alongside absolute change, and state the limitation plainly in the final report. A well-matched comparison group is far better than a single-arm before-and-after, which credits the platform for seasonality and secular trends.

Q: When should a pilot be stopped rather than extended? A: Stop when integration cannot deliver the minimum viable data set reliably, when providers disengage despite enablement, when exception volume grows rather than stabilizes, or when the primary endpoint shows no movement against the comparison group with no identified fixable cause. Extend only with named conditions, an owner, and a new decision date.