A payer-sponsored artificial intelligence (AI) referral management pilot should be designed as a multi-organization operations experiment, not a software installation. Choose two to four provider groups, fix a measurable baseline, define data-sharing boundaries and patient-choice safeguards up front, run 12 weeks with weekly operating reviews, and pre-commit the thresholds that decide scale or no-scale.
Quick answer
- Sponsor at the plan, execute inside provider workflow.
- Pick cohorts by referral volume, specialty mix, and integration readiness — not enthusiasm alone.
- Baseline before go-live; a pilot without a pre-period is a demo.
- Use matched comparison groups where randomization is impractical.
- Publish scale thresholds in week zero so results are judged, not negotiated.
This is deliberately narrower than a general implementation roadmap: the hard parts of a payer pilot are cross-organizational — data boundaries, provider enablement, and attribution across entities you do not employ.
Why payer pilots fail
- No baseline. Completion and leakage were never measured the same way before go-live.
- Volunteer bias. The most engaged group is the least representative one.
- Ambiguous ownership. Plan network teams, provider operations, and IT each assume another party owns exceptions.
- Data friction. Coverage and network files arrive late, so matching runs on stale data and providers lose trust in week two.
- Moving goalposts. Success is defined after results appear.
- Governance as an afterthought. Compliance review lands in week ten and stalls scale.
12-week pilot plan
| Phase | Weeks | Objectives | Exit criteria |
|---|---|---|---|
| Design and governance | 0–2 | Charter, cohorts, evaluation design, data-sharing boundaries, AI governance sign-off | Signed charter and approved data flows |
| Data and integration | 2–4 | Coverage/network files, provider directory, EHR interfaces, test transactions | End-to-end test referral completes in a test environment |
| Baseline lock | 4–5 | Pre-period metrics computed identically for pilot and comparison groups | Baseline report accepted by all parties |
| Enablement | 5–6 | Role-based training, exception playbooks, escalation paths, patient-choice scripting | Every role has a named owner and a runbook |
| Live operations wave 1 | 6–9 | Two specialties per group, weekly operating review | Weekly KPI pack produced without manual heroics |
| Expansion wave 2 | 9–11 | Add specialties or sites, tune matching and outreach | Stable exception volume and provider satisfaction |
| Evaluation and decision | 11–12 | Final measurement, scorecard, scale/no-scale decision | Documented decision with conditions |
Sponsor and RACI
| Activity | Plan sponsor | Plan network ops | Provider operations | Plan IT / data | Compliance / privacy | Vendor |
|---|---|---|---|---|---|---|
| Charter and scale thresholds | A | R | C | C | C | I |
| Cohort selection | A | R | C | C | I | C |
| Data-sharing boundaries | C | C | C | R | A | C |
| Coverage and network data delivery | I | C | I | A/R | C | C |
| EHR integration | I | I | A | C | I | R |
| Provider enablement | I | C | A/R | I | I | R |
| Exception operations | I | C | A/R | I | I | C |
| AI governance and monitoring | C | C | C | C | A | R |
| Measurement and reporting | A | R | C | R | I | C |
A = accountable, R = responsible, C = consulted, I = informed. One accountable name per row, or the row will not happen.
Minimum viable data set
| Data element | Source | Purpose |
|---|---|---|
| Member eligibility, plan, product | Plan | Network verification at point of referral |
| Network participation by provider, location, product | Plan | In-network determination |
| Provider directory with specialty, access, language | Plan and provider | Matching inputs |
| Referral orders with specialty and clinical context | Provider EHR | Workflow trigger |
| Authorization requirements and status | Plan and utilization management | Dependency management |
| Appointment and visit events | Provider and platform | Scheduling and completion signals |
| Specialist result documents | Provider | Loop closure |
| Claims or encounter extracts (pre-period and pilot) | Plan | Completion confirmation and leakage |
Define refresh cadence and file-level acceptance tests. Stale network files are the single most common reason providers stop trusting recommendations.
Data-sharing boundaries and safeguards
- Share the minimum necessary to operate the workflow; document the purpose for each element.
- Keep contracted rate detail out of clinician-facing screens; express cost as relative tiers where appropriate.
- Segregate data by provider organization so one group's performance is not exposed to another without agreement.
- Present recommendations as ranked options with visible rationale and one-click override.
- Record patient preference as a matching input and log every override reason.
- Confirm business associate agreements, access controls, audit logging, and retention before the first live referral.
- Align model documentation, monitoring, and incident response to a recognized framework such as the NIST AI Risk Management Framework. See the companion AI governance checklist and our integration and security practices. This is operational guidance, not legal advice.
Pre-pilot baselines
Compute for the 6–12 months before go-live, using one written definition per metric:
- In-network referral completion rate.
- Referral-to-scheduled rate and referral-to-completed rate.
- Time to appointment by specialty.
- Out-of-network completed care rate and associated cost.
- Authorization turnaround and first-pass approval rate.
- Result return and acknowledgment rate.
- Staff touches per referral, sampled by time study.
- No-show rate at destination.
Recompute the same metrics for the comparison group over the same period. If a metric cannot be computed pre-period, it cannot be a scale threshold.
Evaluation design
Preferred: randomize at the site or clinic level within participating groups, so patient mix and referral patterns are balanced.
Practical alternative: matched comparison groups selected on pre-period referral volume, specialty mix, payer mix, geography, and baseline completion rate. Report both absolute change and difference-in-differences against the comparison group.
Guardrails: pre-register the primary endpoint (in-network referral completion rate is the usual choice), name secondary endpoints, define the analysis population, and specify how exclusions are handled — clinically justified out-of-network referrals, patient-choice overrides, and cases with incomplete claims runout.
Claims lag: allow at least 60–90 days of runout before final completion figures. Report interim results as provisional and label them as such.
Exception workflows
Every pilot generates exceptions; the question is whether they are visible. Define at minimum:
| Exception | Trigger | Owner | Target action |
|---|---|---|---|
| No in-network match found | Matching returns no qualifying destination | Plan network ops | Access remediation or documented exception |
| Authorization pending beyond threshold | Days since submission exceeds standard | Provider operations | Escalate to utilization management |
| Patient unreachable | Outreach attempts exhausted | Provider operations | Alternate channel or care team follow-up |
| Scheduling beyond access standard | No appointment within target window | Plan network ops | Widen radius, alternate destination, or eConsult |
| Result not returned | Days post-visit exceeds threshold | Provider operations | Direct retrieval from destination |
| Patient declines recommendation | Override recorded | Provider operations | Honor choice, log reason, monitor pattern |
Weekly and final KPIs
Weekly operating pack: referrals routed, share with in-network destination selected, referral-to-scheduled rate, median days to appointment, authorization pending count and aging, exception queue volume and aging, override count and reasons, provider-reported friction items.
Final evaluation: in-network referral completion rate, leakage rate, time to appointment, authorization turnaround and first-pass rate, result return and acknowledgment rate, staff touches per referral, no-show rate, and access equity segmentation by language, geography, and coverage type.
Illustrative scale thresholds
The figures below are illustrative examples only — they are not benchmarks, not commitments, and not observed results. Set your own thresholds from your baseline before go-live.
| Metric | Illustrative threshold |
|---|---|
| In-network referral completion rate | +5 to +10 percentage points versus baseline and comparison group |
| Referral-to-scheduled rate | +10 percentage points |
| Median time to appointment | −20% |
| Authorization turnaround | −25% |
| Result return rate | ≥ 85% within 30 days of visit |
| Staff touches per referral | −30% |
| Provider satisfaction | Majority would continue using the workflow |
| Governance | Zero unresolved high-severity findings |
Scale decision scorecard
| Dimension | Weight | Evidence |
|---|---|---|
| Completion and leakage impact | 30% | Primary endpoint versus comparison group |
| Access improvement | 15% | Time to appointment, scheduling within standard |
| Provider adoption and satisfaction | 15% | Usage rates, override patterns, interviews |
| Operational efficiency | 15% | Staff touches, exception aging |
| Integration durability | 10% | Interface stability, data freshness, defect rate |
| Governance and safety | 10% | Monitoring results, audit completeness, incidents |
| Economic case | 5% | Modeled cost avoidance against total cost of ownership |
Score, then apply one of three decisions: scale, extend with named conditions, or stop. Write the decision down with its evidence.
Where ReferralPoint fits
ReferralPoint's public documentation describes a network-aware orchestration layer: IdealMATCH™ for insurance- and network-aware specialist selection across clinical need, network status, cost, quality, access, geography, language and social needs, and patient preference; IntelligentDATA™ for the underlying provider and network data; NetworkMANAGEMENT™ for network performance and tiering; Auto PriorAUTH™ for authorization automation; Auto ReferralCOORDINATOR™ for patient outreach, scheduling, and coordination; and Auto 360° VISIBILITY™ for end-to-end status and loop closure. That combination maps to the pilot stages above, which is why we scope pilots by workflow stage rather than by module count. Clinician and patient choice remain intact. See solutions for payers, our product family, the comparison hub, and case studies. Integration methods vary by EHR and version — verify during procurement.
Sources and methodology
Capability statements reflect ReferralPoint's official product, solution, and integration pages as of publication. Program design draws on primary sources: CMS materials on value-based care, the CMS Closing the Referral Loop electronic clinical quality measure, CMS-0057-F interoperability and prior authorization provisions, HL7 Da Vinci Prior Authorization Support implementation guidance, and the NIST AI Risk Management Framework for AI governance structure. All thresholds are labeled illustrative. No customer outcomes, pricing, or competitor performance claims are asserted.
Frequently asked questions
Q: How should a payer pilot AI-driven referral management? A: Sponsor it centrally, run it inside provider workflow, and treat it as an experiment. Select two to four representative provider groups, lock a pre-period baseline computed identically for pilot and comparison groups, define data-sharing boundaries and patient-choice safeguards before go-live, operate 12 weeks with weekly reviews, and evaluate against thresholds published in week zero.
Q: How long should the pilot run? A: Twelve weeks of live operations is usually the minimum that produces stable signal, with two weeks of design and two to three weeks of data and integration work before it. Add 60 to 90 days of claims runout before final completion and leakage figures. Shorter pilots measure novelty; longer ones without decision gates tend to drift.
Q: How many provider groups should participate? A: Two to four. One group cannot separate platform effect from local operating quirks, and more than four multiplies integration and enablement work before you have learned anything. Choose on referral volume, specialty mix, integration readiness, and willingness to share baseline data — not on enthusiasm alone, which biases results.
Q: What is the single most important metric? A: In-network referral completion rate, because it only improves when every upstream step actually happened. Report it against a matched comparison group rather than against the baseline alone, and pair it with time to appointment, authorization turnaround, and staff touches so you can attribute the change to access, administration, or automation.
Q: How do we protect patient choice in a payer-sponsored pilot? A: Present ranked recommendations with visible rationale, keep override available at one click, capture the override reason, and treat patient preference as a matching input rather than an exception. Monitor override rates and patterns weekly. Include patient-choice safeguards in the pilot charter and review them with compliance before the first live referral.
Q: What if we cannot randomize? A: Use matched comparison groups selected on pre-period volume, specialty mix, payer mix, geography, and baseline completion. Report difference-in-differences alongside absolute change, and state the limitation plainly in the final report. A well-matched comparison group is far better than a single-arm before-and-after, which credits the platform for seasonality and secular trends.
Q: When should a pilot be stopped rather than extended? A: Stop when integration cannot deliver the minimum viable data set reliably, when providers disengage despite enablement, when exception volume grows rather than stabilizes, or when the primary endpoint shows no movement against the comparison group with no identified fixable cause. Extend only with named conditions, an owner, and a new decision date.



