Halcy Research · Health systems · Study companion

How AI recommends health systems, line by line

We asked an AI model with web search for care recommendations 1,800 times, across 10 service lines and 30 US metros, then compared the local providers it recommended with each market’s supply of physicians.

Pre-registered test · verdict

Contradicted

The hypothesis, as pre-registered: “Health systems hold a real structural advantage in AI recommendations, but it thins out badly in the service lines where patients choose themselves.”

In the five tested lines, the pre-registered test rules out the hypothesis’s shortfall: systems’ share of AI-recommended local physicians and groups was not 10 or more points below their share of local physicians, and the estimate sits above supply. The hypothesis’s “thins out badly where patients choose” does not hold on this measure.

Estimate
+9.46points
95% interval
2.04 to 12.05
Shortfall threshold
−10points

This page is the data companion to our article on how AI recommends health systems. It carries the methods, every sensitivity and the results for each service line. Download the results (CSV, 300 rows).

Answers
1,800
Metros
30
Service lines
10
Model
gpt-5.6-luna, web search
Collected
Oct 4, 2026

Published Last updated Analysis commit 9f750fc

The pre-registered test

The shortfall the test looked for did not appear

The estimate is +9.46 points, with a 95% interval of 2.04 to 12.05. The rule fixed before the data called the hypothesis contradicted if the interval’s lower bound sat above −10, and it does.

The gap compares two shares in each market: systems’ share of the local physicians and groups the AI recommended, and systems’ share of local physicians in that line’s specialties. Four sensitivities sit beside the estimate. One of them, which leaves out recommendations resolved by web verification, does not rule out the shortfall. Two of the five intervals include zero, so the data do not show that AI recommends system clinicians above their supply share either.

Pre-registered estimate and sensitivities, gap in percentage points with 95% intervals

Pre-registered estimateFive tested lines, each weighted equally

+9.462.04 to 12.05Contradicted the verdict

Unknowns at the cell’s supply ratePre-registered · 8 unknown recommendations assigned

+9.421.99 to 12.00Contradicted above zero

Web-verified recommendations left outPre-data amendment · 692 left out; maternity has no eligible market and drops out

−5.68−12.42 to 0.98Inconclusive includes zero

Maternity left outFour lines · defined after coding, before the gap was computed

+4.14−0.19 to 8.42Contradicted includes zero

Answer-named recommendations counted as systemAn upper bound · 54 of 370 reassigned · defined after coding, before the gap was computed

+13.085.84 to 15.95Contradicted above zero

Each run aloneRuns one, two and three · maternity out of every run · no interval

−1.65 · +5.60 · +4.80Point estimates only

Points are estimates. Lines are 95% percentile intervals from 10,000 market-cluster bootstrap resamples drawn within tiers. The verdict comes from the pre-registered estimate alone. The sensitivities are reported beside it, and none of them changes the rule.

The five tested lines

Mean gap by tested line, with eligible and excluded markets
Tested lineMean gap, pointsMarkets eligibleExcluded
Primary care+0.81282
Maternity †+30.73129
Joint replacement+18.88273
Colonoscopy+7.63273
Vasectomy−10.73246
Estimate, equal-weight mean of the five+9.462.04 to 12.05

† One eligible market, Philadelphia, whose 4 recommendations were all system fetal-medicine units; 94.3% of local maternity recommendations named a hospital, so 29 of 30 markets had too few clinician recommendations to test. Per-line means are descriptive. Only the five-line mean is tested. A market is eligible for a line when it has at least 10 local physicians in the line’s specialties and at least 3 resolved recommended physicians or groups. In a resample that does not draw Philadelphia, maternity leaves the mean, which is one reason the interval sits unevenly around the estimate.

Answers vary from run to run. Only 11.08% of the distinct names recommended in a market and line appeared in all three runs, and 31.01% in two or more. Computed on each run alone, the estimate is −1.65, +5.60 and +4.80 points.

What AI recommends

Who gets recommended, line by line

In the three referred lines, systems held 93.1% of resolved local recommendations, averaged across lines. In the five tested lines the average was 66.7%, from 99.0% in maternity to 39.5% in vasectomy.

Hospitals’ share of local recommendations runs from 94.3% in maternity to 3.0% in vasectomy. It falls by tier: 35.6% in the ten largest metros, 21.2% in mid-size metros and 11.4% in small metros. Raw system shares partly reflect who employs local physicians, which is why the test above compares recommendations with local supply.

Results by service line, grouped by role in the study
Service lineWho was recommended locallySystem shareGap vs local supply, pointsSystem not named
Tested · self-directed lines · 52,079 local recommendations · 31.3% hospitals · system share 66.7%, averaged across lines
MaternityTested318 local recommendations94.3% hospitals · 3.5% physicians and groups99.0%308 of 311+30.731 market †10.0%1 of 10
Joint replacementTested475 local recommendations39.4% hospitals · 59.2% physicians and groups68.0%323 of 475+18.8827 markets2.1%3 of 140
ColonoscopyTested360 local recommendations24.2% hospitals · 61.7% physicians and groups · 14.2% other, mostly surgery centers68.9%244 of 354+7.6327 markets4.7%6 of 129
Primary careTested624 local recommendations10.7% hospitals · 89.1% physicians and groups58.1%362 of 623+0.8128 markets6.8%20 of 295
VasectomyTested302 local recommendations3.0% hospitals · 96.7% physicians and groups39.5%118 of 299−10.7324 markets2.7%3 of 110
Referred lines · descriptive, no test · 3706 local recommendations · 20.8% hospitals · system share 93.1%, averaged across lines
Heart surgeryDescriptive217 local recommendations30.9% hospitals · 68.2% physicians and groups99.1%213 of 215+19.6314 markets6.8%10 of 146
Brain tumor surgeryDescriptive244 local recommendations18.9% hospitals · 73.0% physicians and groups · 8.2% ER or urgent care94.2%210 of 223+22.9519 markets9.0%15 of 166
Breast cancerDescriptive245 local recommendations13.9% hospitals · 86.1% physicians and groups86.1%211 of 245+20.5927 markets6.1%11 of 180
Self-directed, outside the test · descriptive · 2797 local recommendations · 6.9% hospitals · kept out of the test because Medicare enrollment under-counts independent clinicians in these lines
PsychiatryDescriptive353 local recommendations10.2% hospitals · 63.7% physicians and groups · 26.1% other, mostly ER or urgent care–not in summary+16.8525 markets10.2%13 of 128
PediatricsDescriptive444 local recommendations4.3% hospitals · 75.5% physicians and groups · 20.3% ER or urgent care–not in summary+15.3821 markets3.9%8 of 206

Who was recommended: affirmative recommendations placed in the market, open-ended prompts, three runs pooled. Within each group, lines run from the most hospital-heavy to the least.

System share: the system-affiliated share of resolved local recommendations, counting hospitals, physicians and groups, pooled over markets. The summary does not compute it for pediatrics or psychiatry.

Gap vs local supply: each line’s mean, over eligible markets, of systems’ share of recommended local physicians and groups minus their share of local physicians. Bars run left or right from zero on a ±35-point scale. Only the five tested lines enter the test; every other gap is descriptive.

System not named: recommended system-affiliated physicians and groups for which neither the answer nor a cited page names the health system.

† One eligible market, Philadelphia.

Download results.csv 300 rows, one per market and line, with every sensitivity

Naming the system

When AI recommends a system’s clinicians, it usually names the system

Of 1,510 recommended local physicians and groups affiliated with a health system, 90 came without the system’s name in the answer or its cited page. That is 6.0%, with a 95% interval of 2.85% to 9.84%.

Share of system-affiliated recommendations without the system named
  • All system-affiliated recommendations6.0%90 of 1,510

By who was recommended

  • Physicians named individually1.4%3 of 218
  • Group practices6.7%87 of 1,292

By market tier

  • The ten largest metros8.4%52 of 616
  • Mid-size metros6.0%24 of 397
  • Small metros2.8%14 of 497
Run three · no system named
“…worth considering if you live near the Upper East Side, Upper West Side, Midtown, Murray Hill, Tribeca, Long Island City, or Brooklyn Heights. It has numerous dedicated primary-care offices and internal-medicine practices.”

resp_038709449803c345006ac1f72d5fdc87d08b655718b5b7eb05

Run two · system named
“…a strong option if you prefer coordinated care through NewYork-Presbyterian or want locations in Manhattan, Brooklyn, or Queens.”

resp_01e91ba173379f3e006ac1f72d272887d0b5c86a63d8ae25a6

Same practice, two runs. New York, primary care. AHRQ links the Weill Cornell Medicine practice group to NewYork-Presbyterian, and NewYork-Presbyterian’s primary care locations page lists providers from Columbia and Weill Cornell Medicine. Across all affirmative recommendations, a health system was named 68.05% of the time under open-ended prompts and 68.79% under “best doctor” prompts.

Outside the metro

In small metros, more answers point outside the metro

In the small metros, 14.38% of affirmative recommendations pointed outside the metro, and 113 of 300 answers included at least one. In the ten largest metros it was 2 of 1,436.

All ten lines, open-ended prompts. This describes where answers point and says nothing about the reason.

Out-of-market share of affirmative recommendations by market tier
TierRecommendationsOutsideShare, 95% intervalAnswers with one or more
The ten largest metros1,43620.14%0% to 0.34%2 of 3000.67%
Mid-size metros1,3051269.66%5.26% to 14.64%75 of 30025.0%
Small metros1,25918114.38%10.85% to 17.43%113 of 30037.67%

Methods

How the study was run

Scope statement

A standardized audit of the OpenAI API using gpt-5.6-luna, the model behind free ChatGPT since August 2026, with web search, across 30 US metropolitan markets in three population tiers, in one collection window. Not the consumer app, not rural or micropolitan markets, no claims about bookings, revenue or system size.

Instrument

Model and settings
gpt-5.6-luna through the OpenAI Responses API, in real time, with the web search tool. Reasoning effort medium, up to 4,096 output tokens and 3 tool calls, approximate user location set to the market’s principal city.
Prompts
Two framings per line, one wording each: an open-ended request for care and a “best doctor” question. Both are listed verbatim below.
Collection
October 4, 2026, in one collection window. Three runs of each prompt in each market: 900 open-ended answers and 900 “best doctor” answers.
Coding
Each provider mention was coded for recommendation status, entity type, whether the answer or a cited page names a health system, affiliation and geography, by a primary model coder against a committed rubric. A second model coder checked a sample (validation below).

Supply benchmark

Local physicians
CMS Doctors and Clinicians National Downloadable File, released September 10, 2026. It lists Medicare-enrolled clinicians, one row per NPI, placed in the metro by practice address.
System affiliation
AHRQ Compendium of U.S. Health Systems, 2022 Group Practice Linkage File, with a name match to AHRQ’s 2023 system and hospital names for groups the file does not list. Hospitals: AHRQ 2023 hospital linkage.
One rule, both sides
The same rule classifies local supply and recommended clinicians. Hospital privileges never count as affiliation. Web checks of official sites resolve only recommended clinicians the rule leaves unknown.
Resolution
1,354 of 1,362 recommended local physicians and groups in the tested lines (99.41%) resolved to system or independent: 618 by the rule and 736 by web verification. Eight stayed unknown.

The test

A cell is one market and one tested line, open-ended prompts, three runs pooled. The gap is systems’ share of recommended local physicians and groups minus their share of local physicians, in percentage points. The estimate is the equal-weight mean over the five lines of each line’s mean gap over eligible markets. Uncertainty comes from 10,000 market-cluster bootstrap resamples drawn within tiers (seed 20261007), as a 95% percentile interval.

The verdict rule was fixed before the data: supported if the estimate is −10 or lower and the interval’s upper bound is below zero, contradicted if the lower bound is above −10, otherwise inconclusive. It was set in code at analysis commit 9f750fc, committed before the first run on the study data. Of 397 physician mentions in the tested lines, 44 were left out for a non-physician credential, and 43 tested cells fell short of eligibility (13 in the ten largest metros, 11 mid-size, 19 small). Every cell and its reason is in the CSV.

Markets

Census Vintage 2024 metropolitan areas, in three tiers of ten.

The ten largest metros

All ten, by 2024 population.

  • New York-Newark-Jersey City, NY-NJ
  • Los Angeles-Long Beach-Anaheim, CA
  • Chicago-Naperville-Elgin, IL-IN
  • Dallas-Fort Worth-Arlington, TX
  • Houston-Pasadena-The Woodlands, TX
  • Miami-Fort Lauderdale-West Palm Beach, FL
  • Washington-Arlington-Alexandria, DC-VA-MD-WV
  • Atlanta-Sandy Springs-Roswell, GA
  • Philadelphia-Camden-Wilmington, PA-NJ-DE-MD
  • Phoenix-Mesa-Chandler, AZ

Mid-size metros

Ten seeded draws from metros of 500,000 to 1.5 million.

  • El Paso, TX
  • Des Moines-West Des Moines, IA
  • Omaha, NE-IA
  • Portland-South Portland, ME
  • Ogden, UT
  • Cape Coral-Fort Myers, FL
  • Modesto, CA
  • Salt Lake City-Murray, UT
  • Kiryas Joel-Poughkeepsie-Newburgh, NY
  • Worcester, MA

Small metros

Ten seeded draws from metros of 100,000 to 350,000 with at least one short-term acute-care hospital.

  • Coeur d’Alene, ID
  • Slidell-Mandeville-Covington, LA
  • Charlottesville, VA
  • Olympia-Lacey-Tumwater, WA
  • Joplin, MO-KS
  • Lafayette-West Lafayette, IN
  • Sumter, SC
  • Jefferson City, MO
  • Bloomington, IL
  • Sebastian-Vero Beach-West Vero Corridor, FL

The ten service lines and their prompts

{city} is the market’s principal city and state, for example “Omaha, Nebraska”. Each line’s supply is the CMS specialty shown under its name.

Service lines, roles and prompts, verbatim
LineRoleOpen-ended prompt“Best doctor” prompt
Primary careFamily practice; internal medicine; general practice (hospitalists excluded)TestedI need a new primary care doctor in {city}. Where should I go?Who is the best primary care doctor in {city}?
MaternityObstetrics/gynecologyTestedI’m pregnant and choosing where to have my baby in {city}. Where should I go?Who is the best OB-GYN in {city}?
Joint replacementOrthopedic surgeryTestedI need a knee replacement. Where should I get it in {city}?Who is the best knee replacement surgeon in {city}?
ColonoscopyGastroenterologyTestedI need a colonoscopy. Where should I get one in {city}?Who is the best gastroenterologist for a colonoscopy in {city}?
VasectomyUrologyTestedI want to get a vasectomy. Where should I go in {city}?Who is the best urologist for a vasectomy in {city}?
Heart surgeryCardiac surgery; thoracic surgeryReferredI need heart valve surgery. Where should I go in {city}?Who is the best heart surgeon in {city}?
Breast cancerHematology/oncology; medical oncology; surgical oncologyReferredI was just diagnosed with breast cancer. Where should I get treated in {city}?Who is the best breast cancer doctor in {city}?
Brain tumor surgeryNeurosurgeryReferredI need surgery for a brain tumor. Where should I go in {city}?Who is the best brain surgeon in {city}?
PediatricsPediatric medicineOutside the testI need a pediatrician for my child in {city}. Where should I go?Who is the best pediatrician in {city}?
PsychiatryPsychiatryOutside the testI need a psychiatrist in {city}. Where should I go?Who is the best psychiatrist in {city}?

Validation

A second model coder independently coded a stratified sample of 150 open-ended responses. The pass rule was κ of at least 0.8 and precision and recall of at least 0.9 for the system class. All three fields passed round one, and none was cut.

Validation results
FieldMeasureResultRule
Recommendation statusCohen’s κ0.8777≥ 0.8 · passed
Health system namedCohen’s κ0.967≥ 0.8 · passed
Affiliation, system class554 resolved mentionsPrecision / recall0.9753 / 0.9621≥ 0.9 each · passed
Affiliation, independent classPrecision / recall0.9318 / 0.8865Reported beside, not in the rule

Limitations

What limits these results, and what they do not show

Limitations

  • The API, not the app. The audit queried the OpenAI API with web search and a city-level location. It does not measure the consumer ChatGPT app.
  • One window, variable answers. Every answer was collected on October 4, 2026. Only 11.08% of distinct recommended names appeared in all three runs.
  • A 2022 affiliation file. Of 100 clinicians and groups the rule classed as independent, web checks found 6 now system-employed according to the system’s own site. The same rule classifies supply and recommendations, so staleness touches both sides.
  • Rule-side identification. The second coder agreed less with rule-resolved affiliations (system precision 0.9339, n 210) than with web-resolved ones (0.9959, n 344). These errors move recommended clinicians toward independent. The answer-named sensitivity, +13.08, is an upper bound on their effect.
  • Medicare enrollment. Supply counts Medicare-enrolled clinicians, so independents who do not bill Medicare are missing, which raises the measured supply share. The bias is largest in pediatrics and psychiatry, which is why they sit outside the test.
  • Web verification. 736 of 1,354 resolved tested-line recommendations were resolved by checking official sites. Leaving them out gives the one estimate whose interval reaches −12.42.
  • Maternity rests on one market. Philadelphia is its only eligible market. Without maternity the estimate is +4.14 (−0.19 to 8.42).

What this study does not show

  • That AI recommends system clinicians above their supply share. Two of the five intervals include zero, and the estimate without web-verified recommendations is negative.
  • That AI under-recommends health systems in the lines where patients choose.
  • A result for any single line. Per-line gaps are descriptive. Only the five-line mean is tested.
  • Why an answer recommends one provider over another. The naming and out-of-market figures describe answers, not reasons, and cited sources are described, not shown to raise visibility.
  • Bookings, revenue or system size. The study measures recommendations only.
  • Other products and places. The consumer app, other AI tools, and rural or micropolitan markets are outside the scope. The mid-size and small metros are seeded draws, and the ten largest metros are all ten.

Disclosure. Halcy is a commercial service that sells search and AI-visibility work to healthcare practices, including independent practices. We designed, ran and paid for this study, and we publish the full results so readers can check every figure on this page.