Halcy Research · Health systems · Study companion
How AI recommends health systems, line by line
We asked an AI model with web search for care recommendations 1,800 times, across 10 service lines and 30 US metros, then compared the local providers it recommended with each market’s supply of physicians.
Pre-registered test · verdict
Contradicted
The hypothesis, as pre-registered: “Health systems hold a real structural advantage in AI recommendations, but it thins out badly in the service lines where patients choose themselves.”
In the five tested lines, the pre-registered test rules out the hypothesis’s shortfall: systems’ share of AI-recommended local physicians and groups was not 10 or more points below their share of local physicians, and the estimate sits above supply. The hypothesis’s “thins out badly where patients choose” does not hold on this measure.
- Estimate
- +9.46points
- 95% interval
- 2.04 to 12.05
- Shortfall threshold
- −10points
This page is the data companion to our article on how AI recommends health systems. It carries the methods, every sensitivity and the results for each service line. Download the results (CSV, 300 rows).
- Answers
- 1,800
- Metros
- 30
- Service lines
- 10
- Model
- gpt-5.6-luna, web search
- Collected
- Oct 4, 2026
Published Last updated Analysis commit 9f750fc
The pre-registered test
The shortfall the test looked for did not appear
The estimate is +9.46 points, with a 95% interval of 2.04 to 12.05. The rule fixed before the data called the hypothesis contradicted if the interval’s lower bound sat above −10, and it does.
The gap compares two shares in each market: systems’ share of the local physicians and groups the AI recommended, and systems’ share of local physicians in that line’s specialties. Four sensitivities sit beside the estimate. One of them, which leaves out recommendations resolved by web verification, does not rule out the shortfall. Two of the five intervals include zero, so the data do not show that AI recommends system clinicians above their supply share either.
Pre-registered estimateFive tested lines, each weighted equally
Unknowns at the cell’s supply ratePre-registered · 8 unknown recommendations assigned
Web-verified recommendations left outPre-data amendment · 692 left out; maternity has no eligible market and drops out
Maternity left outFour lines · defined after coding, before the gap was computed
Answer-named recommendations counted as systemAn upper bound · 54 of 370 reassigned · defined after coding, before the gap was computed
Each run aloneRuns one, two and three · maternity out of every run · no interval
Points are estimates. Lines are 95% percentile intervals from 10,000 market-cluster bootstrap resamples drawn within tiers. The verdict comes from the pre-registered estimate alone. The sensitivities are reported beside it, and none of them changes the rule.
The five tested lines
| Tested line | Mean gap, points | Markets eligible | Excluded |
|---|---|---|---|
| Primary care | +0.81 | 28 | 2 |
| Maternity † | +30.73 | 1 | 29 |
| Joint replacement | +18.88 | 27 | 3 |
| Colonoscopy | +7.63 | 27 | 3 |
| Vasectomy | −10.73 | 24 | 6 |
| Estimate, equal-weight mean of the five | +9.462.04 to 12.05 | ||
† One eligible market, Philadelphia, whose 4 recommendations were all system fetal-medicine units; 94.3% of local maternity recommendations named a hospital, so 29 of 30 markets had too few clinician recommendations to test. Per-line means are descriptive. Only the five-line mean is tested. A market is eligible for a line when it has at least 10 local physicians in the line’s specialties and at least 3 resolved recommended physicians or groups. In a resample that does not draw Philadelphia, maternity leaves the mean, which is one reason the interval sits unevenly around the estimate.
Answers vary from run to run. Only 11.08% of the distinct names recommended in a market and line appeared in all three runs, and 31.01% in two or more. Computed on each run alone, the estimate is −1.65, +5.60 and +4.80 points.
What AI recommends
Who gets recommended, line by line
In the three referred lines, systems held 93.1% of resolved local recommendations, averaged across lines. In the five tested lines the average was 66.7%, from 99.0% in maternity to 39.5% in vasectomy.
Hospitals’ share of local recommendations runs from 94.3% in maternity to 3.0% in vasectomy. It falls by tier: 35.6% in the ten largest metros, 21.2% in mid-size metros and 11.4% in small metros. Raw system shares partly reflect who employs local physicians, which is why the test above compares recommendations with local supply.
| Service line | Who was recommended locally | System share | Gap vs local supply, points | System not named |
|---|---|---|---|---|
| Tested · self-directed lines · 52,079 local recommendations · 31.3% hospitals · system share 66.7%, averaged across lines | ||||
| MaternityTested318 local recommendations | 94.3% hospitals · 3.5% physicians and groups | 99.0%308 of 311 | +30.731 market † | 10.0%1 of 10 |
| Joint replacementTested475 local recommendations | 39.4% hospitals · 59.2% physicians and groups | 68.0%323 of 475 | +18.8827 markets | 2.1%3 of 140 |
| ColonoscopyTested360 local recommendations | 24.2% hospitals · 61.7% physicians and groups · 14.2% other, mostly surgery centers | 68.9%244 of 354 | +7.6327 markets | 4.7%6 of 129 |
| Primary careTested624 local recommendations | 10.7% hospitals · 89.1% physicians and groups | 58.1%362 of 623 | +0.8128 markets | 6.8%20 of 295 |
| VasectomyTested302 local recommendations | 3.0% hospitals · 96.7% physicians and groups | 39.5%118 of 299 | −10.7324 markets | 2.7%3 of 110 |
| Referred lines · descriptive, no test · 3706 local recommendations · 20.8% hospitals · system share 93.1%, averaged across lines | ||||
| Heart surgeryDescriptive217 local recommendations | 30.9% hospitals · 68.2% physicians and groups | 99.1%213 of 215 | +19.6314 markets | 6.8%10 of 146 |
| Brain tumor surgeryDescriptive244 local recommendations | 18.9% hospitals · 73.0% physicians and groups · 8.2% ER or urgent care | 94.2%210 of 223 | +22.9519 markets | 9.0%15 of 166 |
| Breast cancerDescriptive245 local recommendations | 13.9% hospitals · 86.1% physicians and groups | 86.1%211 of 245 | +20.5927 markets | 6.1%11 of 180 |
| Self-directed, outside the test · descriptive · 2797 local recommendations · 6.9% hospitals · kept out of the test because Medicare enrollment under-counts independent clinicians in these lines | ||||
| PsychiatryDescriptive353 local recommendations | 10.2% hospitals · 63.7% physicians and groups · 26.1% other, mostly ER or urgent care | –not in summary | +16.8525 markets | 10.2%13 of 128 |
| PediatricsDescriptive444 local recommendations | 4.3% hospitals · 75.5% physicians and groups · 20.3% ER or urgent care | –not in summary | +15.3821 markets | 3.9%8 of 206 |
Who was recommended: affirmative recommendations placed in the market, open-ended prompts, three runs pooled. Within each group, lines run from the most hospital-heavy to the least.
System share: the system-affiliated share of resolved local recommendations, counting hospitals, physicians and groups, pooled over markets. The summary does not compute it for pediatrics or psychiatry.
Gap vs local supply: each line’s mean, over eligible markets, of systems’ share of recommended local physicians and groups minus their share of local physicians. Bars run left or right from zero on a ±35-point scale. Only the five tested lines enter the test; every other gap is descriptive.
System not named: recommended system-affiliated physicians and groups for which neither the answer nor a cited page names the health system.
† One eligible market, Philadelphia.
Naming the system
When AI recommends a system’s clinicians, it usually names the system
Of 1,510 recommended local physicians and groups affiliated with a health system, 90 came without the system’s name in the answer or its cited page. That is 6.0%, with a 95% interval of 2.85% to 9.84%.
“…worth considering if you live near the Upper East Side, Upper West Side, Midtown, Murray Hill, Tribeca, Long Island City, or Brooklyn Heights. It has numerous dedicated primary-care offices and internal-medicine practices.”
resp_038709449803c345006ac1f72d5fdc87d08b655718b5b7eb05
“…a strong option if you prefer coordinated care through NewYork-Presbyterian or want locations in Manhattan, Brooklyn, or Queens.”
resp_01e91ba173379f3e006ac1f72d272887d0b5c86a63d8ae25a6
Same practice, two runs. New York, primary care. AHRQ links the Weill Cornell Medicine practice group to NewYork-Presbyterian, and NewYork-Presbyterian’s primary care locations page lists providers from Columbia and Weill Cornell Medicine. Across all affirmative recommendations, a health system was named 68.05% of the time under open-ended prompts and 68.79% under “best doctor” prompts.
Outside the metro
In small metros, more answers point outside the metro
In the small metros, 14.38% of affirmative recommendations pointed outside the metro, and 113 of 300 answers included at least one. In the ten largest metros it was 2 of 1,436.
All ten lines, open-ended prompts. This describes where answers point and says nothing about the reason.
| Tier | Recommendations | Outside | Share, 95% interval | Answers with one or more |
|---|---|---|---|---|
| The ten largest metros | 1,436 | 2 | 0.14%0% to 0.34% | 2 of 3000.67% |
| Mid-size metros | 1,305 | 126 | 9.66%5.26% to 14.64% | 75 of 30025.0% |
| Small metros | 1,259 | 181 | 14.38%10.85% to 17.43% | 113 of 30037.67% |
Methods
How the study was run
Scope statement
A standardized audit of the OpenAI API using gpt-5.6-luna, the model behind free ChatGPT since August 2026, with web search, across 30 US metropolitan markets in three population tiers, in one collection window. Not the consumer app, not rural or micropolitan markets, no claims about bookings, revenue or system size.
Instrument
- Model and settings
- gpt-5.6-luna through the OpenAI Responses API, in real time, with the web search tool. Reasoning effort medium, up to 4,096 output tokens and 3 tool calls, approximate user location set to the market’s principal city.
- Prompts
- Two framings per line, one wording each: an open-ended request for care and a “best doctor” question. Both are listed verbatim below.
- Collection
- October 4, 2026, in one collection window. Three runs of each prompt in each market: 900 open-ended answers and 900 “best doctor” answers.
- Coding
- Each provider mention was coded for recommendation status, entity type, whether the answer or a cited page names a health system, affiliation and geography, by a primary model coder against a committed rubric. A second model coder checked a sample (validation below).
Supply benchmark
- Local physicians
- CMS Doctors and Clinicians National Downloadable File, released September 10, 2026. It lists Medicare-enrolled clinicians, one row per NPI, placed in the metro by practice address.
- System affiliation
- AHRQ Compendium of U.S. Health Systems, 2022 Group Practice Linkage File, with a name match to AHRQ’s 2023 system and hospital names for groups the file does not list. Hospitals: AHRQ 2023 hospital linkage.
- One rule, both sides
- The same rule classifies local supply and recommended clinicians. Hospital privileges never count as affiliation. Web checks of official sites resolve only recommended clinicians the rule leaves unknown.
- Resolution
- 1,354 of 1,362 recommended local physicians and groups in the tested lines (99.41%) resolved to system or independent: 618 by the rule and 736 by web verification. Eight stayed unknown.
The test
A cell is one market and one tested line, open-ended prompts, three runs pooled. The gap is systems’ share of recommended local physicians and groups minus their share of local physicians, in percentage points. The estimate is the equal-weight mean over the five lines of each line’s mean gap over eligible markets. Uncertainty comes from 10,000 market-cluster bootstrap resamples drawn within tiers (seed 20261007), as a 95% percentile interval.
The verdict rule was fixed before the data: supported if the estimate is −10 or lower and the interval’s upper bound is below zero, contradicted if the lower bound is above −10, otherwise inconclusive. It was set in code at analysis commit 9f750fc, committed before the first run on the study data. Of 397 physician mentions in the tested lines, 44 were left out for a non-physician credential, and 43 tested cells fell short of eligibility (13 in the ten largest metros, 11 mid-size, 19 small). Every cell and its reason is in the CSV.
Markets
Census Vintage 2024 metropolitan areas, in three tiers of ten.
The ten largest metros
All ten, by 2024 population.
- New York-Newark-Jersey City, NY-NJ
- Los Angeles-Long Beach-Anaheim, CA
- Chicago-Naperville-Elgin, IL-IN
- Dallas-Fort Worth-Arlington, TX
- Houston-Pasadena-The Woodlands, TX
- Miami-Fort Lauderdale-West Palm Beach, FL
- Washington-Arlington-Alexandria, DC-VA-MD-WV
- Atlanta-Sandy Springs-Roswell, GA
- Philadelphia-Camden-Wilmington, PA-NJ-DE-MD
- Phoenix-Mesa-Chandler, AZ
Mid-size metros
Ten seeded draws from metros of 500,000 to 1.5 million.
- El Paso, TX
- Des Moines-West Des Moines, IA
- Omaha, NE-IA
- Portland-South Portland, ME
- Ogden, UT
- Cape Coral-Fort Myers, FL
- Modesto, CA
- Salt Lake City-Murray, UT
- Kiryas Joel-Poughkeepsie-Newburgh, NY
- Worcester, MA
Small metros
Ten seeded draws from metros of 100,000 to 350,000 with at least one short-term acute-care hospital.
- Coeur d’Alene, ID
- Slidell-Mandeville-Covington, LA
- Charlottesville, VA
- Olympia-Lacey-Tumwater, WA
- Joplin, MO-KS
- Lafayette-West Lafayette, IN
- Sumter, SC
- Jefferson City, MO
- Bloomington, IL
- Sebastian-Vero Beach-West Vero Corridor, FL
The ten service lines and their prompts
{city} is the market’s principal city and state, for example “Omaha, Nebraska”. Each line’s supply is the CMS specialty shown under its name.
| Line | Role | Open-ended prompt | “Best doctor” prompt |
|---|---|---|---|
| Primary careFamily practice; internal medicine; general practice (hospitalists excluded) | Tested | I need a new primary care doctor in {city}. Where should I go? | Who is the best primary care doctor in {city}? |
| MaternityObstetrics/gynecology | Tested | I’m pregnant and choosing where to have my baby in {city}. Where should I go? | Who is the best OB-GYN in {city}? |
| Joint replacementOrthopedic surgery | Tested | I need a knee replacement. Where should I get it in {city}? | Who is the best knee replacement surgeon in {city}? |
| ColonoscopyGastroenterology | Tested | I need a colonoscopy. Where should I get one in {city}? | Who is the best gastroenterologist for a colonoscopy in {city}? |
| VasectomyUrology | Tested | I want to get a vasectomy. Where should I go in {city}? | Who is the best urologist for a vasectomy in {city}? |
| Heart surgeryCardiac surgery; thoracic surgery | Referred | I need heart valve surgery. Where should I go in {city}? | Who is the best heart surgeon in {city}? |
| Breast cancerHematology/oncology; medical oncology; surgical oncology | Referred | I was just diagnosed with breast cancer. Where should I get treated in {city}? | Who is the best breast cancer doctor in {city}? |
| Brain tumor surgeryNeurosurgery | Referred | I need surgery for a brain tumor. Where should I go in {city}? | Who is the best brain surgeon in {city}? |
| PediatricsPediatric medicine | Outside the test | I need a pediatrician for my child in {city}. Where should I go? | Who is the best pediatrician in {city}? |
| PsychiatryPsychiatry | Outside the test | I need a psychiatrist in {city}. Where should I go? | Who is the best psychiatrist in {city}? |
Validation
A second model coder independently coded a stratified sample of 150 open-ended responses. The pass rule was κ of at least 0.8 and precision and recall of at least 0.9 for the system class. All three fields passed round one, and none was cut.
| Field | Measure | Result | Rule |
|---|---|---|---|
| Recommendation status | Cohen’s κ | 0.8777 | ≥ 0.8 · passed |
| Health system named | Cohen’s κ | 0.967 | ≥ 0.8 · passed |
| Affiliation, system class554 resolved mentions | Precision / recall | 0.9753 / 0.9621 | ≥ 0.9 each · passed |
| Affiliation, independent class | Precision / recall | 0.9318 / 0.8865 | Reported beside, not in the rule |
Limitations
What limits these results, and what they do not show
Limitations
- The API, not the app. The audit queried the OpenAI API with web search and a city-level location. It does not measure the consumer ChatGPT app.
- One window, variable answers. Every answer was collected on October 4, 2026. Only 11.08% of distinct recommended names appeared in all three runs.
- A 2022 affiliation file. Of 100 clinicians and groups the rule classed as independent, web checks found 6 now system-employed according to the system’s own site. The same rule classifies supply and recommendations, so staleness touches both sides.
- Rule-side identification. The second coder agreed less with rule-resolved affiliations (system precision 0.9339, n 210) than with web-resolved ones (0.9959, n 344). These errors move recommended clinicians toward independent. The answer-named sensitivity, +13.08, is an upper bound on their effect.
- Medicare enrollment. Supply counts Medicare-enrolled clinicians, so independents who do not bill Medicare are missing, which raises the measured supply share. The bias is largest in pediatrics and psychiatry, which is why they sit outside the test.
- Web verification. 736 of 1,354 resolved tested-line recommendations were resolved by checking official sites. Leaving them out gives the one estimate whose interval reaches −12.42.
- Maternity rests on one market. Philadelphia is its only eligible market. Without maternity the estimate is +4.14 (−0.19 to 8.42).
What this study does not show
- That AI recommends system clinicians above their supply share. Two of the five intervals include zero, and the estimate without web-verified recommendations is negative.
- That AI under-recommends health systems in the lines where patients choose.
- A result for any single line. Per-line gaps are descriptive. Only the five-line mean is tested.
- Why an answer recommends one provider over another. The naming and out-of-market figures describe answers, not reasons, and cited sources are described, not shown to raise visibility.
- Bookings, revenue or system size. The study measures recommendations only.
- Other products and places. The consumer app, other AI tools, and rural or micropolitan markets are outside the scope. The mid-size and small metros are seeded draws, and the ten largest metros are all ten.
Disclosure. Halcy is a commercial service that sells search and AI-visibility work to healthcare practices, including independent practices. We designed, ran and paid for this study, and we publish the full results so readers can check every figure on this page.