Halcy Research · Dental findings · Part 2 companion

What separates the dental practices ChatGPT cites from the ones it doesn’t

In a matched comparison of 540 dental practice websites, 33.6% of cited sites had at least one audited page of 1,000 words or more, against 19.3% of matched sites absent from the answers. The comparison did not establish an association between citation and FAQ markup or visible Q&A blocks. Credential markup was too rare to assess.

This companion to the second Dental Economics article provides the study methods and comparison tables. It uses the same June–July 2026 corpus as the dental citation dataset.

Cited sites
271
Matched uncited
269
Matched on
20 metros × 7 areas of dentistry
Audited
July 2026

The comparison

Five on-site features, compared across the two groups

Cited sites carried a long page about 1.7 times as often as matched uncited sites (33.6% against 19.3%; odds ratio 2.14 after matching). None of the other four differences reached significance.

Each row is the share of sites in each group carrying the feature. The odds ratio and p-value come from a Cochran–Mantel–Haenszel test stratified on metro × area of dentistry, so the groups are compared within matched markets before the strata are pooled.

Cited, difference held after matchingCited, no significant differenceMatched uncited
Long contentAt least one audited page of 1,000 words or more; treatment/condition page or homepage fallback
33.6%19.3%
OR 2.14 · p = .0002 · +14.2 pp
FAQPage schemaFAQPage markup carrying five or more patient-language questions with answers
8.9%6.7%
OR 1.46 · p = .33 · not significant
On-page Q&A blocksThree or more visible question-and-answer pairs or accordion panels
6.6%5.6%
OR 1.35 · p = .55 · not significant
Structured intro blockOpening text naming a credential, a numeric metric and a distinguishing phrase
3.0%0.7%
OR 4.82 · p = .08 · not significant
Dentist credential schemaPerson or Physician markup carrying a board certification and the body that issued it
0.4%0.0%
Too rare to estimate · 1 site vs 0
Any of the fiveCarries at least one of the features above
38.7%26.0%
OR 1.94 · p = .0008 · +12.7 pp

Shared scale, 0–40%. Odds ratios and p-values are Cochran–Mantel–Haenszel, stratified on metro × area of dentistry; risk differences in percentage points (95% CI for long content 6.9 to 21.6; for the composite 4.9 to 20.5). “Not significant” means the matched test did not reach p < .05.

2.14×

the odds of a long page, cited versus matched uncited. Only long content showed a significant difference after matching. This does not set a word-count target: two thirds of cited sites did not meet it either, and most sites in both groups carried none of the five features (61% of cited, 74% of uncited).

FAQPage schema, on-page Q&A and credential markup were carried at rates that do not separate the groups after matching. Neither group maintained a page-per-procedure structure: the median number of treatment or condition pages the crawl found was zero on both sides.

FAQPage schema appeared on 8.9% of cited sites versus 6.7% of uncited sites: 24 cited adopters against 18. The audit measured presence, not quality, and did not establish an association with citation.

Page length and schema types

Cited sites had longer pages and more schema types

Each row shows the median for each group. Schema.org types are labels in a website’s code for things such as a dentist or a business.

Measure, median per siteCitedUncitedTest
Words on the longest fetched pageAll page text with scripts and styles removed, navigation and footer included. A looser count than the 1,000-word classifier, which reads body content only, so the two are not on one scale1,184905Mann–Whitney U · p < .0001
Distinct schema.org typesDistinct types used in the site’s schema.org markup42Mann–Whitney U · p = .0001
Treatment or condition pages foundPages the crawl identified as describing a specific service or condition00Mann–Whitney U · p = .52 · not significant

Cited n = 271, matched uncited n = 269. The word count is the longest of the pages fetched for each site, so it says nothing about the median page.

Within the 271 cited sites, neither word count nor feature count tracked how often a site was cited (Spearman |ρ| ≤ 0.03), but that sample is the 300 most-cited domains, so the range is restricted by construction and the test says little either way. Whether more content earns more citations once a practice is already cited is a question this study does not answer.

Website availability

About one in six matched uncited practices had no independent website

Among the 347 uncited practices drawn for the comparison, 15.8% had no independent website of their own.

  • 84.1% · Own website292 practices with an independent site; 269 of them returned a usable audit and form the comparison group
  • 4.6% · Third-party page only16 practices: 12 on a DSO or group site, 4 on an association listing
  • 11.2% · No findable website39 practices with no site Google Places could resolve

Every cited practice has its own website by construction: the cited group was drawn from practice-owned domains. The comparison group was resolved through Google Places, and a practice whose only presence was a group site or an association profile entered the count as that page.

Check your own page

Check one treatment page in ten minutes

Open the treatment page patients ask about most and review it using these four checks. The first three are patient-communication checks, not signals this study tested; the fourth is one live query, not the corpus.

  1. Does it answer the question patients ask before booking, in ordinary page text rather than an accordion, PDF or image?
  2. How many words does it carry? The study’s threshold was 1,000 on at least one page; two thirds of cited sites did not meet it.
  3. Is there a named clinician who reviewed it, and a date that reflects a real review?
  4. With web search on, ask ChatGPT “who does [treatment] in [city]?” Does this page appear among the sources?

Prefer a report? Run the free visibility scan and we report the on-site signals from this study as we find them on your pages. No account needed.

Common questions

Questions about this comparison

Want to discuss the findings?

Talk to us for an audit of your website and visibility on Google and ChatGPT. We’ll show you where competitors are ahead and what to work on first.

Does this mean every page needs 1,000 words?

No. One third of cited sites met the threshold on at least one page and one fifth of uncited sites did; two thirds of cited practices did not meet it. The study did not test the effect of adding words.

Is FAQ schema a waste of time?

The study cannot say that. FAQPage markup was slightly more common on cited sites (8.9% versus 6.7%) and the difference was not significant with 24 cited adopters. The audit checked whether the markup was present, not whether it accurately described the page.

Why match on market and specialty but not practice size or age?

Metro and area of dentistry are observable for every practice in NPPES; size and age are not. Matching makes the groups comparable by metro and area of dentistry. Differences in practice size or resources could still explain the association.

What does “uncited” mean here?

Absent from the 13,152 citation events in our corpus: 1,405 dental questions across 20 metros and seven areas of dentistry. A practice ChatGPT cites on questions we did not ask, or in a metro we did not sample, would count as uncited.

Does more content earn more citations once a practice is already cited?

Unknown. Within the cited group neither word count nor feature count tracked citation frequency, but the group is the 300 most-cited domains, so the range is restricted and the test is uninformative. The finding is about the difference between absent and cited practices, not about ranking among the cited.

Methods and caveats

Study methods and limitations

Both groups were matched by market and area of dentistry and audited with the same crawler and classifiers.

Cited groupn = 271

Source
The 300 practice-owned domains cited most often across the 13,152 citation events in the dental corpus
Selection
Independent practices only; hospital, DSO and directory domains excluded
Audited
271 of 300 returned a usable crawl; the rest blocked or failed
Website
Live by construction

Comparison groupn = 269

Source
Dental organizations in the National Plan and Provider Enumeration System (NPPES), entity type 2
Selection
Matched to the cited group on metro × area of dentistry; any domain appearing anywhere in the citation corpus excluded
Audited
347 drawn; 292 with an independent website, found through Google Places; 269 returned a usable crawl
Website
Google Places lookup; see website availability

The audit read treatment and condition pages where the crawl found them and fell back to the homepage where it did not, which was most sites in both groups. Each feature was tested with a Cochran–Mantel–Haenszel test stratified on metro × area of dentistry, pooling only strata with sites in both groups: 60 informative strata for long content, 62 for the composite, 9 for the structured intro block and 1 for the credential feature.

Design

Cited practices were identified from the 13,152 citation events in the June–July 2026 dental corpus: 1,405 questions, GPT-5.5 with web search through the OpenAI API, 20 US metros, seven areas of dentistry. The 300 most-cited practice-owned domains were crawled and 271 returned an audit. Controls were drawn from NPPES, matched on metro × area of dentistry, resolved to websites through Google Places, and excluded if their domain appeared anywhere in the corpus. Both groups were audited with the same classifiers: a long page is 1,000 or more words of visible text after navigation, header, footer and scripts are stripped, read from treatment and condition pages where the crawl found them and from the homepage otherwise. Effect sizes are Cochran–Mantel–Haenszel odds ratios stratified on the matching variables; continuous measures use Mann–Whitney U.

Limitations

The comparison is observational. Matching does not control for practice size, age or spend. The cited group’s websites are live by construction while the comparison group’s are best resolutions, a conservative asymmetry for structure comparisons. The crawl was shallow, and for most sites in both groups the long-content test fell back to the homepage. “Uncited” means absent from this corpus. The results cover one platform in the US during the study period; they do not show changes over time. The association does not establish that adding content causes citations or that longer pages receive more citations once a practice is cited.

Disclosure. Halcy is a commercial service; the practices we work with pay for these deployments. We ran this corpus for our own product development and publish it openly so administrators can audit their own AI presence regardless of any vendor relationship.