How the grades work
Rubric v0, scoring rules v0.1. The full rule definitions and change history live in the public code repository.
In plain words
Every insurer selling plans on HealthCare.gov must publish its full list of doctors and therapists as a public data file and keep it current. Two federal rules stack up here, both enforced by the Centers for Medicare and Medicaid Services (CMS). 45 CFR 156.230(b) requires an up-to-date, accurate and complete provider directory. 45 CFR 156.230(c) is the machine-readable half, requiring issuers in the Federally-facilitated Exchange to publish that same information in a format HHS specifies. The monthly update cadence for these files comes from CMS's 2017 Letter to Issuers (pages 52 and 53), still operative by reference in later years, rather than from the regulation itself. We download every one of those files each month and archive them.
Then we check every entry for problems a patient would hit: phone numbers like 999999999, update dates from the year 1900, providers listed hundreds of miles from the plan's counties, and provider ID numbers the federal government has deactivated. Every problem we find becomes one row of evidence, tied to an archived copy of the insurer's file. The rows add up to a grade for each plan in each county.
Two limits. The grades are computed from files, not phone calls. A good grade means the published list is consistent and current, not that an office will answer. And the score never judges care quality or any individual clinician. It judges the insurer's published data.
The ten checks
Problems with the whole file
Dead file. The web address the insurer gave CMS returns nothing, or most of the files it points to can't be read. In practice the X grade is assigned by outcome rather than by cause. Any insurer that CMS lists as selling plans, but that produced no scorable provider records this month, is marked X for unauditable instead of scored, whether the fetch failed or the file came back unreadable. An F means the published content is bad. An X means there was nothing to check.
Browser-only access. Some insurers' files load in a web browser but refuse automated tools, which defeats the purpose of a machine-readable file. We note this on the insurer's page. It does not lower the score, because the content, once fetched, is complete.
Problems with individual entries, read straight from the file
Placeholder values. A contact field that carries filler instead of contact information. Four kinds count, and the full list matters because it is the heaviest-weighted check on the site:
- Phone. Blank, fewer than 10 digits, or the same digit repeated (999999999, 0000000000).
- ZIP. 99999 or 00000, or any value that is not 5 or 9 digits long.
- Address. Blank, or literally "null", "n/a", "na", or "unknown".
- Update date. Before 2014, or dated later than the day we fetched the file.
An entry like that is unusable exactly as published. Placeholder phones, ZIPs and addresses carry the heaviest penalty weight on the site, tied with a missing or malformed NPI. Placeholder dates carry somewhat less.
Old update dates. Each entry carries a date saying when it was last updated. We flag entries older than 180 days, and more heavily past 365. One caveat matters here. Most files stamp one shared date on every entry, so the date behaves like a file-generation timestamp. An old date is a real warning sign. A fresh date proves nothing about any single entry, and we never present a fresh file as verified.
One shared phone number and nothing else. Every phone number on the entry is a number that appears on at least 1% of the file's entries that carry a usable phone number, or 50 such entries, whichever is larger. We identify whose number it is where we can, because a provider group's own booking line is a smaller problem than a number that reaches nobody. We also check whether the federal registry lists a different direct phone for that provider.
Many addresses on one provider. One individual provider listed at more than 10 street addresses at once, or an organization at more than 25. For individuals we skip specialties that legitimately cover many sites, including radiology, pathology, anesthesiology, emergency medicine, and hospitalists. That exclusion applies to individuals only, because an organization at 25 addresses is a different claim.
Out-of-area listings. A provider attached to a plan even though none of the provider's addresses fall inside any county the plan covers, or any neighboring county. These entries add to the provider count without adding anyone a member can reach. They are excluded from county counts and scored at the plan level.
Missing accepting-patients answer. Insurers are required to say, for each individual provider they list, whether that provider accepts new patients. We flag entries that leave it blank or say unknown. Facilities and groups are allowed to omit this field, and we never count that against them.
Entries that disagree with the federal provider registry
Nearly every provider who bills insurance has a National Provider Identifier (NPI), a 10-digit number issued by the federal government and tracked in a public registry called NPPES. Providers who bill no insurance are not required to have one. CMS publishes a monthly file of deactivated NPIs, listing every NPI deactivated since 2005 and its deactivation date.
Registry status. We flag entries whose NPI is missing, fails its checksum, or appears on CMS's deactivation report while carrying no active record in the monthly registry file. NPIs are deactivated for many reasons, and CMS's own published reasons are death, disbandment, fraud, and other. The flag says one thing only. The insurer's current directory carries an identifier CMS deactivated, as of a dated report.
Specialty disagreement. The file lists a provider under a mental health specialty while the registry records only unrelated fields, such as physical therapy. One of the two databases is wrong about this provider. We do not know which one. The flag states both facts.
How checks become a grade
Every grade attaches to one plan in one county. We grade the mental health list separately from the all-provider list, and both appear side by side.
Each listed provider starts clean. Each problem adds a penalty, but penalties on one entry don't simply add up. A badly broken entry counts once, at the weight of its worst problem plus a fraction of the rest. One broken record is never counted four times.
The in-county entries carry 80% of the grade. The plan's out-of-area rate carries the other 20%, because those entries never appear in any county count and would otherwise be invisible. Inside that second block the rate is multiplied by 0.8 before it scores, so a plan whose every listing is out of area loses 16 points rather than the full 20. We would rather disclose the constant than let a reader derive it and find it undocumented.
A registry disagreement can lower a score by at most 20 points. That cap covers the two checks that rest on comparing the file against a federal registry snapshot, meaning deactivated NPIs and specialty disagreements. It does not cover a missing or malformed NPI, because those are read straight off the insurer's own file and need no registry to establish. The registry is itself self-reported and lags reality, so we hold the cap until insurers have had one monthly notification and correction round to respond.
We also recompute every score with all penalty weights set equal, as a check that our chosen weights are not driving the story. Where the two versions land in different letter grades, the plan's page shows both. The full data download carries both scores for every plan and county.
Grade bands: A is 90 to 100, B is 80 to 89, C is 70 to 79, D is 55 to 69, F is below 55, and X means unauditable. These cutoffs are our own. They are not yet tied to any official standard or benchmark. Once we see how scores are distributed across all plans, we will adjust the cutoffs and publish exactly what changed.
Small lists get special handling. A county list with 10 to 29 providers is graded with a low-sample badge. A list with fewer than 10 is not graded at all; the small count itself is published as the finding.
On the smallest graded lists the flag rate carries a wide confidence interval, so the letter grade is informative while the exact number is not. Where that interval spans more than 25 points, the data download marks the row numeric_suppressed and the score should not be quoted. We show the letter grade on those cells and never a per-cell number anywhere on this site. Be aware that plan-level and insurer-level averages do include those cells, because dropping them would bias an average toward the largest counties.
Three evidence labels
Every flag carries one of three labels, so a reader always knows how strong the evidence is.
E1: read straight off the insurer's file. Anyone holding the same file gets the same answer, with no judgment call in between. Example, a phone number of 999999999.
E2: read off the insurer's file, through a cutoff we chose. Example: we call an update date stale after 180 days. Every cutoff is published, and the table below shows how the counts move at other cutoffs.
E3: the file disagrees with a federal registry snapshot. Either database could be the stale one. These flags always carry the registry date and the note that identifiers are deactivated for many innocent reasons.
The cutoffs at other settings
When findings publish on October 26, 2026, this section will show how each count moves at looser and stricter cutoffs.
Known biases
Most files stamp one shared date on every entry. A fresh date therefore proves nothing, and freshness never earns credit.
The file format lets facilities and groups omit some fields. We never count those omissions as defects. Early in this project we nearly published a wrong number by missing that; the correction is part of the public record.
A few hosting companies publish files for many insurers at once. A bug at one host can lower scores for several insurers at the same time, so every insurer's page names its hosting company.
When one file serves several insurers, each record counts against only its own insurer. This is common: 110 of the 185 insurers we audit share at least one file with another insurer, and one file is published for 24 of them. Every provider record lists the plans it applies to, and we attribute the record to the insurer whose plan it names. A record you did not list is never counted against you, even when it sits in a file you publish. Records that name no plan we can resolve are attributed to no insurer and score nothing.
The federal registry is not ground truth. It is another self-reported database with its own lag. That is exactly why registry disagreements are capped at 20 points.
How we treat insurers fairly
- Every rule has a version number, and every published flag records which version produced it.
- Every downloaded file is archived with its web address, download time, HTTP status, content type, the server's own ETag and Last-Modified, and a cryptographic fingerprint (SHA-256). That index is published as source_manifest.csv.gz on the data page, so every fingerprint cited in an evidence row resolves to a specific address at a specific moment. We keep the archived bytes and send any file on request. We do not republish the archive in bulk, because those files carry named providers' contact details and this project identifies providers by federal ID instead.
- At least 14 days before publication, each insurer's technical contact on file with CMS receives its flag export. Every record-level finding ships in full. The one exception is out-of-area listings, which run to hundreds of millions of rows across all issuers, so those ship as per-plan totals plus a 50,000-row sample. An insurer wanting its complete out-of-area rows can ask, and we will send them.
- Disputes sent to contact@ghostnetworkwatch.org or filed through the public correction tracker are published verbatim alongside the flags they dispute. The one exception is material we will not host, meaning personal or health information, credentials, or abuse. We redact that and say in the thread that we did, rather than editing quietly.
- When a later monthly crawl no longer reproduces a flag, it will be marked resolved with the date. History is retained, not deleted.
- Crawls run monthly, aligned to the federal registry's monthly release. Every page shows its snapshot date, and pages older than 60 days say so.