Methodology
Eight steps, each one auditable.
Find, enrich, understand, score — and at every step, one commitment: nothing is narrated that wasn't returned by a source or read off a page we can link to. Everything below exists to make that commitment mechanical rather than a promise.
- 01
Geography that doesn't dead-end
Name a city, town, ZIP, county, metro, or state and it resolves to state FIPS, every county the place touches, place code, ZCTA, tract, block group, and CBSA before any dataset is queried. Resolution falls back across the Census Geocoder, TIGERweb, the FCC Area API, and place-name matching, and results are cached. Each adapter declares the grain it speaks; the fetch layer translates your place into that grain — rolling ZIPs up to a county, fanning a metro out to member counties. Where a query had to be widened, the panel says so.
- 02
Source selection with a paper trail
Every catalog entry is scored for relevance against your stated goal, with a one-line justification. A fetch budget caps what actually executes, with per-publisher concurrency limits. Sources that were ranked but not pulled stay visible as 'considered, not pulled' with the reason — and one click pulls them anyway. Declined sources carry a reason just as selected ones do.
- 03
The Explain panel
Every source segment scores 0–100, and the score opens rather than asserts. It shows what was used — rows, geography, vintage, live versus retained, the exact query — and exactly why points were deducted.
- Widened geography
- Coarser grain than the decision unit
- Older vintage
- Suppressed values
- Restated units
- Stale capture
- Thin row count
- Upstream failure
- 04
The accuracy guardrail
The narration step receives the fetched rows and nothing else — no access to the original question's assumptions, no tool access — so there is no channel through which an invented figure could enter. No coverage is stated as a gap, by name: 'CMS Part D has no rows for Ada County at this grain.' Suppressed small-cell values are excluded from totals, never zeroed. Every number carries source, geography level, and as-of date.
- 05
One metric layer, reconciled
HCRIS, MA penetration, physician utilization, and Medicaid SDUD are restated into the same canonical KPIs — spend, service volume, capacity, population — with ratios recomputed from summed numerators and denominators rather than averaged. Units are switchable (dollars, thousands, millions; percent, per 1,000, per 100,000) and each metric is badged 'recomputed on switch' or 'rescaled only.' Where two sources can be compared but not added, the dashboard says so instead of summing them.
- 06
Enrichment reads pages, it does not remember them
Once the universe exists, each target is enriched from its own public website: services, levels of care, locations, phone numbers, payers listed, hiring activity, and named decision-makers with titles. Every extracted field is stored as a value plus the exact page URL and the fetch date, or it is stored as null. Nothing is filled from a model's prior knowledge, and a fact that isn't on the page is shown as 'not published' — absence of evidence is displayed as absence of evidence, never as a negative fact about the organisation.
- 07
Gaps and changes, both evidenced on two sides
A gap statement names the demand signal and the page proving the capability was not mentioned — 'large depression population, no interventional service identified on their site.' Change detection re-reads a target and diffs versions: services added or removed, locations opened, decision-makers replaced, registry moves in NPPES. Each change carries the before, the after, and the URL that shows it.
- 08
The fit score is a visible sum
Each target gets a 0–100 score composed of published volume and utilization, payer mix viability, opportunity gap, web-intelligence completeness, whether a named contact was found, and how recently the target changed. The breakdown is shown on the row and in the export, so a thin enrichment result visibly lowers the score rather than silently padding it. A slot is reserved for a commercial claims component; nothing today implies we have one.
- 09
Siting decisions score counties, not names
The Site Selector runs the same discipline one level up. A nationwide catalog of demand, supply, payer and demographic signals is weighted for the facility type you described, then scored in two passes across every county that clears the population floor. Each signal becomes a percentile among counties that actually have data — a county with no data contributes nothing rather than being scored as zero — and the weighted contribution of each signal is printed. Medicare rate context travels with every county: paid per beneficiary, standardized per beneficiary, the geographic rate index (paid ÷ standardized × 100), inpatient payment per user and MA penetration, all arithmetic on published CMS county rows. Rates are always shown; they only carry scoring weight when you switch that on. Commercial and Medicaid rates are not published at county grain, so the results panel names that gap and states what it does to rank confidence instead of filling it in.
- 10
A market is re-read, not remembered
Saved markets are re-checked on a weekly cadence and diffed against the previous capture: providers added or gone, service lines opened, sites moved, decision-makers replaced. Watched targets and competitor alerts surface the same diffs for named organisations. Every entry in the digest carries the before value, the after value, and the page that proves it — a change with no evidence URL is not reported.
Every run has a summary page tracing each step, timing, and row count.
Run oneBuild the list for one county and see whether the names are right.
Ask in plain language. RepVector finds the universe, enriches each target from its own website, scores it 0–100, and hands you a ranked call list with phones, addresses and named contacts.
3 free runs — no card required.
