How accurate is the label?
Every score on this site is inferred from public data. This page measures how often that inference matches what the assessing authority recorded for the same building — the only check that asks whether the output describes the world rather than whether the code does what it says.
How to read it. Baseline is what the label infers everywhere: a modelled structure record plus the census tract's year-built distribution. With assessor adds an observed record from the assessing authority where one resolves. The number that matters is the last table in each section — a year-built error that moves no letter is not a defect anyone can see; one that crosses a code-era boundary is.
Measured so far: Cook County, Illinois, Washington, DC, Washington, DC — condominiums. Each adapter is measured against its own assessor, and the sections are not comparable to each other — different housing stock, different record-keeping, different sample.
Cook County, Illinois
Method. 220 addresses sampled across Cook County Assessor (Open Data) (assessment year 2026, fetched 2026-08-25); a further 5 sampled addresses could not be geocoded or scored and are excluded from every rate below. Scope: all residential improvement records. Each is scored from the address alone, with no construction details supplied, and compared against that jurisdiction's own assessor record.
Field accuracy
| Field | Coverage baseline | Coverage w/ assessor |
Exact baseline | Exact w/ assessor |
Median error baseline | Median error w/ assessor | Truth rows |
|---|---|---|---|---|---|---|---|
| year_built | 100.0% | 100.0% | 1.4% | 74.9% | 12.0 | 0.0 | 215 |
| sqft | 100.0% | 100.0% | 24.0% | 78.8% | 234.5 | 0.0 | 104 |
| stories | 71.2% | 93.5% | 62.0% | 93.1% | 0.0 | 0.0 | 170 |
| construction | 100.0% | 100.0% | 45.8% | 85.5% | — | — | 214 |
| foundation | 100.0% | 100.0% | 44.2% | 86.0% | — | — | 215 |
| condition | 100.0% | 100.0% | 98.6% | 100.0% | — | — | 215 |
Year built, by tolerance. The single field the rest of the construction profile leans on hardest, so the near-misses are worth seeing rather than collapsing into one median: within ±5 years 27.4% of the time at baseline and 83.7% with the assessor; within ±10 years 45.6% and 86.5%.
Does the reader see a different grade?
| Dimension | Baseline | With assessor | n |
|---|---|---|---|
| durability | 36.3% | 9.3% | 215 |
| energy | 28.4% | 5.6% | 215 |
| resilience | 0.9% | 0.9% | 215 |
| environmental | 23.3% | 7.0% | 215 |
| building_axis | 29.8% | 7.9% | 215 |
Assessor lookups resolved for 73.2% of the sample · benchmark digest 88604c464be53fb6.
Washington, DC
Method. 218 addresses sampled across DC Office of Tax and Revenue (Open Data) (assessment year current, fetched 2026-08-25). Scope: non-condominium homes only (condos are ~36% of DC's CAMA stock). Drawn from 220 assessor rows; 2 were not in the parcel layer. Each is scored from the address alone, with no construction details supplied, and compared against that jurisdiction's own assessor record.
Field accuracy
| Field | Coverage baseline | Coverage w/ assessor |
Exact baseline | Exact w/ assessor |
Median error baseline | Median error w/ assessor | Truth rows |
|---|---|---|---|---|---|---|---|
| year_built | 100.0% | 100.0% | 2.8% | 91.7% | 18.5 | 0.0 | 218 |
| sqft | 100.0% | 100.0% | 10.4% | 92.5% | 298.0 | 0.0 | 67 |
| stories | 66.7% | 95.4% | 66.2% | 100.0% | 0.0 | 0.0 | 195 |
| construction | 100.0% | 100.0% | 50.2% | 93.8% | — | — | 209 |
| foundation | — | — | — | — | — | — | 0 |
| condition | 100.0% | 100.0% | 40.4% | 95.0% | — | — | 218 |
Year built, by tolerance. The single field the rest of the construction profile leans on hardest, so the near-misses are worth seeing rather than collapsing into one median: within ±5 years 12.4% of the time at baseline and 92.7% with the assessor; within ±10 years 25.2% and 94.0%.
Does the reader see a different grade?
| Dimension | Baseline | With assessor | n |
|---|---|---|---|
| durability | 50.5% | 4.6% | 218 |
| energy | 32.6% | 4.6% | 218 |
| resilience | 8.3% | 0.9% | 218 |
| environmental | 23.4% | 2.3% | 218 |
| building_axis | 60.6% | 5.5% | 218 |
Assessor lookups resolved for 91.3% of the sample · benchmark digest 37c36ab1949c3db2.
Washington, DC — condominiums
Method. 212 addresses sampled across DC Office of Tax and Revenue (Open Data) (assessment year current, fetched 2026-08-25); a further 1 sampled address could not be geocoded or scored and is excluded from every rate below. Scope: condominium units only (61,329 of DC's 170,602 CAMA records). Drawn from 220 assessor rows; 8 had no active unit record to place them. Each is scored from the address alone, with no construction details supplied, and compared against that jurisdiction's own assessor record.
Field accuracy
| Field | Coverage baseline | Coverage w/ assessor |
Exact baseline | Exact w/ assessor |
Median error baseline | Median error w/ assessor | Truth rows |
|---|---|---|---|---|---|---|---|
| year_built | 100.0% | 100.0% | 1.9% | 97.6% | 22.0 | 0.0 | 211 |
| sqft | 100.0% | 100.0% | 0.0% | 10.4% | 473.2 | 293.9 | 211 |
| stories | — | — | — | — | — | — | 0 |
| construction | — | — | — | — | — | — | 0 |
| foundation | — | — | — | — | — | — | 0 |
| condition | — | — | — | — | — | — | 0 |
Year built, by tolerance. The single field the rest of the construction profile leans on hardest, so the near-misses are worth seeing rather than collapsing into one median: within ±5 years 17.1% of the time at baseline and 97.6% with the assessor; within ±10 years 28.9% and 97.6%.
Does the reader see a different grade?
| Dimension | Baseline | With assessor | n |
|---|---|---|---|
| durability | 39.3% | 0.9% | 211 |
| energy | 45.5% | 16.6% | 211 |
| resilience | 14.2% | 0.0% | 211 |
| environmental | 44.5% | 23.2% | 211 |
| building_axis | 64.5% | 23.7% | 211 |
Assessor lookups resolved for 96.7% of the sample · benchmark digest a9e8366b11ebce95.
What this does and does not establish
- These are the jurisdictions with an adapter, not a national sample. Each figure describes one place's housing stock and record-keeping. Nothing here supports a claim about anywhere else, and the two sections should not be averaged into one.
- Washington, DC excludes condominiums, which are about 36% of its assessor's residential stock (61,329 condo records against 109,273 others). DC keeps them in a separate table keyed by unit, and a unit-level identifier does not appear in the parcel geometry at all — so no coordinate can pick one unit out of a building. The DC figures therefore describe non-condo homes, and the adapter returns nothing for a condo rather than guessing.
- The assessor's own record is treated as truth. It can be stale or wrong; it is the best available reference, not a survey.
- Addresses where the assessor lookup does not resolve fall back to the baseline, so the “with assessor” column includes them. It is the end-to-end number a visitor would experience, not the adapter's accuracy on the rows it answers.
- The benchmarks are fetched on demand and not committed: neither source grants an explicit right to redistribute a dataset. Re-running months later samples a refreshed roll, so each section's digest and date are recorded to make that visible.
- The categorical fields are a weaker test than the numeric ones. Wall material, foundation and condition are translated out of the assessor's vocabulary into the label's by the same table on both sides of the comparison, so their “exact” rates in the assessor column largely measure whether the right parcel was found — not whether the translation is right. One entry is knowingly lossy in each source: Cook's single Masonry category and DC's Brick/Stone are both read as brick, the label's brick/block/stone distinction being finer than either. Year built and floor area carry no such circularity; they are numbers, compared as numbers.
- Rows are not graded on every field. The Truth rows column is each field's own denominator, and it is not always the full sample. Floor area is the clearest case: an assessor records the whole building's area while the label's figure is per dwelling unit, so on a multi-unit parcel the two are different quantities and the row is excluded rather than scored as a miss. A rate is over the rows in that column, not over every address sampled.
- Condition is close to a constant. Most sampled parcels carry one dominant grade in each source, so a high agreement rate on that row reflects the distribution of the source far more than the label's skill, and should not be read as one.
- DC records no basement type, so foundation is never observed there and its row stays at the baseline. That is a gap in the source, not a failure of the lookup.
- The reference profile is not wholly observed. Where the assessor records nothing for a field, the truth arm falls back to the same modelled inputs the other two arms use, so a grade attributed to “true attributes” is built from the assessor's facts plus those fallbacks. It is the best available reference, and on the rows with an incomplete record it understates the distance between the arms rather than overstating it.
Informational purposes only. This label is a modeled estimate built from public data. It is not an inspection, appraisal, survey, or insurance quote, and it is not legal, financial, insurance, engineering, or real estate advice. It describes what the model expects of a home like this one at this location; it cannot tell you the condition, safety, value, or insurability of any particular property. Verify anything you would act on with a qualified professional. Provided as is, without warranty.
Generated
2026-08-25. Regenerate with
python scripts/measure_accuracy.py --jurisdiction <name>; each
run replaces only its own section.