Valuation accuracy
How close are our value estimates to real sales?
Every property value we show is a real comparable-sales estimate —
not a cost markup. Here's how it scores against actual recorded sale prices,
validated on a holdout the model never saw.
In plain English: we predict the price of homes that then actually sold, and publish how far off we were.
Raleigh metro (Wake, NC)4.9%
median error vs. real sale price
91% within ±20% · 1,500 holdout sales (median on the high+medium-confidence subset) · 98% coverage
Phoenix metro (Maricopa, AZ)6.6%
median error vs. real sale price
89% within ±20% · 1,500 holdout sales (median on the high+medium-confidence subset) · 98% coverage
Tampa metro (Hillsborough, FL)7.7%
median error vs. real sale price
85% within ±20% · 1,500 holdout sales (median on the high+medium-confidence subset) · 96% coverage
Pinal5.9%
median error vs. real sale price
88% within ±20% · 1,500 holdout sales (median on the high+medium-confidence subset) · 96% coverage
Omaha metro (Douglas, NE)9.2%
median error vs. real sale price
81% within ±20% · 1,500 holdout sales (median on the high+medium-confidence subset) · 94% coverage
Lincoln metro (Lancaster, NE)8.7%
median error vs. real sale price
82% within ±20% · 1,500 holdout sales (median on the high+medium-confidence subset) · 94% coverage
Omaha metro (Sarpy, NE)7.9%
median error vs. real sale price
85% within ±20% · 1,500 holdout sales (median on the high+medium-confidence subset) · 96% coverage
See what it says about YOUR property Validated on a temporal holdout: we predict each recorded 2023+ sale from comparable sales in the two years before it (never the sale itself), then compare to what it actually sold for. Only markets clearing the bar (median error ≤10%, coverage ≥85%, ≥75% within ±20%) use comp-based values; the rest fall back to a flagged assessed estimate.
Full error distribution
Not just the typical miss — the whole shape. RMSE above MAE means a few large misses pull the tail; bias is the signed mean (+ runs high, − runs low — ours runs close to zero). High+medium-confidence subset.
| Market | p10 | p25 | Median | p75 | p90 | MAE | RMSE | Bias |
|---|
| Raleigh metro (Wake, NC) | 0.9% | 2.0% | 4.9% | 10.6% | 18.6% | 13.5% | 75.2% | +6.9% |
| Phoenix metro (Maricopa, AZ) | 1.2% | 3.1% | 6.6% | 12.8% | 20.7% | 9.8% | 15.1% | +1.4% |
| Tampa metro (Hillsborough, FL) | 1.2% | 3.1% | 7.7% | 14.8% | 24.6% | 12.4% | 26.6% | +4.0% |
| Pinal | 0.9% | 2.4% | 5.9% | 12.4% | 22.7% | 11.4% | 28.8% | +2.5% |
| Omaha metro (Douglas, NE) | 1.4% | 3.9% | 9.2% | 16.7% | 32.7% | 21.9% | 69.5% | +10.9% |
| Lincoln metro (Lancaster, NE) | 1.7% | 4.1% | 8.7% | 16.3% | 27.1% | 16.6% | 50.5% | +7.6% |
| Omaha metro (Sarpy, NE) | 1.3% | 3.6% | 7.9% | 14.8% | 24.8% | 14.6% | 43.5% | +5.7% |
Is the range calibrated? (measured out-of-sample)
We don't just claim a range — we calibrate it and then measure whether it holds. The displayed value range is a conformal ~80% prediction interval: the band is set from the model's own historical error, per market and confidence tier. The table is how often that range actually contained the eventual sale price on a held-out test split the calibration never saw — so the coverage is out-of-sample, not circular. Targeting 80%; this is the calibration work most vendors never disclose.
| Market | Range contained the sale (out-of-sample) | n (test) |
|---|
| Raleigh metro (Wake, NC) | 81% | 736 |
| Phoenix metro (Maricopa, AZ) | 79% | 738 |
| Tampa metro (Hillsborough, FL) | 81% | 723 |
| Pinal | 80% | 718 |
| Omaha metro (Douglas, NE) | 80% | 704 |
| Lincoln metro (Lancaster, NE) | 80% | 702 |
| Omaha metro (Sarpy, NE) | 81% | 719 |
Rent-estimate accuracy
Rents feed the cap-rate / DSCR / hold economics, so we measure them the same way: leave-one-out vs n=1,561 assessor-recorded complexes — predict each Phoenix complex's gross potential rent from every OTHER complex's median $/sqft/yr, then compare to what the assessor actually has on the rent roll. Ground truth is the current recorded roll (not signed leases) and the error is measured in Phoenix — applying it to markets without recorded rents is a labeled assumption.
Multifamily rent (gross potential rent)12.3%
median error vs. the assessor-recorded rent roll
68.9% within ±20% · bias +7.0% (over-predicts the typical complex) · n=1,561 Phoenix complexes
SFR rent (asking-rent consistency twin)10.4%
median gap vs. asking rents of active Phoenix SFR listings — asking ≠ signed lease, and the estimator can see active listings, so this is a consistency check, NOT signed-lease accuracy
74.0% within ±20% · n=100
Reproducible: PYTHONPATH=. python3 scripts/backtest_rent_accuracy.py · as of 2026-07-14 · also in the machine-readable JSON →
Teardown-ranker accuracy
The teardown screen is RANKED by a classifier trained on permit-confirmed demolitions; this measures how well that ranking finds the real ones (leakage-safe 5-fold out-of-fold). Scope honesty: labels cover the jurisdictions with usable permit feeds (Gilbert/Mesa/Scottsdale/Tempe + a one-shot Shovels pull) — Phoenix proper publishes no usable feed, so recall vs ALL Maricopa demolitions is unknowable from this data. This grades the ranking, not any dollar value.
Teardown ranker (recall@250, Scottsdale)19%
of permit-confirmed demolitions surfaced in the top 250 ranked parcels — vs 12% for the static heuristic baseline
OOF AUC 0.8018 · 476,457 old-SFR candidates · 800 labeled demolitions
Reproducible: PYTHONPATH=. python3 scripts/teardown_classifier.py emit · as of 2026-07-13 · also in the machine-readable JSON →
Recommended-offer calibration
Every arm’s-length, single-parcel Maricopa sale since 2023 was re-underwritten with the same leakage guards as the value backtest (comps strictly before the sale, never the sale itself), and the engine’s recommended max offer was compared to the recorded price. Read it as CALIBRATION, not profit: it shows where the recommended ceiling sat relative to the market that day — realized rehab scope, exit prices, and hold outcomes are not observed here.
Recommended offer vs recorded price (median)0.93x
IQR 0.86–0.99 across 18,852 recorded Maricopa arm’s-length sales (2023+) — the served max offer at a 15% return-on-cost target sits deliberately below market
Share of real deals winnable at our offer23%
P(offer ≥ price) while holding the full return target — 17% in 2023 rising to 34% in 2026 as the market cooled
Reproducible: PYTHONPATH=. python3 scripts/measure_bid_validation.py · as of 2026-07-10 · two seeds agree to 3 decimals · also in the machine-readable JSON →
Weekly outcome flywheel
accumulating — value loop engages when sales land. Every week we freeze the exact economics we showed — leakage-safe — then grade them against real recorded sales and permits from county records, so the model sharpens on its own track record rather than on assumptions. Permits grade as they're issued (0 redevelopment permits matched on the cohort so far — live, within weeks). Value self-improvement waits on realized SALES (0 so far; the county sale feed lags first-seen by months) and engages at 100 realized sales. The immutable record is real and accumulating; no value recalibration is claimed until then.
13,410
predictions frozen at first-seen, immutably
9
counties · since 2026-06-04
0
outcomes graded vs. county reality (0 via permits)
Retrospective signal check (leakage-safe backtest — not the live cohort): across 1,401,779 Maricopa parcels, the top-scored decile was 3.04× more likely to actually be redeveloped (a demolition or new-build permit issued after the as-of date) than the base rate — scored only from time-stable structural attributes as of 2024-06-01, so the score cannot encode the outcome. The backtest is evidence the signal predicts redevelopment, while the live cohort matures.
Reproducible: python scripts/backtest_comps_avm.py (seed 1729, temporal holdout). Source output/AVM_BACKTEST.md · content hash b9a44b95ab1a · machine-readable JSON →
Honest limits: estimates use sqft, age,
use, and location (no beds/baths/condition), so a renovated and a tired
same-size home can look alike — that's why every estimate ships with a
confidence tier and a range. See the full methodology →
See the teardown-signal backtest →
— the leakage-safe retrospective study showing the redevelopment signal predicts real demolitions
(2.95× top-decile lift).