What would our rules have predicted—and what happened?
We can inspect outcomes ending in 2025 today. We also keep a separate, unchanged forecast for 2026–2027 so it can be tested against future observations.
A separate market-conditions experiment now tests next-quarter contract time, with adjustable models, historical misses and frozen October–December 2026 predictions. The annual appreciation record below retains its original rules and dates.
Explore the models and your own weighted scenario →
The year ending September 2025
Apply the same rules at a July 31, 2024 feature cutoff; measure home-value growth from September 30, 2024 to September 30, 2025. This preserves the original 12-month horizon and reporting gap. It is not a January–December 2025 test.
Result: the selected county baseline had an average absolute error of 1.72 percentage points. It predicted no city advantage. Some cities grew faster than the county and others grew more slowly.
For example, Burlingame exceeded the county by 4.41 percentage points, while Daly City trailed it by 3.24. The baseline missed both differences.
| City | Selected prediction | Observed difference | Absolute error |
|---|---|---|---|
| Burlingame | +0.00 | +4.41 | 4.41 |
| Daly City | +0.00 | -3.24 | 3.24 |
| Millbrae | +0.00 | +2.30 | 2.30 |
| Pacifica | +0.00 | -2.29 | 2.29 |
| Redwood City | +0.00 | +0.69 | 0.69 |
| San Bruno | +0.00 | +0.61 | 0.61 |
| San Mateo | +0.00 | +0.01 | 0.01 |
| South San Francisco | +0.00 | -0.24 | 0.24 |
This is a retrospective reconstruction using revised historical data, calculated in 2026. These were not predictions issued in 2024. This year was already part of the original six-year holdout; surfacing it does not create another independent test.
Training used only outcomes ending by the feature cutoff; the latest training outcome was 2023-09-30. Historical publication dates are unknown, so this does not prove that every input was available then.
Every model, including the misses
The selection rule was fixed before the original experiment. No candidate earned selection on the earlier validation years. A better result in this single later year cannot change that decision.
| Rule | Role | Average absolute error (pp) | Direction calls correct |
|---|---|---|---|
| Assessment-growth model | Unselected candidate | 1.713 | 5 / 8 |
| Combined model | Unselected candidate | 1.618 | 6 / 8 |
| Match the county | Selected baseline | 1.722 | No direction predicted |
| Historical average | Other baseline | 1.720 | 5 / 8 |
| Momentum model | Unselected candidate | 1.738 | 5 / 8 |
| Repeat recent relative growth | Other baseline | 1.615 | 6 / 8 |
| Relative-value model | Unselected candidate | 1.653 | 5 / 8 |
Across the full holdout, the assessment-growth candidate improved average error by only 2.9% versus the county baseline; it did not clear the original 10% requirement and was not selected. The broader employment, amenities and housing-delivery thesis remains unproven. These annual models use price and assessment features; adding jobs charts did not turn them into a test of that full mechanism.
Reproduction and source record
The independent replay matches all 56 retained city/model predictions. The original protocol, model selection, historical results and forward forecast are preserved.
Download all predictions, errors, timing and source references (JSON). Read the hypothesis and limits.
Original historical-prediction file SHA-256: cdc374651219c1300949d393257939dec375b4a4100eecb504ff5e5fa986832a. Review prepared 2026-09-15.
The separate future test
Issued September 15, 2026 using July 2026 inputs. Outcome window: September 30, 2026–September 30, 2027. All 56 city/model predictions remain frozen; the selected rule still predicts zero excess appreciation versus the county.
Outcome pending. We will score the original predictions once both endpoint cohorts are captured and the full horizon has ended.
Evidence capture
A daily capture checks the official Zillow city and county files. Each changed source is retained privately with its retrieval time, original bytes and hash. We preserve the first complete eight-city-plus-county cohort our process observes for each forecast endpoint. Later revisions remain separate.
- Last successful check (UTC)
- 2026-09-16T03:31:33.890Z
- Latest month shared by all nine geographies
- 2026-08-31
- Last attempt (UTC)
- 2026-09-16T03:31:13.348Z
- 2026-09-30: Waiting for a complete observed cohort.
- 2027-09-30: Waiting for a complete observed cohort.
“First observed” is what our capture process can establish. It is not a claim that we obtained the publisher’s first release. HTTP modification dates are retained as metadata, not treated as publication dates. Failed downloads or changed source structure do not update accepted evidence. Capturing data does not update the city comparison tables or change model selection.
Download capture status and any completed evaluation (JSON).