Research accountability

What would our rules have predicted—and what happened?

We can inspect outcomes ending in 2025 today. We also keep a separate, unchanged forecast for 2026–2027 so it can be tested against future observations.

A separate market-conditions experiment now tests next-quarter contract time, with adjustable models, historical misses and frozen October–December 2026 predictions. The annual appreciation record below retains its original rules and dates.

Explore the models and your own weighted scenario →

The year ending September 2025

Apply the same rules at a July 31, 2024 feature cutoff; measure home-value growth from September 30, 2024 to September 30, 2025. This preserves the original 12-month horizon and reporting gap. It is not a January–December 2025 test.

Result: the selected county baseline had an average absolute error of 1.72 percentage points. It predicted no city advantage. Some cities grew faster than the county and others grew more slowly.

For example, Burlingame exceeded the county by 4.41 percentage points, while Daly City trailed it by 3.24. The baseline missed both differences.

City appreciation minus county appreciation, in percentage points. Positive means faster growth than the county, not necessarily rising prices.
CitySelected predictionObserved differenceAbsolute error
Burlingame+0.00+4.414.41
Daly City+0.00-3.243.24
Millbrae+0.00+2.302.30
Pacifica+0.00-2.292.29
Redwood City+0.00+0.690.69
San Bruno+0.00+0.610.61
San Mateo+0.00+0.010.01
South San Francisco+0.00-0.240.24

This is a retrospective reconstruction using revised historical data, calculated in 2026. These were not predictions issued in 2024. This year was already part of the original six-year holdout; surfacing it does not create another independent test.

Training used only outcomes ending by the feature cutoff; the latest training outcome was 2023-09-30. Historical publication dates are unknown, so this does not prove that every input was available then.

Every model, including the misses

The selection rule was fixed before the original experiment. No candidate earned selection on the earlier validation years. A better result in this single later year cannot change that decision.

Same eight cities and September 2024–September 2025 window. Lower average error is better. This is one shared regional episode.
RuleRoleAverage absolute error (pp)Direction calls correct
Assessment-growth modelUnselected candidate1.7135 / 8
Combined modelUnselected candidate1.6186 / 8
Match the countySelected baseline1.722No direction predicted
Historical averageOther baseline1.7205 / 8
Momentum modelUnselected candidate1.7385 / 8
Repeat recent relative growthOther baseline1.6156 / 8
Relative-value modelUnselected candidate1.6535 / 8

Across the full holdout, the assessment-growth candidate improved average error by only 2.9% versus the county baseline; it did not clear the original 10% requirement and was not selected. The broader employment, amenities and housing-delivery thesis remains unproven. These annual models use price and assessment features; adding jobs charts did not turn them into a test of that full mechanism.

Reproduction and source record

The independent replay matches all 56 retained city/model predictions. The original protocol, model selection, historical results and forward forecast are preserved.

Download all predictions, errors, timing and source references (JSON). Read the hypothesis and limits.

Original historical-prediction file SHA-256: cdc374651219c1300949d393257939dec375b4a4100eecb504ff5e5fa986832a. Review prepared 2026-09-15.

The separate future test

Issued September 15, 2026 using July 2026 inputs. Outcome window: September 30, 2026–September 30, 2027. All 56 city/model predictions remain frozen; the selected rule still predicts zero excess appreciation versus the county.

Outcome pending. We will score the original predictions once both endpoint cohorts are captured and the full horizon has ended.

Evidence capture

A daily capture checks the official Zillow city and county files. Each changed source is retained privately with its retrieval time, original bytes and hash. We preserve the first complete eight-city-plus-county cohort our process observes for each forecast endpoint. Later revisions remain separate.

Last successful check (UTC)
2026-09-16T03:31:33.890Z
Latest month shared by all nine geographies
2026-08-31
Last attempt (UTC)
2026-09-16T03:31:13.348Z
  • 2026-09-30: Waiting for a complete observed cohort.
  • 2027-09-30: Waiting for a complete observed cohort.

“First observed” is what our capture process can establish. It is not a claim that we obtained the publisher’s first release. HTTP modification dates are retained as metadata, not treated as publication dates. Failed downloads or changed source structure do not update accepted evidence. Capturing data does not update the city comparison tables or change model selection.

Download capture status and any completed evaluation (JSON).