What retrospective closed-loan testing reveals about AVM performance that prospective holdout sets miss.
When an AVM vendor publishes accuracy statistics, they typically come from one of two methodologies: prospective holdout testing, where a portion of transactions is withheld from model training and the model's predictions are compared to actual outcomes; or retrospective closed-loan testing, where the model is run on past transactions for which the final sale price is already known. The two approaches measure different things, and for lender use cases, the retrospective approach is substantially more informative.
What Prospective Testing Measures
A prospective holdout study builds the model on transactions through some cutoff date, holds out a validation set, and reports accuracy metrics on the held-out set. If the model was built on 2020-2023 transactions and validated on a 2024 holdout, the accuracy statistics reflect how well the model generalizes to the distribution of properties in the holdout set -- which is the same distribution as the training data, drawn from the same time window.
The limitation is that prospective holdout accuracy is optimized by the model selection process. The modeler runs multiple candidate models, chooses the one with the best validation performance, and reports that number. There is no requirement that the reported accuracy be achieved on future data in a different market condition. When market conditions shift -- a rate environment change, a geographic correction, a supply shock -- the model's actual performance on new transactions may diverge substantially from the reported holdout accuracy.
What Retrospective Testing Measures
A retrospective closed-loan study takes the model as deployed and runs it on a historical dataset of completed transactions where the actual sale price is already known. The study is retrospective in the sense that the ground truth is already available, which eliminates survivorship bias from the analysis: every transaction in the study completed, so the accuracy is measured on the full population rather than only on the transactions where the model was confident enough to produce an estimate.
For lenders, the relevant accuracy question is not "how well does this model perform on the average property in the training distribution?" It is "how well does this model perform on the specific types of properties that appear in our loan pipeline, in our geographies, at the transaction sizes we underwrite?" A retrospective study on a lender's own closed-loan portfolio answers that question directly.
Designing a Meaningful Retrospective Study
A useful retrospective accuracy study for AVM evaluation has several design requirements. First, the study transactions should match the use case: if the lender primarily originates purchase loans in the $300,000-$600,000 range in specific counties, the study transactions should reflect that distribution rather than the vendor's proprietary holdout set which may skew toward different property types or price points.
Second, the study should use the correct accuracy metric for the use case. Median absolute error (MAE) as a percentage of sale price is the right metric for individual-transaction lending decisions, because it is interpretable in terms of the risk the lender is taking on any given loan. Root mean squared error (RMSE) penalizes large errors more heavily, which is useful for detecting catastrophic misses but less useful for characterizing typical performance. A vendor who reports RMSE but not MAE is potentially obscuring systematic moderate errors behind the metric structure.
Third, the study should report the full error distribution, not just the central tendency. Knowing that median error is 4.2% is useful; knowing that 91% of estimates fall within 7% of final sale price is equally useful and tells the lender something about tail risk. A small percentage of estimates with large errors may be acceptable for some use cases and unacceptable for others.
Plotgleam's Accuracy Methodology
Plotgleam's published accuracy figures come from a retrospective closed-loan study: we ran the Plotgleam valuation engine on 830 closed loans from Q3 and Q4 of 2024, across seven geographies in Colorado, where the final sale price was already known. We compared the Plotgleam estimate to the final sale price for each transaction and computed the full error distribution.
The median absolute error across the study was 4.2% of final sale price. 68% of estimates fell within 3% of final sale price. 91% fell within 7%. The full methodology documentation, including the property type distribution of the study transactions, the geographic breakdown, and the treatment of distressed sales and non-arms-length transactions, is available to prospective lender partners on request.
We publish these numbers with the study design disclosed because we believe lenders deserve to evaluate the methodology before relying on the output. An accuracy claim without a disclosed study design is a marketing assertion, not an analytical finding. The right response to a vendor's accuracy claim is always to ask: what transactions were in the study, what methodology was used, and can I run an equivalent study on my own loan population?
Ongoing Accuracy Monitoring
Retrospective testing is not a one-time exercise. A model that performed well in 2024 may perform differently in 2026 if market conditions have shifted, if the geographic coverage has expanded into less familiar markets, or if the underlying data sources have changed. Accuracy monitoring needs to be an ongoing process, not a baseline measurement.
Lenders who integrate an AVM tool into their workflow should build in a regular audit cadence: sample a defined percentage of closed loans, run the AVM model against the final sale price, and track whether accuracy metrics remain within the bounds that informed the original adoption decision. If they do not, the workflow needs to be adjusted. This is not unique to AVM tools; it is the standard ongoing monitoring obligation that applies to any model used in credit decision-making.
Interpreting Accuracy Results by Property Segment
Aggregate accuracy statistics mask important variation by property type, price tier, and geography. A median absolute error of 4.2% across 830 transactions might conceal a 2.8% MAE on suburban single-family properties and a 7.1% MAE on mountain resort properties with thin comparable sales. For a lender whose portfolio is concentrated in the mountain resort corridor, the aggregate number is misleading about the actual risk exposure.
When reviewing an accuracy study -- whether the vendor's published numbers or a lender's own retrospective audit -- request the segment-level breakdown in addition to the aggregate. The segments that matter are the ones that match the lender's own origination mix: the property types, price tiers, and geographies where the tool will actually be used. Accuracy outside those segments is informative background; accuracy within them is the decision-relevant number.
Plotgleam segments accuracy reporting by geography and property type in the methodology documentation available to prospective lender partners. The goal is to provide lenders with the information they need to assess fit with their specific loan population rather than requiring them to infer from headline numbers. A tool that is the right fit for a community bank originating primarily suburban purchase loans in the Denver metro may or may not be the right fit for a credit union with significant mountain-community origination volume. Honest accuracy segmentation makes that determination possible before the tool is deployed rather than after.
Request a demo and run a live report on a property from your own market during the walkthrough.