All projects
Data & ML pipelines
An automated valuation model agents were willing to quote from
A valuation model rebuilt around the pipeline feeding it, with freshness and drift monitoring that tells the team when a number has stopped being trustworthy.
- Client
- European property marketplace
- Industry
- Real estate
- Service
- Data & ML pipelines
- Year
- 2026
The challenge
Listing data arrived from eleven sources in eleven shapes, was cleaned by hand in spreadsheets, and reached the model weeks late. Agents had quietly stopped quoting the automated estimate because it was wrong often enough to be embarrassing, which made the whole feature decorative.
What we did
- Replaced the manual clean-up with orchestrated ingestion per source, each with its own schema contract and quarantine for rows that fail it
- Modelled the warehouse in dbt so every figure has tested lineage back to the source row that produced it
- Split training and inference into separate pipelines with versioned datasets, so a model can be reproduced months later
- Added freshness, drift and prediction-interval monitoring, and taught the product to withhold an estimate rather than show a bad one
- Published confidence alongside every valuation so agents could see when to trust it
Outcome
Weeks to under an hour
Data latency
Any past valuation, re-derivable
Reproducibility
Back in the quoting flow
Agent adoption
- Python
- Airflow
- dbt
- Snowflake
- MLflow
- scikit-learn
- Terraform