Illustrative example. This scenario reflects the kind of project we deliver. It does not describe a named client, and details have been generalised.
Challenge
The company builds analytics tools for property investors and advisers in Singapore. Its product depended on a view of asking rents and resale prices across the market, drawn from public listings on several property portals. An earlier in-house collection had become difficult to maintain. Listings changed layout frequently, the same unit often appeared on several portals under different descriptions, and fields such as floor area and tenure were recorded inconsistently.
The company also wanted to reduce the personal data it held. The earlier collection had captured agent names and phone numbers by default, although the product never used them.
Approach
We scoped the project around the analytics the product needed: asking price, rent, floor area, price per square foot, property type, tenure, district, development name, number of bedrooms and listing dates. Agent contact details were excluded at collection, so they never entered the pipeline. Agency names were retained only where needed to support de-duplication, then dropped from the delivered dataset.
De-duplication was the core technical problem. We combined development name, block and unit attributes, floor area, bedroom count and asking price within a tolerance to identify the same property listed on different portals. Development names were normalised against a reference list, since the same condominium can be spelt several ways. Each cluster of duplicates became a single record with links to its source listings.
Floor areas were converted to both square feet and square metres, and prices normalised to price per square foot. Listings were mapped to planning areas and postal districts to support the client's geographic analysis. Source terms and request rates were reviewed for each portal before collection began.
What we delivered
- A daily feed of new, changed and withdrawn listings, delivered in Parquet to the client's cloud storage.
- A de-duplicated property view, with each record linked to the underlying portal listings.
- Normalised floor area, price per square foot, tenure and property type fields.
- Listing history, so price reductions and time on market could be calculated.
- A data dictionary and a weekly quality summary covering completeness and duplicate rates.
| Field | Example |
|---|---|
| Development | Development A, District 10 (normalised name) |
| Planning area | Bukit Timah |
| Property type | Condominium, 3 bedrooms |
| Floor area | 1,184 sq ft (110 sq m) |
| Asking price | SGD 2,480,000 |
| Price per sq ft | SGD 2,095 |
| Sources | Portal A, Portal C |
Outcome
The client's engineers stopped maintaining collection code and redirected their time to product features. Analysts now work from a single, de-duplicated view of the market rather than reconciling overlapping portal data by hand, which made time-on-market and price-change analyses more reliable. Removing agent personal data at source simplified the company's privacy review and reduced the data it had to secure.
The feed has since become the foundation of a new rental analytics module in the client's product.
