Data Collection Ethics for Engineering Teams
A practical framework for collection work: public versus personal data, proportionality, retention limits, robots and terms, and a pre-launch checklist.
11 articles
A practical framework for collection work: public versus personal data, proportionality, retention limits, robots and terms, and a pre-launch checklist.
How content classification works, why categories affect ad placement and filtering, how misclassification happens, and how to verify a category locally.
Classified page structure, deduplicating reposted listings, per-city scoping and the polite request rates that keep a classifieds collector running.
The Places API versus collecting from Maps pages: what each returns, the real cost trade-off, and why the address location changes local rankings.
Collecting news articles and metadata across outlets: using RSS where it exists, extracting content, deduplicating coverage, and normalising publish times.
Screen scraping reads a rendered interface rather than a structured feed. Where it came from, where it still applies, and why it is the least durable option.
Scraping acquires data, mining analyses it. How the two fit together in a pipeline, and why the distinction decides what you instrument and where proxies sit.
When an official API is available and sufficient, use it. How the two compare on reliability, cost, limits and legal footing, plus the hybrid approach.
A grounded introduction to web scraping: what it is, how it differs from an API, the basic anatomy of a scraper, and where the legal and ethical lines sit.
Property portals localise heavily and throttle scrapers. How to collect listings per market, handle rich media, and normalise records that vary by region.
Collecting pricing, assortment and availability data across regions without hitting localised blocks, and how to design a sampling plan that holds up.
Clean, pre-filtered residential and mobile proxies, sign up and send your first request in minutes.