ATLASResearch
methods
ACROSS FAMILIES/ Data collection

Web scraping

Extract online records reproducibly

Web scraping programmatically extracts information from web pages or publicly available interfaces into a research dataset. A credible workflow defines the target population, extraction rules, timestamps and handling of changing pages. It is collection infrastructure rather than an analysis method. Permission, privacy, provenance and platform-specific selection matter even when the content is publicly visible.

WHEN IT FITS

Use when required information is online and permitted programmatic collection is feasible. Prefer documented APIs when available and verify that page coverage and extraction quality fit the research question.

Strengths

  • Collects repeated detailed observations at scale
  • Can create reproducible timestamped collection pipelines

Limitations

  • Site changes break extraction and comparability
  • Visible pages may be a selective or personalised subset

Know the boundary

A scraped dataset is not automatically a representative sample of offline activity.

USED ACROSS
Banking & financeBusiness & MBAComputer science