Methodology

A transparent, repeatable process. Here is exactly what happens between your brief and your dataset.

Five stages

Each stage produces a record you can inspect.

1. Scope

Research question, region, source types, date range, fields and delivery format are written down and agreed.

2. Source review

Terms, robots rules, copyright and database rights, and personal data risk are assessed per source.

3. Collection

A polite collector fetches public pages at low rates and stores raw snapshots with timestamps.

4. Processing

Extraction, cleaning, deduplication and standardisation, all scripted so they can be re-run.

5. Delivery

Data, datasheet, codebook and log summary are transferred securely.

Quality checks

Completeness

We compare what was collected with what each source lists, and report gaps.

Consistency

Formats for dates, places and categories are validated against the codebook.

Duplicates

Near-duplicate pages and repeated records are detected and flagged or removed.

Spot checks

A sample of records is manually compared with the live source page.

Provenance

Each record links back to its source URL and collection time.

How we behave as a collector

Identified

Our collector uses a clear user agent with a contact address, so site owners can reach us.

Rate-limited

Requests are spaced out and scheduled to avoid load on small servers.

Rule-following

Robots directives and clear opt-out signals are honoured. Where an API or bulk download exists, we prefer it.

Stoppable

If a site owner asks us to stop, we stop and remove their content from ongoing collection.

Limits of web data. Websites change, disappear and are never a full picture of a region. We describe these limits in each datasheet so your conclusions stay honest.

Want to see a sample datasheet?

Ask us and we will explain what one looks like for your topic.

Ask a question