Methodology
A transparent, repeatable process. Here is exactly what happens between your brief and your dataset.
Five stages
Each stage produces a record you can inspect.
1. Scope
Research question, region, source types, date range, fields and delivery format are written down and agreed.
2. Source review
Terms, robots rules, copyright and database rights, and personal data risk are assessed per source.
3. Collection
A polite collector fetches public pages at low rates and stores raw snapshots with timestamps.
4. Processing
Extraction, cleaning, deduplication and standardisation, all scripted so they can be re-run.
5. Delivery
Data, datasheet, codebook and log summary are transferred securely.
Quality checks
Completeness
We compare what was collected with what each source lists, and report gaps.
Consistency
Formats for dates, places and categories are validated against the codebook.
Duplicates
Near-duplicate pages and repeated records are detected and flagged or removed.
Spot checks
A sample of records is manually compared with the live source page.
Provenance
Each record links back to its source URL and collection time.
How we behave as a collector
Identified
Our collector uses a clear user agent with a contact address, so site owners can reach us.
Rate-limited
Requests are spaced out and scheduled to avoid load on small servers.
Rule-following
Robots directives and clear opt-out signals are honoured. Where an API or bulk download exists, we prefer it.
Stoppable
If a site owner asks us to stop, we stop and remove their content from ongoing collection.
Want to see a sample datasheet?
Ask us and we will explain what one looks like for your topic.
Ask a question