Public web data ยท Academic & research organisations ยท UK

Regional web data,
gathered openly and responsibly.

Kabuuga collects publicly available information from websites in the regions you study, then cleans, documents and delivers it in formats your team can analyse straight away.

Public web pagesOpen government portalsRegional news archivesBusiness directoriesEvent listingsPublic noticesOpen datasetsTourism and heritage sitesJob boardsCultural institutions

What we do

Researchers often need web data from a specific region but lack the time or tooling to collect it well. We do that part, and we show our working.

๐ŸŒ

Regional collection

We gather publicly accessible pages and open data tied to a place: a country, county, city or cross-border area you define.

๐Ÿงน

Cleaning and structuring

Messy HTML becomes tidy tables: duplicates removed, dates and places standardised, languages tagged.

๐Ÿ“‘

Full documentation

Every dataset ships with a datasheet: sources, dates collected, method, known gaps and licence notes.

๐Ÿ”

Repeatable snapshots

One-off pulls or scheduled re-collection, so you can study change over months or years.

๐Ÿ›ก๏ธ

Compliance first

We respect site rules, avoid private areas and treat personal data with UK GDPR in mind.

๐Ÿ“ฆ

Analysis-ready delivery

CSV, JSON, Parquet or spreadsheet files, plus a codebook so your analysts start on day one.

How a project runs

A clear sequence, from your research question to a documented dataset.

Scope

We agree your question, region, sources, time frame and output format.

Check

We review each source's terms, robots rules and data protection risk.

Collect

Gentle, rate-limited collection of public pages only, with logs kept.

Clean

Deduplicate, standardise and quality-check against your codebook.

Deliver

Secure transfer of data, datasheet and method notes.

Who we work with

Universities

Departments and research groups in social science, geography, economics, linguistics, public policy and digital humanities.

Research institutes

Independent and public-interest bodies who need dependable regional evidence.

Libraries and archives

Teams preserving and studying the regional web as a cultural record.

Student and doctoral projects

Supervised projects that need a well-scoped, documented dataset.

Our promises

The rules we hold ourselves to on every project.

Public only

No logins, paywalls, private groups or bypassing of access controls.

Polite by design

We identify our collector, honour robots rules and keep request rates low.

Transparent

You see exactly where data came from and when. Nothing is a black box.

Privacy aware

Personal data is avoided or minimised, and never collected without a lawful basis.

Quick answers

Is collecting public web data legal?

It depends on the source, the data and how it is used. We assess each project against site terms, copyright and database rights, and UK GDPR where personal data could appear. We are not a law firm, so your institution's legal or ethics team should confirm suitability for your use.

Do you collect personal data?

By default, no. We design collections around organisations, places, documents and aggregate information. If your research needs personal data, we discuss the lawful basis, minimisation and your ethics approval first.

Can you cover any region?

We work best where sources are openly available and we can review them properly. We will tell you early if a region or source is unsuitable.

How long does a project take?

Small, well-defined collections can take days. Larger multi-source projects take longer, mostly because of the source review and quality checks. We give a plan after scoping.

Have a regional research question?

Tell us the region, the sources you have in mind and what you want to find out.

Contact Kabuuga