Regional web data,
gathered openly and responsibly.
Kabuuga collects publicly available information from websites in the regions you study, then cleans, documents and delivers it in formats your team can analyse straight away.
What we do
Researchers often need web data from a specific region but lack the time or tooling to collect it well. We do that part, and we show our working.
Regional collection
We gather publicly accessible pages and open data tied to a place: a country, county, city or cross-border area you define.
Cleaning and structuring
Messy HTML becomes tidy tables: duplicates removed, dates and places standardised, languages tagged.
Full documentation
Every dataset ships with a datasheet: sources, dates collected, method, known gaps and licence notes.
Repeatable snapshots
One-off pulls or scheduled re-collection, so you can study change over months or years.
Compliance first
We respect site rules, avoid private areas and treat personal data with UK GDPR in mind.
Analysis-ready delivery
CSV, JSON, Parquet or spreadsheet files, plus a codebook so your analysts start on day one.
How a project runs
A clear sequence, from your research question to a documented dataset.
Scope
We agree your question, region, sources, time frame and output format.
Check
We review each source's terms, robots rules and data protection risk.
Collect
Gentle, rate-limited collection of public pages only, with logs kept.
Clean
Deduplicate, standardise and quality-check against your codebook.
Deliver
Secure transfer of data, datasheet and method notes.
Who we work with
Universities
Departments and research groups in social science, geography, economics, linguistics, public policy and digital humanities.
Research institutes
Independent and public-interest bodies who need dependable regional evidence.
Libraries and archives
Teams preserving and studying the regional web as a cultural record.
Student and doctoral projects
Supervised projects that need a well-scoped, documented dataset.
Our promises
The rules we hold ourselves to on every project.
Public only
No logins, paywalls, private groups or bypassing of access controls.
Polite by design
We identify our collector, honour robots rules and keep request rates low.
Transparent
You see exactly where data came from and when. Nothing is a black box.
Privacy aware
Personal data is avoided or minimised, and never collected without a lawful basis.
Quick answers
Is collecting public web data legal?
It depends on the source, the data and how it is used. We assess each project against site terms, copyright and database rights, and UK GDPR where personal data could appear. We are not a law firm, so your institution's legal or ethics team should confirm suitability for your use.
Do you collect personal data?
By default, no. We design collections around organisations, places, documents and aggregate information. If your research needs personal data, we discuss the lawful basis, minimisation and your ethics approval first.
Can you cover any region?
We work best where sources are openly available and we can review them properly. We will tell you early if a region or source is unsuitable.
How long does a project take?
Small, well-defined collections can take days. Larger multi-source projects take longer, mostly because of the source review and quality checks. We give a plan after scoping.
Have a regional research question?
Tell us the region, the sources you have in mind and what you want to find out.
Contact Kabuuga