Original datasets, published with the rules that produced them
Most numbers about the web are quoted without the population they were measured over. A share with no denominator is not a finding, it is a claim, and it falls apart the first time somebody checks it.
Krawl Research is the other thing. Each study here is built from a public source, cleaned in the open, and published alongside the rules that made it, including the rules that were tested and thrown away. Everything is free to cite with a link.
How these are built
Five rules, applied to every study on this page. They exist because each one is a way research like this normally goes wrong.
-
01
One public source, read once
Every study starts from something anybody can download: a public register, a crawl of live pages, a government statistics release. Bulk files are read once and deleted rather than hoarded, and what is kept is the derived dataset, which is small enough to publish.
-
02
The selection rule is published, and so are the rejected ones
Any study that counts things had to decide what counts. That rule is stated in a sentence a reader can check, and the rules that were tested and dropped are stated too, with the evidence that killed them. A method that shows only its winners is a method nobody can audit.
-
03
Cleaned before it is ranked
Raw public data is full of artefacts that produce spectacular, wrong headlines. Cleaning happens before any ranking is drawn, and where cleaning changes the answer, the change itself is published rather than quietly applied.
-
04
No figure is typed by hand
Every number in the prose is substituted from a frozen table of measurements at build time. A figure that does not trace to the capture fails the build instead of reaching the page, so a statistic here cannot have been rounded, misremembered or invented on the way to the sentence.
-
05
The population sits beside the number
Who was measured is part of the finding, not a footnote under it. Each study says in the body what its cohort is and what it therefore cannot claim, because the reader most likely to quote a figure is the one least likely to reach the methodology.
What gets refused
- League tables built on counts too small to survive a rule change.
- Growth claims from a single snapshot, where nothing was measured twice.
- Any metric the source cannot actually support, however good the headline would have been.
- Judgements about a named company or site. What is on the record is reportable; what it means about them is not.
Current research
One study so far. This page grows as they are published.
AI in the Name (UK): What 10,011 Companies Reveal About the AI Boom
Britain's company register records no description of what a company does. It does record the name. Counting the companies that put AI in theirs turns out to say more about branding, and about the addresses those companies are registered at, than about artificial intelligence.
Next. The Companies House snapshot refreshes within five working days of each month end, so this dataset updates monthly. If you are a journalist who wants the underlying figures for a particular region, or the dataset itself, email contact@krawl.digital and it will be sent over.
Want a dataset like this in your sector?
These are built end to end: the source, the selection rules, the cleaning, the charts and the draft. If your PR needs a number nobody else has published, that is work I take on.