There is one thing that always happens the moment you publish a dashboard with percentages: people quote them. They pull them out of context, drop them into a presentation, forward them on. And the caveats, if they live in an appendix or a footnote, are lost on the first forward.

When I built the AI Safety Incident Tracker, this worried me more than the technical side. The dataset has 23 incidents. That is enough to see patterns — which failure modes dominate, which domains hurt most — and completely insufficient to estimate base rates or compare rates across sectors. If someone posts "39% of AI incidents cause serious harm" without that context, my work has produced misinformation with my name on it.

What I did: the methodology tab sits at the same level as the overview tab. It is not a small link at the bottom. It explains what the dataset measures, how the incidents were selected, what each level of the severity scale means, and which questions cannot be answered with this data.

On top of that, the severity rubric was written before anything was classified. It is a detail that looks bureaucratic and is not: if you score first and define the scale afterwards, the scale ends up justifying the scores you already gave. The order matters.

On traceability: every row links to its primary source. Not to an article summarising the news — to the source. It is the rule that admits the fewest exceptions in the whole project, because a claim that cannot be traced back to its origin is not a data point, it is a rumour with formatting.

The lesson I take from it: declaring the limits of your analysis is not hedging or false modesty. It is the part that makes the rest usable. An analysis that only reveals its weaknesses when someone challenges them was never ready to support a decision.