There is a great deal of talk about hypothetical AI risks and very little, with data, about how already-deployed systems fail. The information exists, but it is scattered across headlines, regulatory reports and press releases, with no common taxonomy and no way to compare it.
A headline is not a data point. To be useful for a decision it needs an explicit classification, a severity scale defined before it is applied, and traceability back to the original source.
01Dataset with a contract
23 incidents documented between 2015 and 2025, across 18 public and private organisations, with 13 columns and a data contract validated on every load.
02Failure taxonomy
Each incident is classified by failure mode — bias, security, hallucination, containment — and by severity on a 1 to 5 scale with a written rubric.
03Dashboard
Four views (overview, deep-dive case, dataset and methodology) with filters by failure type, period, severity and application domain.
04Deep-dive case
Root cause analysis of the model supply chain between OpenAI and Hugging Face (2023-2024): four incidents with a common cause and seven controls that would have prevented them.
05Methodology in the open
The rubric, the criteria and the dataset's limitations are published inside the dashboard itself, before anyone quotes a figure.
The decision that matters most
Publishing the limitations inside the product, not in an appendix. The methodology tab explains what the dataset measures and what it does not, and it sits at the same level as the charts.
The reason is practical: as soon as a dashboard shows percentages, people quote them. If the caveats do not travel with the figure, they are lost on the first forward.
23 incidents are not a statistically representative sample of AI failure, and the dataset says so explicitly. It is useful for recognising patterns and failure modes, not for estimating base rates or comparing rates across sectors.
That the difference between a headline aggregator and an analysis tool lies almost entirely in the boring work: defining the taxonomy before classifying, writing the rubric before scoring, and declaring the limits before someone else discovers them.