Methodology

What the public numbers count, how they are built, and the cases where we show no number rather than a bad one. The statistics page itself is not live yet; this went up first so the method is on record before there is a chart to defend.

Where the data comes from

Ransomware groups publish posts naming organisations they say they have compromised. Several projects monitor those sites and republish what they see; we read those projects rather than the sites. We do not scrape leak sites ourselves, we never download anything a group has published, and we never interact with their infrastructure.

The pipeline reads more than one feed, because each monitors a different set of sites and parses them differently, and any one of them can stop without notice. A post seen by two feeds is recorded once, with a note that two feeds carried it.

Which feeds were actually read for a given period, and when each last answered, is shown on the statistics page next to the numbers built from them. The table at the foot of this page is the register of sources and the attribution each one requires — not a claim that every one of them fed the figure you are looking at.

What a claim is

A claim is one post on one leak site in which a group names an organisation. That is the whole of it. It is a thing a criminal group published, observed by a feed, recorded by us.

What a claim is not:

  • Not something we observed happening. We were not there, we have no visibility inside anybody's network, and nothing on our side corroborates a post.
  • Not necessarily true. Groups exaggerate, recycle old material, and republish another group's work after a rebrand. Some posts name an organisation that was reached through a supplier rather than directly. Some name the wrong legal entity, or a parent company, or a brand that belongs to somebody else entirely.
  • Not a measure of what happened in a period. A post appears when a group decides to publish, which is typically weeks after the event and only if negotiation failed. Organisations that paid quietly are, by construction, absent.
  • Not a census. The counts cover the sites our feeds watch, for as long as they were watching. A group with no site, or one nobody is monitoring, contributes nothing and leaves no gap you can see.

So the counts describe leak-site publishing behaviour. They are a reasonable proxy for which groups are busy and who they are aiming at, and a poor proxy for anything else.

How claims are counted

  • Deduplication. Posts with the same normalized organisation name from the same group within 14 days of each other are one claim, whichever feed each arrived from. The earliest is kept and the rest are linked to it, so a group reposting does not read as a second organisation.
  • Feed count. Each claim records how many feeds carried it. Two or three is better evidence that the post existed than one is — it says nothing about whether the post was true.
  • Aggregation. Public numbers are counts per day, country, sector and group. Nothing else is in the aggregate tables: they have no column that could hold an organisation name.

Country means the organisation, not the event

The country on a claim is the one the source attributed, and sources attribute the organisation's headquarters. It is not where anything took place, not where the data sat, and not where the group is. A multinational headquartered in one country counts once, against that country, however its subsidiaries are spread.

Where no country was attributed, the claim is counted under WORLD rather than dropped. That keeps the country breakdown honest: the unattributed claims are visible as their own bucket instead of quietly vanishing from every view and from the totals.

Sector is mapped, and sometimes unknown

Sources describe sectors in free text — the same industry arrives as a dozen spellings. Each is mapped to a fixed 20-item taxonomy through a synonym table.

Anything the table does not cover becomes unknown and goes to a review queue for a person to map. It is never guessed and never silently dropped. A rising unknown bucket means our table is behind, not that a new industry is being targeted.

Groups, aliases and rebrands

Groups rename themselves, split, and reappear under new branding, and the feeds use whichever name the site was using at the time. Known aliases resolve to one group record, and a rebrand is linked to its predecessor, so history does not reset to zero the day a site changes its name. Where the link is not established, the two read as two groups — which is a known way for these counts to understate a long-running operation.

Thin cells: the floor of five

Below five claims in a country-and-sector cell over 90 days, we show the count and nothing derived from it. No trend line, no percentage, no "up 200%" — which is what one claim becoming three looks like when a percentage is applied to a number that small.

In that case the view falls back to the sector worldwide and says on the page that it has done so. The floor applies to every percentage we publish, not only to trends.

Two states that look alike and are not: no claims in a cell, and too few claims to say anything. Only the first is quiet, neither is reassuring, and the page always says which one you are looking at.

A quiet week and a broken feed look identical

This is the failure mode that matters most on a page of counts. If a feed stops answering and we keep drawing the chart, the line goes down and reads as good news.

What we do about it:

  • Every source shows when it was last read successfully, next to the numbers built from it.
  • A source that failed is labelled as failed for that period. It is never rendered as a zero, and its absence is never averaged away.
  • A recent period is still filling up. Groups publish in bursts and feeds lag, so the last few days are marked as partial rather than presented as settled.

Dates on the exploitation timeline

Leak-site posts do not say which vulnerability was used, so there is no honest way to measure "days from disclosure to ransomware exploitation" from this data. Any single number claiming to be that measurement is inference wearing a lab coat. What we build instead is a timeline of four dates, each of which has to name the source it came from. Two of the four are defined and blank on every CVE we draw, and the timeline says which is which on every row rather than leaving a reader to guess:

  • Patch available — when the vendor made a fix available. Defined, and blank on every CVE. No source for it is wired up. The NVD record gives us a CVE publication date and links to vendor advisories, and neither of those is the date a fix shipped, so nothing here stands in for it.
  • KEV listing — the date CISA added the CVE to the Known Exploited Vulnerabilities catalog, read from the catalog's own dateAdded field. Populated. This is the only one of the four we hold today. Where an entry reaches us carrying a date we cannot read, the CVE is shown as listed with its listing date unknown, rather than given a date we chose.
  • Public exploit code — the first dated public proof-of-concept. Defined, and blank on every CVE. No source for it is wired up either.
  • First ransomware attribution — the earliest date a credible published source tied that CVE to a named ransomware group. Defined, and blank on every CVE today. These are entered by hand from published advisories and none has been entered yet. It is a date about reporting, not about first use: it can only be as early as somebody writing it down.

On the signed-in view there is one further marker, and it is the only other date we currently hold: when we first recorded that CVE on your own surface. That one is yours, read from your own findings rather than from any catalog, and it is a lower bound — it starts when our observation started, not when your exposure did.

Where a date is unknown it is shown as unknown, with the reason for the blank beside it, and the reasons are kept apart: a date nobody holds for any CVE is a different statement from a date our sources hold and this CVE has none of. We do not interpolate one, and we do not substitute a nearby date. Any median we publish states the number of CVEs it was computed over, because a median of four is a sentence, not a statistic.

The two timeline statistics, and what they do not say

From those dates we publish exactly two figures, both by quarter, and both with the number they were computed over printed next to them.

  • Median days from KEV listing to first ransomware attribution, over the CVEs in that quarter where both dates are known — with that count beside the median.
  • The share of CVEs added to KEV in the quarter that now carry a ransomware attribution in our knowledge base — with the number of CVEs we hold with a KEV listing date in that quarter as the denominator. That is what we read, not what CISA published: if the KEV feed stops, the newest quarter's cohort is short and the share moves in the flattering direction, which is why every response carries feed health beside the figure. An entry counts here whether or not anybody has dated it, which is why this figure can be real while the median above is empty.

A CVE belongs to the quarter CISA listed it in, for both figures — so the two are read against the same cohort, and the median for a quarter is a statement about the vulnerabilities added to KEV in it rather than about attributions published in it.

The two dates being subtracted are not equally solid, and the difference is the whole reason this section exists:

DateWhere it comes fromHow much weight it bears
KEV listingThe dateAdded field of CISA's Known Exploited Vulnerabilities catalog.High. A date a government agency published, unchanged by us.
First ransomware attributionOur own knowledge base: the earliest date on which a published source we hold reported a named extortion group using that vulnerability. Every entry carries at least one source URL and a named reviewer.Medium — it depends on when a source chose to publish. The group was using the vulnerability before anybody wrote it up, so this is the date reporting caught up, not the date exploitation began.

So the median is a gap between two publication events. What it does not say:

  • Not how long you have before a ransomware group uses a vulnerability. Neither endpoint is the start of exploitation. Both are moments at which somebody wrote something down.
  • Not a measure of a group's speed. A shorter gap can mean attention arrived sooner rather than that anybody moved faster, and the gap can be negative: a group is often reported using a vulnerability before CISA lists it. Those negative gaps stay in the median rather than being clipped to zero, which would say the two dates coincided.
  • Not a census of ransomware vulnerability use. It covers the CVEs our knowledge base has an entry for, which is a hand-curated set built from public advisories.

What the number is today: nothing, and it says so. The attribution dates are entered by hand from published advisories, and none has been entered yet — the ATT&CK data we seeded the knowledge base from carries no first-observed date, and the importer refuses to invent one. So the sample size is zero, every quarter reads "no CVEs where both dates are known", and the count of CVEs added to KEV in that quarter is printed beside it so the difference between nothing happened and we cannot measure thisis visible rather than inferred. When dates are entered, the sample size moves in public along with the figure.

Two other dates belong on a timeline and are shown as unknown for now, because no source for them is wired up: when the vendor made a fix available, and when exploit code first became public. We would rather leave a gap than fill it with a nearby date — treating the CVE publication date as the patch date would manufacture the very number this page exists to be careful about.

CISA's ransomware flag is counted separately and never timed. The KEV catalog marks some entries as known to have been used in ransomware campaigns. That flag names no group and carries no date. We report it as its own count, labelled as CISA's assertion, and it never enters the median or the share — a date-less flag can only join a days-elapsed figure if somebody invents a date for it, and several hundred entries carry that flag, so the fabricated statistic would look impressively well-sampled.

The floor of five applies here too, and the two figures apply it to two different numbers. The median is suppressed below five CVEs in that quarter where both dates are known — not below five CVEs in the quarter, which is a much larger number and a different statement. The share is suppressed below five CVEs listed by CISA in the quarter, which is its denominator. A zero share over a large denominator is published, because "none of these carries an attribution in our knowledge base" is a statement about our knowledge base that a reader is entitled to.

What is deliberately absent

  • Organisation names. No name from a leak-site post appears on any public surface here, in any form. Privacy sets out what we hold, why, and who can read it.
  • Any lookup by company name. There is no search, filter or parameter, at any tier, that takes an organisation name and answers whether it appears in this data.
  • Probabilities. We do not publish odds that a particular organisation will be targeted. There is no defensible way to compute one from these inputs.

Sources and attribution

SourceUsed forTerms and attribution
RansomLookLeak-site posts and the group list. One of two live feeds, so a scraper going dark shows as a gap rather than as a quiet week.CC BY 4.0 — free to use, including commercially, with attribution. This credit is that attribution.
ransomware.liveLeak-site posts, and the only feed that carries a country or a sector — so every country figure on the statistics page rests on this one source.API PRO tier, which permits business use. The unauthenticated API is personal use only and we do not read it.
RansomwatchA 16,000-claim historical corpus, ingested once. The project is archived and published nothing after June 2025, so it contributes nothing current and is no longer polled.Unlicense — public domain. Credited because the history it left still counts.
CISA Known Exploited VulnerabilitiesThe KEV listing date on the exploitation timeline, and the ransomware-campaign flag.US Government work, public domain.
FIRST EPSSThe EPSS score we store against a CVE: the modelled probability that the vulnerability is exploited somewhere in the world within the next 30 days. It is a statement about the vulnerability and never about any reader or organisation. We keep it on the vulnerability record and, on the timeline, alongside the KEV listing date. No page here displays it.Free with attribution requested: EPSS data is produced by the Exploit Prediction Scoring System (EPSS), a FIRST.org special interest group.
NVDCVSS scores, affected-product identifiers, summaries and vendor advisory links on a vulnerability record. It supplies no date on the exploitation timeline.Public domain. This product uses the NVD API but is not endorsed or certified by the NVD.
MITRE ATT&CKTechnique identifiers on initial-access vectors attributed to a group.© 2026 The MITRE Corporation. This work is reproduced and distributed with the permission of The MITRE Corporation.
MaxMind GeoLite2Picking a default country for a first-time visitor. Country level, not stored.This product includes GeoLite2 data created by MaxMind, available from https://www.maxmind.com.

Counts reflect claims published on extortion leak sites. Claims are made by criminal groups, are frequently unverified, and may be duplicated, exaggerated or false.