Skip to content

Methodology

How a campaign gets onto this site

What we take from other people’s published research, what we refuse to take, how often we get it wrong, and how every figure on the Index is computed.

By Antonio RaduLast reviewed 2026-08-06

The numbers, as of the last rebuild

Published campaigns
44
Each reviewed by a human before publication.
Indicators on record
1,359
Each backed by a verbatim string in an archived source.
Audited entries
0
No sample has been drawn yet. The design is five a week, re-checked blind.
Measured error rate
not yet measured
No audit cycle has been reviewed yet.

No extraction attempt has yet been refused by the constraint in §3. The count is published whether it is zero or thousands: a rising number means the constraint is working.

1. Where the data comes from

Public research, published by named people and organisations, that anybody can read for themselves: vendor blogs, CERT advisories, independent researchers. Every publisher appears on /sources, and every campaign page carries the publisher, the date and a followed link to the original above the fold. A source has to meet three conditions:

  • It is public and citable. Nothing arrives from a private channel, a paid feed, or a tip that cannot be pointed at.
  • Its licence permits what we do. We do not ingest feeds whose terms forbid derivative works or redistribution, even where the underlying facts are not copyrightable: a contract is not defeated by a copyright argument. That costs coverage, and buys the right to release everything here under CC BY 4.0.
  • It can be archived. We keep a private copy of the source text, never republished, whose only job is to answer “did this string really appear in that report?” years after the original goes offline.

2. We extract fields, never prose

A language model reads the source and returns structured values only: indicator strings and the sentence each appeared in, first and last observed dates, target operating systems, technique and lure labels, malware families, execution mechanisms. It is not asked for a summary and is not permitted to emit one.

The summary on a campaign page is composed from those fields, in our own sentences, capped at 150 words, and a check rejects any summary sharing an eight-word sequence with the source. Facts are not copyrightable; the paragraphs somebody wrote about them are, and that distinction is the entire reason this site can exist without asking permission. Quotation is allowed where the exact wording carries the meaning: at most 25 words, in quotation marks, attributed inline, under the editorial policy.

Labels are recorded, never inferred — a target platform never derived from an indicator, an execution mechanism never from a command string — and an alias list at ingest folds four vendors’ names for one kit into one label.

3. The verbatim constraint

Every link between an indicator and a source report carries the exact string as printed in that report, including whatever defanging the author used. A database trigger checks that the string is present in the archived copy of the source; if it is not, the insert is refused. Not flagged for review, not published with a lower confidence score — refused, at the storage layer, before the row can exist.

Refused attempts are counted and the total is printed at the top of this page. The point is not that the extractor is perfect; it is that a hallucinated indicator cannot reach a published page, because no code path writes one. If rows are being dropped the extractor is wrong and gets fixed; the constraint does not get relaxed.

So every indicator here traces to a sentence in a document a named publisher put their name on. If you follow the source link and cannot find the string, that is a defect — §9 is how to report it.

4. Who controls the host

Most ClickFix lure pages are not on infrastructure the attacker bought but on somebody’s hacked content-management system: a restaurant, a charity, a local business with no idea. Publishing that domain next to the word malicious is how a reference site becomes a small disaster for a stranger. So every host carries an ownership assessment, and it decides what the site may do with it:

Ownership values and their consequences
AssessmentMeaningIn the blocklist export
Attacker-controlledRegistered or stood up by the operator for this activity.yes
Compromised legitimate siteA real site with a real owner, serving attacker content without consent.never
Shared hosting platformA platform where blocking the domain blocks everyone on it.never
Shared infrastructureInfrastructure used by many parties, some of them legitimate.never
Not yet assessedNobody has looked yet. The default, and it shows no verdict at all.never

Unassessed is deliberately silent: a grey “unknown” badge still implies somebody looked and had doubts. It is also excluded from every export, not only the blocklist files, and can never become indexable — both database rules, because the ownership value is part of the eligibility expression.

5. Later observations, including the negative ones

One pass has run, and it asked the registries rather than the hosts. It recorded 328 dated findings, covering 166 of the 1,359 indicators on record. More names than that were asked about; the ones that produced no finding are the subject of the next paragraph. A registration lookup goes to the registry that issued the name; it does not contact the host, and it never reaches a compromised third party’s server. §8 sets out what that contact never becomes.

A registry answering “no such domain” is a finding, and it is recorded. A registry answering that a registration still exists is not a finding about the host, and nothing is recorded for it — a name can be registered and serving nothing, and this site does not hold the evidence to tell those apart. The indicators with no recorded state are not indicators we believe are live. They are indicators nobody has established anything about.

Negative results are stored exactly like positive ones, which is the point: a record that discards its negatives can never say when something stopped. Checks run on a small budget, oldest-checked-first, so a gap in the record means nobody looked, not that nothing changed.

Every last-observed date on this site is the date a publisher last reported the host, not the date anybody last saw it alive, and a re-check does not move it. A host confirmed gone was not seen on the day we confirmed it — writing that date in would inflate every lifetime computed from the column, and the longest-dead names by the most. Where we looked, and what we found, is recorded separately and shown on the indicator’s own page.

6. The blind weekly audit

No cycle has run yet, which is why the error rate above reads not yet measured. The definition: five already-published entries sampled at random each week and re-checked against their sources by a human, with the original assessment hidden. Each is recorded as correct, wrong or unclear, and the published rate is wrong ÷ reviewed over reviewed samples only — an unreviewed backlog cannot flatter it, because unreviewed samples are not in the denominator.

The cheaper alternative is to count how many proposed entries a human rejected before publication. It is self-deceiving: the review queue measures what we caught, and cannot by construction contain one of the errors that got through — the only errors a reader is exposed to. Auditing published output is the one version of this figure that can embarrass us, which is the point.

Entries found wrong are corrected and logged publicly on /corrections. The rate is not smoothed, not annualised, and not restated when it moves the wrong way; at five samples a week one verdict moves it by whole percentage points, so treat an early number as an order of magnitude.

7. How each Index figure is computed

The ClickFix Index publishes six figures from this record. The sixth is our own error rate, defined in §6. A figure there that is not defined here is a defect.

New ClickFix domains observed per week

Each domain is assigned to the ISO week (Monday to Sunday, UTC) of its first observed date and counted once: four vendors covering one host is one domain. Weeks with no new domains are plotted as zeros. The window spans at most 52 weeks and ends at the most recent week containing an observation, not at today — an empty tail would read as “no domains were reported” when it means “nothing has been ingested since”. Domains with no first observed date, or first observed before the window opens, are counted under the chart rather than dropped.

Read it as reporting, not activity. A quarter of decline is a question worth asking; a fortnight of it is a conference season.

Median observed domain lifetime

For each domain holding both a first and a last observed date, the span is last observed − first observed in whole calendar days; the figure is the median of those spans, and percentiles use the nearest rank without interpolation. Zero-day spans — reported once, never seen again — are included: excluding them would raise the figure by choosing the flattering definition. Domains only, since an address is reassigned rather than retired and a hash has no lifetime. A second version covers only hosts since recorded as not resolving, unregistered, parked or suspended, sinkholed, nearer a true lifetime on a smaller sample.

Read it as a lower bound. A host still alive has no end date, so its span is “so far”; nobody watches continuously, so a domain that went dark on a Tuesday is recorded as dying the day somebody next looked; and a domain enters the record when it was first written about, possibly long after it went up. Quote it with its sample size as “median observed lifetime”, never as “ClickFix domains last N days” — one is a claim about the world, the other about a record.

The lure-theme, target-OS and execution-mechanism mixes

For each label, the number of distinct published campaigns carrying it, over the number carrying at least one label in that dimension — not over the whole corpus, which would understate every theme by however much of it is unlabelled. Campaigns carry more than one label, so shares do not sum to 100 and the Index prints the denominator beside every table. Where labels come from is §2.

Read them as published attention. Research skews towards campaigns that are large, novel, or aimed at a vendor’s own customers.

Sampling, and citing a figure

The weekly series and both lifetime figures come from one sample of domain records, read most-recently-first-observed and capped at 20,000 rows; below the cap that is the whole population. Above it the sample becomes a recency window: older weeks fall out and the median tilts towards recently reported domains, the ones least likely to have finished. The mixes use every published campaign, unsampled.

Figures are revisable downwards: an indicator withdrawn after a dispute leaves the record and every count moves with it, so a citation needs the “as of” date the Index prints — a database column, never the time the page was built. A change to a definition here changes what the numbers mean, and is logged on /corrections.

8. What we deliberately do not do

  • No detonation, and no rendering. The checks in §5 are the only contact this project will ever make with a listed host: a DNS lookup, a registration lookup, and a request read only far enough to tell a live host from a parking page. Nothing is executed and no markup is kept.
  • No screenshots of lure pages. A convincing picture of a fake verification prompt is a working piece of social engineering; we describe the pattern in words.
  • No command strings in runnable form. A command is described, never presented as something to copy. There is no copy button anywhere on this site.
  • No first-party discovery. We do not crawl for new lures; everything traces back to somebody else’s published work, credited. The ceiling on our coverage is whatever the research community chose to write about, through a publisher set that is narrow and English-language.
  • No actor attribution. Where a source names an actor we record that the source named it. We do not adjudicate between sources that disagree.

9. When we are wrong

Researchers whose work is summarised here can have anything corrected within five working days, no questions asked, by writing to corrections@clickfixreport.com. Owners of listed sites should follow the dispute process. Every correction that changes what a page asserts is logged publicly, with a date and a reason. We do not quietly edit pages.