Skip to content

Source-protection aggregation

The ten S-classes are not a list of ten independent admonitions stapled onto tool cards. They are one editorial system. Taken together, they describe how a verification workflow can go wrong on the source-protection side, and they set out the responses a working desk needs ready when each kind of wrong-going is visible. The T6 source-protection tree handles the in-case routing: it is the page a fact-checker opens with a current case in hand. This page is the layer behind T6. Each S-class is set out here as a kind of threat, with the editorial-position statement behind it stated plainly and the connection between operational mitigation and editorial framing made visible.

The framing is consequential. A toolkit that treats source protection as procedural advice (run this check, click this option, do not upload this file) carries less weight than one that treats source protection as editorial policy. The toolkit takes the second route. A workflow that produces a verification at the cost of exposing a source is not a verification; it is a security incident with a verification result attached. Every S-class on this page is one expression of that single position. Triggers and mitigations vary; the underlying claim does not.

How the ten classes sit together

Ten S-classes, four functional groupings.

The first grouping covers the source-and-content interface. S1 (source-identifying upload risk), S3 (graphic, sexual, child-safety, or abuse material) and S4 (doxxing, harassment, or vulnerable-community targeting) name three different ways a file can carry exposure with it. S1 is about the source whose identity travels with the artefact. S3 is about the subject of harmful content whose dignity travels with the artefact. S4 is about the targets whose retaliation risk travels with any republication or platform reporting.

The second grouping sits on the legal-and-state environment. S2 (state-linked or legally sensitive investigation) and S5 (private or encrypted group collection) name two ways a verification workflow can intersect with state authority. S2 is about the topic that brings state attention to the work; S5 is about the collection method that the state already polices independently of topic.

The third grouping is the detector-and-claim architecture. S6 (detector-only accusation) and S8 (liar's-dividend risk) are editorial twins. S6 prevents over-claiming AI from a detector signal; S8 prevents the political abuse of "could be AI" as a way to dismiss real evidence. They sit on opposite sides of the same problem: the gap between what a detector tells you and what is editorially defensible to publish.

The fourth grouping concerns workflow infrastructure. S7 (malware, APK, phishing, payment, or ID-harvesting), S9 (cross-border data transfer or vendor retention) and S10 (staff safety and trauma) name three categories of harm that travel through the tools and the people doing the work, not through the content or the subjects. S7 protects the verifier from the artefact. S9 protects the source from the verifier's vendor stack. S10 protects the verifier from the work itself.

Taken together, the ten classes cover the surface where the verification chain can leak. The mapping is not one-to-one between tool and S-class; in practice, three or four classes often fire on the same case. A politically sensitive WhatsApp source in Laos triggers S1 (the source's voice is on the audio) plus S2 (the topic is Lao state criticism) plus S5 (the artefact came from a private group) plus S9 (any cloud-hosted detector pass routes the file across borders) before the verification work even starts. The case has to clear all four before tool upload or third-party contact, and that clearance is sequential; it is not collapsible into one check.

S1 – Source-identifying upload risk

S1 is the first and most-fired class. It binds whenever a file carries information that identifies a person who would be harmed by exposure: the face on the video, the voice on the audio, the GPS coordinates in the EXIF, the private handle visible in the upper corner, the phone number, the whistleblower's voice, the victim's identity, the private group name, the original file from a vulnerable source.

The threat model is straightforward. A working fact-checker reaches for a detector or a forensics tool that runs in the cloud, uploads the source file as part of the standard workflow, and in doing so transmits the identifying information to a third-party vendor's servers. From that moment, the vendor's retention policy sits between the source and exposure. The toolkit's editorial position is that this transmission is not neutral. Even where the vendor's published retention policy is short and the vendor is reputable, the transmission has happened: the file is on someone else's server. In surveillance-risk countries, that residency can carry legal consequences the source did not consent to.

Mitigation runs in four steps. Classify the file as public, sensitive or source-identifying before any upload. For source-identifying material, route through tools that do not transmit the source file to vendor-controlled servers: reverse-image search on a frame already published, EXIF tooling that runs locally, OCR running on the verifier's machine, archive capture that targets the public URL and not the file. If a detector verdict is operationally necessary on identifying material, strip identifying context locally first: crop to remove backgrounds with location signals, blur faces or text that identifies the source, redact audio to remove identifiable voice characteristics where the detector accepts redacted input. If the upload happened by mistake, the response is documented: request vendor deletion through the support channel, disclose the upload to the source, and document the chain of custody in the post-publication record.

Thirteen tool cards render the S1 admonition directly. InVID-WeVerify, Hiya Loccus, Reality Defender, FotoForensics, Sensity, Hive AI, Deepware Scanner, TrueMedia / Georgetown, TruFor, GeoSpy, OpenAI Whisper in hosted mode, Google Cloud Translation on sensitive content and Google Pinpoint on sensitive documents all carry the admonition. The offline alternatives sit at 1B.4 with Sherloq and ExifTool as the standard routes.

The editorial-position statement on S1 is the one that anchors the rest of the section: a workflow that produces a verification at the cost of exposing a source is not a verification. It is a security incident with a verification result attached. S1 is the most common place where the choice between those two outcomes is made.

S2 – State-linked or legally sensitive investigation

S2 binds on verification work where the topic brings state authority into the case independently of how the work is done. Content concerning police, military, monarchy, ruling party, cybercrime or fake-news laws, protest movements, red-tagging, national security or authoritarian contexts: the trigger is the subject, not the artefact.

The threat model varies country by country. In Laos, Decree 327 makes online criticism of the government and the party prosecutable; the country-legal-context page records the operational implications. Thailand's Article 112 produces the sharpest S2 surface in the toolkit on monarchy-adjacent material. The Philippine anti-terror frame interleaves with red-tagging and converts online monitoring into offline danger. In Sri Lanka, the OSA enforcement pattern and the PTA legacy bind in parallel. Malaysia carries layered exposure across CMA Section 233 oscillation, the Online Safety Act 2025, and the 3R enforcement frame. In Indonesia, UU ITE risk is reduced after the April 2025 Constitutional Court rulings but remains operational on personal-identity material involving named officials.

Mitigation is consultative, not protocol-driven. The required action is to consult Digital Safety, legal counsel or the editor before outreach, publication, platform reporting or contact with state-linked actors. The S2 routing is binary at first-line triage (it fires or it does not), and when it fires the editor-approval gate is mandatory. The country pages carry the per-jurisdiction operational guidance, the country-legal-context page carries the comparative layer, and the threat-models page reads the surveillance-state actor pattern across the six jurisdictions.

Sebenarnya AIFA, Cofact Thailand, Information Tracer, Maltego and Media Cloud carry the S2 admonition directly. The wider tool ecosystem references S2 in prose where regional case work touches state-linked actors. Individual cards render the admonition sparingly because the routing is dominantly topic-driven; the systematic-firing patterns sit at the country-page and country-legal-context layer.

The editorial-position statement on S2 is that the verifier's legal exposure is a workflow input, not a workflow output. A verification that has to be unpublished, taken down or defended against state action after the fact has failed at the source-protection layer earlier in the chain. The consultative mitigation is upstream of the verification work; it is not a post-hoc legal check.

S3 – Graphic, sexual, child-safety, or abuse material

S3 binds on content including sexual imagery, minors, abuse, corpses, graphic violence, torture or non-consensual intimate imagery. The trigger is the nature of the content; the response is at the newsroom protocol layer, not at the tool-card layer.

The threat model has two parts. One is the subject of the harmful content, whose dignity is implicated by every viewing and every replication. The other is the verifier, whose well-being is implicated by repeated exposure to material that would harm any viewer. The first is a moral question; the second is an operational one. S10 (staff safety and trauma) sits adjacent and routes to the same newsroom-protocol layer for the operational side.

Mitigation runs through newsroom and legal protocol, not through tool-side configuration. Do not upload to general verification tools. Follow newsroom protocol, legal counsel and platform protocols. Minimise viewing. Protect staff well-being. Preserve only what policy allows. The specific tool-level admonition currently sits on Auto Archiver for the cross-jurisdiction archive route on conflict-and-violence footage; the wider 1B / 1C image-and-video detector cards reference S3 in prose where conflict and violence verification is in scope.

The editorial-position statement on S3 is workflow-level: the toolkit does not provide tool-side mitigations that substitute for newsroom protocol. The right place for an S3 response is the standing-operating-procedure document the newsroom maintains, with named editors authorised to view and named protocols for handover. Tool-side checklists are useful at the boundary; they do not replace the protocol layer. The operational-checklists page carries the pre-upload and pre-publication steps that apply to S3 cases at the workflow boundary.

S4 – Doxxing, harassment, or vulnerable-community targeting

S4 binds on content identifying activists, journalists, ethnic or religious minorities, LGBTQ+ people, migrants, witnesses, alleged criminals or private citizens. The trigger is the identity-exposure surface in the content itself: the verification work risks deepening the original harm if it replicates identifying material in the debunk.

The threat model is that fact-check publications can themselves be vectors of harm. A debunk that identifies the original poster, names the targeted community in a way that surfaces it to new audiences, or repeats identifiable slurs in the verification text propagates the exposure the original content created. The secondary-harm risk is the dominant editorial question on Filipino red-tagging-adjacent cases, Malaysian 3R-coded material, Sri Lankan communal-memory rumour, and Indonesian ethno-religious content.

Mitigation runs through the publication-craft layer. Redact identifiers in the debunk. Avoid repeating slurs or addresses. Assess the retaliation-risk profile on named targets. Consider a quiet response or a tipline-only response in place of a public debunk where the public-debunk amplifies harm. The decision is editorial, not procedural; the toolkit's role is to make the question visible at the pre-publication stage.

Information Tracer, Maltego, Sinar Project iMAP and CIB Mango Tree carry the S4 admonition directly. The coordinated-operation tools surface it because the attribution work itself can produce identifying material the verification then has to handle responsibly. The wider Pillar 1 institutional analysis layer carries the S4 framing in prose.

The editorial-position statement on S4 is that the public-debunk format is not the only response option. A quiet response, a tipline-only response or a partner-mediated response is the right route on cases where public amplification would deepen harm. The toolkit's T7 tipline routing tree carries the response-routing logic.

S5 – Private or encrypted group collection

S5 binds on evidence collected from WhatsApp, LINE, private Telegram, closed Facebook Groups or private Messenger / Viber chats. The trigger is the platform-of-origin and the consent question: the content came from a space that has membership conditions, and the membership conditions imply a privacy expectation the verification work has to respect.

Two threat-model questions sit behind S5. One is the source pool: scraping or infiltrating a private group exposes the people in the group to identification, which the group's privacy expectation says they did not consent to. The second is legal: in jurisdictions where unauthorised access provisions apply to private digital spaces, the act of infiltration is itself prosecutable independently of what the verifier does with the content. Thailand's Computer Crime Act has been used in cases adjacent to this surface. In Malaysia, the Malaysiakini CMS-access incident shows the regulator-side dimension of the same problem applied to newsroom systems.

Mitigation runs through consent and the tipline route. Use consented submissions and tiplines only. Do not scrape, infiltrate or expose group members without an explicit organisational protocol. The tipline cards already carry this discipline in their workflow text; the S5 admonition reinforces it. On a politically sensitive WhatsApp source the S5 routing combines with S1 (the source identifier travels in the file) and S2 (the topic is state-linked) to produce the standing posture of a high-S country case.

Meedan Check, Cofact Thailand, MAFINDO Kalimasada and Sebenarnya AIFA carry the S5 admonition. Tipline architecture is the operational substitute for infiltration. The platform pages on WhatsApp, LINE and Telegram carry the platform-specific S5 patterns.

The editorial-position statement on S5 is that the consent boundary is non-negotiable. Verification work that breaches the consent boundary to access otherwise private material is not legitimate verification work in the toolkit's framing, even where the content reached is materially newsworthy. The right route is the consented tipline submission, or partner-mediated routing through an organisation that holds documented consent.

S6 – Detector-only accusation

S6 is the operational expression of Architectural Anchor 2 (two non-detector signals required for any strong public claim) and Anchor 3 (multi-detector counts as one signal class). It binds whenever the only evidence for an AI-generated, deepfake, voice-clone, bot or coordinated-inauthentic-behaviour claim is one or more automated tool scores.

The threat model is the false-confidence pattern that runs through the detector class. A working fact-checker on deadline gets a 92% AI-generated reading from a detector, treats it as a signal sufficient to publish, and is wrong. The Brawner / "Dark Eagle" case is the worked instance: Hive returned a 79.3% verdict on a related case from the same actor pool, against a vendor 98% headline. The reading was correct in direction but materially below the marketing claim, and the toolkit's content policy on vendor-accuracy wrapping is built against that gap.

Mitigation is editorial, not technical. Stop. Reframe the claim as "tool flagged for review." Seek independent evidence, or escalate to T5 source-disagreement and Anchor-3-reset routing. The S6 admonition is rendered globally on T6 itself instead of on each detector card individually, because it binds on every detector in the toolkit equally and rendering it sixty-five times would dilute the gate. The T5 escalation tree carries the Anchor-3-reset branch that enforces the gate operationally.

The editorial-position statement on S6 is that the detector class is not a publishable signal class on its own. The toolkit's framing is that detectors are useful as one input class among the four pillars (provenance, source-history, behaviour, cautious-detector); they become misleading the moment a verifier treats them as the load-bearing signal. The vendor-accuracy-wrapping discipline in the tool cards is the operational reminder. The country-page worked cases (Doc Willie Ong as the Pillar 1 ladder running with non-detector signals alongside, Anutin / Mauerberger as the provenance-first non-detector case) are the operational expressions.

S7 – Malware, APK, phishing, payment, or ID-harvesting

S7 binds on content directing users to APKs, payment platforms, "registration" forms, WhatsApp or Telegram-routed claims, aid-application platforms, investment claims or credential collection. The trigger is the verification-pathway question: the verifier's ordinary first move (open the link, install the file, register on the platform) is itself the harmful action.

The threat model sits inside the scam-economy operational form. The artefact's verification pathway is weaponised; clicking through the verification route is what the artefact wants the verifier to do. On Indonesian and Malaysian Ramadan-aid scam content, the routing target is a WhatsApp form that collects ID information. Thai voice-scam content routes through a phone-call response the scam wants the verifier to initiate. On Sri Lankan Tamil-stream financial-fraud content, the Telegram.apk distribution pattern documented in Fact Crescendo Sri Lanka's stream installs malware on the verifier's device.

Mitigation runs through the institutional security layer, not through tool-side configuration. Do not click. Do not install on work or personal devices. Preserve the link safely (URL screenshot with timestamp, archive capture targeting the URL without following the link, sandboxed-VM examination if institutional security capacity allows). Escalate as scam or impersonation through the T7 tipline routing tree into platform reporting and through technical or security support. The platform pages on WhatsApp, Telegram and LINE carry the platform-specific S7 patterns.

S7 fires at the T7 platform-escalation step in T6's architecture. The threat-models page carries the scam-economy actor-pattern reading. The country pages record the documented regional patterns (Indonesia and Malaysia for deepfake-aid scams, Thailand for voice-scams, Sri Lanka for the Tamil-stream financial-fraud and Telegram.apk distribution, the Philippines for impersonation funnels).

The editorial-position statement on S7 is that verification capacity and security capacity are not the same competence. A working desk's standard pre-upload protocol does not include link-click sandboxing as a default. The right route is escalation to institutional security capacity or to partner organisations with the sandboxing infrastructure; the toolkit names that boundary directly so a verifier does not improvise security tooling under deadline pressure.

S8 – Liar's-dividend risk

S8 is the editorial twin of S6. Where S6 prevents over-claiming AI from a detector signal, S8 prevents the political abuse of "could be AI" as a way to dismiss real evidence. The trigger is a powerful actor's claim that real evidence is AI-generated, or the verifier's own temptation to label something fake without independent proof.

The threat model is the political-economy form of the detector-uncertainty problem. A piece of real evidence (a recording of corruption, a video of violence, a leaked document) becomes inconvenient. A powerful actor claims the evidence is AI-generated. The detector class is unreliable enough at the field-reading layer that the claim is hard to disprove with a single detector pass. The result is that real evidence loses public weight even where its provenance is solid.

Mitigation is editorial framing. Verify source, context and provenance through non-detector signals. Say "we cannot verify AI manipulation" when evidence is insufficient; do not say "fake." The framing matters because the language a verifier uses propagates into wider discourse; "fake" is what the liar's-dividend claim wanted the verifier to say. The toolkit's framing on uncertainty is that explicit non-verification is the operational form of editorial honesty, and that it does more public-information work than a falsely confident claim in either direction.

S8 binds on every fact-check format the toolkit produces. It is rendered globally on T6, not on individual cards. The editorial position is that the provenance pillar (C2PA Content Credentials, SynthID watermark identification, original-file chain-of-custody) sits structurally above the detector pillar precisely because the detector pillar is vulnerable to liar's-dividend manipulation. The Thailand country page Anutin / Mauerberger SynthID case is the worked example of provenance-first verification that closes the question more cleanly than a detector pass would.

The editorial-position statement on S8 is that the toolkit's response to liar's-dividend pressure is documentary, not rhetorical. Archive the original. Surface the provenance signal. Cite the source-history pillar's non-detector signals. The detector class is the last resort, and a public uncertainty statement is better than a manufactured certainty.

S9 – Cross-border data transfer or vendor retention

S9 binds whenever hosted tools require upload of sensitive content to proprietary APIs or foreign servers. The trigger is the vendor-jurisdiction question: the file leaves the verifier's machine, transits to a server in a jurisdiction the source did not consent to, and is retained for some period there.

Two layers sit behind the S9 threat model. One is data residency: the file is on someone else's server, and the laws of that jurisdiction govern access requests, retention rules and disclosure protocols. The other is vendor-retention specifics: the documented retention window for the file before deletion. The InVID-WeVerify deepfake tab, which ships frames to CERTH with a 30-day retention window, is the most explicitly documented vendor-retention policy among the toolkit's reviewed tools; the toolkit treats it as the reference case for naming retention windows when they exist.

The vendor-jurisdiction map for the tools in the toolkit is asymmetric. US-jurisdiction servers carry Hive AI, Reality Defender, Hiya Loccus, TrueMedia / Georgetown, Sensity, Deepware Scanner, Resemble, Pindrop, Winston, Copyleaks, Logically and Full Fact AI. EU-jurisdiction servers carry CERTH (via the InVID-WeVerify deepfake tab) and the vera.ai consortium. The asymmetry matters operationally: a Lao source file uploaded to a US-hosted detector is subject to US-jurisdiction subpoena and disclosure rules; the same file uploaded to a vera.ai EU-hosted detector is subject to the EU data-protection framework. Neither resolves the source's exposure question, but the legal-and-disclosure characteristics differ.

Mitigation runs through the data-policy layer. Check organisational data policy. Check vendor retention terms. Check source consent. Prefer local or on-device tools, or trusted expert channels, on identifying material. The S9 admonition is rendered on InVID-WeVerify (the 30-day CERTH window reference case), Hive AI, Reality Defender, Hiya Loccus, TrueMedia / Georgetown, Sensity and Deepware Scanner.

The editorial-position statement on S9 is the access-barrier framing the toolkit carries across the wider methodology layer. Cloud-hosted detector tools sit at progressively higher access-barriers (institutional pricing on Sensity, enterprise-only deployment on Graphika, IFCN-gating on Meta Content Library), and the data-jurisdiction question runs alongside the access question. The toolkit handles the gradient through the repeated "cite as external source" redirect pattern (a Graphika-pattern Quickstart redirect) for tools that are structurally inaccessible to the toolkit's audience. The same gradient applies to S9: data jurisdiction is part of access-barrier framing, not a separate technical question. The methodology editorial-patterns page carries the integrated framing.

S10 – Staff safety and trauma

S10 binds on repeated viewing or listening of harmful content, harassment backlash, threats, graphic content or coordinated abuse after publication. The trigger is the verifier's well-being, not the content's properties or the source's exposure surface.

The threat model is operational. Detector and forensics work on conflict footage requires repeated viewing of material that would harm any viewer; the verifier sees it many times in a way the public sees it once. Audio-clone verification on violent-threat material runs through the same exposure pattern. Coordinated harassment after publication on red-tagging-adjacent cases or 3R-coded debunks targets the verifier specifically. The Tempo pig's-head and decapitated-rats threats documented by CPJ in March 2025, the Andrie Yunus March 2026 acid attack, and the red-tagging patterns recorded in CMFR / NUJP data through April 2025 are the operational anchors.

Mitigation is workflow-level, not tool-level. Limit exposure. Rotate reviewers. Document threats. Secure accounts and devices. Activate the newsroom safety protocol. S10 lives in the Digital Safety section as a workflow-level override; it does not render as a per-card admonition because the response sits at the desk-and-newsroom layer, not at the per-tool layer. 1C and 2C tool cards that require repeated viewing of harmful content cross-link to this section's S10 entry.

The editorial-position statement on S10 is that staff well-being is a verification-workflow input, not an output. A desk that produces high-quality verification at the cost of trauma-loading its reviewers has failed at the workflow design. The right route is rotation, exposure caps, peer-support infrastructure and the newsroom safety protocol's standing operation, with the operational-checklists page carrying the post-publication monitoring steps that apply at the boundary.

Reading the ten classes together: practice on a high-S country case

A working illustration ties the framing back to operational practice. A Lao political video reaches a regional desk via diaspora contact. The video shows a named state critic posting from inside Laos; the audio is in Lao; the artefact has been forwarded through a WhatsApp group with twenty members; the case is being handled through a Bangkok-based regional outlet because the diaspora reporter does not want their name on the verification.

The S-classes fire in combination. S1 binds on the named state critic's face and voice in the artefact. S2 fires on the topic, since Lao political criticism sits inside Decree 327's prosecutable surface. S5 fires on the WhatsApp-group collection method. S9 binds because any cloud-hosted detector pass would route the source file to a US-or-EU jurisdiction server. S10 sits on both the diaspora reporter and the regional outlet's verifier, who carry post-publication retaliation exposure between them. The editorial position is binding on every layer.

The mitigation workflow is sequential. Archive the artefact to a cross-jurisdiction location through auto-archiver before any tool runs. Strip identifying context on the local machine: crop frames, redact audio where possible, document the redaction chain. Route the verification through offline tools on the Bangkok side, with Sherloq for image-and-metadata work and SEA-LION v4 locally deployable for any Lao-language LLM-assisted analysis. Use the Lao country page diaspora-and-regional-partner routing pattern for the editorial weight on publication. Publish from the Bangkok outlet's byline with the diaspora reporter not named. Activate post-publication monitoring through the operational-checklists page S10 routing for both the diaspora reporter and the Bangkok verifier.

The case is the operational expression of the editorial system the ten S-classes constitute. T6 routes the case in real time; this page is the layer that lets a desk understand why the routing is what it is.

Cross-references

Sources

  • T6 source-protection tree (the canonical routing reference)
  • Tool cards rendering S1 through S9 admonitions (per-class card listings on T6)
  • Architectural Anchors – Architectural Anchors 2 (two non-detector signals) and 3 (multi-detector wrapped as one class); vendor accuracy wrapping and honest-gap policy
  • DW Innovation. Synthetic Audio Detectors Put to the Test. Deutsche Welle, 29 September 2025. innovation.dw.com/articles/synthetic-audio-detectors-tested.
  • Electronic Frontier Foundation. Surveillance Self-Defense. EFF, 2025. ssd.eff.org.
  • Country pages (Indonesia, Laos, Malaysia, Philippines, Sri Lanka, Thailand) – regional case anchors operationalising the S-classes