How we chose tools¶
The toolkit ships with sixty-five tools across twenty-two outline subcategories. The discovery space that produced those sixty-five was an inventory of more than two hundred and fifty candidates compiled across six parallel research streams, covering Sri Lanka's Sinhala / Tamil information environment, independent accuracy benchmarks across detector classes, regulatory updates across the six focus countries, and field reconnaissance on tool categories. The selection that produced sixty-five from that base is what this page narrates.
The framing problem behind the selection is worth stating before any criteria. The toolkit is an operational triaging-and-counter-disinformation reference for fact-checkers, OSINT practitioners, and civil-society staff in six focus countries: Indonesia, Laos, Malaysia, Philippines, Sri Lanka, and Thailand. It is not a tool catalogue. It is not a vendor directory. It is not an academic survey. A reader who opens the toolkit has a case in hand and needs a verifiable handle on how to triage, verify, or counter the disinformation in that case under field conditions that include surveillance risk, low bandwidth, mobile-only access, and the operator-identity considerations particular to each of the six jurisdictions.
That framing problem dictates the selection. A tool that is technically interesting but operationally unusable by the toolkit's audience is not in the shortlist. A tool that is widely deployed but carries unwrapped vendor claims is not in the shortlist. A tool whose limitations have been softened in vendor documentation past the point of editorial honesty is not in the shortlist. A tool that would be operationally useful but does not exist for the relevant language or region surfaces as a named gap, not as a substituted alternative. The toolkit's relevance is conditional on these constraints, and the selection's defensibility depends on naming them.
The seven criterion groups¶
Selection runs through a seven-group framework. A condensed evaluative-rubric version for external readers sits on the selection-criteria page. The seven groups are:
Group A – Relevance to outline. The tool maps to a primary outline subcategory at a defensible tier and serves a user need named in the five-lens critique (fact-checker on deadline, OSINT methodologist, disinformation researcher, newsroom editor, regional civil-society organiser). A1 (outline mapping) and A2 (tier appropriateness) are pass/fail gates. A3 (verified user need) weights ranking within a cell.
Group B – SEA fitness. UI accessibility for trained SEA fact-checkers (English-only is sufficient under Decision 7); content language coverage for tools whose core function is natural language processing; language-agnostic declaration for tools whose function is not language-dependent (image, video, network, metadata, provenance); documented regional use as a bonus signal. The seven SEA languages the toolkit tracks are Bahasa Indonesia, Lao, Malay, Filipino / Tagalog, Sinhala, Tamil, and Thai.
Group C – Reliability and honesty. Vendor accuracy wrapping (C1, paired in the Independent Accuracy admonition on every detector card); limitation faithfulness (C2, verbatim from primary sources, not softened); active maintenance evidence within the last eighteen months (C3); no documented hallucination patterns or systematic regional content failure that would invalidate the tool's output, or where such failure is documented, an explicit caveat on the card (C4).
Group D – Accessibility. Cost realistic for SEA newsroom budgets, scored on a one-to-five scale where five is free with no quota and one is enterprise pricing in the fifty-thousand to two-hundred-thousand US dollar per year range; bandwidth and device feasibility relative to tier; install and access friction appropriate to tier; no critical privacy risks for restrictive media environments without a documented mitigation pathway.
Group E – Anti-criteria. Five exclusion rules: vendor claims unvalidated and high false-positive rate (E1); abandoned without responsive maintainer (E2); critical privacy risk with no mitigation (E3); demonstrable failure on regional content without weak-signal framing (E4); better alternative in the same outline subcategory (E5). Group E runs as a gate before the rest of the framework. A tool that triggers any of E1 through E5 is auto-excluded unless an explicit override is documented on the tool card.
Group F – Tier-specific functional. Each tier has a different functional bar tied to the canonical five-thirty-one-hundred-twenty-minute ladder. First-Line Triage tools must be usable in five minutes on a phone with no install or trivial account creation; Professional Verification tools must be usable in thirty minutes at desk with browser-extension or API access acceptable; Institutional-Level Analysis tools may require two-hour-plus workflows, paper-backed methodology, and evidence-grade output.
Group G – Architectural alignment. The three anchors documented on the architectural-anchors page govern the toolkit's framing of the detector class and the signal architecture. G1 covers pillar service across provenance, source-history, behaviour, and cautious-detector; detector-only tools must carry the detector-as-weak-signal caveat. G2 covers signal class declaration on every tool card. G3 covers multi-detector wrapping under Anchor 3, where a single product aggregating multiple detection models is counted as one signal class, not as a multiplier of signals.
A candidate runs through Group A first as a gate, then Group E as an exclusion gate, then Groups B, C, D, and F as scoring and tier checks, then Group G as the architectural-alignment review. Most candidates exit at Group A (no outline match) or Group E (anti-criterion triggered). The candidates that survive are scored, ranked within their outline subcategory, and selected per the density rules in Decision 8.
The architectural-anchors page sets out the three anchors as operational logic the selection encodes. The editorial-patterns page codifies the five recurring patterns the selection produces inside the shipped cards. The two pages run alongside the criterion framework: the criteria are the gate that lets a tool through; the anchors are the discipline the tool ships under; the patterns are the editorial response the card prose embeds.
The seven content policies¶
Five content policies were locked at the end of the initial scoping work. Two more emerged as operational patterns through the cards themselves. The seven together are non-negotiable for any tool card the toolkit ships.
The honest-gap policy for Lao and Sinhala / Tamil is the first. Languages are never masked by fallback to multilingual models. Every detection section carries an explicit block stating what coverage exists for Lao and for Sinhala / Tamil specifically, even where the answer is "none, falls back to XLM-R via Alegre, validate locally before trusting." No silent inheritance. The pairing between Lao and Sinhala / Tamil is intentional; Sri Lankan readers are not held to a higher coverage standard than Lao readers.
The vendor accuracy wrapping rule is the second. Any accuracy claim in a tool card appears as a pair: vendor X percent against independent Y percent, or an explicit "no independent SEA-specific benchmark identified as of [date]" sentence with the vendor's headline recorded for transparency. Without the pair or the no-independent-benchmark sentence, the tool does not enter the toolkit. The Sensity 3-leg wrapping on the editorial-patterns page extends C1 into three legs (vendor headline plus field reading plus category baseline) where the field-reading evidence exists.
The limitation faithfulness rule is the third. The Limitations field on every tool card uses the source's specific language verbatim, or a paraphrase that does not soften. If the source says "fails on Asian faces," the card says exactly that. If the source says "thirty-two percent false-negative rate on compressed audio," that exact figure goes in. If two sources give different numbers (the GPTZero Stanford 2023 sixty-one percent non-native FPR against Perkins 2024's twenty-six and four-tenths percent), both numbers stay; merging is itself a failure mode.
The SEA-language coverage policy is the fourth, applied conditionally per Decision 7. A tool without explicit support for at least one of the seven SEA languages either does not enter the toolkit or enters with an explicit "English / global-only – validate before regional use" disclaimer. Decision 7 carved out the conditionality: language coverage applies to tools whose core function is natural-language processing, not to image, video, network, metadata, or provenance tools where language is technologically irrelevant.
The reliable-triad emphasis is the fifth. Pillar 1 is tactical: detectors are framed as ad-hoc tools, not solutions. Pillar 2 is strategic: provenance, tipline networks, and manual verification carry the durable load. The toolkit does not pretend detectors solve the problem. Section intros across both pillars reinforce the framing; the pillar-1 narrative opens with the four-pillar workflow, not with a detector inventory.
The source-protection-first policy is the sixth, codified during Digital Safety drafting. Verification work that produces a verdict at the cost of exposing a source is not verification work; it is a security incident with a verification result attached. The ten S-classes (S1 through S10) operationalise the policy across every tool card with cloud-upload risk, every country page with surveillance-environment legal exposure, and the T6 source-protection decision tree.
The audit-trail policy is the seventh. Override flags rendered as visible card text. Decisions documented as they are taken, not reconstructed after the fact. The audit trail is what makes the toolkit contestable by a partner or reviewer who disagrees with a selection.
The three architectural anchors¶
The three anchors govern the toolkit's framing of the detector class and the signal architecture. The architectural-anchors page sets them out as embedded operational logic in their own right. In selection terms, they are the discipline the candidate must clear at Group G before its card ships.
Anchor 1: the toolkit is not a directory of AI detectors. The workflow runs on four pillars – provenance, source-history, behaviour, and cautious-detector – and the detector pillar is the last resort, not the first move. Every tool card identifies which pillar or pillars it serves; detector-only tools carry the detector-as-weak-signal caveat explicitly. The selection enforces the anchor at Group G1.
Anchor 2: at least two non-detector signals are required before any strong public claim. This is editorial policy, not advisory. Every decision tree's terminal "publish" node enforces it. Every tool card identifies whether its output is a detector signal or a non-detector signal so the editorial layer can apply Anchor 2 operationally. The selection enforces the anchor at G2.
Anchor 3: detector plus detector is one signal class, not multiple. Running an image through three detectors and getting three "AI" verdicts counts as one signal, not three. The toolkit hard-codes this against false confidence from multi-detector consensus. Multi-detector products (Sensity's multilayer analysis, Reality Defender's massive ensembles, Hive's multimodal stack) count once each. The selection enforces the anchor at G3.
The detector-to-non-detector ratio in the shipped sixty-five is approximately twelve percent pure detector and eighty-eight percent non-detector or mixed-with-declaration signals. The ratio is by design. The toolkit's centre of gravity is non-detector work: tiplines, claim-extraction, transcription, reverse-image, network analysis, frameworks, prebunking, provenance verification. The detector class is wrapped, framed, and consistently positioned as one input class among the four pillars.
The eight foundational decisions¶
Eight decisions were locked as the structural decisions the rest of the toolkit applies. Each is briefly named here for context; the change log page records the rationale for each.
Decision 1 – Canonical time-band ladder of five, thirty, and one hundred and twenty minutes. First-Line Triage is what a fact-checker can do in five minutes on a phone. Professional Verification is what is feasible in thirty minutes at desk with browser tools. Institutional-Level Analysis is two hours minimum, often half a day. The ladder anchors tier appropriateness at A2 and tier-specific functional checks at Group F.
Decision 2 – ABCDE / DISARM split. ABCDE is the surface analytical vocabulary for public-facing work. DISARM is the institutional annex for inter-organisation reporting and researcher use. Both ship in the toolkit, in clearly distinct contexts. 2C.3 cards carry the framework-reference register the split produces.
Decision 3 – Inter-tool conflict resolution. Lives in two places: a Conflict-resolution behaviour field on every detector and mixed-signal tool card stating "if this tool disagrees with another, its evidence weight is X because Y"; and a dedicated T5 escalation tree node handling "two tools disagree."
Decision 4 – Regional suffix scheme. Country suffixes (-id, -ms, -thai, -ph, -si-ta, -lao) and platform suffixes (-wa, -line, -tiktok, -fb, -fbg, -telegram, -youtube, -x) used across decision trees, tool cards, regional case studies, and edge-case mappings. Carries operational specificity into the cross-link layer.
Decision 5 – Consolidated gap catalogue as single gap source. All gap-driven outline additions and drafting priorities trace to the consolidated gap research, not to any earlier separate gap lists.
Decision 6 – Detector skepticism hard-coded at three levels: architectural (Anchor 1), editorial policy (Anchor 2), and operational safety override (the S1-through-S10 system that fires regardless of branch logic). Soft positioning is rejected. A fact-checker on deadline will use a detector verdict if not told plainly it is insufficient, and the toolkit tells the reader plainly across every layer.
Decision 7 – SEA language coverage applies conditionally. UI accessibility is English-sufficient for trained fact-checkers (with one named exception, Sebenarnya AIFA, for public-facing lay use). Content language coverage is mandatory only for NLP-class tools. Image, video, provenance, network, and metadata tools are language-agnostic. Documented regional use is a bonus, not a gate. The honest-gap policy applies to content the reader needs to verify, not to tools' UI.
Decision 8 – Density target of two to three tools per outline subcategory by default, three to four where the landscape is genuinely rich (CIB with thirty-eight inventory entries; claim-extraction with forty-two), one to two where the landscape is genuinely thin (Lao-capable tools; audio detection in SEA languages). Resulting target range of sixty-five to eighty-five tools across the twenty-two cells; final shortlist of sixty-five emerges from the application of density rules, not as a pre-set number to optimise toward.
The decisions interact. Decision 1 (time-band ladder) and Decision 8 (density) interact at every cell. Decision 6 (detector skepticism) and Decision 7 (conditional language coverage) interact across the audio-detection cells. Decision 2 (ABCDE / DISARM split) and Decision 4 (regional suffix scheme) interact across the cross-link layer of the 2C.3 framework references. The interactions are part of what makes the framework operational at the cell level and not just at the corpus level.
What was excluded and why¶
A toolkit shaped by what it excludes carries different weight than a toolkit shaped only by what it includes. The exclusions break into four classes worth naming explicitly.
The first exclusion class is anti-criterion triggers. Roughly thirty candidates exited via E5 (better alternative in the same outline subcategory) across cells where the landscape is dense and a better-evidenced tool sat in the same role. Four candidates exited via E1 (vendor claims unvalidated against documented high false-positive rate): Illuminarty (sixty-seven and four-tenths percent false positive on human art per independent benchmark); Copyleaks (fifty percent false positive on small human-control samples per independent benchmark); Intel FakeCatcher (vendor ninety-six percent with no independent test); Smodin (vendor-only multilingual claim with no independent SEA validation). E2 triggered on TrueMedia.org at the first pass (shut down 14 January 2025) but a Georgetown McCourt School revival was confirmed in May 2026 and the tool re-entered the shortlist at 1B.3 as escalation-option with the closed-beta disclaimer.
The second exclusion class is structural unsuitability. Generic newsroom-productivity AI tools with no AI-disinformation use exited at A1 (no outline match). Academic detector papers without deployable artefacts exited at A2 (tier inappropriate; or the artefact was a benchmark dataset and not a tool). Directory-type entries (Bellingcat OSINT Toolkit, OSINT Framework, osintmap) exited as indexes, not tools.
The third exclusion class is regional unsuitability. Tools whose only documented evaluation showed failure on regional content without weak-signal framing exited at E4. Tools whose vendor "multilingual" marketing claim was contradicted by independent audit — Hiya / Loccus on the DW Innovation November 2025 cross-language reliability audit; AASIST / RawGAT-ST with cross-language performance to Bahasa / Thai / Tagalog / Sinhala / Tamil / Lao untested — entered the toolkit with the audit verbatim and the detector-as-weak-signal framing binding, instead of entering on the marketing claim alone.
The fourth exclusion class is structural absence. Tool categories that the discovery space simply does not contain at field-deployable quality: Lao-language voice-clone detection; Lao-first independent fact-checker; pre-trained SEA-language news-reliability scoring for Sinhala and Lao; Tamil claim-extraction adapted to the Sri Lankan corpus. These exit not as exclusions but as named gaps. The honest-gap pattern is what the toolkit ships in place of the missing tool, and the relevant country pages and cell pins record the operational implications.
The four exclusion classes together account for the gap between the inventory of more than two hundred and fifty and the shortlist of sixty-five. Six tools appear in multiple cells under the cross-cell scheme, which produces seventy-two per-cell entries from the sixty-five unique tools. The cross-cell tools are InVID-WeVerify across 1A.2, 1B.1, and 2A.2; Hive AI across 1A.1 and 1A.3; ExifTool across 1B.1 and 1B.4; Sherloq across 1B.1 and 1B.4; Content Credentials Verify across 1A.1 and 1A.4; Google Pinpoint across 2A.1 and 2B.3. The cross-cell pattern reflects the multi-pillar service profile of those six tools; each is a non-detector multi-function tool that covers more than one outline subcategory at the cell level.
What the selection produces¶
The shipped sixty-five tools sit inside the editorial system the editorial-patterns page codifies. The selection process described above is the gate; the cards themselves are where the discipline shows. A reader who walks through any of the sixty-five cards encounters the seven content policies in operation: vendor wrapping in the Independent Accuracy admonition, verbatim limitations in the Limitations admonition, source-protection routing in the Privacy and threat model section, country and platform applicability per Decision 4, conflict-resolution behaviour on every detector and mixed-signal card.
The selection is contestable. A partner who reads the selection criteria, the architectural anchors, and the editorial patterns together has the operational basis to argue for the inclusion of a tool the toolkit excluded, or for the exclusion of a tool the toolkit included. The change log records the structural decisions that shaped the toolkit's current form. The contestability is part of what makes the toolkit a living document and not a single editorial pronouncement. New tools enter the discovery space every month; the framework above is what the next maintainer applies to decide whether they enter the shortlist.
The selection is also auditable. A reader who wants to reconstruct why a given tool entered or did not enter the shortlist can walk the seven criterion groups, seven content policies, three architectural anchors, and eight foundational decisions together — they are the codified record of what the selection process applied. The current shortlist is the product of that framework applied across the inventory.
The selection is, finally, regional. The six focus countries are not interchangeable. Indonesia has the most developed civil-society fact-check ecosystem in the region; Laos has the thinnest tool landscape and the most restrictive legal environment; Malaysia has the most layered state and civil-society co-existence pattern; the Philippines has the most institutionalised fact-check coalition (#FactsFirstPH); Sri Lanka has the most asymmetric language ecosystem (Sinhala stronger than Tamil at the tool layer); Thailand has the most platform-specific tool landscape (LINE-centric). The selection respects those differences. A Lao-relevant tool that would be sub-par in a Philippine context can still ship at 2B.3 with the honest-gap framing applied. A Philippine-specific tool that would be irrelevant in Laos still ships at 2B.1 with the country-applicability table making the regional scope explicit. The toolkit's regional specificity is its main differentiator from globally-framed counter-disinformation references, and the selection is what produces that specificity at the tool-card layer.