Skip to content

Selection criteria

The seven criterion groups on this page are the operational rubric the toolkit applies to candidate tools. This page condenses the framework for an external reader who has a new tool in attention and wants a structured first read on whether it belongs in their workflow.

The rubric is evaluative, not descriptive. A descriptive checklist would say "the toolkit prefers free tools, multilingual interfaces, and active maintenance." That is true but uselessly thin. An evaluative rubric says how the toolkit weighs evidence: which sub-criteria are pass / fail gates, which are scoring, which are weightings, which are exclusion rules, which carry caveat-pass or mitigation-pass routes. The condensed version below preserves the gate-then-score-then-tier-then-architecture sequence the full framework runs. A reader who walks a new tool through the seven groups in that sequence produces a structured assessment that can be defended against another reader's read on the same tool.

Two reader profiles use this page. A working fact-checker who hears about a new SEA-language deepfake detector at a regional convening wants a fifteen-minute structured assessment to decide whether to spend an afternoon piloting it. A methodology partner who is extending the toolkit to a different region or use case wants the rubric as a basis they can adapt to their own context. Both profiles share the same operational handle: apply the seven groups in sequence, produce a verdict, document the override flags for the verdict's caveat-pass or mitigation-pass routes.

Group A – Relevance to outline

The toolkit's value is judged against its approved outline. Tools that are interesting but unmatched to a subcategory belong in a maintenance backlog, not in the shortlist.

A1 maps the candidate to a primary IPMR outline subcategory. Pass / fail. Without a primary mapping the tool cannot enter the shortlist. Secondary mappings are encouraged but never substitute for a primary mapping. A2 checks tier appropriateness against the canonical five-thirty-one-hundred-twenty-minute ladder. A GPU-only model recommended at First-Line Triage fails; the same tool may pass at Institutional-Level Analysis. A3 weights ranking based on whether the tool serves a need named in the five-lens critique (fact-checker on deadline, OSINT methodologist, disinformation researcher, newsroom editor, regional civil-society organiser) or in a Critical / High gap catalogue entry. High weight if named in the five-lens critique or a Critical / High gap; medium if named in a HighxMedium or MediumxHigh gap; low otherwise.

A working application: an external reader encounters a new browser-extension reverse-image search tool. A1 maps it to 2A.2 (reverse-image / geolocation enhanced) as a primary candidate. A2 checks tier appropriateness: browser extension passes at Professional Verification under the thirty-minute desk-tool framing. A3 weights it as medium-to-high if it pairs with InVID-WeVerify's reverse-image module on the source-history pillar. Group A produces a verdict of "passable for further evaluation" but not yet a shortlist place.

Group B – SEA fitness (per Decision 7)

Decision 7 ended the blanket "must support a SEA language" rule the toolkit started with. Language coverage applies conditionally to what the tool processes.

B1 covers UI accessibility. The tool's interface is in a language a trained SEA fact-checker can use; English-only UI passes by default for the trained audience. One named exception: tools intended for public-facing lay use (a citizen-facing tipline chatbot in Malaysia, for example) require local UI for the country they target. B2 covers content-language coverage and applies only to NLP-class tools whose core function is to process natural language: AI-text detection, voice-clone detection, NLP-based claim extraction, transcription, translation, sentiment and hate-speech classification. For these, the tool must document support for at least one of the seven SEA languages (Bahasa Indonesia, Lao, Malay, Filipino / Tagalog, Sinhala, Tamil, Thai) or enter with an explicit "English / global-only – validate before regional use" disclaimer. B3 declares language-agnostic status for tools whose function is not language-dependent: image detection, video deepfake detection, content provenance, reverse-image search, network and CIB analysis, archiving, metadata forensics. B2 does not apply; the auditor marks B2 as "N / A – language-agnostic" in the tool card. B4 is a one-to-three scoring bonus on documented regional use: 1 = no documented SEA use; 2 = documented SEA use in one or two countries; 3 = documented across three or more focus countries or named in a SEA fact-checker workflow recommendation.

A working application: the same browser-extension reverse-image search tool from the Group A example. B1 passes (English UI sufficient for the trained audience). B2 is N / A because reverse-image search is language-agnostic per B3. B4 scores 2 if the tool has documented use in two of the six focus countries or 1 if it has none. The verdict carries forward without language-related blockers.

A framework-level commitment sits alongside the four sub-criteria: paired Lao / Sinhala / Tamil honest gap. Where Lao or Sinhala / Tamil coverage is structurally absent in a cell, the toolkit names the absence instead of masking it through multilingual fallback. The commitment applies to the toolkit's editorial layer (section intros, country pages, decision-tree leaves), not to individual tool eligibility. See editorial-patterns.md Pattern 5.

Group C – Reliability and honesty

Reliability is editorial, not technical. A tool can be unreliable in field conditions and still pass the toolkit's reliability bar if its limitations are faithfully represented and its claims are wrapped. The bar exists because IPMR readers will take the toolkit's framing as load-bearing.

C1 covers vendor accuracy wrapping. Any numeric accuracy claim on the tool card appears paired: "vendor X percent / independent Y percent" or an explicit "no independent SEA-specific benchmark identified as of [date]" sentence with the vendor's headline recorded for transparency. No pair, no entry. The editorial-patterns page extends C1 into the three-leg Sensity pattern where field-reading evidence exists; the third leg (category baseline) is recommended where Deepfake-Eval-2024 or equivalent pool-level evidence is available. C2 covers limitation faithfulness. The Limitations field on the tool card uses the source's specific language verbatim or a paraphrase that does not soften; if the source says "fails on Asian faces," the card says exactly that. Where two sources give different numbers, both stay; merging is itself a failure mode. C3 covers active maintenance evidence within the last eighteen months: a release, repo commit, vendor changelog, or clear support channel. Status statements like "actively maintained" without evidence are insufficient. C4 covers documented hallucination patterns or systematic regional content failure: a tool with such failure passes only if a faithful caveat accompanies the recommendation. Without the caveat, the tool is excluded.

A working application: a new AI-text detector claims ninety-five percent accuracy on a vendor-curated benchmark. C1 fails unless an independent test is locatable; the toolkit documents the Stanford 2023 sixty-one percent non-native FPR for GPTZero and the Perkins 2024 twenty-six-point-four percent figure for a comparable detector. The reader checking C1 looks for analogous independent evidence on the new detector. If none is locatable, the tool either enters with the "no independent SEA-specific benchmark identified" sentence or does not enter at all. C2 then requires that any limitation the vendor disclosed (or any independent finding the reader located) ships in the Limitations field verbatim. C3 requires an eighteen-month maintenance window. C4 fires if the new detector has a documented systematic failure on regional content; the reader checks the available independent audits for the relevant evidence.

Group D – Accessibility

SEA fact-check desks operate on small budgets, mobile devices, unstable bandwidth, and often inside legal regimes where cloud uploads are themselves a source-protection risk. A tool that works only with a fifty-thousand US dollar licence, a desktop GPU, or unrestricted cross-border data transfer fails the field, regardless of how accurate it is.

D1 scores cost on a one-to-five scale: 5 = free with no quota for ordinary use; 4 = free tier sufficient for non-edge use, paid for surge; 3 = low-cost paid (up to approximately fifty US dollars per seat per month); 2 = mid-cost paid, no NGO path; 1 = enterprise-only at fifty-thousand US dollars or more per year with no documented free or grant access. D2 covers bandwidth and device feasibility relative to tier: Android phone for First-Line Triage, mid-range laptop with intermittent broadband for Professional Verification, desktop with stable connectivity for Institutional. No GPU requirement below Institutional. D3 covers install and access friction relative to tier: First-Line Triage = no install or trivial; Professional Verification = browser extension, account, or API key acceptable; Institutional = full install and technical setup acceptable. D4 covers critical privacy risks for restrictive media environments. The tool does not require uploading source-identifying content to cloud servers, does not transfer data cross-border without disclosure, and does not encode creator identity into outputs without warning. Tools with these risks may still pass if the toolkit binds them with a documented mitigation (local pre-strip, consent rule, local-tools-first instruction). The Sensity card is the worked instance of D4 mitigation-pass via the S1 source-protection sub-routine.

A working application: an enterprise CIB platform marketed at "regional newsrooms" with pricing on application. D1 scores 1 or 2 depending on whether an NGO path is documented; without one, D1 = 1 and the access-barrier framing pattern applies. D2 passes at Institutional tier. D3 passes at Institutional. D4 requires checking whether the platform pre-strips identifying material before processing or whether cloud upload is required. If cloud upload is required, the S1 / S9 mitigation pathway is documented or the tool fails D4.

Group E – Anti-criteria (auto-exclude unless overridden)

Anti-criteria are exclusion rules. A tool that triggers any of these five is excluded from the shortlist by default. Override requires explicit documentation, not a judgement call.

E1 fires on vendor claims that are unvalidated AND show documented high false-positive rates. Either condition alone does not exclude; the conjunction does. E2 fires on abandoned tools: last commit / release more than eighteen months ago with no responsive maintainer for open-source tools; no vendor activity (no changelog, no support response, no published release) for hosted tools. E3 fires on critical privacy / safety risks in the SEA context with no documented mitigation pathway. E4 fires on demonstrable failures on regional content (named SEA language, regional face / voice population, regional script) without weak-signal framing as the override. E5 fires when another tool covers the same use case at the same tier with stronger scores across Groups A through D, and the candidate adds nothing the alternative lacks.

Override format: "Anti-criterion triggered: [Ex]. Override rationale: [reason]. Mitigation: [mitigation binding into the tool card]." Override entries are recorded on the tool card itself, so the rationale ships with the card the reader sees.

The shortlist's worked overrides are instructive. GPTZero is admitted at 1B.2 under E1's "rescued via independent-evidence-already-cited" route because the Stanford sixty-one percent and the Perkins twenty-six-point-four percent independent findings sit alongside the vendor claim; the wrapping pair fills C1. TrueMedia.org is admitted at 1B.3 under E2's "revival confirmed" route because Georgetown McCourt School's May 2026 revival was verified; the closed-beta-disclaimer flag ships in the access section. Sensity is admitted at 1C.2 under E5's "rescued as primary" route because of the Rappler / #FactsFirstPH documented deployment and the Deepfake-Eval-2024 anonymised-pool position. Graphika is admitted at 1C.1 under E5's "rescued as escalation-option" route because of unique multi-platform coverage breadth, with the frontline-no-access disclaimer rendering the access boundary visible.

Group F – Tier-specific functional

Each tier has a different functional bar tied to the canonical time-band ladder.

F1 covers First-Line Triage tools. Must be usable in five minutes on a phone, with no install and no account (or trivially fast account creation), and free or freemium-sufficient for non-edge use. F2 covers Professional Verification tools. Must be usable in thirty minutes at desk; browser extension or API access acceptable; account or API key acceptable; integrates with existing verification workflow patterns. F3 covers Institutional-Level Analysis tools. Two-hour-plus workflows acceptable; requires technical capacity (Python, ML, data engineering); paper-backed or peer-reviewed methodology; evidence-grade or audit-trail output.

A tool can pass F at one tier and fail at another; that re-tiers the tool, it does not exclude it. The reader who is evaluating a new tool should apply F at the proposed tier first; if the tool fails F at the proposed tier but passes at a different one, the verdict is "rebase to the tier where F passes," not "fail outright."

Group G – Architectural alignment

The three architectural anchors govern the toolkit's framing of the detector class and the signal architecture. The architectural-anchors page sets them out in operational detail. Group G is the per-tool check that the candidate complies with the anchors.

G1 covers Anchor 1: the tool serves at least one of the four pillars (provenance, source-history, behaviour, cautious-detector). Detector-only tools must carry the detector-as-weak-signal caveat explicitly on the card. Without the caveat the tool fails G1. G2 covers Anchor 2: the tool's output is classifiable as either a detector signal or a non-detector signal, and the classification is explicit on the card. Tools with mixed output declare which class predominates and how the uncertainty is presented to the reader. G3 covers Anchor 3: multi-detector products are scored as one signal class, not as a multiplier of detector signals. Declarative bookkeeping on the card; the workflow section names the rule.

A working application: a new multimodal AI detector that processes image, audio, and video together and returns three confidence scores. G1 requires the detector-as-weak-signal caveat. G2 declares the output as a detector signal class. G3 binds the three modality scores as one detector signal class under Anchor 3; the workflow section says so explicitly. The reader who skipped G3 and treated the three scores as three independent confirmations has violated Anchor 3, and the resulting publishable claim is editorially indefensible under Anchor 2.

Application sequence

The framework runs in a fixed order so the result is reproducible and reviewable. Step 1: Group A gate. Confirm primary outline mapping (A1) at a defensible tier (A2). Step 2: Group E gate. Run the five anti-criteria; if any triggers, the default outcome is exclusion unless an explicit override is documented with evidence. Step 3: Group B, C, D scoring. Tools that survive Steps 1 and 2 are scored across SEA fitness, reliability, and accessibility. Step 4: Group F tier check. Confirm the candidate's functional profile passes the tier bar of the cell it occupies; if it fails F at the proposed tier, re-tier and re-run A2 or fail. Step 5: Group G architectural-alignment review. G1 caveats and G2 signal-class declarations are written into the tool-card draft; G3 bookkeeping binds how the tool counts in evidence-weighting downstream.

Tie-breaking inside an outline subcategory. When two tools both pass A, B, C, D, F, and G in the same tier-cell, ranking uses, in order: A3 weighting, B4 score, D1 score, G1 multi-pillar breadth. The first axis on which the tools differ resolves the tie. If they remain tied, E5 (better alternative) is invoked: a documented comparison selects one as primary and parks the other in the backlog.

A worked example – a hypothetical new SEA-language deepfake detector

A regional convening introduces a new tool: a 2026-launched browser-based image and video deepfake detector with documented support for Bahasa Indonesia, Thai, and Filipino interface menus, vendor-claimed ninety-four percent accuracy, free tier with one hundred scans per day, and Jakarta-based vendor hosting.

Group A. A1 maps the tool to 1A.1 (image) and 1A.2 (video) as primary candidates, with possible secondary at 1B.1 if it integrates as an InVID-style plugin. Pass. A2 checks tier appropriateness: browser-based with no install, free tier, fast scan times. Passes at First-Line Triage and Professional Verification. A3 weights medium: the tool serves the fact-checker-on-deadline lens at 1A.1 and the OSINT methodologist lens at 1A.2 if integration paths exist.

Group E. E1 requires independent evidence on the ninety-four percent claim. The reader checks independent benchmark evidence or equivalent; if no independent test is locatable, the tool can still enter with the "no independent SEA-specific benchmark identified" sentence. E2 (abandoned) does not fire on a 2026 launch. E3 (critical privacy) requires checking whether the tool uploads source files to vendor cloud and what retention is documented; if Jakarta-based hosting carries Indonesia's UU ITE exposure for sensitive newsroom material, the reader documents the S9 cross-border-jurisdiction concern as a mitigation-pass route via the country-legal-context page. E4 (regional content failure) requires checking documented evaluation on the three claimed-supported SEA languages; absence of evaluation in the available independent audits is a flag, and the tool either enters with the weak-signal framing binding or fails E4. E5 (better alternative) requires comparing against Hive AI, ImageWhisperer, InVID-WeVerify, and Deepware Scanner. If the new tool offers genuinely new regional-language coverage and the existing four do not match it, the tool clears E5. If it does not, the tool is excluded as redundant.

Group B. B1 passes (English UI option). B2 applies if the tool processes text alongside the image and video; if the tool is purely visual, B2 is N / A and B3 declares it language-agnostic for those modalities. B4 scores 2 if regional deployment is documented in one or two countries (Indonesia and the Philippines plausibly given vendor location), 3 if documented across three or more.

Group C. C1 is the binding gate. The reader pairs the vendor ninety-four percent against an independent reference (Deepfake-Eval-2024 anonymised-pool ceiling of zero point seven eight if no tool-specific number exists) or applies the "no independent SEA-specific benchmark identified" sentence. C2 requires verbatim limitations from the vendor docs plus any independent findings the reader located. C3 confirms the 2026 launch is recent enough to satisfy the eighteen-month maintenance window. C4 checks for documented systematic regional content failure; on a new tool the answer is usually "not yet documented," which is itself a caveat the card should carry.

Group D. D1 scores 4 (free tier sufficient for non-edge use, paid for surge if quotas apply). D2 passes at First-Line Triage (browser, mobile-friendly). D3 passes at First-Line Triage (no install, no account or trivial). D4 fires the S1 / S9 mitigation pathway because Jakarta-based hosting requires cross-border consideration for non-Indonesian source material. The S1 source-protection sub-routine binds on the card.

Group F. F1 passes if the five-minute phone use case holds. F2 passes if a thirty-minute desk integration with InVID-WeVerify or another professional-verification tool is possible.

Group G. G1 requires the detector-as-weak-signal caveat. G2 declares the tool's output as a detector signal class. G3 binds the multi-modal scores as one signal class.

The verdict. The tool enters the shortlist at 1A.1 and 1A.2 as alternative entries (with Hive AI as the existing primary at 1A.1 and InVID-WeVerify as the existing primary at 1A.2), conditional on the C1 wrapping pair, the C2 verbatim limitations, the D4 S1 / S9 mitigation-pass, the G1 detector-as-weak-signal caveat, and the B4-bonus regional-deployment evidence in two-or-more focus countries. Override flags documented on the tool card. If any of these conditions are unmet at evaluation time, the tool is held in the backlog until the missing condition is satisfied.

The example shows how the framework actually operates on a new case. It is conservative but not exclusionary. The C1 wrapping pair and the G1 caveat are non-negotiable; the B4 bonus shapes ranking but does not gate inclusion; the D4 mitigation-pass route preserves the tool's accessibility framing while binding the source-protection routing into the card prose. The same sequence applies to any new tool entering a regional partner's attention. The editorial-patterns page carries the rendered output: the five patterns the cards embed once the criterion gate is cleared.