Sri Lanka¶
Sri Lanka has real multilingual fact-check capacity. Watchdog (OCCRP-funded, hybrid OSINT and editorial, around thirteen named staff with Sinhala and Tamil capacity), Hashtag Generation (around eight publicly named fact-check staff, registered nonprofit, donor-funded), FactCheck.lk (Verité Research-operated, IFCN status currently expired), Fact Crescendo Sri Lanka (Meta third-party partner, 3,000-plus fact-checks over four years), FactSeeker (SLPI-affiliated, tri-lingual), and Citizen Fact Check (transparent on funding but with unclear current throughput) make up the operational landscape. The language-tech and research layer rests on LIRNEasia, the University of Moratuwa's National Languages Processing Centre, and the University of Colombo School of Computing's Language Technology Research Laboratory.
Sinhala and Tamil sit on opposite sides of an asymmetric coverage pattern that this page documents directly. Sinhala has more locally developed Sri Lanka-specific language-tech assets: the MisinformationCorpusSinhala, SinLlama, the SLTK tokenizer, and the LIRNEasia–Watchdog Dissect prototype. Tamil leans on pan-Tamil and India-based resources, with explicit Sri Lanka-adaptation gaps. The Online Safety Act 2024 and the Prevention of Terrorism Act legacy shape the legal environment. The fact that no Sri Lankan operator currently holds active IFCN signatory status closes Meta Content Library access. This page documents the 2024–2025 election cycle response, the DRI 2025 chatbots-and-misinformation study, the asymmetric language-tech stack, and the operational routing that the Sri Lankan ecosystem has worked out under those conditions.
Information environment¶
The Sri Lankan fact-check ecosystem is denser than a quick scan suggests. On the strongest reading of the 2024–2026 evidence, six locally based operators are active or recently active. FactCheck.lk (Verité Research, launched 2018, Sinhala and Tamil archives). Hashtag Generation (registered nonprofit, founded 2015, fact-check arm matured from 2021 onwards). Watchdog (OCCRP-funded hybrid OSINT collective, founded April 2019 after the Easter Sunday bombings). Fact Crescendo Sri Lanka (Meta third-party partner, tri-lingual). FactSeeker (SLPI-affiliated, tri-lingual). And Citizen Fact Check, with uncertain current operational intensity per LIRNEasia research. AFP Sri Lanka and Newschecker Sri Lanka work on Sri Lankan misinformation with international parent organisations, not as locally based entities. Dissect is a tool and not a fact-check organisation; the LIRNEasia–Watchdog–Appendix partnership built it as an AI prototype for journalists and fact-checkers.
Three strengths run through the most-documented material. The fact-check operational layer is multilingual and donor-supported: Hashtag Generation's Citizen Reporters' Network brings together critically engaged youth from diverse ethnic, geographical, and linguistic backgrounds, and Watchdog's editorial and engineering capacity supports investigations through both Sinhala and Tamil. The civil-society capacity layer extends into media-literacy and digital-rights training (Hashtag's documented 600-plus individuals trained over three years, including Tamil and Upcountry Tamil communities; SLPI's repeated workshops on basic media literacy in Sinhala and Tamil). The research-and-tooling layer rests on LIRNEasia's MisinformationCorpusSinhala (3,576 Sinhala documents annotated CREDIBLE / FALSE / PARTIAL / UNCERTAIN), the University of Moratuwa's SinLlama LLM (continual pre-training on Llama-3-8B over a 10-million-sentence Sinhala corpus), and the broader UCSC / UoM stack (Sinhala POS Tagger, Sinhala-and-Tamil NER, ThamizhiUDp and ThamizhiPOSt for Tamil, the Sri Lanka NLP catalogue).
The tipline layer is the real structural gap. Per LIRNEasia research, Sri Lanka's intake infrastructure is "real but thin": a patchwork of web forms, email, phone numbers, and WhatsApp distribution and update groups. No documented Sri Lanka-based automated or hybrid multilingual tipline runs purpose-built verification-backend infrastructure at the Meedan Check level. Fact Crescendo Sri Lanka has the most visible public intake through its Submit For Fact-Check workflow and active WhatsApp distribution channels. Hashtag Generation's Fact Check Claim page accepts uploaded screenshots with a Keep me Anonymous option useful for communal-sensitivity rumours. FactSeeker uses email, a public phone number, and separate Sinhala / Tamil / English WhatsApp groups (broadcast, not clearly two-way tiplines). The 2024–2025 election cycle showed local actors absorbing both classical and AI-intermediary threats; the operational ceiling sat where intake volume, Tamil localisation, and AI-synthetic scale outran available staff time.
Platform dominance shapes the work. Facebook remains central to public discourse. WhatsApp carries most personal messaging and most viral rumour propagation. Telegram is operationally important on the fraud-and-scam side, especially the malicious.apk distribution patterns documented in Fact Crescendo Sri Lanka's Tamil stream and the broader cross-banking scam ecosystem. TikTok has grown sharply in election cycles. Cross-language propagation is common: a Sinhala-language rumour reaches Tamil-speaking communities through translation, often by partisan actors. The Tamil version then propagates in different platform ecologies, and the verification work has to track both to close the case.
State institutions appear in the verification ecosystem chiefly as sources, subjects, and amplifiers of corrective communication, and not as publicly visible equivalents of civil-society fact-check desks. The Election Commission shows up as a verification anchor; Hashtag debunked a false Sinhala claim about Election Commission recruitment by reference to formal public procedures. The police and CID appear as part of the evidentiary chain on financial-fraud cases. Ministries (the Ministry of Finance in particular) are routine targets of FactCheck.lk's verification work. There is no publicly documented multilingual misinformation-response unit operated by the state at the scale civil-society work runs. The toolkit's editorial position records this without prescribing what should change.
Documented cases¶
Watchdog and Hashtag Generation on the 2024–2025 election cycle¶
Across the September 2024 presidential election and the 2025 local government elections, Hashtag Generation published a comparative analysis of ethnonationalist hate speech and post-election disinformation. Watchdog ran multilingual OSINT investigations through its hybrid editorial-engineering team. The work covered communal-sensitivity content (Mahaviru-related and Thesawalamai-law-related debunks in Hashtag's archive), election-rumour spikes, AI-generated political content surfacing in late 2025, and the DRI 2025 study that tested major chatbots in English, Sinhala, and Tamil on election-related questions and found systematic misinformation and bias risks.
The case shows the Sri Lankan ecosystem in operational form. Watchdog's Dissect prototype carried part of the Sinhala claim-extraction load through 2025; the LIRNEasia 2025 usability and scalability report describes the second project phase as focused on usability, scalability, and refinement before public launch. Hashtag's Citizen Reporters' Network brought community sensing into the workflow at the intake layer, with the Keep me Anonymous option operational on communal-memory rumours. Fact Crescendo Sri Lanka ran active Sinhala and Tamil streams alongside, with documented fact-check output volume that materially exceeds what Sri Lanka-only local competitors produce. The 2025 cycle was also the first Sri Lankan election cycle where AI intermediaries (and not only viral posts and forged PDFs) entered the verified threat model. The DRI study is the citable evidence base.
Three lessons sit inside the case. One: the multilingual capacity exists, asymmetrically. Sinhala has Dissect, MisinformationCorpusSinhala, and SinLlama; Tamil has Fact Crescendo's stream, Hashtag's Tamil staff, and X-CLAIM with the Sri Lanka adaptation gap. The verification work has to read both sides and accept that the tooling will lean Sinhala-stronger. Two: the tipline layer is the operational ceiling. Intake volume on a single major case can outrun manual processing; the route in practice is to combine Hashtag's anonymised submission, Fact Crescendo's WhatsApp distribution, and Watchdog's OSINT investigation, and not to expect a single tipline backend to carry the load. Three: AI intermediaries are now part of the threat model alongside artefact-level deepfake work. Chatbots themselves can produce election misinformation, and the verification workflow has to include the source-of-the-claim question alongside the authenticity-of-the-artefact question.
The decision-tree path the case demonstrates is T7 tipline routing into T1 image triage and T3 audio triage for AI-generated political content. The 2B.2 AI claim extraction cell routes through Dissect for Sinhala and through X-CLAIM for Tamil (with the SL-adaptation caveat applied). The T6 source-protection tree S2 sub-section fires on communal-sensitivity material and the S5 sub-section on named-source-on-camera communal cases.
Dissect, LIRNEasia, and the Sinhala language-tech stack¶
Dissect is not a fact-check organisation. The LIRNEasia research file is direct on this. It is a tool, a prototype built by LIRNEasia in partnership with Appendix / Watchdog Sri Lanka. The 2025 LIRNEasia usability and scalability report frames it as an AI tool for journalists and fact-checkers, with the second project phase oriented at usability, scalability, and refinement before public launch. The tool sits at 2B.2 AI claim extraction as the only Sri Lanka-specific Sinhala NLP claim tool in the toolkit shortlist. It carries the LIRNEasia Sinhala language-tech stack into the verification workflow.
The case in this country page is the structural one. Dissect is the productive expression of Sri Lanka's locally developed Sinhala stack. The MisinformationCorpusSinhala provides the training and evaluation substrate. SinLlama provides the LLM layer for summarisation and retrieval augmentation. The SLTK tokenizer (released as version 1.0.0 on 25 March 2025 on PyPI) provides the lexical infrastructure. The Sri Lanka NLP catalogue carries the broader resource-discovery surface. For a Sinhala-language claim entering a Sri Lankan fact-check operation, the route is Dissect for claim extraction, with Meedan Alegre at 2B.2 as the multilingual fallback when material crosses into Tamil or English. Yudistira-style Bahasa work is explicitly absent: Bahasa coverage is for Indonesia, not for cross-language Sri Lankan work.
The case also shows the Sinhala/Tamil asymmetry directly. The Sri Lanka language-tech stack has no equivalent Sri Lankan Tamil LLM. There is no equivalent Tamil misinformation corpus tagged on Sri Lankan named entities and communal-memory politics. There is no Tamil tokenizer release equivalent to SLTK from a Sri Lankan institution. Tamil work in Sri Lanka leans on the pan-Tamil and India-based resources documented in LIRNEasia research (AnaadiAI TamilTokenizer, the Tamil_Hate_Speech Hugging Face dataset, the Multilingual Fact-Checking using LLMs workshop paper benchmarking Tamil among other languages). None of these are trained on or curated for Sri Lankan Tamil political discourse, Sri Lankan named entities, or island-specific communal narratives. The X-CLAIM tool card carries the Sri Lanka adaptation gap directly.
DRI 2025 chatbots-and-misinformation study and the AI-intermediary threat¶
Democracy Reporting International released the 2025 study titled "Biased by Design? Chatbots and Misinformation in Sri Lanka's 2025 Local Elections" as the first publicly documented Sri Lanka-specific evidence that AI intermediaries themselves are a misinformation vector, separate from the artefact-level deepfake threat. The study tested major chatbots in English, Sinhala, and Tamil on election-related questions and found systematic misinformation and bias risks across the three languages.
The case matters here because it changes the verification workflow in a documented way. Before 2025, the Sri Lankan threat model centred on viral posts (forged PDFs, fabricated statement screenshots, communal-narrative video clips) and on the broader rumour ecology around elections, economic-crisis aftermath, and Easter Sunday anniversaries. After the DRI study, the threat model includes the AI-intermediary vector: a user asks a chatbot an election question, the chatbot returns a misinformation-shaped answer, the user propagates that answer through social channels as if it were a verified fact. The verification workflow now includes asking where a claim came from, with chatbot output as a documented possible source.
For a fact-checker reading the case, the handle is that LLM hallucination at the consumer-chatbot layer is now a documented threat in Sri Lanka, not a hypothetical one. The toolkit's 2B.3 LLMs in fact-check workflows framing applies: every LLM output is suggestive, never a signal class on its own. The case is also the citable evidence base for the wider regional argument the toolkit makes in the methodology layer: chatbots are vectors of misinformation, not only outputs to fact-check. The decision-tree implication is that T5 escalation on a textual claim now includes the chatbot-as-source question, with Google Fact Check Explorer at 1B.5 carrying the existing-fact-check lookup that often closes the case quickly.
The IFCN access barrier and the Meta Content Library gap¶
The Sri Lankan fact-check ecosystem currently has no active IFCN signatory. FactCheck.lk's IFCN status is marked Expired in the IFCN profile. Hashtag Generation, Watchdog, and FactSeeker are not IFCN-accredited on the evidence reviewed in LIRNEasia research. Fact Crescendo Sri Lanka sits inside the wider Fact Crescendo network with IFCN status at the network level; the Sri Lanka-specific operational independence question is partly answered by the network's certification and partly remains open.
The operational consequence is that Meta Content Library, the access-controlled research surface for Facebook and Instagram content, is structurally unavailable to most Sri Lankan fact-check operators. The toolkit treats Meta Content Library as a Graphika-pattern access-barrier tool with the "cite as external source" Quickstart redirect. The Sri Lanka case is the clearest worked instance of why the redirect pattern matters: the access surface is real, but it is closed to the operators who would use it most productively.
The case sits in the methodology editorial-patterns layer as part of the access-barrier framing running across the toolkit. For a working Sri Lankan fact-checker, the route around the gap is to combine Sinar Project iMAP-style independent platform monitoring (where applicable), the Hashtag and Watchdog OSINT and editorial capacity, and Fact Crescendo network access where the case warrants institutional escalation. The gap is not closable from the Sri Lankan side. It is structural, and the toolkit names it directly instead of papering it over.
Language paths¶
Sri Lanka's language-coverage picture in the toolkit stack is the most asymmetric of any country in the six focus set. Sinhala has the deepest local stack. Dissect at 2B.2 handles claim extraction. Meedan Alegre sits as the multilingual fallback (XLM-R coverage at documented performance for Sinhala). The broader LIRNEasia / Moratuwa / UCSC research stack (MisinformationCorpusSinhala, SinLlama, SLTK, Sinhala POS / NER) sits outside the toolkit shortlist as specialist resources. SEA-LION coverage is limited; Sinhala is partial in the major SEA LMs, and the toolkit pairs this absence with the Lao gap under Decision 7. Sinhala-language audio has limited Whisper coverage (less production-quality than Bahasa or Tagalog); the audio-detector class at 1B.3 has no Sinhala-specific benchmark.
Tamil has documented operational capacity but thinner Sri Lanka-specific tooling. X-CLAIM is pan-Tamil and India-trained, useful for general Tamil claim work but with the explicit Sri Lanka adaptation gap recorded on the card (the tool is not trained on Sri Lankan named entities and communal-memory politics). Meedan Alegre provides the multilingual fallback for cross-language work. The wider pan-Tamil / India-based resources (AnaadiAI TamilTokenizer, Tamil_Hate_Speech, Mozilla Common Voice Tamil at 425.32 hours / 979 speakers as of March 2026) provide preprocessing and a broader-Tamil substrate but carry the same Sri Lanka adaptation caveat. The University of Moratuwa's Sinhala-and-Tamil NER and ThamizhiUDp / ThamizhiPOSt provide Tamil parsing and POS-tagging support at the local level, but again sit outside the toolkit shortlist as specialist research resources.
The Sinhala/Tamil pairing under Decision 7 binds the toolkit to treating both with equal rigour. The operational read is that Tamil fact-check work in Sri Lanka needs more human review, more source-list curation, and more care with named entities, dialectal nuance, and communal-memory politics than Sinhala work needs. Sinhala work is better positioned for local automation experiments but still requires human oversight, because the country's most dangerous content often mixes misinformation with identity signalling, sarcasm, visual manipulation, and coded political speech that no current tool surfaces reliably.
For Tamil work, the handle is that the toolkit's tools are operational substrate, not verification authority. Desk-tier verification on a Sri Lankan Tamil claim runs through Hashtag's or Watchdog's Tamil-capable staff. The toolkit's tools handle the standard authenticity work (image and video triage, archive capture, metadata extraction). The language-specific claim work stays largely human-led.
Legal and threat context¶
The Online Safety Act No. 9 of 2024 (OSA) is the central legal instrument shaping Sri Lankan online speech. The Act was certified 1 February 2024 and establishes an Online Safety Commission with broad powers over "prohibited statements," online accounts, and "online locations" used for prohibited purposes. The Cabinet approved an amendment-committee process in February 2025; the law remained in force through 2025–2026 with the amendment process described as ongoing in September 2025 reporting. The first arrest under the OSA was made in February 2024, and the Online Safety Commission was not appointed at the time of that arrest, which is itself part of the operational pattern. Access Now, CPJ, and more than fifty other organisations called in January 2024 for the bill to be withdrawn. GNI's March 2025 year-in-review characterised the Act as one of the most sweeping threats to freedom of expression and privacy in any jurisdiction.
The Prevention of Terrorism Act (PTA) legacy sits alongside the OSA. RSF in September 2024 called for repeal of both the PTA and the OSA, noting that both are used against journalists. In April 2025, the Centre for Policy Alternatives warned that the government was continuing to use the PTA for conduct with no apparent connection to terrorism, including arresting a youth over anti-Israel stickers. Human Rights Watch in January 2026 said the proposed Protection of the State from Terrorism Act risked reproducing PTA abuses. Groundviews described the December 2025 draft as retaining wide executive powers and strong overlap between ordinary offences and terrorism. The legal-risk question for Sri Lankan verification work runs alongside the verification question, not after it.
The prosecutions pattern in 2024–2026 affects both Sinhala-language and Tamil-language journalism, and disproportionately affects Tamil journalists. CPJ reported in August 2025 that counter-terrorism police summoned Tamil photojournalist Kanapathipillai Kumanan, who had been documenting mass graves in the north; IFEX relayed a joint appeal demanding an end to the harassment. In March 2026, authorities detained Lanka-e-News editor Sandaruwan Senadheera after he arrived at Colombo airport. The airport-detention pattern is operationally important. The threat surface includes border control on returning journalists, not only inland investigation.
For T6 source-protection, the S2 sub-section (state-linked or legally sensitive investigation) fires on Sri Lankan verification work touching the OSA enforcement pattern, the PTA legacy, north-and-east community reporting, or any case where counter-terrorism framing is plausible. The S5 sub-section (private group infiltration and identifying-material caution) sharpens on cases involving Tamil community sources in the north and east, where the documented physical-and-digital surveillance pattern (per RSF UN reporting September 2024, Tamil Guardian February 2025) interleaves with the OSA's centralised authority over accounts, content, and "online locations." The practical implication in regional legal-context research is that staff travelling to the north and east should use travel-clean devices where possible, store sensitive archives off-device, and establish immediate check-in procedures.
The surveillance environment also drives the choice of forensics tooling. Sherloq at 1B.4 is the offline forensics route when the source file cannot leave the verifier's machine. The toolkit treats this offline desktop pattern as mandatory mitigation for Sri Lanka's surveillance-environment threat model under the Decision 7 paired honest-gap policy. The auto-archiver cross-jurisdiction archive route at 2A.3 handles the archive workflow where retention outside Sri Lankan jurisdiction is part of the source-protection posture.
Operational routing¶
If a Sinhala-language claim or artefact reaches a Watchdog, Hashtag, or FactCheck.lk desk, the first-minute routing is through Dissect for claim extraction at 2B.2, with Meedan Alegre as the multilingual fallback if the case crosses into Tamil or English. Pair the claim-extraction read with reverse-image work via InVID-WeVerify at 1B.1, T1 image triage for any image artefact, and Google Fact Check Explorer at 1B.5 for the existing-fact-check lookup.
If a Tamil-language claim or artefact reaches a Fact Crescendo Sri Lanka, FactSeeker, or Hashtag desk, route the artefact-level work through the same Pillar 1 stack as Sinhala work: InVID-WeVerify, T1 image triage, Google Fact Check Explorer. Treat the Tamil claim-extraction layer as human-led, with X-CLAIM as a preprocessing aid carrying the Sri Lanka adaptation caveat. The named-entity verification step on Sri Lankan Tamil content is operational; pan-Tamil tools will not reliably identify Sri Lankan named entities, and the verification benefits from human Tamil-competent review against community context.
When a case touches the OSA enforcement environment, the PTA legacy, north-and-east community reporting, or counter-terrorism-adjacent framing, the T6 source-protection tree S2 and S5 routing fires before the verification workflow proceeds. Pre-upload review for cloud-hosted tools is operational. Sherloq at 1B.4 is the offline forensics route. auto-archiver at 2A.3 carries the cross-jurisdiction retention pattern.
For institutional-tier escalation that would benefit from Meta Content Library access, the route is via Fact Crescendo Sri Lanka's network-level connection, or via a partner organisation outside Sri Lanka that holds the IFCN signatory status Sri Lankan operators do not currently carry. The gap is structural, and the toolkit's framing on it is direct.
On an AI-intermediary case (a claim originating from a chatbot and not from a viral artefact), the workflow now includes the source-of-the-claim question alongside the authenticity question. The DRI 2025 study is the citable evidence base. The 2B.3 LLMs in fact-check workflows framing applies. The existing-fact-check lookup through Google Fact Check Explorer closes a meaningful proportion of cases at the 1B.5 layer before deeper verification work is needed.
Cross-references¶
- 2B.2 AI claim extraction – Dissect for Sinhala, X-CLAIM for Tamil (with SL-adaptation gap), Alegre for multilingual fallback
- 1B.1 multi-tool plugins, 1B.4 metadata / ELA forensics, 1B.5 news-reliability scoring – the Pillar 1 ladder for Sri Lankan desk work
- 2A.3 archiving – auto-archiver cross-jurisdiction route for OSA / PTA environment
- 2B.3 LLMs in fact-check workflows – DRI 2025 chatbot-as-vector framing
- T1 image triage, T6 source-protection, T7 tipline routing – decision-tree routing for Sri Lankan work
- Laos country page – the paired-gap country under Decision 7 (Sri Lanka has fact-check capacity, Laos has neither at scale)
- Indonesia country page – Pillar 2 tipline contrast (Sri Lanka has the real gap; Indonesia anchors the operational form)
- Methodology editorial-patterns – Meta Content Library access-barrier framing
- Digital Safety — Country Legal Context – OSA enforcement, PTA legacy, north-and-east surveillance pattern, airport-detention pattern
- Facebook, WhatsApp, Telegram, TikTok – platform-level context for Sri Lankan operations
Sources¶
- LIRNEasia. MisinformationCorpusSinhala. LIRNEasia, 2025. github.com/LIRNEasia/MisinformationCorpusSinhala. (3,576-document Sinhala misinformation corpus; SinLlama on Llama-3-8B over 10M-sentence Sinhala corpus; SLTK v1.0.0 language toolkit.)
- LIRNEasia. Misinformation and Language Resources. LIRNEasia, 2025. lirneasia.net/themes/misinformation-and-language-resources. (Sinhala / Tamil NLP asymmetry; Mozilla Common Voice Tamil 425.32 hours versus Sinhala 0.12 hours; Watchdog / Dissect stylometric pipeline.)
- Democracy Reporting International (DRI). AI Chatbots and Misinformation in Sri Lanka. DRI, 2025. democracy-reporting.org. (DRI 2025 chatbots study; 2024–2025 election-cycle disinformation cases; Hashtag Generation, FactCheck.lk, Fact Crescendo SL ecosystem.)
- Human Rights Watch. Sri Lanka: New Terrorism Law Threatens Rights. HRW, January 2026. hrw.org. (Online Safety Act No. 9 of 2024; Protection of the State from Terrorism Act warning; PTA legacy; Kumanan / Senadheera prosecutions; airport-detention and north-and-east surveillance patterns.)
- Centre for Policy Alternatives. Statement on the Online Safety Act Amendment. CPA, April 2025. cpalanka.org. (Cabinet amendment-committee approval February 2025; GNI year-in-review; OSA digital-surveillance dimension.)
- change-log — paired Lao / Sinhala-Tamil honest-gap policy (Decision 7).
- Tool cards: Dissect, X-CLAIM, Meedan Alegre, InVID-WeVerify, Sherloq, Google Fact Check Explorer, Google Cloud Translation, Meta Content Library, AI Disinfo Hub, SEA-LION, auto-archiver, Sinar Project iMAP