Lost in Translation: How AI Summarization Tools Misrepresent Urdu-Language News Archives for Researchers

dc.audienceAudience::IFLA Publications
dc.congressWLICIFLA WLIC 2026 - Busan, South Korea
dc.contributor.authorKhanum, Almas
dc.contributor.authorBashir, Faiza
dc.coverage.spatialPakistan
dc.date.accessioned2026-07-31T08:20:26Z
dc.date.available2026-07-31T08:20:26Z
dc.date.issued2026-07-30
dc.description.abstractThere is an increasing tendency for South Asian scholars to use AI-based summarization tools in order to sift through the massive amounts of newspaper archives written in Urdu – vernacular papers that provide unique insights into political events, religious discussion, and sociocultural developments in Pakistan from 1947 onwards. However, there has been very little consideration of whether these summarization tools can be considered effective when working with non-Western, resource-poor languages. The current study seeks to critically analyze the distortion of the Urdu-language newspaper archives created by modern AI summarization software, focusing specifically on three thematic categories: religion, regional politics, and civil-military interactions. For that reason, a purposeful sample of 120 Urdu-language newspaper articles published from 1947 until now will be analyzed using GPT-4o, Gemini 1.5 Pro, and LLaMA 3.The four types of epistemic distortions noted included: (i) omission of politically and religiously relevant information; (ii) cultural elimination of linguistically encoded meaning; (iii) factual distortion; and (iv) reversal of evaluative position, in which the evaluative position taken in the source text is negated in the translated version. Pre-1990 papers experienced much higher rates of distortion due to the cumulative effects of OCR errors and the absence of early Urdu journalism materials from AI training datasets. These are not technical problems alone but stem from structural inequalities in AI development, in which the dominance of English-language data sets leads to the systematic marginalization of non-Western knowledge systems. This paper takes Urdu journalism as a case study in the exclusion of low-resource languages from AI. Keywords: Urdu-language archives; Low-resource languages; Epistemic ince; Low-resource language processing; AI summarization;
dc.identifier.urihttps://2026.ifla.org
dc.identifier.urihttps://repository.ifla.org/handle/20.500.14598/7189
dc.language.isoeng
dc.publisherInternational Federation of Library Associations and Institutions (IFLA)
dc.relation.ispartofseriesWorld Library and Information Congress (WLIC); 2026 - Busan, South Korea - Libraries Powering Transformation
dc.rights.holderKhanum, Almas
dc.rights.holderBashir, Faiza
dc.rights.licenseCC BY 4.0
dc.rights.urihttps://creativecommons.org/licenses/by/4.0/
dc.subjectArchives
dc.subjectLanguages
dc.subjectArtificial intelligence
dc.titleLost in Translation: How AI Summarization Tools Misrepresent Urdu-Language News Archives for Researchers
dc.typePublication
ifla.UnitHeadquarters::Headquarters

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
182-khanum-en.pdf
Size:
496.99 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
2.28 KB
Format:
Item-specific license agreed upon to submission
Description: