Lost in Translation: How AI Summarization Tools Misrepresent Urdu-Language News Archives for Researchers
Loading...
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
International Federation of Library Associations and Institutions (IFLA)
Abstract
There is an increasing tendency for South Asian scholars to use AI-based summarization tools in order to sift through the massive amounts of newspaper archives written in Urdu – vernacular papers that provide unique insights into political events, religious discussion, and sociocultural developments in Pakistan from 1947 onwards. However, there has been very little consideration of whether these summarization tools can be considered effective when working with non-Western, resource-poor languages. The current study seeks to critically analyze the distortion of the Urdu-language newspaper archives created by modern AI summarization software, focusing specifically on three thematic categories: religion, regional politics, and civil-military interactions. For that reason, a purposeful sample of 120 Urdu-language newspaper articles published from 1947 until now will be analyzed using GPT-4o, Gemini 1.5 Pro, and LLaMA 3.The four types of epistemic distortions noted included: (i) omission of politically and religiously relevant information; (ii) cultural elimination of linguistically encoded meaning; (iii) factual distortion; and (iv) reversal of evaluative position, in which the evaluative position taken in the source text is negated in the translated version. Pre-1990 papers experienced much higher rates of distortion due to the cumulative effects of OCR errors and the absence of early Urdu journalism materials from AI training datasets. These are not technical problems alone but stem from structural inequalities in AI development, in which the dominance of English-language data sets leads to the systematic marginalization of non-Western knowledge systems. This paper takes Urdu journalism as a case study in the exclusion of low-resource languages from AI.
Keywords: Urdu-language archives; Low-resource languages; Epistemic ince; Low-resource language processing; AI summarization;