Using Named Entity Recognition for Automatic Indexing
Loading...
Date
Authors
Goh, Rachael
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Automatic indexing has been used by libraries to index their collections for many years. Recent technological advances have allowed for a refinement of automatic indexing, providing quicker and more accurate results.
In 2016, the National Library Board (NLB) embarked on a Named Entity Recognition (NER) project which leveraged on natural language processing techniques to extract names of entities such as people, organisations and places that are found in documents such as articles. Articles from NLB’s Infopedia (eresources.nlb.gov.sg/infopedia) and HistorySG (eresources.nlb.gov.sg/history) were identified for extraction. This has since enabled users to search for related resources in NLB’s vast collections.
In late 2017, the NER project was extended to include extraction of topics mentioned in articles and metadata records. This automated indexing was done using NLB’s controlled vocabularies such as Events, Historical Events, Programmes, Legal Acts, Awards, Time, and SingHeritage. Data from other agencies, National Archives of Singapore (NAS) and National Heritage Board (NHB) collections were also included. A smaller scope of the project includes running the documents against external vocabularies such as GeoNames and Wikidata.
This paper will discuss the process, the method used to evaluate the accuracy of the named entities extracted, the NER issues discovered across the collections, and the challenges faced in ensuring that the extraction process is improved after verification. ...