Corpus-driven Bantu Lexicography Part 2: Lemmatisation and Rulers for Lusoga

Gilles-Maurice de Schryver, Minah Nabirye


This article is the second in a trilogy that deals with corpus-driven Bantu lexicography, which is illustrated for Lusoga. The focus here is on the macrostructure and in particular on the building of a lemmatised frequency list directly within a dictionary-writing system. The programming code for the parts of the lemmatisation that may be automated is included as addenda. A second focus is on the embedded part-of-speech and alphabetical rulers, for which it is shown how these may be used to plan the actual compilation of the dictionary entries.


Bantu; Lusoga; corpus lexicography; lemmatisation; lemmatised frequency list; part-of-speech ruler; alphabetical ruler; multidimensional lexicographic ruler; dictionary planning; dictionary-writing system; TLex; TshwaneLex

Full Text:




  • There are currently no refbacks.

ISSN 2224-0039 (online); ISSN 1684-4904 (print)

Creative Commons License CC BY 4.0

Powered by OJS and hosted by Stellenbosch University Library and Information Service since 2011.


This journal is hosted by the SU LIS on request of the journal owner/editor. The SU LIS takes no responsibility for the content published within this journal, and disclaim all liability arising out of the use of or inability to use the information contained herein. We assume no responsibility, and shall not be liable for any breaches of agreement with other publishers/hosts.

SUNJournals Help