South African tool helps low-resource languages enter the age of artificial intelligence
Read more
IOL
iol.co.za

South African tool helps low-resource languages enter the age of artificial intelligence

Researchers from South Africa have developed a software tool designed to address one of the major challenges facing African languages in the digital age, and it is now being used by scientists worldwide.

This tool, named TextAugment, is created to assist developers in building language technologies when they lack the vast amount of digital text required to train artificial intelligence systems. The open-source software was developed by Professor Vukosi Marivate, Director of the African Institute for Data Science and AI (AfriDSAI) and holder of the Absa UP Chair of Data Science at the University of Pretoria, as well as researcher Tshephisho Sefara. Currently, the tool has registered over 286,000 downloads and is being applied to work with languages such as Swahili, Arabic, and Uzbek.

The dissemination of this software was marked on July 16th when Marivate and the TextAugment team received the first NSTF-SADiLaR Research Software Award for human language technologies. This award is given for creating open-source software that helps researchers generate synthetic training data for languages with insufficient digital resources.

This recognition is part of a series of recent awards for Marivate. On May 19th, the computer scientist, born in Ga-Rankuwa, received the Mapungubwe Silver Order for his contribution to data science, artificial intelligence, and natural language processing. The presidency highly praised his work, noting 'his outstanding contribution to data science, artificial intelligence (AI), and natural language processing (NLP), which has significantly advanced both national and continental technological capabilities.'

However, his work goes beyond simply teaching machines language comprehension. It raises a broader question: what happens if the technologies shaping the future cannot understand the languages spoken by millions of people? He emphasized that 'the mere existence of AI is not enough for it to truly impact people'; suitable conditions must be created to amplify the positive aspects and minimize the negative consequences of any technology.

African languages in AI research are classified as 'low-resource' languages. In Marivate's view, this does not mean there are few speakers of such languages, but rather that they lack the digital data and tools necessary for machines to process them effectively. He explained that 'one of the biggest obstacles to training robust machine learning models for low-resource languages (such as isiZulu, Sesotho, or Yoruba) is the lack of data. Traditional AI models require huge amounts of clean, labeled data to perform well. If such data is absent, these languages are effectively excluded from the modern AI revolution.'

The Data Gap Excluding African Languages from AI

This gap becomes particularly evident when people interact with online systems. For example, sentiment analysis systems can determine whether an online review is positive, negative, or neutral. But systems developed predominantly for English cannot automatically provide the same service for African languages.

Marivate noted: 'In South Africa, we have 12 official languages, and we are not actually serving most of our people. To interact with all these online systems, you are forced to write in English because that is how the processing happens.'

A language can have millions of speakers while remaining poorly represented online. This is where one of Marivate's most frequently used research tools comes into play. TextAugment was created to help researchers overcome the scarcity of high-quality language data. The tool works by synthetically augmenting existing datasets, including replacing words in sentences with synonyms while attempting to preserve the original meaning.

This technology supports applications for translation, sentiment analysis, summarization, and other language processing tasks. It also helps researchers create tools such as spell checkers and speech processing systems. Ultimately, TextAugment was released as open-source software so that researchers outside Marivate's team could use it. Since then, it has been used by researchers working with languages such as Swahili, Uzbek, and Arabic.

Creating AI That Understands More Than Just Words

However, increasing the volume of language data for AI models is only part of the problem. African languages carry cultural meanings, idioms, and proverbs that cannot always be translated literally. Marivate acknowledges that some augmentation methods have limitations, as replacing a word with a synonym does not necessarily account for the context in which that word is used.

He stated: 'Writing in English and then translating to isiXhosa is not the same as a person writing in isiXhosa, because they write in a way that touches upon culture, geography, and much else that is inherent.'

For Marivate, African language technologies should not be reduced to taking English-language systems and forcing them to output African words. The technology must be capable of interacting with the cultural context in which these languages are used. Therefore, his research group is studying how AI models understand idioms and proverbs in African languages, including isiZulu and Tshivenda.

The issue extends beyond how AI processes individual words. A broader question about whose languages and knowledge are represented in AI was also addressed by the Vice-Chancellor of the University of Ghana, Professor Nana Aba Appiah Amfo, during the University of Warwick Distinction Lecture in 2026. In a lecture titled 'Whose Language Matters? African Voices, Knowledge Systems, and the Future of AI,' Amfo explored the implications of language exclusion in new technologies and challenged common assumptions about whose knowledge is represented in AI systems.

According to her, the absence of African languages in digital ecosystems is not merely a technical limitation, but a matter of representation and inclusion. She stated: 'When a language is absent from the digital corpus, it is not just a translation problem. It is a visibility problem. It is a knowledge problem. And ultimately, it is a matter of justice.'

It is easy to frame the development of AI in Africa as something that began with the emergence of ChatGPT. Marivate rejects this narrative. He points to an African AI ecosystem that has been developing for over a decade. The Deep Learning Indaba conference began in 2016 as a gathering for people involved in African research, development, and innovation in AI. Its 2026 meeting in Lagos hosted about 1000 people. The Masakhane Research Fund, which emerged from this broader community in 2019, according to Marivate, has grown to over 3000 researchers and innovators working on African languages. Marivate stressed: 'The African AI ecosystem did not wait. They did not wait; these organizations started a decade ago.'

Investment is needed directly into the languages themselves, including the development of digital tools, research potential, and education. Governments and private companies must also support locally developed technologies instead of simply importing or reselling systems created elsewhere. He argues that without procurement and investment supporting African tech companies and researchers, a sustainable ecosystem cannot emerge. This is especially critical because poorly functioning AI systems can have consequences that go beyond inconvenience. Marivate warns that if an AI service works well in English but provides unreliable results in an African language, users may mistakenly assume the technology is equally capable across all languages, when it is not, which can lead to genuinely harmful experiences.

The next phase of his work involves assessing how well existing large language models actually perform on African languages. He acknowledged that Africa remains heavily reliant on foreign-developed technologies. However, he believes that building local research capacity will allow the continent to anticipate technological changes rather than just react to them. He concluded: 'We must build ourselves to learn how to do it better.'

TextAugment initially emerged as a response to the limited digital representation of African languages. Its adoption demonstrates a broader possibility: when African researchers create solutions tailored to the realities of the continent, they can produce technologies of global value.

Popular