MANSA IS NOW LIVE

African Languages Lab research recognised at ACL

Our paper, "The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLP," was selected for an oral presentation at ACL 2026. Sheriff Issaka, our founder and head of research, presented the work on Monday, 6 July.

Oral presentations at ACL go to a small fraction of accepted papers, and acceptance itself is competitive. For work on African languages, which has historically been under-represented at the major NLP venues, the selection matters beyond our own team.

African Languages Lab research recognised at ACL

The paper introduces the largest validated multimodal dataset for African languages assembled to date: 40 languages, 19 billion text tokens and 12,000 hours of aligned speech. Every part of it was collected and validated by native and fluent speakers rather than scraped, which is what makes it usable for training rather than simply large.

The more consequential finding is about where the bottleneck actually sits. For African languages, data quality rather than model size is what limits performance. You can add parameters indefinitely, but a model trained on thin, mistranslated or unvalidated data inherits those flaws. That result reframes the problem for anyone working in this space: the work is not primarily a modelling problem, it is a collection problem.

The paper was written with collaborators at UCLA, Georgia Tech, the University of Wisconsin-Madison, the University of Cape Coast, Carleton University, Stetson University, Northwestern University in Qatar, Cornell, Soka University of America and Columbia.

Read the paper

Create a free website with Framer, the website builder loved by startups, designers and agencies.