Meet the winners of the WAXAL speech recognition challenge
Speech comes first. Long before anyone learns to read, they learn to talk, and across Africa the spoken word remains the main way people share knowledge, do business and care for their families. Millions speak Lingala, Shona and hundreds of other languages fluently without ever having had the chance to read or write them.
That should make voice the easiest way for them to use technology. Ask a question, hear an answer. Yet most AI systems today cannot understand these languages, so the people who stand to gain the most from talking to their phones are the ones currently shut out.
We built WAXAL to help change that. It is an open-source collection of speech data covering 32 African languages, and we recently worked with Zindi to put it to the test. The WAXAL ASR Challenge invited data scientists to use it to build systems that can accurately understand and transcribe spoken African languages, starting with Lingala and Shona.
The response went well beyond what we planned for. 1,462 innovators from 100 countries entered, 42 of them African nations. Between them they submitted 8,837 solutions and put more than 24,000 hours into the problem. When we met participants at the Deep Learning Indaba, many told us how much it meant to work on a challenge built around African languages and the everyday situations people use them in.
We're thrilled to announce the winners.
Meet the innovators
The top three scores were remarkably close, and each winner got there by a different route. Together, their approaches show how much creative thinking this problem rewards.
1st Place: Team Pasketti (Kazakhstan and Canada)
Roman Solovyev and enes3774 started from the idea that no single AI model gets everything right. So they built eight and made them work together. Each model makes its own kind of mistake, one dropping a word, another swapping it for the wrong one. Blended together, the models could check each other's work and settle on the most accurate transcription.
2nd Place: Alban Nyantudre (Burkina Faso)
For Alban Nyantudre, a machine learning engineer, the challenge sat close to home. He already builds speech tools for his mother tongue, Mooré. His system listens to the same audio clip in several different ways, then picks the reading it trusts most. He put the case for this work simply: "Plenty of people I grew up around speak their language perfectly but cannot read or write it. That is why speech is the only realistic way for those people to use technology at all."
3rd Place: Abdourahamane Ide Salifou (Niger)
Abdourahamane, an AI engineering student, began by studying how each language actually sounds. In Lingala, words run into one another, and a computer struggles to tell where one ends and the next begins. He trained his system to wait and take in the flow of speech before deciding on the words. Shona has clearer word boundaries, so there he leaned on the surrounding context to guide the prediction.
What comes next
Give innovators good open data and they will build technology that serves everyone. This challenge proved it. The methods developed here point toward a future where a farmer can check crop prices, a parent can find health advice and a student can learn, all by speaking to a phone in the language they grew up with.
Thank you to Zindi, to every participant, and to the wider data science community for the time, skill and care you brought to this.