Sharing our work, learning from Africa’s AI community, and rethinking how we evaluate AI at Deep Learning Indaba 2026.
At Deep Learning Indaba 2026 in Lagos, Kitala AI and YUX Design joined Microsoft Research Africa to facilitate “Beyond Benchmarks: Scaling Multi Turn Participatory AI Evaluations in Health and Education,” a hands on workshop exploring how researchers can evaluate AI in ways that better reflect how people actually use these systems.
The workshop was a key part of our participation, but our contributions extended beyond it. Our AI Engineer Ertony Bashil presented our work on culturally grounded maternal health evaluation during Research in Africa Days, while Oche David Ankeli joined a mentorship session with emerging AI practitioners. Throughout the week, we also attended keynotes, technical workshops, and research sessions that challenged us to think more deeply about what meaningful AI progress should look like across African contexts.
Across these experiences, one question kept coming back:
How do we know when an AI system actually works for the people it is supposed to serve?

Our team at Deep Learning Indaba 2026 in Lagos, Nigeria.
Taking AI evaluation beyond benchmarks
Traditional benchmarks remain an important part of AI development. They help researchers compare models, track technical progress, and identify strengths and weaknesses at scale.
But people do not experience AI as a benchmark.
Real conversations unfold over time. People switch between languages, leave out important information, change their questions, use local expressions, and bring their own cultural expectations and lived experiences into an interaction.
This was the gap we wanted participants to explore during our Beyond Benchmarks workshop.
Jointly facilitated by Kitala AI, YUX Design, and Microsoft Research Africa, the session brought researchers and practitioners together to explore participatory and human centred approaches to evaluating large language models.

Facilitating the Beyond Benchmarks workshop at Deep Learning Indaba 2026.
Instead of focusing only on whether a model produced the expected answer, we asked participants to consider questions such as: Who is using the system? What are they trying to accomplish? What could go wrong for this particular user? And what should we measure to understand whether the interaction actually worked?
Participants worked together to create realistic multi turn scenarios, select human-centred evaluation metrics, interact with an LLM, evaluate its responses, and collectively examine the results.

Participants working in groups to explore multi turn scenarios and human centred approaches to AI evaluation.
The workshop moved from understanding the limitations of existing evaluation approaches to scenario creation, technical setup, live LLM testing, and collective analysis.
The goal was not to argue that benchmarks should be replaced.
It was to demonstrate what becomes visible when technical evaluation is complemented by human experience, context, and participation.
Bringing our research into the conversation
The week also gave us an opportunity to share some of the research behind the questions we brought into our workshop.
During the Research in Africa Days session, our AI Engineer Ertony Bashil presented SenMH QA, Senegal Maternal Health QA, a culturally grounded dataset developed to evaluate how well large language models understand maternal and reproductive health in Senegal.
The work addresses an important challenge in health AI.
A response can be medically accurate while still missing the cultural and social realities that influence how people understand health information and make decisions.
SenMH QA therefore looks beyond general model performance to examine topics including maternal risks, family planning, healthcare access, and culturally influenced health decisions.
The question behind the work reflects much of what we were discussing throughout Indaba:
Can an AI system understand not only the question being asked, but also the context surrounding it?

Ertony Bashil presenting SenMH QA during Research in Africa Days at Deep Learning Indaba 2026.
Sharing knowledge with the next generation
Our participation was also an opportunity to share what we have learned with others building AI across the continent.
During the Breakfast Mentorship Session, Oche David Ankeli, our AI Engineer, joined aspiring AI practitioners for a conversation about building AI that is rooted in African communities.
Participants raised questions about creating culturally relevant datasets, working with African languages, responsible data collection, compute limitations, moving from academia into industry, and the skills companies value in emerging AI researchers and engineers.
What stood out was how quickly the conversation moved beyond simply asking how to build more powerful models.
Young practitioners were asking where data should come from, how communities should participate in AI development, and how technology can be built around the realities of the people expected to use it.
These are exactly the kinds of questions Africa's growing AI ecosystem needs.

Group photo with emerging AI practitioners during the Breakfast Mentorship Session.
Learning from the wider African AI community
While we came to Lagos to contribute our own work, we also came to learn.
Across the week, our team attended conversations on African language models, responsible and sovereign AI, AI safety, governance, multilingual evaluation, and the future of AI research on the continent.
Prof. Vukosi Marivate’s keynote challenged the community to think more carefully about what our benchmarks actually measure, particularly when evaluating African languages.
A session by Dr. David Ifeoluwa Adelani explored the progress and continuing challenges involved in building language models for African languages, while conversations led by Dr. Margaret Mitchell and Blessing Ogbuokiri raised broader questions about responsibility, governance, safety, and whose interests AI ultimately serves.
We left these sessions with a stronger appreciation for something that also sits at the centre of our work at Kitala AI: building AI for African contexts requires us to think just as carefully about evaluation as we do about model development.
It is not enough to ask whether a model is becoming more capable.
We also need to understand who benefits from that capability, where it fails, and whether our methods for measuring progress reflect the languages, cultures, and realities of the people these technologies are meant to serve.

Researchers presenting their work and exchanging ideas during the poster sessions.
What we are taking forward
Our time in Lagos gave us the opportunity to contribute, test ideas, learn from others, and connect our work to a much larger conversation about the future of AI in Africa.
Our workshop demonstrated how participatory evaluation can complement traditional benchmarks. Ertony’s presentation showed why culturally grounded datasets matter when evaluating AI in sensitive areas such as maternal health. Oche’s mentorship session reminded us that building Africa’s AI ecosystem also means investing in the researchers and engineers who will shape what comes next.
The wider conversations throughout the week challenged us to keep asking difficult questions about language, culture, safety, responsibility, and who gets to define what good AI looks like.
We arrived in Lagos asking how we can evaluate AI beyond technical performance, and we left even more convinced that the answer begins with people. The question is no longer simply, “How well does this model perform?” We also need to ask, “For whom does it perform well, in what context, and according to whose definition of success?” These are the questions that continue to shape how we approach AI evaluation at Kitala AI and YUX.

Members of Microsoft Research Africa alongside our team member, Oluchi Audu,