Beyond English-Led AI: How Local Languages Could Unlock New Markets, Jobs and Public Services

AI’s rapid growth risks deepening inequality as thousands of underrepresented languages remain excluded from digital services, markets and innovation. UNDP and Current AI urge governments, development partners and businesses to invest across the full Data-to-AI value chain, ensuring local communities gain both access to AI and a fair share of its economic value.

Beyond English-Led AI: How Local Languages Could Unlock New Markets, Jobs and Public Services
Representative Image.

Artificial intelligence is spreading quickly across economies and public institutions, but its benefits remain highly uneven. A report by the United Nations Development Programme (UNDP) and Current AI, drawing on two years of experience from UNDP's Local Language Accelerator, warns that communities speaking underrepresented languages risk being left behind. The programme has worked across 11 countries, while UNDP's wider AI Landscape Assessments across 26 countries show that the biggest barriers are not only technical. Weak governance, limited infrastructure, fragmented government responsibilities, poor procurement systems and inadequate funding often prevent promising AI projects from reaching people.

Thousands of Languages, but AI Speaks for Only a Few

The world has roughly 7,000 languages, yet most have little representation in AI research and development. The gap is also becoming self-reinforcing: languages with more digital data attract more researchers, producing more models and generating still more data and investment.

English can gain as many as 50,000 new models annually, while languages such as Congolese-Swahili, Wolaytta, Kwangali and Kuanyama average only 4.15 new models a year.

Africa demonstrates the imbalance clearly. A review of 884 research papers on African Language AI found that research output quadrupled over five years, reaching more than 200 papers in 2023. However, over 70% of the African languages covered appeared in fewer than five papers. Swahili, Hausa, Yoruba and Amharic attracted much of the attention despite Africa having more than 2,000 languages, including at least 75 spoken by over one million people.

The inequality has economic consequences. AI increasingly influences access to banking, healthcare, education, agriculture and government services. In Dholuo, for example, a word that takes one token in English can take four or five tokens, increasing the cost of AI services and weakening the commercial incentive to develop applications for underserved users.

The Missing Link Between AI Pilots and Public Services

UNDP proposes a five-stage Data-to-AI Value Chain: data collection, data curation, model training, deployment, and impact and reuse. Its central message for governments and development partners is straightforward: collecting language data is not enough.

Funding often supports datasets and prototypes but stops before financing computing infrastructure, evaluation, deployment and long-term operations. Once grants expire, promising tools can disappear without ever becoming permanent public services.

Governments can change this through procurement. Requiring local-language capability and independent performance benchmarks in public AI contracts would encourage companies to invest in underserved languages while giving domestic startups access to government markets.

Governments also possess valuable language resources in school curricula, broadcast archives, parliamentary debates, civil-registration documents and other public records. Auditing and responsibly releasing appropriate datasets, with privacy and legal safeguards, could turn previously locked information into infrastructure for domestic AI development.

Local Languages Could Unlock New Digital Markets

The economic opportunity extends far beyond translation. Local-language AI could expand access to agricultural advice, financial services, healthcare information, education and government programmes, particularly for rural, elderly and non-literate populations.

The Democratic Republic of the Congo provides an important example. UNDP's #Elikya initiative is developing AI resources for Lingala, spoken by more than 20 million people in the DRC and over 40 million across Central Africa. The programme has produced a linguist-validated corpus containing one million parallel Lingala-French sentences while developing community-driven voice-data collection.

Such resources could eventually underpin voice-enabled agricultural advisory systems, health information lines and digital public services. Reusable datasets and infrastructure can also reduce costs for future startups and researchers, rather than forcing each new project to rebuild the same foundations.

For development partners, this means moving beyond short-term pilot financing. Investment should support shared computing infrastructure, local technical skills, governance systems, model evaluation, deployment and impact measurement. Regional infrastructure could further reduce costs by allowing countries with similar challenges to share resources.

Turning Language Data Into Value for Communities

The private sector has a substantial opportunity in underserved language markets. Fintech, agritech, health technology, education platforms, telecommunications providers and government technology suppliers could reach millions of new users through voice and local-language interfaces.

But the report identifies an equally significant risk: communities can supply voices, stories and cultural knowledge while receiving little of the economic value subsequently created. UNDP therefore calls for Free, Prior and Informed Consent, community-centred licensing, transparent data provenance and benefit-sharing mechanisms. It also highlights fair pay and safer conditions for annotators performing essential transcription, labelling and validation work.

To address these weaknesses, UNDP and Current AI propose Every Language Matters, organised around five priorities: data and community sovereignty, locally governed AI infrastructure and model commons, ecosystem and capacity-building, public-service deployment, and measurement and accountability.

For policymakers, the priority is to establish clear institutional responsibility, unlock appropriate public datasets, reform procurement, finance deployment and measure results. Development partners should finance complete AI ecosystems rather than isolated experiments, while private companies need business models that recognise community rights and return value to those supplying language resources.

The larger development question is therefore no longer simply whether AI can understand thousands of languages. It is whether countries can build the institutions, markets and capabilities needed to turn their linguistic and cultural resources into locally retained economic and social value, rather than supplying the raw material for an AI economy whose benefits are captured elsewhere.

  • FIRST PUBLISHED IN:
  • Devdiscourse
Give Feedback

Use this form for editorial or site feedback. We usually reply within 2 to 3 working days.

By submitting, you agree that we may use your email address to respond.