Artificial intelligence has the potential to become a powerful tool for people around the world, but its usefulness depends heavily on the data it is trained on, and that data is not equally available across languages.
A recent World Bank Data Blog published on August 21, 2026, by Daniel Gerszon Mahler, Senior Economist, Development Data Group, World Bank, examines how well AI represents different languages online and how uneven online representation can affect AI systems.
Mahler points to the uneven availability of data on the internet as a key concern, noting that some voices and information are far harder to find than others.
Languages with a strong online presence, such as English, have much more data available for AI training. English alone accounts for nearly half of all global URLs.
In contrast, many languages spoken in low- and middle-income countries have far less online content, even though they are spoken by hundreds of millions of people.
The chart shows a clear pattern.
.
Languages spoken in high-income countries tend to have a much stronger online presence, while languages predominantly spoken in low- and lower-middle-income countries are consistently underrepresented online.
This imbalance can affect AI systems because they have less data to learn from in these languages. As a result, AI models may be less accurate and lack the accuracy and nuance needed to be useful to people in these regions.
The gap could further marginalize lower-income countries, particularly in places where access to information and knowledge is already limited. Mahler’s analysis argues that closing this gap will require deliberate investment, including expanding text data in underrepresented languages and supporting local AI development capacity.
“AI can become a truly universal technology, but it requires not leaving the languages often spoken in poorer regions of the world behind,” Mahler writes.
The issue is not entirely new. A May 2026 World Bank analysis, “Inequalities in Use of and Exposure to Artificial Intelligence“, had also highlighted the digital divide in language as part of broader inequalities in AI.
It noted that languages primarily spoken in low- and middle-income countries have a significant deficit of online data.
Also Read: Don’t Chase Frontier AI Models, World Bank Tells Low & Middle Income Countries






