Learning Area | Interprefy

Fair Multilingual AI: Why The EU MMLU Benchmark Matters for AI Live Translation

Written by Dayana Abuin Rios | July 30, 2026

Most conversations about AI language performance still centre on English. A model scores well on a standard benchmark, gets praised as state-of-the-art, and is rolled out for global use. But a growing body of evidence, most recently the European Commission's new EU MMLU dataset, shows that strong English performance says very little about how a model handles French, Hungarian or Maltese. For anyone relying on AI live translation at international events, that gap is not an academic footnote. It is the difference between a delegate understanding a keynote and missing it entirely. 

This matters because the benchmarks used to evaluate large language models (LLMs) shape which systems get deployed and trusted. If those benchmarks are built almost entirely in English, they cannot tell us how a model performs everywhere else, including in the languages spoken by the audiences that AI live translation is meant to serve. 

What the EU MMLU Benchmark Is 

EU MMLU is a new multilingual benchmarking dataset released by the European Commission's Directorate-General for Translation (DG Translation). It builds on Massive Multitask Language Understanding (MMLU), one of the most widely used frameworks for evaluating LLMs, which tests models across thousands of multiple-choice questions spanning 57 subjects, from science and law to ethics and public affairs.

EU MMLU adapts this framework specifically for the EU context. It focuses on 7 subject areas chosen for their relevance to European institutions and currently covers 16 EU official languages, including Croatian, Czech, Dutch, French, German, Greek, Hungarian, Irish, Italian, Lithuanian, Polish, Portuguese, Romanian, Slovak and Slovenian, with more planned. Rather than simply translating existing English questions, the project set out to preserve the same meaning, difficulty and testing value across every language version, so that a score in Polish means the same thing as a score in French.

Why a Model Can Excel in English and Stumble Elsewhere 

Most AI evaluation datasets were built in English and reflect English-speaking educational, cultural and societal contexts. A model trained and tested predominantly on English data can appear highly capable while its actual performance in other languages goes largely unmeasured.

This is precisely the gap DG Translation set out to close. As the EU MMLU announcement puts it, a model may score well in English while underperforming significantly in French, Hungarian or Maltese. The underlying issue is not that models cannot process other languages at all. It is that standard benchmarks rarely test them rigorously enough to reveal where they fall short, so weaknesses in grammar, nuance or subject-specific terminology go unnoticed until a real audience is depending on the output.

For AI live translation specifically, this uneven performance has direct consequences. A system might handle straightforward English source material well, only to introduce errors, awkward phrasing or dropped meaning once that same content is rendered into a less-tested target language, precisely the moment when accuracy matters most to a listener following along in real time.

 

How Linguistic and Cultural Bias Shows Up in Live Translation

Language performance gaps are only part of the picture. DG Translation has also published quality criteria for multilingual benchmarking that go beyond translation accuracy to test how well an AI model reflects EU values and handles culturally specific content, including idioms, humour, references, date and number formats, and differences in expected tone or politeness.

These are exactly the elements that make live translation genuinely difficult. A conference speaker who uses a regional idiom, a culturally specific reference, or a level of formality tied to local convention is relying on more than word-for-word conversion. An AI system without exposure to this kind of nuance in a given language may translate the literal words correctly while losing the intended meaning or tone entirely. At a multilingual event, that kind of gap can change how a message lands with part of the audience, even when nobody notices anything is technically wrong.

Why Human-Reviewed Multilingual Data Still Matters

One of the more notable aspects of EU MMLU is its methodology. Unlike most multilingual benchmarks, which rely primarily on machine translation to generate non-English test questions, EU MMLU used a human-centred approach. DG Translation worked with nearly 250 student translators and project managers from the European Master's in Translation (EMT) network, drawn from 21 universities across Europe, to translate and revise more than 1,000 benchmark questions.

This distinction matters more than it might first appear. A benchmark built through machine translation risks inheriting the same weaknesses it is meant to expose, testing a model's output against another model's output rather than against a genuine human standard. Human review anchors the evaluation in how people actually use language, which is also the standard that matters most for live translation at real events, where the audience is human and the tolerance for subtle mistranslation is low.

What Fairer AI Evaluation Could Mean for Multilingual Events  

DG Translation frames its long-term ambition clearly: establishing a new standard for multilingual AI evaluation in Europe, so that AI systems perform consistently across languages and cultures rather than primarily in English-speaking environments. If that ambition is realised, the practical benefits extend well beyond research labs.

For organisations running international conferences, institutional meetings or global corporate events, better multilingual benchmarking means clearer visibility into which AI systems can be trusted for which languages, rather than a mid-event discovery that a particular tool underperforms outside English. It supports more informed decisions about where AI-only live translation is reliable, and where a hybrid approach combining AI with human interpretation better protects meaning and nuance for a specific language or audience.

Benchmarks like EU MMLU will not resolve every gap in multilingual AI overnight, but they represent a meaningful shift in how performance gets measured and communicated. As these evaluation standards mature, they give event organisers, institutions and communication teams a clearer basis for choosing AI live translation solutions suited to genuinely multilingual audiences. To see how Interprefy approaches AI live translation across languages, explore our AI speech translation platform.