AI-MT Quality Evaluation: MQM Deep Dive - Part 2
AI-MT-mqm-error-categories
How Are AI-MT Errors Classified? — The MQM Category Breakdown
# Overview: The 10 MQM Error Categories
Q. I've heard MQM has many error types. What's the top-level structure?
A. MQM organizes errors into 10 top-level categories, each with one or more sub-types. Here is the full overview:
| # | Category | Core Question |
|---|---|---|
| 1 | Accuracy | Is the source meaning faithfully conveyed? |
| 2 | Terminology | Are specialized terms used correctly? |
| 3 | Linguistic Conventions | Are grammar, spelling, and punctuation correct? |
| 4 | Style & Fluency | Does it read naturally and fit the audience? |
| 5 | Consistency | Is the same concept rendered the same way throughout? |
| 6 | Locale Conventions | Is date/currency/unit formatting localized? |
| 7 | Compliance | Does it follow the client style guide and project instructions? |
| 8 | Format & Non-Translatables | Are tags, code, and layout handled correctly? |
| 9 | Source Error | Is the problem in the source text, not the translation? |
| 10 | Other | Anything that doesn't fit the above categories? |
1. Accuracy — The Most Critical Category
Q. Accuracy seems like the most important category. What sub-types does it have?
A. Accuracy is indeed the most impactful category in AI-MT evaluation — it's where the most dangerous errors lurk. It has five sub-types:
1.1 Mistranslation
The translation does not accurately reflect the meaning of the source. This includes distorted or ambiguous meaning, overly literal renderings that lose the original intent, and the use of non-existent or foreign words in the target text.
1.2 Omission
Content present in the source is missing from the translation — such as omitted sections, incomplete sentences, or dropped clauses. Note: missing punctuation falls under 3.3; missing tags or placeholders fall under 8.1.
1.3 Addition
Content not present in the source has been added to the translation. This includes AI hallucination — where the model invents plausible-sounding but fabricated content. In production environments, hallucinated additions are almost always Major or Critical.
1.4 Untranslated / Overtranslation
Source language content has been left untranslated when it should have been rendered in the target language, or conversely, items that should remain as-is (variables, brand names, etc.) have been translated. Note: if a term is specified as 'keep in English' in the glossary, a violation is classified as 2.1 Glossary Compliance, not 1.4.
1.5 Numbers & Symbols
Verifiable elements such as numbers, dates, or trademark symbols (©, ®, ™) do not match the source. Especially dangerous in legal and medical contexts.
2. Terminology
Q. How is a Terminology error different from a Mistranslation? They seem similar.
A. The key distinction is whether the term in question is governed by an approved glossary or termbase:
| Scenario | Correct Classification |
|---|---|
| Wrong word used — term IS in the client glossary | Terminology |
| Wrong word used — term is NOT in the glossary | Mistranslation |
| Glossary term used inconsistently across the document | Term Inconsistency |
| Non-glossary expression used inconsistently | Internal Consistency |
2.1 Glossary Compliance
The translation conflicts with an approved glossary or termbase — including incorrect spelling or case. This also covers cases where a term flagged 'keep in source language' was translated, or vice versa.
2.2 Industry-Standard Terminology
The translation does not follow widely accepted industry-standard terminology or recognized third-party product terminology (e.g., WHO medical terms, official software UI strings).
2.3 Term Inconsistency
A glossary-listed term is rendered inconsistently across the document — appearing in one form in one segment and another form elsewhere.
3. Linguistic Conventions
Q. What counts as a Linguistic Conventions error? Is this just grammar?
A. Linguistic Conventions covers the mechanical rules of the target language — grammar, spelling, punctuation, and capitalization. Four sub-types:
3.1 Grammar / Syntax
Violations of the target language's grammatical or syntactic rules — subject-verb disagreement, incorrect case, wrong word order, etc.
3.2 Spelling / Typos
Spelling errors and typos. For Asian languages, this also includes incorrect spacing rules.
3.3 Punctuation
Incorrect, missing, or spurious punctuation marks. Always classify here — unless the punctuation error actually breaks the grammar of the sentence, in which case reclassify as 3.1.
3.4 Capitalization / Hyphenation
Incorrect capitalization, accent marks, or hyphenation per target language norms. Reclassify to 3.1 only if the error breaks the grammar.
4. Style & Fluency
Q. What if the grammar is correct but the text still feels unnatural? Where does that go?
A. This is exactly what the Style & Fluency category captures — grammatically correct but still problematic output, which is one of AI-MT's most common failure modes.
4.1 Fluency / Naturalness
The text is grammatically correct but reads as stiff, overly literal, or non-idiomatic. This includes 'translationese' — the robotic, word-for-word style that plagues AI-MT output and hurts readability.
4.2 Tone / Register
The tone or register is inappropriate for the target audience or content type. Examples: formal documentation rendered in casual language, a children's game tutorial written in stiff legalese, or inconsistent mixing of formal and informal address within a document.
5. Consistency
Q. What's the difference between Consistency and Terminology errors?
A. The dividing line is the glossary:
| Inconsistency Type | Classification |
|---|---|
| A glossary term rendered differently across segments | Terminology — Term Inconsistency |
| A non-glossary word/phrase rendered differently across segments | Internal Consistency |
| A previously approved TM match ignored in favor of a new translation | External Consistency |
| Formal/informal register mixed within a single text | Style — Tone / Register |
6. Locale Conventions
Q. Is localization format the same as style? How does this differ from Compliance?
A. Locale Conventions covers locale-specific formatting norms — how numbers, dates, currencies, and addresses should be presented in the target market. It is distinct from Compliance (category 7), which covers adherence to client-provided style guides and project instructions.
| Error Type | Category |
|---|---|
| Date written as MM/DD/YYYY instead of the target locale's DD.MM.YYYY | Locale Conventions |
| Metric units not converted to imperial for a US audience | Locale Conventions |
| Cultural reference (legal title, proverb) not adapted for target audience | Cultural Adaptation |
| Client's preferred comma usage style not followed | Style Guide (Compliance) |
7. Compliance
Q. What falls under Compliance? Is this the same as Style & Fluency?
A. No — Compliance is about following client-provided project documents, while Style & Fluency is about inherent linguistic quality. Two sub-types:
7.1 Style Guide
The translation does not follow the client-provided style guide for this project. Locale-specific formatting issues should be classified under category 6 instead.
7.2 Project Instructions & Feedback
The translation does not follow the translation brief or Q&A file, or previously confirmed query answers and QA feedback have not been applied. This is a Major-level error in most frameworks.
8. Format & Non-Translatables
Q. Category 8 seems very technical. Why does it matter in translation evaluation?
A. Translation files frequently contain tags, placeholders, code snippets, URLs, and HTML entities that the translator must pass through unchanged. AI-MT systems often mistakenly translate, move, or drop these elements — causing software to malfunction.
8.1 Tags, Placeholders & Code
Formatting tags, URLs, code, parameters, HTML entities, or placeholders have been translated, modified, repositioned, deleted, or spuriously added. Example: {hero_name} translated as 'hero name' instead of being left as the variable.
8.2 Length, Spacing & Encoding
String length limit violations, incorrect whitespace, broken characters, or wrong file encoding (UTF-8, Unicode, etc.).
8.3 Layout & Typography
Incorrect bold/underline/italic/all-caps formatting, wrong font or language code, GUI clipping due to unadjusted control size, incorrect numbering in the final formatted file.
9. Source Error — Not a Translation Error
Q. What is a Source Error, and why is it included in MQM?
A. During evaluation, reviewers sometimes discover that the real problem lies not in the translation but in the source text itself. Category 9 exists so that this can be recorded without penalizing the translation.
Source Errors include: ambiguous or incomplete source text, source text that is incomprehensible or factually incorrect, and segmentation issues that interfere with translation quality. Source Errors are ALWAYS classified as Neutral — they do not affect the translation score.
Boundary Rules: How to Classify Ambiguous Cases
Q. Some error types seem to overlap. Are there rules for borderline cases?
A. Yes — consistent boundary rules are essential for inter-rater reliability. Here are the most important ones:
| Borderline Scenario | Correct Classification |
|---|---|
| Punctuation error that doesn't affect grammar | Punctuation |
| Punctuation error that breaks the sentence's grammar | Grammar |
| Wrong word — term is in the approved glossary | Terminology |
| Wrong word — term is not in the glossary | Mistranslation |
| Glossary term used inconsistently | Term Inconsistency |
| Non-glossary expression used inconsistently | Internal Consistency |
| Formal/informal register mixed inside a single text | Tone / Register |
| Same concept rendered differently across segments | Internal Consistency |
| Tag, placeholder, or URL missing or modified | Omission or Addition |
| Date/currency format not localized | Locale Conventions — NOT Compliance |
| Content left in source language (should be translated) | Untranslated |
| Glossary term specified 'keep in English' — was translated | Glossary Compliance |
Explore more insights
View All Articles
Connecting People, Through Language
Professional language services for global success
Partners
Trusted by leading brands worldwide











































