top of page
Aiqara - AI Powered qa translation.jpg

What is Aiqara?

Aiqara is an AI-powered translation QA solution that helps brands and teams protect clarity, consistency, and trust across languages. It reviews translated content, highlights potential issues, and provides scores and explanations to support a more informed QA process.

Translation Quality Evaluation in the AI Era: Standards, Frameworks and the Role of Human Review

  • Jul 28
  • 16 min read

Translation quality assurance has changed significantly over the past few years.

For a long time, many localization teams relied on relatively simple models: identify an error, assign it a category, decide whether it is minor or major, calculate a score, and determine whether the translation passes or fails.

That basic structure still matters. But today’s localization workflows are more complex.

Teams may be evaluating:


  • fully human translations;

  • translation memory matches;

  • machine-translated content;

  • post-edited machine translation;

  • AI-generated suggestions;

  • content reviewed by linguists;

  • high-volume multilingual material that cannot all receive the same level of human attention.


This has changed the central question in language quality assurance.

The question is no longer only:

Is this translation good?

It is also:

Good for which purpose, audience, market, risk level and workflow?

A quality framework helps teams answer that question consistently. AI can support the process, but it cannot replace the need to define what quality means for a particular project.


Translation Quality Evaluation in the AI Era

This article looks at the main standards and frameworks used in translation quality evaluation, how they differ, and how AI-assisted tools such as Aiqara can support a structured QA process while keeping human judgement in the loop.


From general LQA models to modern quality evaluation


Earlier approaches to language quality assurance often focused heavily on error counting and penalty scoring. Frameworks such as the LISA QA Model, SAE J2450, TAUS DQF and MQM gave organizations more structured ways to classify translation problems and compare quality results.

Those models still form part of the history of translation quality management, but the industry has moved further.


Three changes are especially important:

  1. Translation output now comes from more than one production method.

  2. Quality expectations increasingly depend on content purpose and risk.

  3. AI can now assist with evaluation, but it can also introduce false positives, inconsistent judgement and overconfidence.


Modern QA therefore needs more than an error checklist. It needs:

  • clear project specifications;

  • relevant error types;

  • defined severity rules;

  • appropriate sampling;

  • qualified evaluators;

  • reliable human oversight;

  • and a way to use automation without treating automated judgement as unquestionable.


ISO 5060:2024 and the translation quality evaluation output


One of the most important recent developments is ISO 5060:2024, Translation services — Evaluation of translation output — General guidance.

ISO 5060 provides guidance for evaluating:

  • human translation output;

  • post-edited machine translation output;

  • unedited machine translation output.


It also discusses evaluator qualifications, evaluator competence and the role of sampling. The standard focuses on analytical evaluation using error types and penalty points that can be configured to produce an error score and an overall quality rating.

That scope matters because modern localization projects rarely use only one production method.

A single project might include:

  • human-translated high-visibility marketing content;

  • translation memory matches for repeated legal text;

  • machine translation for large support libraries;

  • AI-generated suggestions for selected segments;

  • and human post-editing for customer-facing content.


A useful evaluation approach must be capable of reviewing all of those outputs without pretending that they were created in the same way or carry the same level of risk.

ISO 5060 does not tell every organization to use one identical scorecard. Instead, it supports a structured analytical approach. That means the organization still has to decide:

  • which error types matter;

  • which severity levels will be used;

  • how many penalty points each severity carries;

  • what constitutes a pass or fail;

  • and how the evaluation sample will be selected.


This is an important principle:

A standard can provide a disciplined evaluation method, but the project still needs its own quality profile.

ISO 5060 is not the same as ISO 17100


ISO standards are often mentioned together, but they do not all cover the same part of the process.

ISO 17100:2015 focuses on the requirements for providing translation services. It covers core processes, resources and other aspects required to deliver translation services that meet applicable specifications.

That means ISO 17100 is primarily about the translation service process.

It addresses questions such as:

  • Are appropriate resources being used?

  • Are the required translation and revision processes defined?

  • Are the people involved suitably qualified?

  • Are client specifications being followed?

  • Is the translation service being delivered through a controlled process?


By comparison, ISO 5060 focuses more directly on evaluating the output itself.

A simple way to distinguish them is:

Standard

Main focus

ISO 17100

Translation-service processes and resources

ISO 5060

Analytical evaluation of translation output

ISO 18587

Full human post-editing of machine translation output

These standards can complement each other, but they should not be treated as interchangeable.


ISO 18587 and machine translation post-editing


ISO 18587:2017 applies specifically to the process of full human post-editing of machine translation output and to the competencies of post-editors. It is intended for translation service providers, clients and post-editors, and applies only to content that has been processed by machine translation.

This matters because post-editing is not simply ordinary translation performed more quickly.


A post-editor may need to deal with:

  • fluent but inaccurate machine output;

  • omissions that are difficult to notice;

  • terminology that looks plausible but is wrong;

  • repeated structural patterns introduced by the engine;

  • inconsistent register;

  • hallucinated content;

  • and wording that is grammatically sound but inappropriate for the market.


For example:

Source: Select Save to keep your changes.

Machine translation output: Select Delete to keep your changes.


The target may be fluent and grammatically correct, but it contains a serious action error.

A quality model that focuses too heavily on surface fluency might miss the real risk. A human evaluator or a properly configured AI-assisted QA system should identify the change from Save to Delete as an accuracy problem with potentially major impact.

ISO 18587 helps define the post-editing process. ISO 5060 helps structure the evaluation of the resulting output.


MQM: a practical framework for analytical evaluation


Multidimensional Quality Metrics, or MQM, is one of the most relevant frameworks for modern translation quality evaluation.

The MQM Council describes MQM as a framework for analytical translation quality evaluation that can be applied to human translation, machine translation and AI-generated translation.

An analytical evaluation looks for specific issues rather than judging the text only as a whole.

The evaluator identifies:

  • the location of the problem;

  • the type of error;

  • the severity of the error;

  • and, where appropriate, the reason or root cause.


The MQM error typology contains seven high-level dimensions, with more detailed subtypes underneath them. For example, Accuracy includes error types such as Addition, Omission and Mistranslation.

A simplified project profile might use categories such as:

  • Accuracy

  • Terminology

  • Linguistic conventions

  • Style

  • Locale conventions

  • Audience appropriateness

  • Design and markup


A team does not necessarily need every available subtype.

That is one of MQM’s strengths: the framework can be tailored to the project.


MQM Core and MQM Full


MQM offers different levels of detail.

MQM Core provides a more manageable set of commonly needed error types. It is useful when an organization wants consistency without creating an overly complex scorecard.

MQM Full provides much greater granularity. It can be useful for:

  • complex research programmes;

  • specialist domains;

  • detailed root-cause analysis;

  • large enterprise quality programmes;

  • or projects where the organization needs to distinguish closely related problem types.

The danger of an overly detailed model is that reviewers may spend more time deciding which label to use than evaluating the actual risk.


For example, imagine this translation:

Source:The device must remain switched off during installation.

Target:The device should remain switched off during installation.


A detailed framework might encourage discussion about whether this is:

  • a mistranslation;

  • a change in modality;

  • a weakened obligation;

  • a technical relationship error;

  • or an instruction-related issue.


That level of detail may be valuable in a technical or safety-critical project.

In a lower-risk project, it may be enough to classify it as:

  • Accuracy

  • Major

The right level of detail depends on what the data will be used for.

If the purpose is simply to determine whether the translation is acceptable, a smaller model may be enough.

If the purpose is to identify recurring engine weaknesses, compare vendors or improve a terminology process, more detailed classification may be worthwhile.


Error type and severity are not the same thing


One of the most important principles in quality evaluation is that error category and severity must be separated.

The category describes what kind of problem occurred.

The severity describes how much the problem matters in context.


Consider these examples:


Example 1: Terminology issue with low impact

Source: Open the settings menu.

Target: Open the options menu.


The project terminology says settings must always be translated using the approved equivalent.

This may be a terminology issue, but if the user still understands the instruction and no functional confusion results, it may be minor.


Example 2: Terminology issue with high impact

Source: Do not disconnect the grounding cable.

Target: Do not disconnect the power cable.


This is also a terminology or accuracy issue, but its impact is much greater. The target identifies the wrong cable and may create a safety risk.

The category alone does not tell us enough.

Severity must reflect:

  • the intended purpose of the content;

  • the likely effect on the user;

  • whether the error changes meaning;

  • whether it creates legal, medical, technical or financial risk;

  • and whether the content remains usable.


Minor, major and critical errors in practice


MQM scoring models use error types and severity levels as inputs to quality scores. Implementers select error types relevant to the project and assign severity values and weights according to their specifications.

The exact definitions should be documented before evaluation begins.

A practical model might look like this:


Minor

A minor issue does not materially change the meaning or prevent use of the content.

Examples:

  • slightly awkward wording;

  • a punctuation issue that does not cause ambiguity;

  • a minor inconsistency in capitalization;

  • an isolated style-guide deviation;

  • a non-preferred synonym that remains clear.

Example:

Source: Your account has been updated successfully.

Target: Your account was updated successfully.

Depending on context, this may be acceptable or at most a minor tense or style issue.


Major

A major issue changes an important part of the meaning, creates confusion, prevents proper use or violates an important project requirement.

Examples:

  • an omitted instruction;

  • an incorrect date or amount;

  • a wrong action;

  • a material mistranslation;

  • inconsistent approved terminology in a critical location;

  • an untranslated segment;

  • a broken placeholder;

  • a tone mismatch that makes regulated or official content inappropriate.

Example:

Source: Select Cancel to leave without saving.

Target: Select Save to leave without saving.

The user is told to take the opposite action. This should not be treated as a small wording problem.


Critical

A critical issue can make the content unusable or create serious harm.

The MQM Council’s example guidance describes critical errors as issues that can render a text unusable and may cause harm to people, equipment or an organization’s reputation. In some scoring models, one critical error can result in an automatic fail.

Examples may include:

  • a medical dosage changed incorrectly;

  • a dangerous safety instruction;

  • a legal obligation reversed;

  • an incorrect allergen statement;

  • a severe financial error;

  • content that exposes confidential information;

  • or a translation that creates substantial reputational or regulatory risk.

Example:

Source: Take one tablet once daily.

Target: Take ten tablets once daily.

This is not simply an accuracy issue. Its severity comes from the possible consequence.


Not every difference is an error


This is where both human reviewers and AI systems can fail.

A translation does not have to mirror the source word for word to be correct.

Consider:

Source: Please enter your email address.

Target: Enter your email.

Depending on the language, interface and context, the target may be completely acceptable.

The following should not automatically be treated as errors:

  • natural restructuring;

  • omission of unnecessary politeness;

  • a shorter UI label;

  • use of a valid local convention;

  • an accepted spelling variant;

  • adaptation for space;

  • or a wording choice that preserves the same user action.


A QA process becomes noisy when it flags every difference from the source.

That is especially dangerous with AI-assisted evaluation because AI systems can produce confident explanations for issues that are not actually wrong.

Good QA must therefore ask:

  • Is there a genuine problem?

  • Does it affect meaning, use, consistency or market appropriateness?

  • Is the form actually prohibited by the project style guide?

  • Could this be a valid linguistic or locale-specific variant?

  • Does the reviewer have enough context to judge it?

Sometimes the right QA decision is to leave correct content alone.


Locale conventions belong in quality evaluation


Many translation issues are not grammatical errors. They are locale-convention issues.

Examples include:

  • decimal commas versus decimal points;

  • date order;

  • currency-symbol placement;

  • quotation marks;

  • spaces before punctuation;

  • full-width versus half-width characters;

  • capitalization of months and days;

  • right-to-left punctuation behaviour;

  • and local forms of address.

Consider a German target:

4.5 kg

The meaning is obvious, but a German-language product may require:

4,5 kg

Or consider Spanish:

Marzo 2026

The month may be incorrectly capitalized if the context does not require a capital:

marzo de 2026

Or Japanese:

(設定)

A project may require the Japanese full-width brackets:

(設定)


However, this does not mean the reviewer should apply one global rule mechanically. The project’s locale, content type, platform behaviour and style guide matter.

The correct question is not:

Does this look different from English?

It is:

Does this follow the target locale and the agreed product style?

Sampling: reviewing less without learning less

Full human evaluation of every segment may be appropriate for:

  • short content;

  • safety-critical material;

  • legal obligations;

  • public-health information;

  • high-visibility campaigns;

  • or content where the cost of an undetected error is high.

For large, lower-risk volumes, full review may not be practical.

ISO 5060 explicitly discusses sampling as part of translation output evaluation.

However, sampling must be designed carefully.

A weak sample can create a false sense of quality.

For example, selecting only:

  • short segments;

  • repetitions;

  • high translation-memory matches;

  • or content from the easiest section

may produce an excellent score while ignoring the parts where errors are more likely.

A more useful sample may consider:

  • content length;

  • risk level;

  • document section;

  • translator or vendor;

  • source complexity;

  • presence of terminology;

  • numbers and placeholders;

  • language direction;

  • and production method.

Sampling should also match the purpose of the evaluation.

A vendor-performance sample may need broad representation.

A release-readiness sample may need to focus on high-impact user journeys.

A machine-translation evaluation may need enough difficult segments to reveal recurring weaknesses.

AI-assisted tools can support this process by identifying potentially risky content, but the sampling logic still needs human design.


How AI can support a structured QA process


AI can add real value to translation QA when it is used as an assistant rather than as an unquestioned authority.

A properly configured AI-assisted QA process can help:

  • compare source and target meaning;

  • identify possible omissions and additions;

  • flag likely mistranslations;

  • check terminology against project expectations;

  • detect inconsistent tone or register;

  • identify suspicious locale conventions;

  • highlight placeholder or tag risks;

  • provide an initial category;

  • suggest severity;

  • and help reviewers prioritize segments.


This is particularly valuable in large projects where human reviewers cannot inspect every segment with the same level of attention.

The useful role of AI is not:

Decide the final truth for every translation.

It is:

Help the reviewer find where closer attention may be valuable.

What AI should not decide on its own

AI should not independently define:

  • what counts as acceptable quality;

  • which categories matter to the client;

  • what severity scale should be used;

  • whether a local variant is permitted;

  • which terms are approved;

  • whether a regulatory risk is critical;

  • or whether a client’s style preference is mandatory.


Those decisions depend on project context.

For example, the same terminology inconsistency may be:

  • minor in an internal draft;

  • major in customer-facing software;

  • critical in a medical device instruction.


AI sees text. It does not automatically know the organization’s contractual, legal or operational risk unless that information has been provided clearly.


The risk of false positives in AI QA


A QA system that flags too much can become as unhelpful as one that misses errors.

False positives create several problems:

  • reviewers stop trusting the output;

  • real issues are hidden among weak flags;

  • review time increases instead of decreasing;

  • translators spend time defending correct choices;

  • and the QA process becomes confrontational.


Common false-positive patterns include:

  • flagging a natural paraphrase as an omission;

  • treating a valid locale variant as incorrect;

  • confusing style preference with mistranslation;

  • objecting to necessary UI compression;

  • interpreting placeholders as natural language;

  • misunderstanding right-to-left display;

  • and assuming punctuation must copy the source exactly.


A useful AI QA system must therefore be configured to reduce noise, not simply maximize the number of flags.


The role of context

Individual segments can be misleading.

Consider:

Source: Continue.

Without context, that could mean:

  • continue to the next screen;

  • resume a paused action;

  • keep reading;

  • proceed with payment;

  • or confirm a process.


The appropriate translation may depend entirely on what came before and what happens next.


Context can include:

  • neighbouring segments;

  • screen title;

  • button function;

  • document section;

  • product terminology;

  • user journey;

  • or previous translations.


This is one reason context-aware QA is more useful than isolated sentence checking.

However, more context is not automatically better. Irrelevant context can distract an evaluation system. The goal is to provide enough surrounding information to support the judgement without overwhelming the model.


From generic QA to project-specific quality profiles


A good quality framework should reflect the content.

A software UI project might prioritize:

  • accuracy;

  • approved terminology;

  • placeholder preservation;

  • consistency;

  • locale conventions;

  • concise wording;

  • and screen fit.


A marketing project might prioritize:

  • tone;

  • audience impact;

  • brand voice;

  • naturalness;

  • cultural adaptation;

  • and persuasive effect.


A medical project might prioritize:

  • dosage;

  • warnings;

  • omissions;

  • unambiguous terminology;

  • numbers;

  • contraindications;

  • and safety-critical instructions.


A legal project might prioritize:

  • obligations;

  • defined terms;

  • dates;

  • names;

  • references;

  • modality;

  • and completeness.


The same scorecard should not be copied blindly across all four.


Example quality profile: software UI

Area

What the reviewer checks

Typical impact

Accuracy

Correct action and meaning

Major if user action changes

Terminology

Approved UI and product terms

Minor or major

Placeholders

Variables and tags preserved

Usually major

Locale conventions

Numbers, dates and punctuation

Minor or major

Style

Consistent voice and register

Usually minor

UI fit

Text fits buttons and screens

Minor or major

Consistency

Same feature named consistently

Minor or major


Example quality profile: medical content


Area

What the reviewer checks

Typical impact

Accuracy

Clinical meaning preserved

Major or critical

Numbers

Dosage, frequency and measurements

Major or critical

Omissions

Warnings and contraindications retained

Critical

Terminology

Approved medical terms

Major

Ambiguity

Instructions cannot be misread

Major or critical

Formatting

Tables, labels and symbols preserved

Major


This is what structured QA should accomplish: not merely produce a score, but reflect actual project risk.


Where older frameworks still fit


Earlier quality models remain useful historically and, in some cases, operationally.


LISA QA Model

The LISA QA Model was influential because it offered a clear error-categorization and penalty approach. It helped normalize the idea that translation quality could be evaluated systematically rather than only through general impressions.

However, it is no longer actively maintained and is less adaptable than modern MQM-based approaches. It is better understood today as an important predecessor rather than the default framework for new programmes.


SAE J2450

SAE J2450 was developed for automotive and technical translation, with a strong focus on terminology, omissions, syntactic issues, word structure and spelling.

It can still be relevant where:

  • technical precision matters;

  • documentation is highly structured;

  • terminology consistency is critical;

  • and teams need a stable, domain-oriented error metric.

Its weakness is that it may be too narrow for modern content types such as marketing, conversational UI, multimedia or culturally adaptive content.


TAUS DQF

TAUS DQF broadened quality evaluation beyond a static scorecard and supported more operational and data-driven quality management.

TAUS describes the DQF error typology as the harmonized DQF-MQM typology.

The history is important because it shows that the industry has gradually moved away from isolated, proprietary taxonomies and toward more interoperable approaches.


How Aiqara fits into this evolution

Aiqara has grown out of this changing quality environment.

The aim is not to replace the quality framework or replace human reviewers.

The aim is to help teams apply quality criteria more efficiently across real localization content.

Aiqara supports structured AI-assisted assessment across areas such as:

  • Accuracy

  • Fluency

  • Terminology

  • Style

  • Formatting

  • Source-text issues


It can assess source and target content, use neighbouring context, provide an explanation, identify possible error types and generate a score for review.

The intended workflow is not:


  1. AI produces a judgement.

  2. The judgement is automatically treated as final.


It is:

  1. The project defines the quality criteria.

  2. Aiqara applies those criteria across the content.

  3. Potential issues are surfaced with explanations.

  4. Human reviewers decide what is valid, what needs correction and what should be dismissed.

  5. The results become part of a more focused review process.


This distinction matters.

Aiqara is not the standard.

It is a tool that can help teams apply their chosen quality approach more consistently and at greater scale.


Why human review remains essential

No quality framework eliminates judgement.

Even well-defined rules require interpretation.

A reviewer may need to decide:

  • whether an omission is intentional;

  • whether a shorter UI translation is appropriate;

  • whether a tone shift is acceptable;

  • whether the target locale allows more than one form;

  • whether a terminology change affects meaning;

  • or whether an apparent formatting issue is caused by the platform rather than the translation.

Native-language expertise is especially important.

As recent Aiqara educational posts have shown, a Chinese example can be typographically correct but use the wrong label for a person’s name. A Japanese punctuation character can be correct at the Unicode level but rendered incorrectly by a platform font. A Korean punctuation form may look unusual while still being explicitly permitted by the official rule.

Those are not reasons to reject AI.

They are reasons to design AI-assisted QA around review, traceability and correction rather than blind acceptance.


A practical model for AI-assisted QA


A reliable workflow could look like this:

Step 1: Define the content purpose

Is the content:

  • informational;

  • transactional;

  • promotional;

  • legal;

  • medical;

  • technical;

  • or internal?


Step 2: Define the quality criteria

Choose only the criteria that matter.

For example:

  • Accuracy

  • Terminology

  • Fluency

  • Locale conventions

  • Formatting

  • Style


Step 3: Define severity

Document what minor, major and critical mean for the project.

Do not leave severity entirely to individual preference.


Step 4: Provide supporting information

Where possible, supply:

  • terminology;

  • style guidance;

  • locale;

  • content type;

  • intended audience;

  • neighbouring context;

  • and product constraints.


Step 5: Run AI-assisted evaluation

Use AI to surface possible issues and prioritize review.


Step 6: Validate the findings

Human reviewers confirm, reject or correct the findings.


Step 7: Learn from the result

Review recurring issues:

  • Which problems were genuine?

  • Which flags were false positives?

  • Which criteria created noise?

  • Which languages need more specific instructions?

  • Which project rules should be clarified?

This makes QA a feedback loop rather than a one-time inspection.


What quality scores can and cannot tell you

Scores are useful because they summarize complex review data.

They can help teams:

  • compare projects;

  • monitor trends;

  • assess vendors;

  • identify recurring weaknesses;

  • prioritize corrective work;

  • and report quality at scale.


But a score is only meaningful if the model behind it is meaningful.

A score of 92 does not automatically tell us:

  • whether the content is safe;

  • whether one critical error exists;

  • whether the sample was representative;

  • whether the evaluator was consistent;

  • or whether the error weights matched project risk.


Two projects can receive the same score for completely different reasons.

One may contain several minor style issues.

The other may contain one serious mistranslation.

That is why a score should be accompanied by:

  • error categories;

  • severity;

  • explanations;

  • sample size;

  • evaluation scope;

  • and, where relevant, critical-issue handling.


The future of translation QA is structured, assisted and human-led

Translation quality evaluation is moving away from two extremes.

The first extreme is purely subjective review:

This sounds good to me.

The second is blind automation:

The system gave it a high score, so it must be acceptable.

Neither is enough.

The more reliable approach combines:

  • structured standards;

  • project-specific criteria;

  • relevant error typologies;

  • calibrated severity;

  • qualified reviewers;

  • intelligent automation;

  • and continuous feedback.

ISO 5060 gives the industry a current framework for analytical translation output evaluation. MQM provides a flexible and detailed way to classify issues. ISO 17100 and ISO 18587 support the translation and post-editing processes around that evaluation. AI tools can help apply these ideas at scale.

But quality still depends on people deciding what matters.

That is the direction Aiqara is built around: using AI to make translation QA more focused and manageable while keeping human expertise at the centre of the final decision.


Sources


Friendly note

If you work with these standards or frameworks and spot anything that needs clarification, please shine a light on it and let us know. Translation quality develops through shared expertise, and we always appreciate thoughtful corrections from people working directly in the field. You can contact us through the Aiqara website.


Read more:

 
 

Request an Aiqara QA pilot

© 2026 by Aiqara. 

bottom of page