Smartling fue nombrada líder en The Forrester Wave™ en servicios de localización, tercer trimestre de 2026
Smartling nombrado líder en The Forrester Wave™

How do you choose MQM translation quality scoring software?

MQM translation quality scoring software is a tool that logs translation errors against the Multidimensional Quality Metrics (MQM) error typology, weights each error by severity, and converts the result into a pass/fail quality score. Choose it on five tests: how faithfully it implements the MQM typology and scoring models, how far severity weights and thresholds can be configured per content type, whether its AI scoring has been calibrated against human reviewers for your locales, whether it produces string-level evidence an auditor or vendor can check, and whether scoring runs inside your translation workflow instead of beside it. Those five tests decide whether a score survives a dispute, which matters more than the feature list.

Last reviewed: October 7, 2026

Why do two MQM tools give the same translation different scores?

Two MQM tools can score the same translation differently because MQM is a framework with configurable parts, not a single fixed formula. The MQM Council, which stewards the framework, publishes an open error typology and several scoring models, and each tool makes its own choices inside them. (For how a single score is calculated, see how translation quality scoring works.)

  • Typology subsets differ. A tool can use the MQM Core or a reduced subset. Smartling's own templates show the range: Simplified MQM groups errors into five categories, while Full MQM uses eight, including Locale Conventions and Design & Markup. Fewer categories make evaluation faster but hide where errors come from.
  • Severity weights differ. Smartling's default weights are Critical 25, Major 5, Minor 1 and Neutral 0. A tool that weights a Major error at 10 will report a lower score for identical work, so a score is only comparable alongside its weights.
  • Raw and calibrated scores differ. With Smartling's default Acceptable Penalty Points of 20, the Raw Quality Score passing threshold is 98 and the Calibrated Quality Score passing threshold is 80. A "95" is a fail on one scale and a strong pass on the other.
  • AI scorers and humans differ. Smartling documents that its LQA Agent tends to score more strictly than a human evaluator. Any AI scorer needs its strictness measured against your reviewers before its numbers drive vendor decisions.
  • Repeated and preferential errors are handled differently. Some schemas count every repeat of the same error; others record repeats at Neutral severity so one mistake does not sink a whole file.

What capabilities should MQM scoring software have?

  • Typology fidelity — The tool should offer the full MQM catalog as well as a streamlined option, so a regulated legal file and a product UI string can be evaluated at the depth each needs. The MQM Council describes the framework as aligned with ISO 5060, which gives procurement a standard to point to.
  • Configurable weights and thresholds per content type — Marketing, legal and product documentation carry different risk, so each needs its own severity weights and pass threshold. In Smartling, a schema is assigned to each workflow step, so different content streams can be scored against different schemas.
  • Locked schemas once in use — Smartling allows weights and Acceptable Penalty Points to be edited only while a schema is a draft. That constraint is a governance feature: published scores cannot be quietly re-weighted after a vendor dispute.
  • Error neutralization — Repeated and preferential errors should be recordable without affecting the score. Smartling lets teams neutralize specific error types while keeping them visible in the LQA Errors & Arbitration report.
  • Calibrated AI scoring — AI scoring should state its accuracy against human reviewers, the locales where it is reliable, and what it cannot detect. Without those limits on the record, an AI score is a guess with a decimal point.
  • Arbitration and string-level evidence — Every error should carry its category, severity and reviewer comment, and vendors should be able to dispute it on the record. This is the evidence an auditor or a procurement reviewer actually reads.
  • Workflow routing — Scores should change what happens next. Routing strings with Critical or Major errors to human review, while letting Minor-only strings proceed, is what turns scoring into a control rather than a report.

MQM scoring settings worth comparing across tools

Ambientación Valor documentado fuente
Default severity weightsCrítico 25, Mayor 5, Menor 1, Neutral 0Smartling Help Center, "LQA: MQM Schema Templates"
Default pass/fail thresholdAcceptable Penalty Points 20 = Raw Quality Score 98 or Calibrated Quality Score 80Smartling Help Center, "LQA: MQM Schema Templates"
MQM schema templates available4: Simplified MQM, Full MQM, MQM with Repeated Error Types, LQA Agent MQM SchemaSmartling Help Center, "LQA: MQM Schema Templates"
AI scoring accuracy vs. a human reviewer90% of the accuracy of a human reviewerSmartling Help Center, "LQA Agent: AI-Powered Quality Assurance"
Recommended AI scoring thresholdAcceptable Penalty Points 100 = MQM passing threshold 90Smartling Help Center, "LQA Agent: AI-Powered Quality Assurance"
Recommended sample size2–8% of translations per monthSmartling Help Center, "LQA Agent: AI-Powered Quality Assurance"
Human review scope with targeted validationReduced by up to 90%Smartling Help Center, "LQA Agent: AI-Powered Quality Assurance"
Standards alignment of the frameworkMQM described as aligned with ISO 5060MQM Council, themqm.org (verified October 7, 2026)

How do you run an MQM software evaluation?

A side-by-side evaluation on your own content is the only reliable way to compare MQM tools, because demo content is chosen to score well.

  1. Write the specifications first — For each content type, agree the evaluation word count, the severity levels and the pass threshold before looking at any tool. The MQM Council puts specifications ahead of annotation for the same reason: a threshold chosen after seeing scores is not a standard.
  2. Build a calibration set — Take a sample of real translations, including known errors, and have your own reviewers annotate it. This becomes the answer key every tool is measured against.
  3. Score the set in each tool and compare by locale — Check whether each tool's scores track your reviewers' scores, which locales drift, and how many errors reviewers reject as false positives.
  4. Test the evidence trail — Export the string-level error report, dispute an error as a vendor would, and confirm the arbitration outcome is recorded. If an auditor could not follow the trail, neither could a procurement reviewer.
  5. Pilot inside a live workflow — Run sampling and routing on real jobs for at least one reporting cycle, and measure reviewer hours and turnaround alongside the scores themselves.

MQM scoring software fits teams that...

  • Manage several translation vendors or language service providers and need one scale to compare them.
  • Must show auditors, regulators or executives documented quality evidence rather than reviewer opinion.
  • Translate legal, medical or financial content where a Critical error has real consequences.
  • Are scaling AI or LLM translation and need quality monitoring to keep pace with volume.
  • Want routing decisions, not just reports, driven by error severity.

When MQM scoring software may not be the right priority

  • Very small or one-off programs where a single trusted reviewer's judgment is sufficient.
  • Transcreation and creative campaigns, where an error typology captures little of what makes the copy work.
  • Programs with no glossary or style guide yet, since terminology and style errors cannot be scored against assets that do not exist.
  • Teams wanting AI-only scoring for non-English source content; check each tool's source-language support first, because Smartling's LQA Agent currently requires English source content.

Evaluation checklist: questions to ask MQM scoring vendors

Which translation management systems use MQM to score translation quality?
Several platforms advertise MQM-based scoring, and Smartling's LQA tools score against MQM-compatible schemas. Ask every vendor to show the schema itself, its categories, weights and threshold, rather than accepting "MQM" as a label.

Can AI score our translations automatically at scale, and how accurate is it?
AI scoring is practical now, but ask for its measured accuracy against human reviewers and the locales where it is reliable. For a wider look at AI scoring tools, see top tools for automated translation quality monitoring at scale.

Can one platform score marketing, legal and product documentation differently?
It should support a separate schema, with its own weights and threshold, for each content stream. One schema for everything either over-penalizes marketing copy or under-penalizes contracts.

How does scoring scale across hundreds of languages and locales?
Human MQM review scales through sampling, not full coverage; ask whether sampling rules run automatically. For AI scoring, ask for the list of locales where scores have been validated against human reviewers.

What evidence will an auditor or a regulated-industry reviewer get?
Ask for a string-level error report with category, severity, reviewer comments and arbitration history. For healthcare and finance, confirm that Critical severity is defined around legal or commercial consequences.

Which security and compliance certifications cover the platform where scoring runs?
Quality scores are produced inside the same platform that stores your content, so the platform's certifications apply. Smartling lists SOC 2, ISO/IEC 27001, ISO/IEC 42001:2023, HITRUST e1 and PCI Level 1 on its Security at Smartling page; ask every vendor for the reports behind any claim.

What proof of ROI and which references can you share?
Ask for references that can show an MQM score trend over time alongside reviewer hours or vendor spend. A premium price is only worth paying if the tool changes routing and vendor decisions, not just the dashboard.

How Smartling approaches MQM quality scoring

Smartling builds MQM scoring into the translation workflow through its Linguistic Quality Assurance (LQA) tools. Teams start from one of four MQM-compatible schema templates (Simplified MQM, Full MQM, MQM with Repeated Error Types, and the LQA Agent MQM Schema), set severity weights and Acceptable Penalty Points while the schema is a draft, and assign the published schema to a workflow step.

LQA Agent, available on Enterprise tier accounts, evaluates translations against that schema automatically, recording errors by category and severity at 90% of the accuracy of a human reviewer. LQA Memory stores validated results and reuses them on later evaluations, reducing false positives as a team's review history grows. A Decision step in a dynamic workflow can then route strings whose highest-severity error is Critical or Major to human review, while Minor-only strings move on.

Results appear in MQM results in Job Details, the LQA Dashboard, the LQA Errors & Arbitration report and the LQA Error Density report, and LQA Suite's automated sampling keeps evaluation running on a regular cadence. Together these give a localization leader a score, the string-level evidence behind it, and a workflow that acts on it.

¿Listo para ver a Smartling en acción?

Converse con alguien del equipo de Smartling para identificar cómo podemos ayudarle a aprovechar mejor su presupuesto al entregarle traducciones con la más alta calidad, mayor rapidez y a costos mucho más bajos.