Home » Blog » Can You Trust AI Translation? A Survey of 77,630 Real Business Documents

Can You Trust AI Translation? A Survey of 77,630 Real Business Documents

When AI Translates Your Business Documents

The supplier agreement is due back Friday, the counterparty reads Romanian, and the AI translator returns a clean, confident paragraph in about four seconds. It looks right. That is the whole problem. Nothing in the output tells you which parts the model was certain about and which parts it guessed at. By the time a clause reads differently to a lawyer in Bucharest than it did to you, the contract is signed.

So we went looking for a better answer than gut feel. What follows is what 77,630 real translations, produced by real people doing real work over 90 days, say about when AI translation can be trusted and when it cannot.

THE SHORT ANSWER: You cannot tell from a single output. Across 77,630 real translations, independent AI models produced meaningfully different versions of the same source text roughly one time in six, and the disagreement clustered in legal and official content. The reliable signal is not fluency. It is whether several independent models converge on the same wording.

What this survey covers

  • How the survey was run, and why it is not a benchmark
  • Finding 1: five out of six translations are too easy to reveal anything
  • Finding 2: even the strongest language pairs leave room to disagree
  • Finding 3: models have a stable style and an unstable accuracy
  • Finding 4: content type sets the quality ceiling
  • Finding 5: the long-document cliff
  • Three real situations, and what each one teaches
  • A five-point checklist for business documents
  • What this survey cannot tell you
  • Frequently asked questions

How the survey was run, and why it is not a benchmark

There was no curated test corpus. The sample is production traffic. Between 8 May and 6 August 2026, roughly 454,000 translations were created by users of an AI translation platform that runs several large language models against the same source text simultaneously. Of those, 77,630 carried a full cross-model comparison verdict, meaning the source was ambiguous enough for the models to actually diverge. Those 77,630 translations span 3,215 language pairs and form the survey sample.

A larger layer sits underneath it: 2,120,449 individually scored translations from the same window, joined to source word count and content category. That layer supports the domain and length findings below.

All of it is what people actually pasted and uploaded during three ordinary months of work. Benchmarks tell you how models behave on problems someone selected. This tells you how they behave on the supplier agreement that had to go out on Friday.

Finding 1: five out of six translations are too easy to reveal anything

Only 17% of translations in the window, the 77,630 in this sample, produced enough divergence between models to compute a comparison verdict at all. The other 83% were short, fragmentary, or unambiguous enough that every model landed in the same place.

That 17% is the sample worth having, because it maps onto the work that carries consequences. Nobody loses a contract over a mistranslated menu item. The moment text becomes long, formal, or technical enough that reasonable models disagree is also the moment the stakes rise.

Finding 2: even the strongest language pairs leave room to disagree

Across the ten highest-volume language pairs, average agreement between models on the same source text ranged from 91.9% down to 83.6%.

Read the bottom of that chart carefully. Japanese-to-English, the pair a technology company is most likely to lean on for board materials or vendor documentation from Asia, sat at 83.6%. Roughly one segment in six came back differently depending on which model produced it. English-to-Spanish, one of the most heavily resourced pairs in existence, was barely better at 83.8%. Volume and maturity do not close the gap.

Finding 3: models have a stable style and an unstable accuracy

Each comparison assigns every model a role based on how its output differs from the others: most formal, most natural, most literal, most concise. Counting those role wins across the top ten pairs produces a clear split.

RoleModel that won most oftenPairs wonHow reliable
Most formalClaude10 of 10Consistent everywhere
Most creativeQwen10 of 10Consistent everywhere
Most distinctive phrasingQwen10 of 10Consistent everywhere
Most naturalChatGPT9 of 10Consistent almost everywhere
Closest to consensusClaude9 of 10Consistent almost everywhere
Most thoroughMistral7 of 10Contested, no clear owner
Most literalMistral5 of 10Contested, no clear owner
Most conciseMistral and Qwen, tied4 of 10 eachContested, no clear owner

The roles that describe register, how formal or natural or inventive a model sounds, are owned outright and behave like personality traits. The roles that describe fidelity, how literal and how thorough a rendering is, have no reliable owner and flip from one language pair to the next. In plain terms: you can predict how a model will sound. You cannot predict, from the model alone, whether it got the meaning right.

Finding 4: content type sets the quality ceiling

Joining scored translations to their content category exposes something more uncomfortable than a leaderboard. Legal and government content had the lowest ceiling in the dataset: the strongest model in that category averaged 7.08 out of 10, and official certificates were barely better at 7.15. Customer service text topped out at 8.66.

Figure 2. Strongest and weakest model performance within each content domain, ordered by ceiling.

The floor matters as much as the ceiling. On official certificates, the weakest model averaged 0.23 out of 10. The same document, handed to two different models, comes back near-usable or near-worthless, and nothing on the screen distinguishes the two.

One note on why the models are unnamed in this chart and the next. The scores were assigned by an automated scorer whose identity could not be confirmed, and a single engine topped every category and every length bucket without exception. That is consistent with a genuinely strong model. It is equally consistent with a scorer sharing a lineage with one of the engines it grades, a documented failure mode in automated evaluation. So the spread is reported here and the leaderboard is not.

Finding 5: the long-document cliff

Every model in the survey performed best on passages of 10 to 60 words and worse on anything longer. Not some models. All of them.

Business documents live on the wrong side of that line. In the same window, documents uploaded by paying users had a median length of 5,795 words, and the longest 10% ran past 90,000. The content most likely to matter is the content furthest outside every model’s best operating range.

Three real situations, and what each one teaches

Case 1: a governing law clause, English to Romanian

Governing law clauses look like the safest thing in a contract. They are short, formulaic, and repeated almost verbatim across thousands of agreements. In this survey they sit where the two worst conditions overlap: legal and government content, the lowest-ceiling category, and short text, where one wrong lexical choice has nowhere to hide.

Take a sentence like “This Agreement shall be governed by and construed in accordance with the laws of England and Wales.” One model reaches for the formal Romanian legal register and produces something a court would recognise. Another produces plain-language phrasing that reads perfectly well and carries a subtly different obligation. Both come back in under five seconds. Both sound authoritative. Nothing in either output flags the divergence.

The cheap way to catch that split is to put the clause through several independent models and look at where they land. If six of seven converge on the same construction and one wanders off, you have learned something a single output could never tell you. If they scatter, that clause needs a lawyer rather than a faster tool. A worked example of exactly that comparison, showing how different engines render the same clause into Romanian, is set out in this breakdown of how to translate a governing law clause.

The lesson: short does not mean safe. Agreement across independent models is the cheapest available proxy for confidence, and taking the rendering the majority produced removes the largest avoidable source of error, which is trusting one model’s guess on language that only looks routine.

Case 2: the 12,000-word supplier agreement

A document of that size is entirely ordinary, roughly twice the median paying upload in this dataset, and it sits far past the 60-word mark where every quality curve turns down.

The instinctive fix makes things worse. Splitting the file into chunks small enough to sit in the sweet spot means each chunk is translated without sight of the others, so a defined term rendered one way on page 3 drifts by page 30. You trade a quality problem for a consistency problem, and consistency problems in contracts are the ones that end up in dispute.

The lesson: long documents need a workflow that keeps the whole file in view and enforces terminology across it, not a manual chunking habit. If your process involves copying paragraphs into a chat window, the length curve is already working against you.

Case 3: Japanese to English board materials

This pair had the lowest cross-model agreement in the survey at 83.6%. It also produced the most decisive single result in the data: one model won the formality role in 63% of comparisons, a commanding margin where role winners typically take 15% to 30%.

It would be easy to read that as “use this model for Japanese.” The literal-accuracy role in the same pair went to a different model entirely, on just 12% of assignments. One model owned the tone. Nobody owned the meaning.

The lesson: a model that reliably nails your register is not the same as a model that reliably nails your content. Register consistency is easy to observe and easy to over-trust.

A five-point checklist for business documents

  1. Treat fluency as noise. Every model in this survey sounded equally sure of itself, whatever it scored.
  2. Classify the content before you translate it. Legal, official, and technical content have measurably lower ceilings and deserve an extra step. Marketing copy does not.
  3. Watch the length. Anything past 60 words is outside every model’s strongest range, and anything past a few thousand words needs a whole-document workflow, not manual chunking.
  4. Compare before you commit. Several independent models converging on the same wording is evidence. One model sounding certain is not.
  5. Escalate what scatters. Model disagreement is a triage signal, telling you which small share of a document actually needs a human.

What this survey cannot tell you

  • The traffic comes from a general-purpose platform, so it skews toward the language pairs and content types its users bring. Enterprise localisation pipelines look different.
  • Agreement is a similarity measure, not a correctness measure. Models can converge and still be wrong together. Consensus reduces idiosyncratic error, not shared error.
  • The 0 to 10 quality figures come from an automated scorer. Read them as directional, not settled. Ninety-two percent of scored translations fell into an uncategorised bucket and were excluded, so Figure 2 describes the categorised minority.
  • There are no human reference translations here. The survey measures convergence and scored quality, not accuracy against a professional gold standard.

The alternatives are not better. Vendor benchmarks are chosen by vendors, and independent research such as Intento’s annual State of Translation Automation report or CSA Research’s market studies works at a level of abstraction that never reaches the individual clause. Intento’s latest edition, covered by the industry publication Slator, finds that legal and technical content is exactly where quality problems surface, which is what the domain findings above show from the other direction.

Frequently asked questions

How many AI models should I compare before trusting a translation?

The value comes from independence rather than volume. Comparing two models tells you only whether they disagree. Comparing six or more tells you where the majority lands, which is the more useful signal when you have to decide what to send.

Does model agreement mean a translation is correct?

No. Agreement means the models did not make different mistakes. They can still make the same one, particularly on rare terminology. Consensus removes the risk of a single model’s idiosyncratic guess. It does not remove the need for human review on high-stakes text.

Which content types need the most caution?

Legal and government content and official certificates had the two lowest quality ceilings, at 7.08 and 7.15 out of 10. Healthcare and technical manuals followed. Marketing and customer service scored highest, so everyday material is the safest to automate.

Why do long documents translate worse than short ones?

Every model peaked on passages of 10 to 60 words. Past that, context has to be held across more text and terminology decisions compound. The drop is consistent across all seven models, which points to a structural limit rather than one vendor’s weakness.

Is any single AI model best for business translation?

No model owned both style and accuracy. Register roles like formality had consistent winners across all ten top language pairs. Fidelity roles like literalness had no consistent winner at all. Picking one model as a standing default optimises for the wrong half of the problem.

5/5 - 1 vote

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top