Skip to content

The Language Routing Contract: Localization Beyond Fluency

Fluent translation does not prove native reasoning, cultural fit or stable permissions. The language routing contract tests all five before trusting a score.

5 min readData, adaptation and alignment

Research area 12: Multilingual, cross-cultural and localization engineering

In brief

  • Fluent translation does not prove native reasoning, cultural fit or safe permissions in another language.
  • The language routing contract gates every task by language, locale and entities before an answer is trusted.
  • Skipping native checks lets safety failures and wasted spend hide behind one passing global average.

Multilingual, cross-cultural and localization engineering proves an AI system works in a language and a culture, not just that it produces fluent text in one. A model can translate cleanly and still misname a local institution, misstate a permission, or answer a question that made sense in English but not in the target language, and one aggregate score hides all of it.

What it is

Localization adapts an interaction to a language, a locale and a culture, and it is broader than translation. Translationese names features of text that has passed through translation, and finding those features in output does not prove a model reasoned in English first, only that the output resembles translated text: native reasoning needs its own test. A refusal-rate comparison across two languages is not interpretable unless the prompts are equally harmful and culturally loaded.

Five capabilities need separate tests, and none validates the others: language proficiency, task reasoning, translation adequacy, local factual knowledge and culturally appropriate interaction. A translationese classifier measures a proxy that can learn genre habits rather than reasoning, so it belongs alongside a native reader's judgment, not in place of it. Original entities, units, dates and names must survive every translation step, and permission boundaries need their own test separate from refusal style, since a polite local-language response can still authorize the wrong action.

Two 2026 studies show both the mechanism and the limits of current evaluation. A study made public August 18, 2026 examined translationese in Gemini-2-Flash and Llama-3.3-70B across five languages and found German classifier flag rates of 60.2 percent and 37.6 percent, a result its own authors do not treat as proof of internal English reasoning, since native checks on 60 sentences per language for German and Spanish diverged from the classifier. SurakshaEval, made public August 8, 2026, covers 2,968 prompts across ten Indic languages plus English and 27 models, and found overall human agreement of 69.2 percent that fell to 44.16 percent on sociocultural judgments, the category where cultural interpretation matters most. A safety average that looks acceptable in aggregate can conceal a rubric barely two in five evaluators agree on for exactly the judgments carrying the most cultural risk.

The architecture

I call the answer the language routing contract: a versioned record attached to a task, not a language label, that forces a routing decision through an explicit gate before any answer is trusted. It exists because picking one model and assuming it behaves the same way in every language it claims to support has no evidence behind it here.

The animation opens on a single node labeled Task in local language, arriving with no routing decision attached yet. A line draws into a second node, The language routing contract, as a dot travels the connecting line, showing the task pass through a check of its language, locale and original entities. From the contract, two paths branch. One leads to a red node labeled Escalated for review, the outcome when a task is ambiguous or carries high stakes. The other leads to a green node labeled Validated automated answer, the outcome when the task clears every check. A traveling dot follows each path before the diagram fades to an empty stage and repeats.
The language routing contract: how a task earns a validated answer

The contract runs in a fixed order. A task first states its language, locale, script and original entities, none discarded once processing begins. It then routes on held-out performance for that workload: native prompting, translate-then-solve, multilingual retrieval when facts only exist in local sources, a language-specific adapter when behavior is narrow enough to justify its upkeep, deterministic conversion for dates and units, or human escalation when the decision depends on local expertise, ambiguous intent or an unverifiable fact. Whichever route runs, the contract checks units, dates, names and authorization scope before the answer is accepted, because a translate-then-solve path compounds error: an illustrative pipeline with a 95 percent chance of preserving meaning, a 90 percent chance of solving the task, and a 97 percent chance of a faithful back-translation succeeds end to end only about 83 percent of the time.

The most common failure is skipping the gate and treating one translated score as global truth. The second is letting a route go unmonitored after a model, tokenizer, glossary or retrieval source updates.

Ownership sits with a named localization owner accountable for the contract across every language shipped, holding the routing decisions, the glossary, and the authority to pull a language back to human escalation without asking permission.

Fund native task construction and native adjudication in every claimed language, and held-out route comparisons priced on translation, retrieval and review, not model tokens alone. Per-language safety testing needs a hard floor: zero violations across 30 prompts still permits an underlying rate near 9.5 percent, and only roughly 300 prompts brings that bound near 1 percent. Stop funding parity claims resting on a translated benchmark with no native check, since SurakshaEval's own agreement rate drops 25 points on the sociocultural judgments that matter most.

Two decisions belong to an executive: which languages get native task construction and adjudication funded, since that budget line turns a claim of supporting dozens of languages into a tested one, and where the human escalation threshold sits, a risk-appetite call a benchmark score cannot make.

What it is worth

Money and risk show up in three places: compute and review spent on a route a matched-cost comparison would have ruled out, incidents where a permission boundary held in the tested language and failed in an untested one, and rework when a translated benchmark score does not predict native performance.

Measure before and after with the same yardstick across languages: native task success, translation-path reliability end to end, per-language safety agreement rather than one global figure, and cost per validated task including translation, retrieval and review. The 2026 evidence supports that translationese detection and native checks can disagree, and that a safety benchmark's agreement rate varies sharply by category. It does not support a matched-cost ranking of native prompting against translation, a settled measure of code-switching cost, or proof that any tested safety boundary holds across every language a deployment will see. No return figure appears in that evidence, and none should be assumed here.

Questions leaders ask

What is the language routing contract, in practical terms?
It is a versioned record attached to a task, not a language label, that fixes the task's language, locale, script and original entities before any routing happens, then sends the task down a validated route: native prompting, translation, retrieval, a specialist adapter, deterministic conversion, or human escalation. Units, dates, names and permission scope are checked before the answer is trusted.
Does a model that avoids translationese in its output prove it reasons natively in that language?
No. A 2026 study examining Gemini-2-Flash and Llama-3.3-70B across five languages found translationese classifier flags that its own authors do not treat as proof of internal English reasoning, and native checks on 60 sentences per language for German and Spanish diverged from what the classifier alone suggested. Fluent output and native reasoning are different properties, and native reasoning needs its own operational test.
Can a global safety average be trusted across every language a product supports?
Not on its own. SurakshaEval, a 2026 Indic safety benchmark covering 2,968 prompts across ten Indic languages, English and 27 models, found overall human agreement of 69.2 percent that fell to 44.16 percent specifically on sociocultural judgments. A single average conceals exactly the category where cultural disagreement runs highest, so report that category separately before trusting the average.
When should a company use a smaller specialist model or a human reviewer instead of one global frontier model?
When a smaller specialist's validated coverage meets the task at lower maintenance cost, when a transformation is well defined enough for deterministic conversion, such as dates and units, or when the decision depends on local expertise, ambiguous intent or a consequential fact that cannot be verified automatically. The available 2026 evidence supports these criteria but stops short of a specific product recommendation or a frontier-versus-specialist cost ranking.
What decision about multilingual AI belongs to an executive rather than a technical team?
Two decisions. Which languages get native task construction and native adjudication funded, since that is what turns a claim of supporting dozens of languages into a tested one rather than a marketing line. And where the human escalation threshold sits, meaning how much ambiguity or cultural stake a task needs before an automated route hands it to a qualified local reviewer, since that is a risk-appetite decision a benchmark score cannot set.

Want this thinking applied to your organization?