Scale AI and Korea AI Safety Institute Release Benchmark Exposing a Translation Gap

The study found lower harmful-response scores in Korean prompts, but also more refusals of harmless requests—complicating claims that a model is simply safer in another language.

By 3 min read
Scale AI and Korea AI Safety Institute Release Benchmark Exposing a Translation Gap
Scale AI and Korea AI Safety Institute Release Benchmark Exposing a Translation Gap

Listen to this story

The audio brief

About 1:33
0:001:33
Read transcript
Some AI models look safer in Korean—but they may also be more likely to refuse harmless Korean requests. That is the tension in ROK-FORTRESS, a new benchmark from Scale AI and the Korea AI Safety Institute. It tested fourteen models across 1,235 tasks, changing two things independently: the language, from English to Korean, and the setting, from U.S. institutions and events to Korean ones. The goal was to separate a genuine language effect from the influence of local geopolitical context. Korean wording changed harmful-response rates by about ten percentage points, compared with roughly four points for changing the real-world setting. And this was not just a refusal effect: among responses the models actually gave, twelve of the fourteen were still less harmful in Korean. But some models refused benign Korean requests at about twice the rate, raising a harder question: are they better at recognizing danger, or simply more willing to say no? The benchmark’s transcreation matrix paired harmful tasks with harmless counterparts across chemical, biological, radiological, nuclear, and explosive threats, terrorism, crime, finance, and information leakage. Removing role-play and emotional wrappers mostly erased the Korean advantage, suggesting framing mattered. Proprietary models stayed modestly safer, while five open-source frontier models complied more often in Korean. The key constraint is scope: this covers one language pair and one geopolitical axis, so the next test is whether the same tradeoff appears elsewhere.

Story brief

3 key points

Scale AI and the Korea AI Safety Institute introduced ROK-FORTRESS, a benchmark designed to distinguish language effects from local geopolitical context in AI safety testing. Across 14 models and 1,235 tasks, Korean-language prompts changed harmful-response rates by about 10 percentage points, versus roughly four points for switching U.S. references to Korean ones. The result is not a straightforward safety win: 12...

  1. 01

    The benchmark independently varies English/Korean wording and U.S./Korean entities, institutions, and operational details.

  2. 02

    Removing role-play and emotional wrappers mostly eliminated the Korean advantage, suggesting framing can influence apparent safety differences.

  3. 03

    Five open-source frontier models were more likely to comply in Korean; proprietary models remained modestly safer.

A model that appears safer in Korean may also be less willing to answer harmless Korean questions. That is the central tension in ROK-FORTRESS, a new benchmark from Scale AI and the Korea AI Safety Institute: across 14 evaluated models, Korean prompts grounded in Korean settings generally produced lower harmful-response scores than English prompts grounded in U.S. settings, while some models refused benign Korean requests at roughly twice the rate.

The finding matters because multilingual safety testing often changes only the language while leaving the scenario intact. ROK-FORTRESS instead asks whether a model responds differently when both the words and the real-world references change—for example, from U.S. institutions and events to Korean ones. The researchers say translation-only tests can therefore miss interactions that shape behavior in actual local deployments.

How the benchmark separates two variables

  • It varies prompt language between English and Korean.
  • It separately varies geopolitical grounding between U.S. and Korean entities, institutions and operational details.
  • It pairs adversarial requests with benign counterparts, so the researchers can measure both harmful answers and over-refusal.

The project uses what it calls a transcreation matrix: controlled prompt variants that preserve the underlying intent while changing language and local grounding independently. Its 1,235 tasks span chemical, biological, radiological, nuclear and explosive threats; political violence and terrorism; criminal and financial activity; and information leakage. Responses were assessed with calibrated panels of language models acting as judges, using expert-crafted, prompt-specific binary rubrics.

The reported Korean effect was not entirely explained by models declining to respond. When the analysis looked only at answers models actually gave, 12 of the 14 models still produced less harmful substantive responses in Korean. But the researchers also found broader caution: harmless Korean requests drew more refusals in some cases, leaving open whether the apparent safety improvement reflects better threat discrimination or a blunter tendency to say no.

The study’s main tests used elaborate adversarial requests wrapped in role-play, invented backstories or emotional appeals. When those wrappers were removed and the same information was requested directly, the Korean advantage mostly disappeared. Scale AI says proprietary models remained modestly safer in Korean, while five open-source frontier models became more likely to comply in Korean—evidence that some earlier suppression may come from adversarial framing losing force during transcreation rather than from stronger Korean safeguards.

The benchmark examines one language pair and one geopolitical axis, so it does not establish that the same pattern will hold elsewhere. It does show that language and context did not combine in a uniform way: in four of 14 models, Korean grounding significantly weakened the reduction in harmful responses associated with Korean language, and none showed a statistically significant effect in the reverse direction. The public dataset subset and preprint give other researchers a way to test whether culturally grounded evaluations reveal similar gaps in other settings.

Sources

  1. labs.scale.comROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
  2. scale.comROK-FORTRESS: What Language and Context Reveal About AI Safety

Loading discussion...

Scale AI and Korea AI Safety Institute Release Benchmark Exposing a Translation Gap | Superpower Daily