On the Systematic Challenges of Culturally Loaded Machine Translation:
Dream of the Red Chamber as the Cultural Lens
July 22, 2026
Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms. Although large language models (LLMs) have enabled MT systems to achieve human-like quality in many scenarios, their ability to handle culturally loaded expressions remains underexplored. In this study, we systematically investigate the challenges posed by culturally loaded translation in LLM-based MT systems. We construct a Chinese-Japanese bilingual dataset from the culturally representative corpus Dream of the Red Chamber1, containing 500 segments across diverse cultural categories. Using a comprehensive evaluation protocol, we reveal three main challenges: (1) task challenges, where frontier LLMs exhibit notable performance gaps and struggle with culturally loaded content; (2) human evaluation challenges, where evaluator backgrounds lead to substantial disagreement in translation judgments; and (3) automatic evaluation challenges, where widely used metrics fail to reliably assess translation quality for this task. These findings may offer valuable insights for culture-oriented translation research in both computational science and linguistics.
Translation aims to convey the source message in the target language through the closest natural equivalent while preserving semantic, stylistic, and pragmatic effects. However, language is inseparable from culture, and cultural priorities shape linguistic expression [1]. Each culture organizes its vocabulary around its own areas of emphasis, leading certain domains of meaning to become increasingly detailed and complex [2]. This culturally shaped development gives rise to a category of lexical items termed culturally loaded words by [3]. Such words are deeply embedded in specific socio-cultural contexts and reflect the traditions, beliefs, and value systems of a community [4], [5].
The translation of culturally loaded expressions is highly contextualized. Its context extends beyond the immediate text to the broader cultural background behind corresponding languages. Differences in cultural focus across languages may lead to semantic mismatches or lexical gaps [6], [7]. Therefore, translators must consider both source and target cultural contexts and balance interpretation with communication. At the interpretive level, they should pursue conceptual equivalence rather than rely on surface correspondence, ensuring that target readers grasp the intended meaning, since literal translation is often insufficient [1]. At the communicative level, these expressions carry substantial cultural weight and should be rendered in ways that preserve, as far as possible, meanings rooted in the source culture [8]. These place high demands on the professional expertise of translators.
On the computer science side, the rapid development of large language models (LLMs) [9]–[11] has enabled machine translation (MT) systems to achieve human-like quality in many scenarios [12]–[14]. However, capturing rich contextual information remains a major challenge for current MT systems [15]–[18], making the culturally loaded translation task particularly difficult for current models due to its inherently contextualized nature. Evaluating such translations adds another layer of difficulty. Translation quality is already subjective, and cultural elements amplify this subjectivity: evaluators from different cultural backgrounds may judge the same output differently based on their cultural familiarity and preferences (see §2.3 for the theoretical basis), introducing unpredictable variance in human evaluation. Furthermore, it remains unclear whether widely used automatic metrics can reliably assess translation quality in these scenarios. Thus, from the task itself to its evaluation, culturally loaded translation presents unique challenges for modern MT research.
In this study, we systematically investigate the above challenges that it poses for LLM-based MT. We construct a Chinese-Japanese bilingual dataset (§3) based on the highly representative cultural corpus of Dream of the Red Chamber. The dataset comprises 500 representative culturally loaded segments, and spans diverse cultural categories, with representative examples shown in Fig.1. We also design a specific evaluation protocol (§4) that incorporates diverse evaluation dimensions and evaluator backgrounds, ensuring the completeness and comprehensiveness of the results. All preliminary model translation qualities are based on human evaluation, which serves as the gold standard for our formal analysis. Based on these, we reveal three comprehensive core challenges:
Task Difficulty (§5.1): Frontier LLMs still underperform on culturally loaded translation, showing substantial gaps both across models and relative to human references. Their performance also varies across cultural categories and is sensitive to category-specific textual characteristics and difficulty. This indicates considerable room for improvement in LLM-based MT systems.
Human Evaluation Disagreement (§5.2): Human evaluation of culturally loaded translation is sensitive to evaluator background, including native cultural environment and expertise. This indicates that evaluating this task requires careful design of evaluator diversity and comprehensiveness; otherwise, systematic bias may arise.
Automatic Evaluation Unreliability (§5.3.0.2): Mainstream automatic metrics struggle to reliably capture model rankings and totally fail to distinguish sample-level quality differences in culturally loaded translation. This highlights the need for more task-aware automatic methods.
In addition, we conduct extended error and translation strategy analyses to better understand model behavior patterns and key bottlenecks in this task. We hope these findings provide valuable insights for both linguistic and computational research.
[3] first coined the term culturally loaded words to describe lexical items that encode society’s traditions, beliefs, and value systems. Similar concepts in translation studies, such as cultural words [6] and culture-specific items [19], generally refer to expressions that are deeply rooted in particular socio-cultural contexts and whose meanings cannot be fully understood without cultural background knowledge [4], [5].
Taking Chinese culture as examples,
UTF8gbsn”布衣”
(bùyī) in ancient China referred to ordinary people and implied a modest lifestyle associated with coarse cloth garments. If translated literally as “coarse clothes,” English readers may find it confusing, since the phrase only refers to a type of fabric and does not convey the social meaning in Chinese. Another example is
UTF8gbsn”鸿雁”
(hóngyàn), a bird image in Chinese ecological culture. Although “swan goose” is its literal zoological equivalent, it fails to capture the cultural meaning. In Chinese tradition,
UTF8gbsn”鸿雁”
often symbolizes a messenger carrying letters. Translating it as “message-bearing swans”3 better conveys this cultural imagery, even though it sacrifices strict lexical equivalence. These examples show that translating culturally loaded expressions requires conveying both meaning and cultural significance.
As the foundational scholar of modern translation studies, Eugene A. Nida pioneered the categorization of culture into five types: ecology, religion, material, linguistics, and society [1], [20]:
Ecology: Natural elements such as flora, fauna, climate, geographical landscapes, and ecological phenomena that carry cultural meaning.
Religion: Spiritual beliefs including ancestor worship, folk superstitions, and mythological concepts that shape worldviews.
Material: Tangible artifacts of daily life like clothing, food, tools, and other physical objects that define a civilization’s material culture.
Linguistics: Culturally embedded expressions like idioms, proverbs, slang, riddles, and fixed phrases that exhibit unique rhetorical patterns.
Society: Social structures, customs, institutions, kinship systems, etiquette norms, and conventions that govern interpersonal relations.
This classification has shown strong theoretical adaptability and continues to inform contemporary translation research [21]–[24]. In this study, we adopt this five-fold taxonomy as the framework for dataset construction.
Cross-cultural translation has long been a central topic across disciplines. Among the most influential frameworks is Venuti’s domestication and foreignization taxonomy [8]. Domestication adapts a text to target-culture norms to enhance readability, whereas foreignization preserves source-culture elements to foreground cultural difference. This framework highlights the translator’s agency and offers a critical perspective for evaluating cross-cultural transfer. It has been widely used to analyze how translators handle culture-specific items by balancing accessibility and authenticity.
Other theoretical perspectives further deepen the understanding of translation as a culturally embedded practice. The dynamic equivalence theory [1] prioritizes equivalent reader response rather than literal correspondence. The cultural turn thoery [25] conceptualizes translation as a form of cultural rewriting shaped by ideological forces; The norms theory [26] emphasizes the socio-cultural constraints that guide translators’ decisions. Together, these frameworks suggest that translation is not a neutral linguistic transfer but a process of cultural negotiation — an important foundation for examining how language models handle culturally loaded translation, which we will also explore in our subsequent experiments.
We used Dream of the Red Chamber as our source corpus. First, widely regarded as the “encyclopedia of Chinese classical culture” [27], [28], it displays exceptional linguistic richness and cultural depth. Second, it covers all five cultural categories in §2.2, enabling comprehensive evaluation across cultural dimensions within a unified textual framework [29]. Third, its long translation history and the extensive scholarship surrounding it, often referred to as “redology” [30], [31], allow findings derived from this corpus to connect with broader research and cross-cultural communication.
Building on Chinese culture as the source, we select Japanese as the target language, as its cultural interpretive distance [25], [26] from Chinese strikes a balance. Unlike translations between Chinese and Western languages, which often require extensive cultural explanation, Chinese and Japanese belong to the East Asian cultural sphere and share elements such as the kanji writing system and related aesthetic traditions. This shared foundation allows many culturally loaded concepts to be transferred more directly, while still requiring careful strategic choices when meanings diverge. For example, the Chinese term ”
UTF8gbsn阴司
” (yīnsī) is rendered in the Japanese translation as
UTF8min「閻魔の庁」
. This substitution draws on shared religious conceptions between Chinese and Japanese cultures, where both traditions envision judgment by King Yama after death. In contrast, Western readers’ imagination of the “nether world” tends to evoke Greek Hades or the Christian purgatory, lacking the distinct East Asian imagery of Yama’s judgment. This intermediate cultural distance makes the Chinese–Japanese pair particularly revealing for studying culturally loaded translation.
| Ecology | Religion | Material | Language | Society | |
|---|---|---|---|---|---|
| Source | 23.40±12.35 | 31.70±16.21 | 24.20±16.69 | 15.20±9.42 | 24.90±24.19 |
| Target | 50.00±25.43 | 64.30±36.92 | 56.50±44.72 | 32.05±19.42 | 51.05±43.33 |
We prioritize target-native translations because conveying culturally loaded meanings in the target language requires generative competence in its linguistic and cultural norms [32], [33]. Mastery of the target culture is thus more consequential than source-culture familiarity, leading us to adopt a translation by a native Japanese speaker.
Japanese translations of Dream of the Red Chamber date back to Mori Kainan in 1892. Over the following century, they evolved from partial renderings into at least 38 full versions. Among these, the full translations by Matsueda Shigeo, Ito Sohei, and Inami Ryoichi are widely regarded as the three pillars of the Japanese tradition [34]. We adopt Ito Sohei’s version for this study. Unlike Inami Ryoichi’s modern Japanese rendering, which prioritizes reader accessibility, Ito’s translation achieves a more balanced approach by preserving source fidelity while providing rich annotations that deepen cultural interpretation [35]. Furthermore, developed over nearly fifty years with four revisions, it has become the most influential and widely studied version in academic research [36]–[38].
We collected the bilingual Chinese-Japanese edition from the state-sponsored Library of Chinese Classics series4, translated by Ito Sohei. This edition was compiled under the auspices of Chinese government agencies and extensively revised by renowned scholars, ensuring its authority and reliability. Also, its format enables direct comparisons between source and target texts (see §9.2), ensuring rigor throughout the data processing.
Currently, only image versions of this edition are available, so OCR processing is required. Four Chinese graduate students conducted the initial annotation, each selecting 480 culturally loaded segments (4 per chapter across all 120 chapters) and categorizing them according to the taxonomy in §2.2. Four Japanese graduate students then performed a second round of screening and revision. To control the workload of subsequent human evaluation, the dataset was further refined to 500 representative segments, balanced across the five cultural categories (100 per category). The entire annotation process took 20 days. These annotators also participated in human evaluations, so their backgrounds are presented in subsequent §4.1.
Text length statistics are shown in Tab.1 using the Qwen3 tokenizer [39]. Overall, target texts are about twice as long as source texts. Among categories, Religion tends to be longer, Linguistics shorter, and the other three are similar. Importantly, there are no marked differences, meaning that category difficulty is not driven by text length.
| Model Name | Abbreviation | Affiliation | Citation |
|---|---|---|---|
| Non-Reasoning Model | |||
| Deepseek-v3 | DS-v3 | Deepseek | [40] |
| Qwen3-235B-A22-Non-Thinking | Qwen3-T | Alibaba | [39] |
| GPT-4.1 | GPT-4.1 | OpenAI | [41] |
| Gemini-2.5-Flash | Gemini-2.5 | [42] | |
| Reasoning Model | |||
| Deepseek-r1 | DS-r1 | Deepseek | [43] |
| Qwen3-235B-A22-Thinking | Qwen3-NT | Alibaba | [39] |
| OpenAI-o4-mini | o4-mini | OpenAI | [44] |
| Claude-Sonnet-4 | Claude-4 | Anthropic | [45] |
| Dimension Name | Abbreviation | Target | Scope |
|---|---|---|---|
| Content Accuracy | Acc. | Content | General |
| Language Fluency | Flu. | Language | General |
| Cultural Appropriateness | Cult. | Content | Culture-specific |
| Native Readability | Read. | Language | Culture-specific |
| Full Dataset | Ecology | Religion | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2-6 (lr)7-11 (lr)12-16 Model Name | Overall | Acc. | Flu. | Cult. | Read. | Overall | Acc. | Flu. | Cult. | Read. | Overall | Acc. | Flu. | Cult. | Read. |
| Reference | 4.27 | 4.35 | 4.24 | 4.28 | 4.11 | 4.33 | 4.38 | 4.27 | 4.32 | 4.16 | 4.23 | 4.33 | 4.29 | 4.24 | 4.09 |
| Deepseek-v3 | 3.92 | 3.67 | 3.92 | 3.67 | |||||||||||
| Qwen3-235B-A22B-Non-Thinking | 3.30 | 3.60 | 3.46 | 3.37 | 3.40 | 3.49 | 3.79 | 3.66 | 3.59 | 3.56 | 3.33 | 3.71 | 3.54 | 3.42 | 3.49 |
| GPT-4.1 | 3.66 | 3.75 | 3.72 | 3.63 | 3.82 | 3.89 | 3.85 | 3.71 | 3.64 | 3.84 | 3.71 | ||||
| Gemini-2.5-Flash | |||||||||||||||
| Deepseek-r1 | 3.37 | 3.51 | 3.53 | 3.33 | 3.48 | 3.51 | 3.60 | 3.61 | 3.38 | 3.60 | 3.34 | 3.56 | 3.49 | 3.25 | 3.47 |
| Qwen3-235B-A22-Thinking | 3.65 | 3.76 | 3.63 | 3.31 | 3.85 | 3.88 | 3.81 | 3.50 | 3.65 | 3.81 | 3.38 | ||||
| OpenAI-o4-mini | |||||||||||||||
| Claude-Sonnet-4 | 3.57 | 3.78 | 3.80 | 3.64 | 3.42 | 3.73 | 3.91 | 3.92 | 3.72 | 3.58 | 3.54 | 3.85 | 3.78 | 3.39 | |
| Average | 3.64 | 3.80 | 3.80 | 3.62 | 3.61 | 3.82 | 3.96 | 3.94 | 3.76 | 3.77 | 3.64 | 3.86 | 3.83 | 3.60 | 3.62 |
| Range (max-min) | 0.59 | 0.49 | 0.71 | 0.50 | 0.63 | 0.60 | 0.60 | 0.72 | 0.67 | 0.59 | 0.57 | 0.50 | 0.74 | 0.58 | 0.53 |
| Material | Linguistics | Society | |||||||||||||
2-6 (lr)7-11 (lr)12-16 Model Name |
Overall | Acc. | Flu. | Cult. | Read. | Overall | Acc. | Flu. | Cult. | Read. | Overall | Acc. | Flu. | Cult. | Read. |
| Reference | 4.39 | 4.39 | 4.33 | 4.34 | 4.17 | 4.17 | 4.29 | 4.16 | 4.22 | 4.05 | 4.23 | 4.36 | 4.15 | 4.28 | 4.08 |
| Deepseek-v3 | 4.12 | 3.88 | 3.73 | 3.51 | 3.42 | ||||||||||
| Qwen3-235B-A22B-Non-Thinking | 3.56 | 3.84 | 3.72 | 3.67 | 3.64 | 3.10 | 3.45 | 3.33 | 3.20 | 3.28 | 3.00 | 3.35 | 3.15 | 3.10 | 3.14 |
| GPT-4.1 | 3.85 | 3.93 | 3.86 | 3.72 | 3.53 | 3.66 | 3.65 | 3.52 | 3.45 | 3.58 | 3.50 | 3.45 | |||
| Gemini-2.5-Flash | |||||||||||||||
| Deepseek-r1 | 3.59 | 3.68 | 3.69 | 3.46 | 3.70 | 3.22 | 3.38 | 3.43 | 3.20 | 3.35 | 3.18 | 3.34 | 3.33 | 3.18 | 3.33 |
| Qwen3-235B-A22-Thinking | 3.92 | 3.85 | 3.56 | 3.47 | 3.64 | 3.49 | 3.21 | 3.37 | 3.65 | 3.47 | 3.40 | 3.10 | |||
| OpenAI-o4-mini | |||||||||||||||
| Claude-Sonnet-4 | 3.79 | 3.98 | 3.97 | 3.77 | 3.64 | 3.43 | 3.66 | 3.71 | 3.51 | 3.32 | 3.38 | 3.61 | 3.60 | 3.48 | 3.27 |
| Average | 3.87 | 4.01 | 3.98 | 3.81 | 3.83 | 3.47 | 3.65 | 3.68 | 3.47 | 3.48 | 3.40 | 3.59 | 3.55 | 3.42 | 3.40 |
| Range (max-min) | 0.57 | 0.51 | 0.71 | 0.63 | 0.57 | 0.61 | 0.43 | 0.67 | 0.47 | 0.62 | 0.68 | 0.42 | 0.73 | 0.55 | 0.66 |
We evaluated eight language models (see Tab.2), selected for two aspects of diversity: (1) they span five organizations and multiple model families, providing broad source coverage; (2) they include both reasoning and non-reasoning paradigms, enabling us to examine the impact of long chain-of-thought (CoT) ability [43], [46] on culturally loaded translation. Implementation details are provided in §10.1. Based on model outputs, we established the following evaluation protocol.
Given the contextual complexity and inherent subjectivity of this task, our evaluation is human-centered, with human judgments serving as the gold standard. We adopt the standard MQM framework [47], a widely used functionalist evaluation approach5, and adapt it to better capture the characteristics of culturally loaded translation. Specifically, we define four evaluation dimensions, shown in Tab.3 (see §11 for detailed explanations and scoring guidelines). The first two assess semantic accuracy and linguistic fluency following MQM. The remaining two focus on cultural aspects, evaluating whether the translation conveys the cultural meaning of the source expression and whether it reads naturally to target-culture readers. In addition, evaluators assign an overall quality score to reflect their holistic judgment.
Our evaluators come from diverse cultural and academic backgrounds and consist of sixteen native speakers, including eight Chinese and eight Japanese. Within each language group, four are graduate students specializing in the other language’s literature, and four are senior academics serving as university professors in the other country’s cultural studies. To ensure evaluation reliability, all student evaluators hold at least JLPT-N1 or HSK-6 certification, or an equivalent qualification. This design yields four groups by language (zh / ja) and academic level (stu. / prof.).
Given the data scale, evaluating all translations would be prohibitively labor-intensive: (8 models + 1 reference) × 500 samples = 4,500 translations, resulting in 22,500 scores per evaluator. Therefore, within each group, four evaluators each review 125 samples (25 per category), collectively covering the full dataset. For Read., scores are assigned only by target-language natives. To avoid bias from prior assumptions about model and human abilities, translation sources were anonymized. Evaluators saw only the source text and nine shuffled translations for each sample. Before the evaluation, pilot assessments verified intra-group agreement. Sampling-based comparisons showed that evaluators within the same group had comparable expertise and low judgment variance. After the evaluation, we conducted follow-up interviews with the evaluators to examine their fine-grained decision criteria and interpretations of some anomalous phenomena, which are needed for analysis in §5.2.
We adopt three types of metrics: lexical-based BLEU [48], semantic-based xCOMET [14], and LLM-as-a-Judge [49]. For the latter, we employ two models: GPT-5 (a non-reasoning model, [50]) and Gemini-3-Pro (a reasoning model, [51]). See §10.2 for the specific prompt.
We first analyze overall model performances on the full dataset. The top three models in overall score are OpenAI-o4-mini (3.89), Gemini-2.5-Flash (3.87), and Deepseek-v3 (3.83). Notably, the first two rank among the top three across all five evaluation dimensions, indicating strong robustness. However, performances vary greatly across models, with overall scores ranging from 3.89 to 3.30. This wide range demonstrates that this task poses a meaningful challenge and effectively differentiates model capabilities. In addition, a considerable gap remains between current models (an average of 3.64) and human references (4.27), whose scores are consistently higher across all dimensions, indicating that culturally loaded translation remains difficult for current LLMs to match human performance.
Another interesting phenomenon is that reasoning and non-reasoning models reveal no systematic advantage for either group. Their average overall scores are nearly identical (3.66 vs. 3.62), and medal counts are similarly balanced (15 vs. 12 medals, with 2 vs. 3 golds). While the top-performing OpenAI-o4-mini employs reasoning, the runner-ups Gemini-2.5-Flash and Deepseek-v3 do not, suggesting that longer CoT are not decisive for our culture-oriented translation tasks.
Next, we analyze model performance across evaluation dimensions. A clear pattern emerges: the two culturally related dimensions, Cult. and Read., consistently receive lower scores than the general two, Acc. and Flu.. This suggests that while current models perform relatively well on general quality, they struggle more with culturally loaded aspects of translation. Moreover, strong performance on general dimensions does not necessarily imply strong cultural competence. For example, OpenAI-o4-mini achieves the highest scores in Acc.(4.00) and Flu.(4.17), yet its Cult.(3.74) and Read.(3.72) only rank third. Conversely, GPT-4.1 attains the highest Read. (3.94) but performs less strongly in general dimensions. These phenomena indicate that cultural competence cannot be inferred directly from general translation quality, revealing a gap between linguistic correctness and deeper cultural adaptation.
Finally, we analyze model performance across cultural categories. The results reveal a clear difficulty hierarchy: Material (3.87) and Ecology (3.82) score the highest, followed by Religion (3.64), while Linguistics (3.47) and Society (3.40) lag significantly behind.
This gap is largely driven by differences in textual characteristics across cultural categories. Material and Ecology texts mainly describe concrete objects, entities, and natural phenomena, which are relatively straightforward for models to translate. Religion texts, although conceptually abstract, often contain narrative elements and real-world descriptions that models can partially handle using general knowledge. In contrast, Linguistics and Society texts are considerably more challenging. Linguistics frequently includes classical idioms, allusions, and elliptical expressions with dense meanings and flexible syntax, making accurate interpretation and translation difficult. Social Culture involves etiquette and honorific expressions that depend on subtle social relationships, requiring pragmatic inference that models often fail to capture, leading to inappropriate tone or meaning.
The inter-group agreement varies across evaluation dimensions. For the first three dimensions, the differences among evaluator groups remain moderate, generally below 0.5 points. In contrast, the Read. dimension shows substantially larger divergence. For several lower-ranked models, the gap between evaluator groups approaches 1 point. We analyze this anomaly in detail below.
Despite these variations, the relative ranking of the eight models remains largely consistent across evaluator groups and aligns well with the overall scores. Among the four groups, ja-prof. shows the closest agreement with the overall ranking, followed by zh-prof.. The two student groups exhibit slightly greater fluctuation, though the differences are not substantial. In addition, agreement tends to decrease as model quality declines, as poorer outputs leave more room for subjective interpretation.
Within the same language group (zh or ja), students consistently score higher than professors. This may reflect differences in evaluation standards and familiarity with culturally loaded texts. Through the post-evaluation interview, we find that when encountering ambiguous expressions, students tend to accept them more readily, which introduces an upward bias. By contrast, professors generally apply stricter criteria when judging translation adequacy.
This divergence is most evident in the Read. dimension, where students and professors show substantial disagreement. Through the post-evaluation interview, we note an important fact: culturally loaded texts are often associated with specific historical or literary registers. Students tend to evaluate readability primarily based on fluency and ease of comprehension. Professors, however, pay closer attention to whether the stylistic register (e.g. classical or modern language) appropriately reflects the cultural and historical context embedded in the source text, even if this reduces immediate readability. In practice, models often produce relatively colloquial outputs that align with students’ reading habits, leading them to score such translations more favorably. Professors, however, may consider these stylistically inappropriate, arguing that culturally significant texts should preserve more register-appropriate expressions. This reflects a broader tension between accessibility and authenticity rather than a purely objective notion of correctness.
This is further evidenced by their scoring of the reference on the Read. dimension: students and professors gave average scores of 3.70 and 4.52, respectively. Notably, the students’ rating is even lower than that of some model outputs, providing additional support for the above conclusions.
Within the same academic group (stu. or prof.), Japanese are slightly stricter than Chinese evaluators on the two general dimensions, Acc. and Flu.. This is expected, as Japanese annotators are native speakers of the target language and can therefore apply more nuanced judgments of translation quality.
In contrast, the relative scoring patterns for the cultural dimension Cult. vary across models — sometimes Japanese evaluators score higher, while in other cases Chinese evaluators do. Through the post-evaluation interview, we find that although both groups value the preservation of source culture and the appropriateness of the target expression, they tend to emphasize different aspects. Chinese evaluators often focus more on faithfulness to the source culture, whereas Japanese evaluators place greater weight on appropriateness within the target cultural context. Because current models often struggle to balance these two aspects simultaneously, translations that favor one side may draw stricter scrutiny from evaluators who prioritize the other, causing noticeable fluctuations in scoring.
Fig.3 compares system-level rankings between human evaluation and automatic metrics. Overall, correlations with human judgments remain limited. xCOMET and two LLM-based judges show modest positive correlations (Kendall’s \(\tau\) between 0.21 and 0.36, Spearman’s \(\rho\) between 0.38 and 0.48), with several models positioned near the diagonal. However, BLEU exhibits almost no correlation (\(\tau<0.1\)), as its reliance on surface-level lexical overlap with reference translations prevents it from capturing semantic and culturally grounded translation quality. These results suggest that while newer metrics offer somewhat more informative signals than BLEU, they still fall short of reliably distinguishing model performance in culturally loaded translation tasks.
Fig.4 illustrates the correlation between human evaluation scores and four automatic metrics across all eight models. Overall, correlations with human judgments remain extremely weak at the sample level for all metrics. Although xCOMET and the LLM judges occasionally show slightly stronger positive correlations than BLEU, the overall alignment remains limited, and the scatter plots display substantial dispersion around the fitted regression lines. In many cases, samples with similar automatic scores receive noticeably different human ratings, while samples with comparable human judgments correspond to a wide range of automatic metric values. This pattern suggests that automatic metrics capture only part of the translation quality signal and fail to reliably reflect fine-grained differences among individual outputs. Consequently, the magnitude of an automatic metric score alone is insufficient to determine whether a specific generated sample is actually better or worse according to human evaluation.
To further explore the model limitations in translating culturally loaded content, we conduct extended analyses of their translation behavior. Due to space limitations, detailed discussions are provided in the appendix, with key conclusions presented here.
We classify translation errors into five types: No Obvious Error, Mistranslation, Overtranslation, Undertranslation, and Omission. Mistranslation is the most frequent error across all models, while higher-performing models exhibit a lower proportion of such errors. The remaining three error types occur relatively infrequently. Notably, reasoning models produce substantially more Overtranslation, yet this tendency does not translate into better overall translation quality.
We focus on the domestication-foreignization framework introduced in §2.3. We find that over 80% of translations adopt a single strategy, with a clear preference for domestication. Under this, models tend to rigidly adapt cultural expressions, often sacrificing source cultural imagery for target-language readability, resulting in notable loss of cultural transmission.
To provide a more intuitive understanding of model translation styles, we present several case studies that illustrate how models underperform across various evaluation dimensions.
Culturally loaded translation has long been a central concern in translation studies, as language serves not only as a means of communication but also as a carrier of culture. In this work, we connect this linguistic topic with modern MT research6. Using a dataset based on Chinese-Japanese translations of Dream of the Red Chamber, we systematically uncover the challenges of culturally loaded translation from three perspectives: the task itself, human evaluation, and automatic evaluation. Our analysis may offer insights for future research on culturally loaded translation in computer science. More broadly, the results suggest that, despite rapid progress in LLMs, MT still has a long way to go, as many researchers have long believed [52]–[54].
The primary limitation of this study is language diversity. Although we have justified in §3.2 the representativeness of the current language pair for studying culturally loaded translation, including additional languages, particularly those with broader speaker populations such as English, would undoubtedly enhance the generalizability of our findings. However, due to various constraints, including limited funding, the lengthy cycle of human evaluation, and the need for culturally proficient native speakers (crowdsourcing is not a viable option for this study), we are currently unable to extend our study to more languages in the near term. As a mitigation, we will open-source our data sources and provide raw multilingual corpus to facilitate further research. We believe that future work can build upon this foundation to yield more valuable insights.
Our study involves engineering experiments and analyses, and does not propose any new architectures or algorithms. Therefore, it does not introduce uncontrollable consequences in practical use. Regarding the research data, we have carefully verified the compliance of both the constructed dataset and the translations generated by LLMs, ensuring no potential risks are involved.
We have confirmed that our research data, including the constructed dataset and the translations generated by LLMs, does not contain any sensitive elements related to politics, violence, or similar topics. We ensure that no uncontrollable impact will be caused to the human evaluators involved in the experiment or to any individuals who may come into contact with the data.
All evaluators were fully informed about the research use of the evaluation data collected from them, and their consent was obtained. Throughout the evaluation process, we consistently respected their autonomy and paid them with mutually agreed-upon remuneration.
AI assistants (ChatGPT and Deepseek) were used exclusively for writing polishing and played no role in any other aspect of this work.
Regarding related work on culturally loaded translation, §2 has already discussed foundational theoretical studies. Here, we focus only on MT research that potentially relates to our work.
MT research has long recognized the challenges of culture-oriented translation [55]. In response, several studies have proposed benchmarks to quantify cultural translation [56]–[59], but these often focus on specific domains or concrete contexts (e.g., food) and rely primarily on automatic metrics. Other work has developed systems for identifying culturally loaded content [60]. [18] shares a similar perspective with ours, namely rethinking the challenges of culturally loaded translation in MT, but focuses specifically on the aspect of paratexts. In contrast, our study adopts a more systematic perspective, offering a comprehensive analysis of culturally loaded translation challenges from the task itself to its evaluation.
Tab.5 - 9 presents dataset examples from five cultural categories.
Source: UTF8gbsn 每日早起,拿上等燕窝一两,冰糖五钱,用银铫子熬出粥来。 |
Target: UTF8gbsn 毎朝、上等の燕窩(海つばめの巣。燕巣)一両分と氷砂糖五銭分とを、銀の銚子で煮いてお粥にこしらえなさい。 |
Source: UTF8gbsn 我不吃六安茶 |
Target: UTF8gbsn 「わしは六安茶(安徽省霍山県—昔、六安郡に属した—産の銘茶)は飲まないのだよ」 |
Source: UTF8gbsn 此刻忽见宝玉笑问道:“宝姐姐,我瞧瞧你的红麝串子? |
Target: UTF8gbsn 宝玉はにこにこしながら、「宝釵お姉さま、あなたの分の赤い香りの香珠、わたしに拝ませてくださいよ」と言い出すのでした。 |
Source: UTF8gbsn 只见晴雯如得了世露一般,一气都灌下去了。 |
Target: UTF8gbsn すると晴雯はさながら甘露でも得たかのように、一気にこれを咽喉に流しこむのでした。 |
Source: UTF8gbsn 刘姥姥道:“这个菜里有毒,俺们那些都成了砒霜了。那怕毒死了,也要吃尽了。” |
Target: UTF8gbsn 劉婆さんは答えて、「このご馳走に毒がはいっておるちゅうことになりますると、てまえどもの料理なんぞは、砒霜(砒素の化合物、毒物)の塊ってえことになっちまいますわい。たとい毒で死のうとも、ご馳走さえ平らげられたら本望でございます」 |
Source: UTF8gbsn 贾珍便命贾琼,责琛、贾璘、贾蔷四个人去陪客,一面吩咐去请钦天监阴阳司来择日。 |
Target: UTF8gbsn に言いつけて客の応対に当たらせることとし、一方では使いをやって欽天監(天文暦法をつかさどる役所)の陰陽司の係(陰陽生、後出。易占をつかさどる)に日柄を見たてさせます。 |
Source: UTF8gbsn 风姐听了这话,便发了兴头,说道:“你是素日知道我的,从来不信什么是阴司地狱报应的。 |
Target: UTF8gbsn 無鳳はこのことばを聞いてつい釣りこまれ、「あなたはかねがねこのわたしという人間をご存じのはず、閻魔の庁だの地獄の応報だのはついぞじたこともないわたしですもの」 |
Source: UTF8gbsn “我昨日叫赖升媳妇出去叫人给宝玉算算命。这先生算得好灵,说要娶了金命的人帮扶他,必要冲冲喜才好。不然只怕保不住。” |
Target: UTF8gbsn わしは昨日、頼昇の家内を街へつかわして宝玉のことを占わせてみたのだが、その占い者の立てた卦というのがぴったりなのだよ—金の性の者を嫁にとって連れ添わせなさい。どうしてもここでひとつ縁起直しせぬことにはいかん。でないと恐らくは生命も保つまい、とこういったそうだ。 |
Source: UTF8gbsn 你舅舅今日斋戒去了 |
Target: UTF8gbsn 伯父さんは今日はお役で斎戒にお出かけになってお留守なの。 |
Source: UTF8gbsn 王子腾那边,仍是一套衣服,一双鞋袜,一百寿桃,一百束上用银丝挂面; |
Target: UTF8gbsn 王子騰のもとからは、例によって衣服一襲・靴下一足・寿桃(誕生祝いに送る小麦粉の桃)百個・宮中用素麵百束が届きます |
Source: UTF8gbsn 这红玉也不梳洗,向镜中胡乱挽了一挽头发,洗了洗手,腰内束了一条汗巾子便来打扫房屋。 |
Target: UTF8gbsn 紅玉とておちおち身づくろいなどしてはおられず、そうそうに鏡に向かって髪をわがね、手を洗い、腰帯をしめるなり、部屋の掃除にやってきました。 |
Source: UTF8gbsn 宝琴披着凫靥裘站在那里笑 |
Target: UTF8gbsn 鳧の毛の装をはおった宝琴がそこに立って笑っています。 |
Source: UTF8gbsn 另换了三四个衣帽周全十七八岁的小厮上来,复抬起轿子,众婆五步下围随,至一垂花门前落下。 |
Target: UTF8gbsn するとこんどは別に三、四人、お仕着せ姿をした十七、八の若党が交替にきて、また輪をかきあげ、老女たちがその囲りをとりまくようにして徒歩でついてゆき、垂花門(正門をはいった中門、二の門。垂花のかざりがあるのでいう)まできて幅をおろしました。 |
Source: UTF8gbsn 散押岁钱荷包金银锞;摆上合欢宴来,男东女西归坐,献屠苏酒、合欢汤、吉祥果、如意糕毕。 |
Target: UTF8gbsn そこで押歳銭(大晦日に長上から子供に取らせるお年玉の金。穴あき銅銭百枚を赤紙で通してある)・中着・金銀の小粒などを分けて取らせます。また合歓宴の宴席を設けて、男は東の、女は西の席につき、屠蘇酒・合歓湯(スープ)・吉祥果(果物)・如意糕(蒸し菓子)を献じ終えました。 |
Source: UTF8gbsn 谋事在人,成事在天 |
Target: UTF8gbsn 人間さまがお膳立て、天道さまがお取りあげ |
Source: UTF8gbsn 那宝玉是个丈八的灯台,照见人家,照不见自家的。 |
Target: UTF8gbsn どだい肝腎の宝玉さまが、『文八(一文八尺)のお灯明台:他人は照らせても、わが身は照らせぬ(「灯台もと暗し」)』。 |
Source: UTF8gbsn 巧媳妇做不出没米的粥来 |
Target: UTF8gbsn 遣繰り上手の嫁さんでも米なしでは粥はできぬ |
Source: UTF8gbsn 偏偏凤姐想出一条偷梁换柱之计 |
Target: UTF8gbsn あいにく鳳ちゃんが替え玉の計略を考え出してくれた |
Source: UTF8gbsn 于是尤氏一行人悄悄的来至窗下,只听里面称三赞四,耍笑之音虽多;又兼着恨五骂六,忿怨之声亦不少。 |
Target: UTF8gbsn かくて尤氏ら一行、足音を忍ばせて窓の下までやってきましたところ、なかではほめたりたたえたりで笑いさんざめく声がしきりにする一方、またわめいたり恨んだりの怒りと憎しみの声も少なくないふう・・・・・ |
Source: UTF8gbsn 我们合家大小登门去磕头。 |
Target: UTF8gbsn わたくしども家中揃ってお宅に伺い叩頭させていただきます |
Source: UTF8gbsn 他是我们这里有名的一个泼皮破落户儿,南省俗谓作辣子。 |
Target: UTF8gbsn これはうちでは聞こえたお転婆の破落戸、江南の方なら俗に『辣子』というやつよ |
Source: UTF8gbsn 他又成了香饽饽了,都抢不到手。 |
Target: UTF8gbsn あのひとはほかほか饅頭(人気者の意)なってしまい、奪い合いでめったに手にははいらないね |
Source: UTF8gbsn 平儿便福下去,宝玉作揖不迭。 |
Target: UTF8gbsn 平児がそこで「万福(女子の敬礼のときのことば)」といってお辞儀をしますと、宝玉は遅れじ揖礼(手を挟いて上下する男子の敬礼法)を返します。 |
Source: UTF8gbsn “孽障!你生气,要打骂人容易,何苦摔那命根子!” |
Target: UTF8gbsn 「この罰あたりめが!おまえ、かんしゃくを起こしたら、人をぶつたりどなりつけたりするのだってたやすいのに、選りに選ってその命の綱も同然の品を投げつけてなんとする?」 |
Fig.5 and 6 present two cases of the original PDF format of our corpus with direct bilingual comparisons.
| Model | Top-\(k\) | Top-\(p\) | Temperature \(T\) |
|---|---|---|---|
| Deepseek-v3 | 20 | 0.95 | 0.6 |
| Deepseek-r1 | 20 | 0.95 | 0.6 |
| Qwen3-235B-A22B-Non-Thinking | 20 | 0.8 | 0.7 |
| Qwen3-235B-A22B-Thinking | 20 | 0.95 | 0.6 |
All model implementations follow default settings. For closed-source models, we do not have access to hyperparameters. For open-source models, specific sampling hyperparameters are listed in Tab.10.
Given the randomness of sampling, theoretically, we should generate multiple responses and report the average. However, due to the high cost of human evaluation, both in time and expense, it is impractical to score every sampled output. To address this, we randomly selected 20 samples, had each model generate eight responses, and scored them separately by our human evaluators. We observed that score variance across samples was minimal, unlike in deterministic reasoning tasks, where outputs can vary significantly [61]–[63]. Under this premise, we use only the first sampled response for evaluation.
Tab.11 presents the prompt used for our llm-as-a-judge evaluations.
| ### Instruction: |
| You are an expert in translation evaluation with native-level proficiency in both source and target languages and deep knowledge of both cultures. You will evaluate a translation based on the following four dimensions. After considering each dimension, provide an **Overall Score** from 0 to 5 that holistically reflects the translation’s quality. |
| — |
| ### Evaluation Dimensions: |
| **1. Content Accuracy** |
| Measures whether the translation faithfully and accurately conveys the meaning of the source text. |
| - High accuracy means no noticeable deviation from the original meaning. |
| - Low accuracy indicates mistranslations, distortions, or loss of core content. |
| **2. Language Fluency** |
| Measures whether the translation conforms to the linguistic norms of the target language, including grammar, vocabulary, spelling, punctuation, and naturalness of expression. |
| - Fluent translations read smoothly and naturally. |
| - Low fluency involves awkward wording, grammatical errors, or obvious translationese. |
| **3. Cultural Appropriateness** |
| Evaluates whether appropriate translation strategies are used when handling cultural elements in the source text, ensuring that cultural connotations are properly conveyed and understandable to target readers. |
| - High cultural appropriateness preserves cultural nuances without causing confusion. |
| - Low appropriateness results in cultural loss, distortion, or misunderstanding. |
| **4. Native Readability** |
| Evaluates whether the translation sounds natural and acceptable to native speakers of the target language and can be easily understood without barriers. |
| - High readability means the text is clear and fully understandable at first reading. |
| - Low readability requires inference, repeated reading, or leaves only a general impression. |
| — |
| ### Overall Scoring Criteria: |
| - **5 - Excellent**: Highly accurate, fluent, culturally appropriate, and reads naturally for native speakers. No noticeable issues. |
| - **4 - Good**: Generally accurate and fluent with minor deviations; cultural meaning largely preserved; occasional awkwardness but easily understood. |
| - **3 - Adequate**: Main content is understandable, but noticeable issues exist in accuracy, fluency, or cultural handling; may require some effort to read. |
| - **2 - Poor**: Significant deviations or errors; cultural meaning poorly conveyed; awkward or hard to follow; only general ideas graspable. |
| - **1 - Very Poor**: Severe distortions or omissions; cultural content missing; highly unnatural or unintelligible. |
| - **0 - Completely Wrong**: Translation is unrelated to the source or empty. |
| — |
| ### Source Text: |
| {source_text} |
| ### Translation: |
| {target_text} |
| ### Evaluation: |
| Provide a **brief justification** (one or two sentences) and then the **Overall Score** in the following format: |
| Justification: [your reasoning] |
| Overall Score: [0-5] |
Tab.12 presents the explanations and scoring guidelines of four specific evaluation dimensions and one overall dimension used in our experiments.
Note: During human evaluation, all evaluators were presented with guidelines in their respective native languages. The English version is provided in this section for ease of presentation.
| Dimension | Definition | Scoring Criteria | |||
|---|---|---|---|---|---|
| Content Accuracy | Measures whether the translation faithfully and accurately conveys the meaning of the source text. | ||||
| 4: Generally accurate with minor distortion. | |||||
| 3: Some deviations exist, but the main content is understandable. | |||||
| 2: Significant deviation; core content difficult to understand. | |||||
| 1: Severely deviates from the source text; core meaning completely lost. | |||||
| Language Fluency | Measures whether the translation conforms to the linguistic norms of the target language, including grammar, vocabulary, spelling, punctuation, and naturalness of expression. | ||||
| 4: Generally fluent with slight unnaturalness. | |||||
| 3: Basically understandable but with noticeable issues. | |||||
| 2: Awkward wording; difficult to understand or obvious translationese. | |||||
| 1: Frequent errors; difficult to understand. | |||||
| Cultural Appropriateness | Evaluates whether appropriate translation strategies are used when handling cultural elements in the source text, ensuring that cultural connotations are properly conveyed and understandable to target readers. | ||||
| 4: Generally fluent and appropriate with minor deviations. | |||||
| 3: Cultural information largely preserved but with some mistranslation. | |||||
| 2: Cultural meaning poorly conveyed, causing misunderstanding. | |||||
| 1: Cultural content missing or seriously distorted. | |||||
| Native Readability | Evaluates whether the translation sounds natural and acceptable to native speakers of the target language and can be easily understood without barriers. | ||||
| 4: Overall natural with minor awkwardness. | |||||
| 3: Comprehension requires inference or repeated reading. | |||||
| 2: Most content obscure; only general idea graspable. | |||||
| 1: Highly unnatural; difficult for native readers. | |||||
| Overall Score | Provides a comprehensive evaluation of the translation’s overall quality by considering all aspects above. | ||||
| 4: Generally good with minor shortcomings. | |||||
| 3: Moderate quality with noticeable issues. | |||||
| 2: Poor overall quality with significant problems. | |||||
| 1: Extremely poor overall performance. |
| Category | Description |
|---|---|
| A. No obvious error | The translation accurately conveys the source content without any issues. |
| The content is incorrectly understood, resulting in a translation that deviates from the original meaning. | |
| Information not present in the source text is added to the translation. | |
| Portions of the source text are left untranslated and directly copied into the target text. | |
| Source information is omitted, resulting in an incomplete translation. |
| Category | Description |
|---|---|
| A. Domestication | Oriented toward the norms of the target culture, adapting the translation to fit the cognitive and cultural expectations of target readers. Strategies include localization, generalization, cultural substitution, functional equivalence, etc. |
| Preserves the cultural characteristics and heterogeneity of the source text, allowing the translation to reflect the original cultural context. Strategies include literal translation, transliteration, semantic borrowing, formal equivalence, etc. | |
| Combine domestication and foreignization to strike a balance between cultural fidelity and reader comprehension. Strategies include literal translation with annotations, retention with explanation, partial literal translation combined with partial free translation, etc. | |
| Translation approaches that do not fall directly under domestication or foreignization, often involving omission or entirely creative renditions. Strategies include omission, innovative translation, etc. |
To better understand error patterns of models, we analyze error distributions across eight models in Tab.7. Errors are categorized into five types: No Obvious Error, Mistranslation, Overtranslation, Undertranslation, and Omission (see Tab.13 for details). Top-performing models, including OpenAI-o4-mini, Gemini-2.5-Flash, and Deepseek-v3, show the highest No Obvious Error rates (55–58%) and lowest Mistranslation (25–36%). In contrast, Deepseek-r1 achieves only 33.8% error-free and 52.7% mistranslation, confirming that misunderstanding source content is the primary accuracy bottleneck.
The remaining three error types are relatively rare, together accounting for under 15% of cases for most models. Undertranslation and Omission remain consistently low across all systems (\(<\)7% and \(<\)5.4% respectively), indicating models rarely refuse to translate or drop content entirely. In addition, Overtranslation shows an interesting pattern: reasoning models exhibit notably higher rates (Qwen3-235B-A22-Thinking 9.5%, Claude-Sonnet-4 8.1%, Gemini-2.5-Flash 8.1%) compared to non-reasoning Deepseek-v3 (2.7%). While longer CoT encourages elaboration, such additions often over-explain cultural elements or introduce irrelevant content, failing to improve outcomes. This confirms that “thinking more” does not guarantee better cultural adaptation — effective translation requires appropriately calibrated elaboration.
We further analyze the translation strategies adopted by the models during the translation process, as shown in Fig.8. Specifically, we draw on the “domestication–foreignization” theory introduced in §2.3 and categorize translation strategies into four types (see Tab.14 for details). Overall, each model employs either domestication or foreignization strategies in over 80% of cases, with mixed or other strategies being relatively infrequent. A clear preference for domestication emerges across models, prioritizing adaptation to target language norms. Notably, reasoning models adopt domestication at higher rates than their non-reasoning counterparts.
Two primary domestication patterns are observed. The first involves Japanese idiomatic expressions, which preserve core meaning and ensure fluency but replace source cultural imagery with target-culture equivalents, resulting in cultural loss. The second relies on generic paraphrasing, which captures only coarse semantics while omitting nuanced components, leading to partial meaning degradation (see §14.1 and 14.4 for a typical case). These findings suggest that while domestication enhances readability and target-language acceptability, it often comes at the cost of source culture preservation — underscoring a key limitation of current models in conveying culturally embedded meanings.
Source:
UTF8gbsn 他又成了香饽饽了,都抢不到手。
Reference:
UTF8min あのひとはほかほか饅頭(人気者の意)なってしまい、奪い合いでめったに手にははいらないね
Deepseek-v3:
UTF8min 彼はまた人気者になって、みんな手に入れようにも手に入れられないんだ。彼はまた人気者になって、みんなが奪い合っている。
Qwen3-235B-A22B-Non-Thinking:
UTF8min 彼はまたもてはやされる存在になり、手に入れようとしても手に入らないほどだ。
GPT-4.1:
UTF8min また引っ張りだこになって、誰も手に入れられないよ。
Gemini-2.5-Flash:
UTF8min 彼はまた引っ張りだこになって、もう手に入らないよ。
Deepseek-r1:
UTF8min 彼はまた引く手あまたになってしまって、もらえないんです。
Qwen-235B-A22B-Thinking:
UTF8min 彼はまたもてはやされる存在となり、誰もが手に入れようとするものの、なかなか手に入らない状態です。
OpenAI-o4-mini:
UTF8min 彼はまた人気者になって、みんな手に入れようにも手に入れられないんだ。
Claude-sonnet-4:
UTF8min 彼はまた人気者になって、取り合いになっているよ。
The expression
UTF8gbsn”香饽饽”
in the original text functions as a key culturally loaded word. Literally, it refers to a “steaming hot bun”, while metaphorically it conveys the cultural meaning of “a highly sought-after person”. Ito renders it as
UTF8min「ほかほか饅頭(人気者の意)」
. This approach represents a classical paradigm in cultural translation: the source-language imagery is preserved through literal translation, while the metaphorical meaning is made explicit through explanatory notes in parentheses. In this way, the translation both maintains the cultural imagery of the source text and ensures comprehensibility for target-language readers.
LLMs tend to prioritize alignment with the linguistic conventions and expression patterns of the target language at the expense of source-language cultural imagery. This preference is realized primarily through two approaches:
The first approach involves replacing the source expression with established Japanese idioms. For example, GPT-4.1 and Gemini-2.5-Flash translate the term as
UTF8min「引っ張りだこ」
, while Deepseek-r1 uses
UTF8min「引く手あまた」
. Although these expressions accurately convey the core meaning of “being highly sought after” and produce fluent, idiomatic output, they essentially constitute cultural substitution. As a result, the culturally specific imagery associated with
UTF8gbsn”饽饽”
, a symbol rooted in Chinese culinary culture, is lost.
The second approach generalizes the expression into more neutral vocabulary. Examples include
UTF8min「人気者」
in the translations by Deepseek-v3 and Claude-4-sonnet, and
UTF8min「もてはやされる存在」
produced by Qwen series. While such renderings convey the attribute of “being popular,” they substantially weaken or even eliminate the dynamic semantic component of “being competed for” (i.e., the sense of rivalry or competition implied in the original expression). Consequently, part of the original semantic richness is lost.
Source:
UTF8gbsn 那宝玉是个丈八的灯台,照见人家,照不见自家的。
Reference:
UTF8min どだい肝腎の宝玉さまが、『丈八(一丈八尺)のお灯明、他人は照らせても、わが身は照らせぬ(「灯台もと暗し」)』。
Deepseek-v3:
UTF8min 宝玉は八尺の灯台で、他人を照らすが、自分の足元は照らさない。
Qwen3-235B-A22B-Non-Thinking:
UTF8min その宝玉というやつは、丈八(じょうはち)の灯台(とうだい)のようなもので、他人(ひと)のことは照らし出すが、自分自身は照らし出せないのだ。
GPT-4.1:
UTF8min あの宝玉は、まるで丈八(じょうはち)の灯台のようなもので、人のことはよく見えるが、自分のことは見えないのです。
Gemini-2.5-Flash:
UTF8min この宝玉(ほうぎょく)は、まさに「灯台下暗し(とうだいもとくらし)」という男だった。
Deepseek-r1:
UTF8min 宝玉は一丈八尺の灯籠台、他人を照らすが、自分の足元は見えぬ
Qwen-235B-A22B-Thinking:
UTF8min 宝玉は丈八(じょうはち)の灯台のようなものだ。他人は照らすが、自分自身は照らせぬ。
OpenAI-o4-mini:
UTF8min あの宝玉は丈八の燈台みたいなもので、人のことはよく照らせても、自分の家(身内)のことはまったく見えない。
Claude-sonnet-4:
UTF8min その宝玉は、他人のことはよく分かるが、自分のことは分からない人なのである。
This case belongs to a citation-type two-part allegorical saying
UTF8gbsn(歇后语)
structure composed of “tenor + vehicle + explanation”. The first half establishes the metaphor, while the second half reveals the intended meaning of the pun. From a translation perspective, the major difficulty of this sentence lies in two aspects: the treatment of the cultural image and the explicit rendering of the metaphorical implication. On the one hand, the translation needs to reproduce the exaggerated metaphorical image of
UTF8gbsn”丈八的灯台”
; on the other hand, it must convey the satirical meaning embedded in the original expression, namely “seeing others’ faults while failing to examine one’s own”.
Ito’s translation adopts a strategy of preserving the source metaphor while providing explanatory support. First, the expression
UTF8min「丈八(一丈八尺)のお灯明」
retains the original imagery, and the term
UTF8gbsn”丈八”
is annotated to ensure comprehension for target-language readers. Second, the latter half is translated literally as
UTF8min「他人は照らせても、わが身は照らせぬ」
, which accurately conveys the meaning of the original while maintaining a formal correspondence with the source structure. Finally, the translator supplements the Japanese proverb
UTF8min「灯台もと暗し」
, further facilitating reader understanding. As a result, the translation both preserves the cultural imagery and makes the metaphorical meaning explicit, achieving a desirable balance between cultural fidelity and reader comprehensibility.
For LLMs, their translations can generally be categorized into the following types:
The first type combines structural preservation with explanatory adaptation. For example, GPT-4.1 adopts a canonical Japanese comparative construction for expressing metaphor and contrast,
UTF8min「まるで…のようなもので、…よく見えるが、…見えないのです」
. The resulting translation is fluent and rhythmic, effectively restoring the syntactic structure of the original sentence.
The second type adopts a literal translation strategy for the cultural image. Examples include deepseek-v3-250324, o4-mini-2025-04-16, deepseek-r1-250528, and the qwen3 series. A common characteristic of these translations is that
UTF8gbsn”丈八”
is directly rendered as
UTF8min「八尺」
,
UTF8min「丈八」
, or
UTF8min「一丈八尺」
. Although such literal translation preserves the cultural features of the source language to the greatest extent, the absence of further explanation or adaptation may reduce readability for Japanese readers unfamiliar with the cultural reference.
The third type adopts a domestication or paraphrasing strategy, as seen in the translations produced by gemini-2.5-flash and claude-sonnet-4-20250514. While these translations achieve high fluency, the cultural imagery of the source text is entirely replaced or erased. From the perspective of cultural transmission, this strategy may fulfill communicative purposes but inevitably results in a loss of cultural information.
Finally, some translations generated by the DeepSeek series contain clear errors. In these cases,
UTF8gbsn”丈八的灯台”
is interpreted as a literal physical object rather than a metaphor, and the metaphorical relationship is therefore not reflected. Moreover,
UTF8gbsn”照不见自家”
is rendered as
UTF8min「自分の足元は見えない」
, which concretizes the expression into a spatial description. Consequently, the original satirical implication—referring to a lack of self-reflection—degenerates into a neutral depiction of physical visibility.
Source:
UTF8gbsn ......只听里面称三赞四,耍笑之音虽多;又兼着恨五骂六,忿怨之声亦不少。
Reference:
UTF8min ......なかでは、ほめたりたたえたりで、笑いさんざめく声がしきりにする一方、またわめいたり恨んだりの怒りと憎しみの声も少なくないふう......
Deepseek-v3:
UTF8min 中では三人を褒め四人を誉め、笑い戯れる声が多いかと思えば、また五人を恨み六人を罵る、怒り怨む声も少なくない。
Qwen3-235B-A22B-Non-Thinking:
UTF8min 中からは三人称えて四人ほめる声、遊び笑う音は多いにあり、また五人を恨み六人を罵る声、怒りと不満の言葉も少なくない。
GPT-4.1:
UTF8min 中からは、「三を褒め四を称える」声や、ふざけ合う音が多く聞こえてきたが、一方では「五を恨み六を罵る」怒りや恨みの声も少なからず混じっていた。
Gemini-2.5-Flash:
UTF8min (30)中からはお世辞を言ったり、褒めそやしたり、ふざけ合ったりする声が多く聞こえたが、その一方で、罵声や恨み言、不平不満の声も少なくなかった。
Deepseek-r1:
UTF8min 中からは賑やかな笑い声とともに、褒めそやす話し声が多く聞こえる一方で、恨み罵る憤りの声も少なからず聞こえてきた。
Qwen-235B-A22B-Thinking:
UTF8min 中をうかがうと、そこには三人称賛し、四人で笑い声を立てている者が多く、また五人を恨み、六人を罵るような怒りや不満の声も少なくなかった。
OpenAI-o4-mini:
UTF8min 中からは絶え間なくおだてるような笑い声が聞こえてきたが、同時に罵倒や恨み言を口にする怒りの声も少なくなかった。
Claude-sonnet-4:
UTF8min 中からは三を褒め四を讃えるような、戯れ笑う声が多く聞こえてくるものの、同時に五を恨み六を罵るような、憤懣怨嗟の声も少なくなかった。
Taking the expression
UTF8gbsn”称三赞四”“恨五骂六”
, a typical culturally loaded word under the linguistic category, as an example, the key challenge in its translation lies not only in correctly interpreting the rhetorical device of numerical vagueness characteristic of Chinese, but also in reproducing the rhythmic vitality and emotional intensity created by the parallel structure of the original expression.
Ito renders
UTF8gbsn”称三赞四”
as
UTF8min「ほめたりたたえたり」
and
UTF8gbsn”恨五骂六”
as
UTF8min「わめいたり恨んだりの怒り」
. This translation employs the Japanese parallel construction
UTF8min「〜たり〜たり」
, which not only accurately conveys the semantic content of the original but also reproduces its rhythmic pattern and emotional tension. In contrast, LLMs generally fall into two categories.
On the one hand, the translations by Deepseek-v3, GPT-4.1, and Claude-sonnet-4 mistakenly interpret the numerals
UTF8gbsn”三、四、五、六”
(“three, four, five, six”), which function rhetorically as vague intensifiers, as literal quantity markers. Consequently, they are rendered as expressions such as
UTF8min「三人を褒め四人を讃える」
(“praising three people and commending four”) or
UTF8min「五人を恨み六人を罵る」
(“hating five people and insulting six”), resulting in a clear distortion of the original meaning.
On the other hand, OpenAI-o4-mini, Deepseek-r1, and Gemini-2.5-Flash, render
UTF8gbsn”称三赞四”
and
UTF8gbsn”恨五骂六”
as expressions such as
UTF8min「お世辞を言う」
,
UTF8min「褒めそやす」
,
UTF8min「罵倒する」
, or
UTF8min「恨み言」
. Although these translations successfully convey the core semantic content of praise and denunciation expressed in the original, they lack the linguistic tension and rhetorical dynamism of the source text, and therefore fail to fully reproduce its tone and emotional force.
Source:
UTF8gbsn 巧媳妇做不出没米的粥来
Reference:
UTF8min 遣繰り上手の嫁さんでも米なしでは粥はできぬ
Deepseek-v3:
UTF8min 巧婦、米なしに粥を作ることはできない。
Qwen3-235B-A22B-Non-Thinking:
UTF8min 巧妇も米がなければ粥は作れない。
GPT-4.1:
UTF8min どんなに腕のいい主婦でも、米がなければお粥は作れない。
Gemini-2.5-Flash:
UTF8min 元手がなければ商売はできない
Deepseek-r1:
UTF8min ない袖は振れぬ
Qwen-235B-A22B-Thinking:
UTF8min 巧婦も米がなければ粥は炊けぬ。
OpenAI-o4-mini:
UTF8min どんなに器用な嫁でも、米のないお粥は作れない。
Claude-sonnet-4:
UTF8min 米なしでは、どんなに腕の良い嫁でも粥は作れない
Overall, the translations produced by LLMs can be broadly divided into two categories.
The first category adopts a domestication strategy aimed at achieving functional equivalence. A representative example is the DeepSeek-r1 translation
UTF8min「ない袖は振れぬ」
. This type of translation completely abandons the concrete imagery in the original text—such as
UTF8gbsn”米”
,
UTF8gbsn”粥”
, and
UTF8gbsn”媳妇”
—and instead aligns directly with a semantically equivalent idiomatic expression in Japanese culture. Although this approach diverges significantly from the original in terms of linguistic form, it succeeds in achieving pragmatic equivalence at the functional level.
The second category adopts a foreignization strategy, preserving the cultural imagery of the source language. For instance, the Qwen3-thinking translation
UTF8min「巧婦も米がなければ粥は炊けぬ」
retains the original imagery of
UTF8gbsn”巧妇”
,
UTF8gbsn”米”
, and
UTF8gbsn”粥”
. Moreover, by employing expressions such as
UTF8min「炊けぬ」
, which carry a strong classical Sino-Japanese stylistic flavor, the translation closely mirrors the structural and rhetorical style of the original Chinese proverb, thereby achieving a dual form of foreignization in both imagery and syntax.
It is worth noting, however, that when foreignization lacks sufficient cultural contextual support, it may lead to semantic deviation. Models such as Gemini-2.5-Flash and GPT-4.1 exhibit varying degrees of strategic fluctuation in their sentence-level translations. For example, Gemini renders the proverb as
UTF8min「元手がなければ商売はできぬ」
(“without capital, business cannot be conducted”). Although the translation maintains the formal structure of a proverb, it introduces a substantial shift in meaning by abstracting the original expression into a commercial context.
Source:
UTF8gbsn 那畸角上我有一个花冢。
Reference:
UTF8min あの隅の方に、わたくし、花塚をこしらえてあります
Deepseek-v3:
UTF8min あの隅っこに私の花塚があります。
Qwen3-235B-A22B-Non-Thinking:
UTF8min その変な角のところに、私はお墓参りをしています。
GPT-4.1:
UTF8min その歪んだ角に、私は一つの花塚(はなづか)を持っている。
Gemini-2.5-Flash:
UTF8min その奇妙な角の上に、私は一つの花塚を建てた。
Deepseek-r1:
UTF8min あの隅っこに私の花の塚があるんだ。
Qwen-235B-A22B-Thinking:
UTF8min あの片隅に、私の花塚があります。
OpenAI-o4-mini:
UTF8min あの歪んだ角に、私はひとつの花の塚を持っている。
Claude-sonnet-4:
UTF8min その隅っこに私は花塚を持っています。
UTF8gbsn”畸角”
originally refers to a spatial notion, meaning a “remote, neglected, or unnoticed corner”. The translation strategies adopted by LLMs models can be broadly divided into two categories.
The first category consists of contextually grounded literary translations. Representative examples include the Reference translation, Qwen3-235B-A22B-Thinking, and the Deepseek series. These models correctly identify the spatial attribute embedded in the character
UTF8gbsn”畸”
within Chinese rhetoric. The Reference translation uses
UTF8min「隅の方」
, while DeepSeek renders it as
UTF8min「隅っこ」
, both of which accurately anchor the expression in physical spatial orientation. Among them, the translation
UTF8min「片隅」
produced by Qwen3-235B-A22B-Thinking stands out as particularly effective. In Japanese lexical intuition,
UTF8min「片隅」
not only denotes a marginal space but also carries a subtle sense of loneliness and isolation, which aesthetically resonates with the imagery of
UTF8gbsn”花冢”
. This alignment suggests that models equipped with stronger reasoning abilities can move beyond literal lexical correspondence and achieve cross-lingual equivalence at the level of scene or atmosphere, rather than merely word meaning.
The second category can be described as moderate translations based on literal semantic equivalence. Models such as GPT-4.1, OpenAI-o4-mini, and Gemini-2.5-Flash exhibit a clear tendency toward dictionary-like literal translation. They interpret the character
UTF8gbsn”畸”
as indicating abnormal shape, rendering the phrase as expressions such as
UTF8min「歪んだ角」
or
UTF8min「奇妙な角」
. Similarly, Qwen3-non-thinking translates it as
UTF8min「変な角」
. Such treatments establish only a superficial lexical correspondence and fail to capture the intended meaning of the original expression.
In summary, the quality of translating
UTF8gbsn”畸角”
largely depends on whether the model recognizes its metaphorical function as an environmental modifier rather than a description of shape. Models capable of detecting contextual cues and employing expressions such as
UTF8min「片隅」
or
UTF8min「隅」
, which align with natural usage in the target language, demonstrate clear advantages over literalist approaches in both translational accuracy and literary aesthetic quality.
The Chinese novel
UTF8gbsn《红楼梦》
, also known as
UTF8gbsn《石头记》
, was first translated as Dream of the Red Chamber by Wang Jizhen (1929). David Hawkes later adopted the alternative title The Story of the Stone in his influential translation (1973–1986). Following broader academic and general convention, we use the former title to refer to the original work.↩︎
Equal Contribution. We also thank Natsuki Oe, Kanon Yamaguchi, and Mouye Weng for coordinating the human study with Japanese volunteers, and Yuya Goto of Mejiro University for his suggestions on parts of the background discussion and case studies in this manuscript.↩︎
Translated by Chinese translator Yuanchong Xu.↩︎
https://zh.wikipedia.org/wiki/
UTF8gbsn大中华文库
The standard MQM framework includes five dimensions: Accuracy, Fluency, Verity, Design, and Internationalization.↩︎
In recent MT research, several culture-oriented translation studies have also emerged; we discuss them in §8.↩︎