Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding


Abstract

Existing AI-assisted oracle bone inscription (OBI) visual recognition and understanding studies mainly focus on character-level, ignoring the long-form textual coherence and contextual dependencies embedded in complete divination charges. Recently, the powerful visual perception capabilities of multimodal large language models (MLLMs) have opened new possibilities for OBI information processing. In this work, we introduce S-OBI, a novel benchmark for evaluating MLLMs in Sentence-level OBI understanding. Instead of using noisy and incomplete rubbings as the visual input, S-OBI synthesizes clear and standardized sentence-level OBI instances through glyph substitution and composition. According to 95 original rubbings with translations that have been identified, corrected, and verified by experts, we replace characters in the original rubbings with corresponding clean glyph samples sourced from existing OBI datasets while preserving the overall inscriptional structure and semantic organization. This mitigates the influence of low-level distortions and enables a more focused evaluation of sentence-level OBI understanding. Based on this, we design semantic matching, semantic slot extraction, and contextual reasoning tasks and obtain 695 question-answer pairs. Experiments reveal the inferiority of contemporary MLLMs on sentence-level OBI understanding. In particular, visual perception errors in unmasked regions propagate through the reasoning chain, leading to erroneous predictions for masked characters, which indicates that sentence-level OBI understanding in current models remains strongly dependent on character-level recognition. Overall, S-OBI provides a diagnostic benchmark for evaluating whether MLLMs can move beyond isolated character recognition toward structured inscription-level understanding.

1 Introduction↩︎

Oracle bone inscriptions (OBIs) are among the earliest mature writing systems in ancient China. They were mainly carved on turtle plastrons and animal bones in the late Shang dynasty and provide important evidence for studying early Chinese language, political organization, religious practices, and social structure [1]. Unlike ordinary text recognition tasks, an OBI record is not simply a collection of isolated characters. A complete divinatory inscription usually follows a relatively structured form and may contain components such as the date, diviner, divination marker, charge, object, action, and sometimes missing or implicit contextual information. Therefore, OBI understanding requires more than recognizing individual glyphs. It also requires the model to integrate glyph perception, character order, semantic roles, and divinatory context to interpret the complete inscription.

Traditional OBI studies mainly rely on expert interpretation. Researchers usually examine published materials, rubbings, hand copies, photographs, dictionaries, and archaeological context to decipher characters and interpret inscriptions [2]. This expert-led process provides reliable and explainable results, but it is time-consuming, highly dependent on specialist knowledge, and difficult to scale. With the development of computer vision and deep learning, OBI processing has gradually shifted from manual consultation to data-driven image recognition [3]. Existing datasets and methods [4][6] have supported tasks such as OBI character detection, classification, retrieval, rejoining, and decipherment, enabling models to learn the correspondence between single-character images and modern Chinese characters or character categories. More recently, multimodal large language models (MLLMs) have been introduced into OBI-related tasks [7][9], making it possible to explore the relation between visual glyph features, semantic knowledge, and cross-modal reasoning.

Figure 1: Examples of OBI appearances covered in constructing the S-OBI. The red and green rectangles are paired original sentential OBI rubbings and the corresponding synthetic OBI sentences, respectively.

However, most existing studies focus on character-level or localized tasks. Their basic evaluation unit is typically a single glyph, with the primary objective being character identification or mapping the glyph to its modern Chinese equivalent. While these tasks are fundamental to OBI recognition and decipherment, they fall short of fully evaluating whether a model comprehends the internal structure of a complete divinatory inscription. For example, a model may recognize several individual characters but fail to identify their semantic roles in the inscription. It may also select a plausible interpretation without correctly recovering the reading order, punctuation boundaries, or missing characters. For a practical OBI understanding system, character recognition alone is not sufficient. The model should also be able to reason over sentence-level visual inputs and combine glyph identity, character order, divinatory structure, and contextual clues.

To address this limitation, we introduce S-OBI, a diagnostic benchmark for sentence-level OBI understanding. Unlike previous studies that mainly focus on isolated characters, S-OBI treats a complete divinatory inscription as the basic evaluation unit. It is built from 95 expert-corrected inscriptions and uses high-quality single-character OBI resources to reconstruct standardized sentence-level OBI images through glyph substitution. Fig. 1 provides an example of the original sentential OBI rubbings, the single-character OBI pool, and the synthetic OBI sentences. This design reduces the influence of low-level visual noise and allows a more focused evaluation of sentence-level understanding. S-OBI contains 695 question-answer instances across three tasks, i.e., semantic matching, semantic slot extraction, and contextual reasoning. These tasks evaluate whether MLLMs can understand the overall meaning of an inscription, recover its structured semantic frame, and reason about punctuation boundaries and masked glyphs from context.

We systematically evaluate representative MLLMs on S-OBI. The results show that current models are still far from reliable sentence-level OBI understanding. Although some models can handle coarse semantic selection or reading-order choices to a certain extent, they struggle with structured semantic recovery, punctuation restoration, and masked glyph completion. These findings suggest that current MLLMs have not yet developed a robust ability to model the structure and context of complete divinatory inscriptions.

The main contributions of this paper are summarized as follows:

  • We propose a new sentence-level OBI understanding task. Different from previous studies that mainly focus on isolated characters, our task treats a complete divinatory inscription as the basic evaluation unit and requires models to jointly consider glyph perception, character order, semantic roles, and divinatory context.

  • We construct S-OBI, a diagnostic benchmark for sentence-level OBI understanding. S-OBI reconstructs standardized sentence-level OBI images through glyph substitution using high-quality single-character OBI resources and provides 695 question-answer instances covering semantic matching, semantic slot extraction, and contextual reasoning.

  • We evaluate 10 MLLMs on sentence-level OBI understanding. The results show that existing models can sometimes handle coarse semantic selection but still struggle with structured semantic recovery, punctuation restoration, and masked glyph completion, indicating that sentence-level OBI understanding remains a challenging problem for current MLLMs.

2 Related Work↩︎

2.1 Oracle Bone Inscription Datasets↩︎

Existing OBI datasets have primarily been developed for character-level recognition [7], classification [10], retrieval, generation [11], and decipherment tasks [1]. Representative datasets, such as OBC306 [4], Oracle-MNIST [5], and HUST-OBC [6], contain large-scale oracle character images with category labels for recognition or decipherment, serving as a common basis for single-glyph evaluation. Another important line of character-level datasets usually pairs cropped OBI glyphs with their modern Chinese counterparts, textual glosses, or evolutionary forms, thereby providing supervised signals for bridging the historical gap between ancient and modern scripts. For example, EVOBC [12] models the evolutionary trajectory from oracle bone script to later Chinese scripts. OracleSem [13] enriches character-level samples with semantic annotations, including pictographic composition, structural organization, and semantic evolution. PictOBI-20k [14] links oracle glyphs with real-world object images and question-answer pairs to evaluate MLLMs on visual decipherment of pictographic OBIs.

However, these evaluation artifacts remain predominantly isolated in character, which mainly answer questions, such as Which character is this? or Which modern Chinese character may this glyph correspond to?, while paying limited attention to inscription-level structure, character order, semantic roles, and contextual reasoning.

2.2 OBI Understanding↩︎

Early OBI understanding was primarily expert-led, with researchers manually examining published materials, rubbings, tracings, photographs, and ancient books to decipher characters and interpret inscriptions. This process laid the foundation for OBI studies, but it was time-consuming, highly dependent on specialist knowledge, and difficult to reproduce at scale [1], [15]. Recently, the rise of deep learning shifted OBI processing from expert-centric manual procedures to data-driven automatic frameworks. Deep neural networks have been widely applied to OBI character recognition, rejoining, classification, and retrieval, and Transformer-based architectures further improved the modeling of structural dependencies in degraded or fragmented glyphs [16][18]. However, these methods are still mainly driven by visual features and often lack explicit modeling of the textual and semantic content of divinatory inscriptions. More recently, large multimodal models have introduced a new paradigm for cross-modal OBI perception and interpretation [19], enabling the field to move beyond isolated character recognition toward contextual decipherment. Representative efforts include OBI-Bench [7], OracleSage [13], OracleFusion [20], and related LMM-based OBI decipherment systems [21], [22]. Nevertheless, current LMMs remain fragile on noisy, original, and cross-modal OBI inputs, and their ability to perform structured sentence-level reasoning over complete divinatory inscriptions is still insufficient. Our proposed S-OBI is designed to address this gap by treating the sentence-level OBI image sequences as the basic evaluation unit.

Figure 2: The construction pipeline of S-OBI. We reconstruct the sentence-level OBI images based on the expert-corrected, punctuated, interpreted multi-character OBI images. Three task tiers are designed to evaluate MLLMs through semantic matching, slot extraction, and order-sensitive contextual understanding.

3 S-OBI Benchmark↩︎

3.1 Task Definition↩︎

S-OBI evaluates whether an MLLM can perceive a sentence-level OBI image and a task prompt to a semantic, structural, or contextual answer. The benchmark is designed around complete inscriptional groups rather than isolated glyphs, which therefore challenges models in combining glyph perception, character order, divination formulas, and sentence-level semantic composition. Each evaluation instance is represented as \[q_i=(X_i,p_i,a_i|t_i),\] where \(i\) is the instance index. \(X_i\) denotes the sentence-level OBI image. \(p_i\) and \(a_i\) denote the task prompt and the expert-corrected ground-truth answer, respectively. \(t_i\) denotes the task type. Given the image-prompt pair \((X_i,p_i)\), an evaluated MLLM \(M\) predicts an answer: \[\hat{a}_i=M(X_i,p_i).\] The prediction \(\hat{a}_i\) is then compared with the ground-truth answer \(a_i\) according to the corresponding task type \(t_i\).

3.2 Data Curation↩︎

3.2.1 Source Content Collection with Expert Verification.↩︎

High-quality original OBI rubbings and reliable translations are the foundation of S-OBI. In this work, we mainly enforce two criteria. First, the oracle bone characters on the rubbing should contain multiple characters, be sufficiently complete, and identifiable. Second, the original OBI rubbings together with their translations should be semantically complete, of a certain length, and relatively uncontroversial. Based on this, we choose an authentic OBI book published by Liu et al. [23], and select 95 sentential OBI rubbings (Fig. 1 (a)) as the basis for S-OBI construction. We invited 5 OBI experts to parse, punctuate, and interpret these raw materials. The whole human verification process lasts over a month. Fig. 2 illustrates the overall construction pipeline.

3.2.2 Glyph Substitution.↩︎

Given the scarcity of original sentence-level oracle bone inscription (OBI) materials, which are insufficient to support diverse evaluation needs, we synthesize various sentence-level OBI image sequences by replacing individual glyphs while strictly preserving the original semantic content. After obtaining the meticulously annotated sequential OBI rubbings, we first extract individual characters from the source images. We then determine their corresponding meanings and map them to appropriate character categories using large-scale OBI datasets, such as OBC306 [4] and HUST-OBC [6]. Within each matched category, we choose clear and reliable glyph samples as substitution candidates. These candidates are then used to generate diverse sentence-level OBI images. The purpose of glyph substitution is not to change the semantic structure of the original inscription but to reduce accidental visual noise and degradation in the original rubbing. By replacing noisy or incomplete glyphs with clearer and more diverse corresponding samples, S-OBI preserves the original inscriptional structure while enabling a more focused evaluation of sentence-level understanding.

3.2.3 Sentence-level OBI reconstruction.↩︎

After obtaining candidate glyphs, we compose them according to the original character order and generate standardized sentence-level OBI images. Meanwhile, we generate controlled variants, e.g., substituted glyph sequences, shuffled character orders, and masked-glyph images, for different tasks, allowing the benchmark to evaluate different abilities while keeping the underlying inscriptional meaning consistent.

3.3 Task Design and Difficulty Tiers↩︎

In this section, we elaborate on the three tasks, i.e., semantic matching, semantic slot extraction, and contextual reasoning in S-OBI (Fig. 2).

3.3.1 T1: Semantic matching.↩︎

This task includes two variants, i.e., meaning selection and semantic order judgment. In the meaning selection task, given a sentence-level OBI image with a complete divination charge and a natural language question concerning its holistic semantic content, the model is required to select the correct answer from multiple candidates. In the semantic order judgment task, the sentence-level OBI image is divided into several semantic units, and the model is asked to select the correct reading order from four candidate orders. This task evaluates not only whether the model can identify the reading order of the inscription, but also whether it can use the semantic meaning of the corresponding characters to recover the correct order from the visual input.

3.3.2 T2: Semantic slot extraction.↩︎

This task requires models to recover a structured divination frame from a sentence-level OBI image. Specifically, the model needs to recognize Oracle bone characters and output the contents of each slot in the structured frame, which assesses both direct identification of OBI glyphs and inference of ambiguous or problematic character types from known context. Compared with single-character recognition, this task evaluates whether an OBI understanding system can use known characters to infer other less recognizable characters within the same inscriptional group.

3.3.3 T3: Contextual reasoning.↩︎

Furthermore, to fully examine the fine-grained contextual reasoning ability, we introduce masked glyph completion and shuffled order recovery. For the first task, we randomly mask 30% of the characters in a sentence-level OBI image and select one masked position as the target. The model is required to choose the corresponding modern Chinese character of the masked glyph from four options. For the second task, we randomly shuffle the character order in a sentence-level OBI image and ask the model to recover the correct reading order from the candidate options. These two subtasks explicitly evaluate the understanding ability of an OBI understanding system at the sentence level and the ability to recover the correct reading order based on contextual and semantic cues.

3.3.4 Difficulty Split.↩︎

We adopt an intuitive division strategy, namely dividing tasks into four difficulty levels according to the number of characters in the sentence-level OBI image [24]. Four subsets with different scales, i.e., samples with \(2 \leq L \leq 5\), \(6 \leq L \leq 8\), \(9 \leq L \leq 10\), and \(L \geq 11\) characters, are collected. We discuss our observations among different difficulty levels in Sec. 4.4, where we reveal the relevance between the length of sequential OBI images and different tasks.

4 Experiments↩︎

4.1 Experimental Setup↩︎

We evaluate ten cutting-edge MLLMs, including proprietary models, GPT-5.2 [25], GPT-4o [26], Gemini 3.1 Pro [27], and Qwen3.6-Plus [28], from OpenAI, Google, and Alibaba, as well as open-source models, DeepSeek-V4-Pro [29], InternVL3.5 [30], Qwen3.6-35B [28], Qwen3.5-9B [31], and Qwen2.5-VL-72B [32] from DeepSeek, Shanghai AI Laboratory, and Alibaba. Apart from the proprietary models and those with a parameter size greater than 72B that are deployed via API, all other models are performed using up to 8 Nvidia RTX4090 24GB. We use the default parameter settings (e.g., temperature, top_k, and top_p) of respective models. We mainly report the overall, task-averaged, and instance-wise accuracy scores in percentage for performance comparison.

4.2 Main Results↩︎

Table 1: Performance comparison on S-OBI. The upper and lower parts are closed-source and open-source MLLMs, respectively. The best and second-best results are in bold and underlined.
Model Overall Avg. T1 T2 T3
GPT-5.2
GPT-4o
Gemini-3.1-Pro 31.4 32.2 59.0
Qwen3.6-Plus-thinking 33.5 34.1 60.9 27.2 14.1
DeepSeek-V4-Pro
Qwen3.6-35B-A3B 13.3
Qwen3.5-9B 25.2
Qwen2.5-VL-72B
InternVL3.5-38B
InternVL3.5-8B

3pt

Current MLLMs are still far from reliable sentence-level OBI understanding. Tab.  1 shows that the best overall score is only 33.5%, achieved by Qwen3.6-Plus-thinking, while Gemini-3.1-Pro ranks second at 31.4%. All remaining models fall below 30% overall. This ceiling indicates that S-OBI is not solved by general OCR-rich visual recognition or broad multimodal instruction following.

The main difficulty lies in moving from plausible meaning selection to structured and contextual interpretation. For the strongest model, performance drops from 60.9% on T1 to 27.2% on T2 and 14.1% on T3. The same pattern appears across model families, i.e., models can often choose a semantically plausible answer, but they struggle to recover the internal divination frame and to use sentence context for punctuation or masked-glyph reasoning.

Open-source models are competitive, but the gap to reliable benchmark performance remains large. Qwen2.5-VL-72B and InternVL3.5-38B are the strongest open-source models, reaching 28.1% and 27.3% overall, respectively. DeepSeek-V4-Pro follows at 25.1%. Their proximity to several closed-source systems suggests that the access regime alone does not explain performance. We speculate that the dominant bottleneck is the lack of sentence-level ancient-script grounding.

4.3 Diagnostic Results↩︎

Table 2: Performance comparison on the subtasks of S-OBI.
Model Avg. Meaning Order Slots Punc. Mask
GPT-5.2
GPT-4o 20.0
Gemini-3.1-Pro 33.5 32.6 14.9
Qwen3.6-Plus-thinking 35.6 32.6 89.2 27.2 12.5 16.7
DeepSeek-V4-Pro 87.8
Qwen3.6-35B-A3B
Qwen3.5-9B 25.2 16.7
Qwen2.5-VL-72B
InternVL3.5-38B 33.7
InternVL3.5-8B
Task-wise Avg.

3pt

Figure 3: Pearson’s correlation between different tasks.

Models capture coarse order more easily than sentence-level meaning. Tab.  2 reports an average score of 79.6% for semantic-unit order selection, but only 22.6% for meaning matching. This contrast is visible even in strong runs. For example, Qwen3.6-Plus-thinking reaches 89.2% on order selection but 32.6% on meaning matching, while DeepSeek-V4-Pro reaches 87.8% on order selection but only 13.7% on meaning matching. This suggests that visible sequence regularities are easier to exploit than inscription-level semantic interpretation.

Structured slot extraction remains a central bottleneck rather than a formatting issue. The average slot score is nearly 20.0%, and the best score is 27.2%. Although several models produce syntactically valid structured outputs, all models obtain 0% complete slot-frame exact match. This gap shows that models often miss at least one essential component of the divination structure, such as the subject, action, target, time, outcome, preface, or charge.

Order-sensitive contextual recovery is especially weak. Punctuation restoration remains below 15% for every model, with Gemini-3.1-Pro achieving the best score at 14.9%. Besides, masked glyph completion is also fragile. The best result is below the nominal 25% chance level of a four-choice task. These failures indicate that current MLLMs can not yet use visible context and inscriptional order robustly when the local glyph evidence is incomplete.

Correlation analysis. Fig. 3 shows Pearson’s correlation between different tasks. We observe moderate interdependence (\(r \in [0.371, 0.564]\)) in macro-level correlations (Fig. 3 (a)), demonstrating the rationality of our task design and the reduction of redundancy. Granular subtask analysis (Fig. 3 (b)) reveals a strong semantic-syntactic coupling (\(r = 0.817\)) between "Meaning" and "Punc.", indicating that LMMs heavily depend on overall semantic decoding to segment unpunctuated ancient scripts. Structural reasoning tasks, such as "Order" and "Slots", align moderately, suggesting shared downstream pathways for grammatical arrangement. Conversely, the "Mask" task acts as an orthogonal outlier (e.g., \(r = -0.142\) with Meaning), proving that masked character restoration prioritizes local visual pattern matching over broad semantic logic. Consequently, robust evaluation frameworks must strategically pair these correlated semantic tasks with orthogonal visual tasks to thoroughly assess both high-level text decipherment and low-level visual robustness.

Table 3: Difficulty-conditioned results. \(\Delta\) denotes ‘Very Hard’ minus ‘Easy’.
Task Profile Easy Medium Hard Very hard \(\Delta\)
T1 Mean +8.8
T1 Qwen3.6-Plus-thinking +3.4
T2 Mean
T2 Qwen3.6-Plus-thinking

3pt

Figure 4: Case studies on semantic matching and contextual reasoning tasks across different difficulties.

4.4 Difficulty Analysis↩︎

Tab.  3 shows that character length affects different task families in opposite ways. T1 does not decrease with length, suggesting that longer inscriptions may expose more formulaic ordering cues. By contrast, T2 degrades as length increases. Specifically, the average score drops by 4.6 points from easy to very hard, and Qwen3.6-Plus-thinking drops by 16.4 points. This supports the interpretation that length is a pressure test for structured role assignment rather than a simple visual-complexity factor. Fig. 4 provides a qualitative comparison between Qwen3.6-plus-thinking, Gemini 3.1 Pro, and GPT-4o.

5 Limitations↩︎

Although we have conducted extensive exploration on S-OBI and obtained important observations, S-OBI has several limitations. First, the current benchmark is compact, with 95 base inscriptions, and is therefore intended as a diagnostic benchmark rather than a large-scale leaderboard. Second, the sentence-level images are reconstructed from aligned glyph images, which simplifies some material properties of original oracle bones. Third, multiple-choice subtasks may still contain answer-choice shortcuts, especially for formulaic order selection.

6 Conclusion↩︎

In this paper, we introduce S-OBI, a novel benchmark for evaluating MLLMs on sentence-level oracle bone inscription image understanding. Unlike single-character OBI resources, S-OBI targets evaluating the inscriptional-group understanding through global meaning matching, semantic slot extraction, punctuation restoration, and masked glyph completion. Experiments on ten mainstream MLLMs show that current systems remain unreliable in sentence-level OBI understanding, with the best overall performance being 33.5%. Subtask analysis further shows that MLLMs perform much better on semantic-unit order selection than on full meaning matching, suggesting reliance on coarse formulaic regularities rather than robust inscriptional reasoning. These findings position sentence-level ancient-script understanding as a distinct and challenging multimodal evaluation problem for language models.

References↩︎

[1]
Chen, Z., Hua, W., Li, J., Zhu, Y., Zhi, X., Liu, Z., Chen, T., Zhang, W., Zhai, G.: Oracle bone inscriptions information processing: A comprehensive survey. npj Heritage Science 14,  220 (2026).
[2]
Keightley, D.N. (ed.): Sources of Shang History: The Oracle-Bone Inscriptions of Bronze Age China. University of California Press, Berkeley (1978).
[3]
Li, J., Chen, Z., Chen, T., Liu, Z., Wang, C.: Obiformer: A fast attentive denoising framework for oracle bone inscriptions. Displays 89, 103059 (2025).
[4]
Huang, S., Wang, H., Liu, Y., Shi, X., Jin, L.: OBC306: A large-scale oracle bone character recognition dataset. In: Proceedings of the International Conference on Document Analysis and Recognition. pp. 681–688 (2019).
[5]
Wang, M., Deng, W.: A dataset of oracle characters for benchmarking machine learning algorithms. Scientific Data 11,  87 (2024).
[6]
Wang, P., Zhang, K., Wang, X., Han, S., Liu, Y., Wan, J., Guan, H., Kuang, Z., Jin, L., Bai, X., Liu, Y.: An open dataset for oracle bone script recognition and decipherment. Scientific Data 11,  976 (2024).
[7]
Chen, Z., Chen, T., Zhang, W., Zhai, G.: OBI-Bench: Can LMMs aid in the study of ancient script on oracle bones? In: Proceedings of the International Conference on Learning Representations (2025).
[8]
Zhang, J., Li, R., Pang, H., Xia, D., Zhu, Z., Zhang, Q., Li, C., Yang, X.: Specializing large models for oracle bone script interpretation via component-grounded multimodal knowledge augmentation. In: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 35221–35236 (2026).
[9]
Zhou, Z., Shi, D., Shi, L., Song, R., Qiu, P., Diao, X., Xu, H.: Acse: An ancient character semantic-aware embedding for large language models. In: Findings of the Association for Computational Linguistics: ACL 2026. pp. 9000–9012 (2026).
[10]
Li, Z., Chen, Z., Yang, Z., Liu, Z., Zhai, G., Chen, T.: Roots: Recognizing oracle bone inscriptions via an organized tree structure (2026).
[11]
Li, J., Chen, Z., Jiang, R., Chen, T., Wang, C., Zhai, G.: Mitigating long-tail distribution in oracle bone inscriptions: Dataset, model, and benchmark. In: Proceedings of the 33rd ACM International Conference on Multimedia. pp. 7729–7738 (2025).
[12]
Guan, H., Wan, J., Liu, Y., Wang, P., Zhang, K., Kuang, Z., Wang, X., Bai, X., Jin, L.: An open dataset for the evolution of oracle bone characters: EVOBC. arXiv preprint arXiv:2401.12467 (2024).
[13]
Jiang, H., Pan, Y., Chen, J., Liu, Z., Zhou, Y., Shu, P., Li, Y., Zhao, H., Mihm, S., Howe, L.C., Liu, T.: OracleSage: Towards unified visual-linguistic understanding of oracle bone scripts through cross-modal knowledge fusion. arXiv preprint arXiv:2411.17837 (2024).
[14]
Chen, Z., Hua, W., Li, J., Deng, L., Du, F., Chen, T., Zhai, G.: PictOBI-20k: Unveiling large multimodal models in visual decipherment for pictographic oracle bone characters. arXiv preprint arXiv:2509.05773 (2025).
[15]
Li, J., Chi, X., Wang, Q., Huang, K., Wang, D.H., Liu, Y., Liu, C.L.: A comprehensive survey of oracle character recognition: Challenges, datasets, methodology, and beyond. Pattern Recognition 169, 111824 (2026).
[16]
Liu, M., Liu, G., Liu, Y., Jiao, Q.: Oracle bone inscriptions recognition based on deep convolutional neural network. Journal of image and graphics 8(4), 114–119 (2020).
[17]
Yue, X., Li, H., Fujikawa, Y., Meng, L.: Dynamic dataset augmentation for deep learning-based oracle bone inscriptions recognition. ACM Journal on Computing and Cultural Heritage 15(4), 1–20 (2022).
[18]
Fujikawa, Y., Li, H., Yue, X., Aravinda, C., Prabhu, G.A., Meng, L.: Recognition of oracle bone inscriptions by using two deep learning models. International Journal of Digital Humanities 5(2), 65–79 (2023).
[19]
Chen, Z., Tian, Y., Sun, Y., Sun, W., Zhang, Z., Lin, W., Zhai, G., Zhang, W.: Just noticeable difference for large multimodal models. arXiv preprint arXiv:2507.00490 (2025).
[20]
Li, C., Ding, Z., Hu, X., Li, B., Luo, D., Wu, A., Wang, C., Wang, C., Jin, T., Shu, S., et al.: Oraclefusion: Assisting the decipherment of oracle bone script with structurally constrained semantic typography. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 19893–19902 (2025).
[21]
Liu, Y., Guan, H., Wang, P., Wang, X., Wan, J., Zhang, K., Zheng, H., Liu, X., Kuang, Z., Yang, H., et al.: Alphaoracle: Oracle bone script decipherment via human-workflow-inspired deep learning. The Innovation (2026).
[22]
Li, C., Ding, Z., Hu, X., Li, B., Luo, D., Peng, X., Jin, T., Liu, Y., Han, S., Yang, J., et al.: Oracleagent: A multimodal reasoning agent for oracle bone script research. arXiv preprint arXiv:2510.26114 (2025).
[23]
Liu, Z., Zhang, D., Takashima, K.i., Zang, K., et al.: A Classified Collection of Oracle Bone Inscriptions with Modern Chinese and English Translations. Guangxi Education Press (11 2005).
[24]
Huang, S., Price, B., Fan, Y., Morse, B.: Beyond single images: A comprehensive benchmark for album-level vision-language understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 38564–38573 (2026).
[25]
OpenAI: Introducing GPT-5.2. https://openai.com/index/introducing-gpt-5-2/(2026), accessed 22 June 2026.
[26]
OpenAI: GPT-4o system card. arXiv preprint arXiv:2410.21276 (2024).
[27]
Google: Gemini model documentation. https://ai.google.dev/gemini-api/docs/models(2026), accessed 22 June 2026.
[28]
Qwen Team: Qwen3.6 model documentation. https://qwen.ai/models(2026), accessed 22 June 2026.
[29]
DeepSeek-AI: DeepSeek-V4-Pro model card. https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro(2026), accessed 22 June 2026.
[30]
Wang, W., Gao, Z., Gu, L., Pu, H., Cui, L., Wei, X., Liu, Z., Jing, L., Ye, S., Shao, J., et al.: InternVL3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265 (2025).
[31]
Qwen Team: Qwen3.5-9B model card. https://huggingface.co/Qwen/Qwen3.5-9B(2026), accessed 22 June 2026.
[32]
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., et al.: Qwen2.5-VL technical report. arXiv preprint arXiv:2502.13923 (2025).