BamiBERT: A New BERT-based Language Model for Vietnamese

Dat Quoc Nguyen\(^1\), Thinh Pham\(^2\), Chi Tran\(^1\), Linh The Nguyen\(^1\)
\(^1\)Qualcomm AI Research
1
\(^2\)Virginia Tech
{datnq, chitran, linhnt}@qti.qualcomm.com, thinhphp@vt.edu


Abstract

In this paper, we introduce BamiBERT, a new BERT-based pre-trained language model for Vietnamese that addresses key limitations of PhoBERT—the current de facto Vietnamese text encoder. Trained from scratch on a 129GB corpus of general-domain Vietnamese text for 20 epochs, BamiBERT supports an extended context length of up to 2048 tokens and operates directly on raw input, eliminating the need for external word segmentation. Across 8 Vietnamese benchmarks, it achieves the best score on 11 of 15 metrics and the second-best on 3 others, setting a new state of the art among "base"-sized Vietnamese encoders and demonstrating strong cross-domain generalization.

1 Introduction↩︎

In today’s LLM-driven era, BERT-based models [1], [2] remain essential for tasks that demand high precision and low latency, such as span labeling, classification, and information retrieval. Their lightweight nature makes them particularly well-suited for resource-constrained applications. Rather than competing with large language models (LLMs), BERT-based models often serve as core components in hybrid systems, delivering strong performance at a fraction of the computational cost while effectively complementing LLMs [3]. As a result, the BERT family continues to evolve, with recent additions such as ModernBERT [4] and NeoBERT [5]. While English benefits from a rich ecosystem of pre-trained BERT-based models, the development of Vietnamese counterparts remains comparatively limited. The multilingual model XLM-RoBERTa [6] achieves competitive performance on a wide range of Vietnamese NLP tasks by leveraging 138GB of CC100 Vietnamese text in its multilingual pre-training corpus. PhoBERT [7] was the first large-scale monolingual BERT model pre-trained from scratch specifically for Vietnamese, using 20GB of text. Subsequent general-domain monolingual models include viBERT and vELECTRA [8], pre-trained on 60GB of Vietnamese text, and ViDeBERTa [9], which uses the same 138GB CC100 Vietnamese corpus as XLM-RoBERTa. More recently, CafeBERT [10] was continually pre-trained from the XLM-RoBERTa “large” model on 18GB of Vietnamese text collected prior to 2021. Several domain-specific monolingual models have also emerged: VnLawBERT [11] for legal text, ViHealthBERT [12] and ViPubmedDeBERTa [13] for health and biomedical text, and ViSoBERT [14] for social media text.

Among these monolingual models, PhoBERT has become the default choice for many Vietnamese NLP tasks thanks to its strong and consistent performance. Since its release, it has gained widespread adoption, with over 200K monthly downloads on HuggingFace and active use across the NLP community, whereas all other Vietnamese monolingual models receive fewer than 5K monthly downloads each. Despite its popularity, PhoBERT has several limitations: it supports only a short maximum context length of 256 subword tokens, and it requires input text to be word-segmented by an external tool prior to processing. These limitations motivate the development of a new Vietnamese BERT-based model that supports a longer context and operates directly on raw text.

In this paper, we introduce BamiBERT—a new pre-trained language model for Vietnamese—trained from scratch on a large corpus of 129GB of uncompressed text for 20 epochs, with an extended maximum context length of 2048 tokens. Unlike PhoBERT, which requires Vietnamese text to be pre-segmented by an external word segmenter, BamiBERT operates directly on raw input text, making it more flexible and easier to integrate into a wider range of downstream applications. Experimental results on 8 Vietnamese benchmark datasets show that BamiBERT delivers state-of-the-art or near-state-of-the-art performance (ranking #1 on 11/15 metrics, #2 on 3/15 metrics, and #3 on the remaining one), demonstrating strong cross-domain generalization. We release BamiBERT at: https://huggingface.co/Qualcomm-AI-Research/BamiBERT.

2 Pre-trained language model BamiBERT↩︎

This section presents how we pre-train our new BERT-based language model from scratch.

2.0.0.1 Architecture:

We pre-train a text encoder, named "BamiBERT",2 from scratch, employing the BERT’s "base" architecture with 12 Transformer block layers [1]. To pre-train BamiBERT, we use the masked language modeling objective [1] and the RoBERTa pre-training approach [2] which optimizes BERT with a dynamic masking strategy and without the next sentence prediction objective. For tokenization, we extend the PhoGPT’s Vietnamese-specific byte-level BPE tokenizer [15] with an additional "<mask>" token, resulting in a final vocabulary of 20481 token types. We set a maximum sequence length of 2048.

2.0.0.2 Pre-training dataset:

We use a clean, 129 GB dataset of uncompressed, general-domain text.

2.0.0.3 Optimization:

The model is optimized using Adam [16]. We use a batch size of 1024 sequence blocks distributed across 8 A100 GPUs (each with 40GB of memory) and a peak learning rate of 0.00015. The pre-training process runs for 20 epochs, with the initial 2 epochs dedicated to warming up the learning rate.

3 Experiments↩︎

3.1 Setup↩︎

Table 1: Statistics of 8 experimental datasets.
Dataset Train. Valid. Test
ViNLI 24,376 3,009 2,991
PhoNER_COVID19 5,027 2,000 3,000
UIT-VSFC (Sentiment) 11,426 1,583 3,166
UIT-VSFC (Topic) 11,426 1,583 3,166
ViSpamReviews 14,306 1,590 3,974
UIT-ViSFD 7,786 1,112 2,224
UIT-ABSA (Hotel) 7,180 795 2,030
UIT-ABSA (Restaurant) 7,028 771 1,938
Table 2: Results of pre-trained "base"-architecture models. \(^\dagger\) denotes results extracted from previous works.
Dataset Metric BamiBERT ViDeBERTa ViSoBERT XLM-RoBERTa PhoBERT
ViNLI (4-label) Accuracy 81.01 61.08 67.70 76.83\(^\dagger\) 78.00\(^\dagger\)
F1 81.15 60.71 67.82 77.01\(^\dagger\) 78.05\(^\dagger\)
PhoNER_COVID19 F1 94.90 94.50\(^\dagger\) 92.90 92.50\(^\dagger\) 94.20\(^\dagger\)
UIT-VSFC (Sentiment) Accuracy 93.86 87.86 93.15 93.56 94.10
F1 83.41 70.74 81.49 82.20 83.27
UIT-VSFC (Topic) Accuracy 89.34 83.94 88.80 89.18 89.24
F1 79.90 66.49 79.86 79.56 79.80
ViSpamReviews Accuracy 90.76 86.21 90.99\(^\dagger\) 90.16\(^\dagger\) 89.83\(^\dagger\)
F1 78.20 67.04 79.06\(^\dagger\) 76.55\(^\dagger\) 76.18\(^\dagger\)
UIT-ViSFD F1 (Detection) 89.14 75.53 88.63 82.73 86.03
F1 (Classification) 84.24 64.20 83.55 65.81 78.76
UIT-ABSA (Hotel) F1 (Detection) 79.99 72.05 79.41 77.70\(^\dagger\) 79.16\(^\dagger\)
F1 (Classification) 72.65 62.97 74.24 71.23\(^\dagger\) 73.73\(^\dagger\)
UIT-ABSA (Restaurant) F1 (Detection) 88.01 73.56 86.86 82.18\(^\dagger\) 86.53\(^\dagger\)
F1 (Classification) 74.89 63.78 73.87 71.58\(^\dagger\) 73.52\(^\dagger\)

0.3em

We conduct experiments to compare our model BamiBERT with the previous strong and public pre-trained "base"-architecture ones for Vietnamese, including: Vietnamese-specific models ViDeBERTa-base [9], ViSoBERT [14] and PhoBERT-base [7] as well as the multilingual XLM-RoBERTa-base [6].3 Here, BamiBERT, ViSoBERT and XLM-RoBERTa take raw texts as input, while ViDeBERTa and PhoBERT are Vietnamese word-level models. That is, a Vietnamese word segmentation tool must be applied to produce word-segmented texts before feeding them to the word-level ViDeBERTa and PhoBERT. For ViDeBERTa and PhoBERT experiments, we utilize the RDRSegmenter component [17] from the VnCoreNLP toolkit [18] for Vietnamese word segmentation.

We employ the following experimental benchmark datasets: ViNLI—a Vietnamese dataset for open-domain natural language inference [19], PhoNER_COVID19—a dataset for recognizing COVID-19 related named entities in Vietnamese [20]; UIT-VSFC (Sentiment) and UIT-VSFC (Topic)—Vietnamese students’ feedback benchmarks for sentiment-based and topic-based classifications [21]; ViSpamReviews—a dataset for spam review detection on Vietnamese e-commerce websites [22]; UIT-ViSFD—a Vietnamese aspect-based sentiment analysis dataset of feedbacks and comments for smartphone e-commerce [23]; and UIT-ABSA (Hotel) and UIT-ABSA (Restaurant)—Vietnamese aspect-based sentiment analysis datasets for hotel and restaurant domains [24]. ViNLI and PhoNER_COVID19 are based on general-domain texts, whereas the remaining benchmarks are derived from social media and forum discussions. See Table 1 for the statistics of these datasets.

For all experimental models, we employ transformers [25] to fine-tune them using the AdamW optimizer [26] and set the batch size to 32. We also perform a grid search on the validation set to select the initial learning rate for AdamW from {1e-5, 2e-5, 5e-5}. We train for 30 epochs on the training set, compute F1 on the validation set after each training epoch, and select the model checkpoint with the best F1 to report final metric scores on the test set.

3.2 Main Results↩︎

Table 2 reports the performance of BamiBERT and four baselines—ViDeBERTa, ViSoBERT, XLM-RoBERTa, and PhoBERT—across eight Vietnamese benchmarks. BamiBERT achieves the best performance on 11 of the 15 evaluation metrics, ranks second on three metrics, and ranks third on the remaining metric, establishing a new state of the art for "base"-sized Vietnamese BERT-based language models.

3.2.0.1 ViNLI

BamiBERT achieves the best performance on both metrics (81.01 Accuracy and 81.15 F1), yielding substantial absolute gains of +3.01 Accuracy and +3.10 F1 over the second-ranked PhoBERT (78.00/78.05). XLM-RoBERTa follows in third place (76.83/77.01), trailing PhoBERT by roughly one point on both metrics. ViSoBERT (67.70/67.82) and ViDeBERTa (61.08/60.71) lag considerably behind, with gaps of 13–20 points relative to BamiBERT.

3.2.0.2 PhoNER_COVID19

BamiBERT obtains the highest F1 score (94.90), outperforming ViDeBERTa by 0.40 points and PhoBERT by 0.70 points. Its advantage widens against the remaining baselines, reaching +2.0 F1 over ViSoBERT (92.90) and +2.4 F1 over XLM-RoBERTa (92.50).

3.2.0.3 UIT–VSFC (Sentiment)

With 93.86 Accuracy and 83.41 F1, BamiBERT ranks among the top-performing models. It is essentially on par with PhoBERT, trailing by 0.24 Accuracy but leading by 0.14 F1. BamiBERT maintains consistent margins over XLM-RoBERTa (+0.30 Accuracy, +1.21 F1) and ViSoBERT (+0.71 Accuracy, +1.92 F1), while ViDeBERTa underperforms substantially on both metrics.

3.2.0.4 UIT–VSFC (Topic)

BamiBERT again leads on both metrics (89.34 Accuracy and 79.90 F1), outperforming PhoBERT by +0.10 on both, XLM-RoBERTa by +0.16/+0.34, and ViSoBERT by +0.54/+0.04. Although the top four models are tightly clustered (within 0.54 Accuracy and 0.34 F1), BamiBERT remains the most consistent. ViDeBERTa, in contrast, trails markedly (83.94/66.49).

3.2.0.5 ViSpamReviews

BamiBERT ranks second overall (90.76 Accuracy and 78.20 F1), narrowly behind ViSoBERT (90.99/79.06) by 0.23 Accuracy and 0.86 F1. Nevertheless, it clearly outperforms XLM-RoBERTa (+0.60 Accuracy, +1.65 F1) and PhoBERT (+0.93 Accuracy, +2.02 F1), while ViDeBERTa lags well behind (86.21/67.04). BamiBERT remains highly competitive at the top tier and continues to surpass other established baselines by clear margins.

3.2.0.6 UIT–ViSFD

BamiBERT delivers the strongest performance on both subtasks. For aspect detection, it attains 89.14 F1, ahead of ViSoBERT (88.63; \(-\)​0.51) and PhoBERT (86.03; \(-\)​3.11), with XLM-RoBERTa (82.73) and ViDeBERTa (75.53) trailing further. For aspect-based sentiment classification, BamiBERT achieves 84.24 F1, again surpassing ViSoBERT (83.55; \(-\)​0.69) and PhoBERT (78.76; \(-\)​5.48). The consistent gains over PhoBERT (+3.11 and +5.48 points) underscore BamiBERT’s robustness across both subtasks.

3.2.0.7 UIT–ABSA (Hotel)

On aspect detection, BamiBERT obtains the highest F1 (79.99), outperforming ViSoBERT (79.41; \(-\)​0.58) and PhoBERT (79.16; \(-\)​0.83), while XLM-RoBERTa (77.70) and ViDeBERTa (72.05) remain less competitive. On aspect-based sentiment classification, however, ViSoBERT takes the lead (74.24 F1), followed by PhoBERT (73.73; \(-\)​0.51) and BamiBERT (72.65; \(-\)​1.59); XLM-RoBERTa (71.23) and ViDeBERTa (62.97) trail by a wide margin.

3.2.0.8 UIT–ABSA (Restaurant)

BamiBERT produces the strongest end-to-end performance, with 88.01 F1 on aspect detection and 74.89 F1 on aspect-based sentiment classification. These results exceed those of ViSoBERT (86.86/\(-\)​1.15 and 73.87/\(-\)​1.02) and PhoBERT (86.53/\(-\)​1.48 and 73.52/\(-\)​1.37), while XLM-RoBERTa (82.18/71.58) and ViDeBERTa (73.56/63.78) lag considerably behind.

4 Discussion↩︎

Overall performance Across 8 Vietnamese benchmarks and different subtasks (Table 2), BamiBERT consistently delivers SOTA or near-SOTA results. The most substantial improvement appears on ViNLI, where BamiBERT surpasses the next-best PhoBERT by +3.01 Accuracy and +3.10 F1, while outpacing ViSoBERT/ViDeBERTa by 13–20 points—demonstrating strong sentence-pair semantics and cue-word sensitivity.

4.0.0.1 Domain effects

Note that PhoNER_COVID19 and ViNLI represent general-domain text, while the remaining benchmarks reflect social media and forum content. BamiBERT exhibits strong cross-domain generalization, outperforming or closely matching the social-media-focused ViSoBERT across both general-domain (e.g., ViNLI, PhoNER_COVID19) and social-domain benchmarks. Its consistent top-tier performance across diverse tasks—NER, span detection, and classification—demonstrates resilience to domain shift and label granularity. This robustness positions BamiBERT as a reliable choice for NLP pipelines operating under domain heterogeneity and distributional uncertainty.

4.0.0.2 Detection vs.classification

In aspect-based sentiment analysis pipelines, BamiBERT often excels in detection (e.g., Hotel: 79.99 F1; Restaurant: 88.01 F1), while its classification performance is strongest in the Restaurant domain (74.89 F1) but trails ViSoBERT/PhoBERT in the Hotel domain (72.65 F1). This pattern suggests complementary strengths: BamiBERT appears particularly effective at span/target localization and boundary-sensitive cues, whereas domain-specific sentiment nuances in the Hotel domain may benefit more from domain-adapted pretraining (ViSoBERT). Combined with the new SOTA results on UIT-ViSFD, this indicates BamiBERT’s robust end-to-end capability for fine-grained social content analysis.

4.0.0.3 Takeaway

BamiBERT delivers SOTA or near-SOTA performance across diverse Vietnamese benchmarks, excelling in NLI, topic classification, and fine-grained sentiment tasks. Its cross-domain stability and strong results on both detection and classification make it a robust default for real-world Vietnamese NLP applications.

5 Conclusion↩︎

In this paper, we have presented BamiBERT, a new pre-trained language model for Vietnamese designed to address key limitations of existing monolingual encoders. Unlike PhoBERT—the current de facto choice of Vietnamese text encoder—BamiBERT is trained from scratch on a 129 GB corpus of general-domain text for 20 epochs, supports an extended maximum context length of 2048 tokens, and operates directly on raw text, removing the dependency on external word segmentation. Experiments on 8 Vietnamese benchmark datasets show that BamiBERT does better than PhoBERT, establishing a new state of the art among “base”-sized Vietnamese encoders with strong cross-domain generalization.

References↩︎

[1]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. . In Proceedings of NAACL, pages 4171–4186.
[2]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. . arXiv preprint, arXiv:1907.11692.
[3]
Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024. . In Proceedings of KDD, page 6491–6501.
[4]
Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Griffin Thomas Adams, Jeremy Howard, and Iacopo Poli. 2025. . In Proceedings of the ACL, pages 2526–2547.
[5]
Lola Le Breton, Quentin Fournier, Mariam El Mezouar, John X. Morris, and Sarath Chandar. 2025. .
[6]
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019. . arXiv preprint, arXiv:1911.02116.
[7]
Dat Quoc Nguyen and Anh Tuan Nguyen. 2020. . In Findings of EMNLP 2020, pages 1037–1042.
[8]
The Viet Bui, Thi Oanh Tran, and Phuong Le-Hong. 2020. . In Proceedings of PACLIC, pages 13–20.
[9]
Cong Dao Tran, Nhut Huy Pham, Anh Tuan Nguyen, Truong Son Hy, and Tu Vu. 2023. . In Findings of EACL 2023, pages 1071–1078.
[10]
Phong Nguyen-Thuan Do, Son Quoc Tran, Phu Gia Hoang, Kiet Van Nguyen, and Ngan Luu-Thuy Nguyen. 2024. . In Findings of NAACL, pages 211–222.
[11]
Chieu-Nguyen Chau, Truong-Son Nguyen, and Le-Minh Nguyen. 2020. . In Proceedings of NICS, pages 298–301.
[12]
Nguyen Minh, Vu Hoang Tran, Vu Hoang, Huy Duc Ta, Trung Huu Bui, and Steven Quoc Hung Truong. 2022. . In Proceedings of LREC, pages 328–337.
[13]
Manh Tran-Tien, Huu-Loi Le, Dang Nhat Minh, T. Tran Khang, Huy-The Vu, and Nguyen Minh-Tien. 2023. . In Proceedings of PACLIC, pages 831–840.
[14]
Nam Nguyen, Thang Phan, Duc-Vu Nguyen, and Kiet Nguyen. 2023. . In Proceedings of EMNLP, pages 5191–5207.
[15]
Dat Quoc Nguyen, Linh The Nguyen, Chi Tran, Dung Ngoc Nguyen, Dinh Phung, and Hung Bui. 2023. . arXiv preprint, arXiv:2311.02945.
[16]
Diederik P. Kingma and Jimmy Ba. 2015. . In Proceedings of ICLR.
[17]
Dat Quoc Nguyen, Dai Quoc Nguyen, Thanh Vu, Mark Dras, and Mark Johnson. 2018. . In Proceedings of LREC 2018, pages 2582–2587.
[18]
Thanh Vu, Dat Quoc Nguyen, Dai Quoc Nguyen, Mark Dras, and Mark Johnson. 2018. nCoreNLP: A Vietnamese natural language processing toolkit. In Proceedings of NAACL: Demonstrations, pages 56–60.
[19]
Tin Van Huynh, Kiet Van Nguyen, and Ngan Luu-Thuy Nguyen. 2022. . In Proceedings of COLING, pages 3858–3872.
[20]
Thinh Hung Truong, Mai Hoang Dao, and Dat Quoc Nguyen. 2021. . In Proceedings of NAACL.
[21]
Kiet Van Nguyen, Vu Duc Nguyen, Phu X. V. Nguyen, Tham T. H. Truong, and Ngan Luu-Thuy Nguyen. 2018. . In Proceedings of KSE, pages 19–24.
[22]
Co Van Dinh, Son T. Luu, and Anh Gia-Tuan Nguyen. 2022. . In Proceedings of ACIIDS, pages 595–607.
[23]
Luong Luc Phan, Phuc Huynh Pham, Kim Thi-Thanh Nguyen, Sieu Khai Huynh, Tham Thi Nguyen, Luan Thanh Nguyen, Tin Van Huynh, and Kiet Van Nguyen. 2021. . In Proceedings of KSEM, pages 647–658.
[24]
Dang Van Thin, Ngan Luu-Thuy Nguyen, Tri Minh Truong, Lac Si Le, and Duy Tin Vo. 2021. . ACM Trans. Asian Low-Resour. Lang. Inf. Process., 20(4).
[25]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. . In Proceedings of EMNLP: System Demonstrations, pages 38–45.
[26]
Ilya Loshchilov and Frank Hutter. 2019. https://openreview.net/forum?id=Bkg6RiCqY7. In Proceedings of ICLR.

  1. Qualcomm Vietnam Company Limited. Qualcomm AI Research is an initiative of Qualcomm Technologies, Inc. This work was completed while all authors were at Movian AI, Vietnam. All datasets and models were downloaded, trained, and evaluated using Movian AI’s resources.↩︎

  2. "Bami" denotes "bánh mì" which is a popular type of sandwich in Vietnam.↩︎

  3. ViDeBERTa, ViSoBERT and XLM-RoBERTa were trained using a maximum sequence length of 512 tokens.↩︎