Model Selection for Homophone Spelling Correction in Low-Resource Languages

A Systematic Review with Khmer Case Study

Authors

  • Seanghort Born LIUM, Le Mans University, France , Institute of Digital Research and Innovation, Cambodia Academy of Digital Technology, Cambodia Author
  • Madeth May LIUM, Le Mans University, France Author
  • Claudine Piau-Toffolon LIUM, Le Mans University, France Author
  • Sebastien Iksal LIUM, Le Mans University, France Author

DOI:

https://doi.org/10.21467/proceedings.8.1.2

Keywords:

spelling correction, model selection, low-resource languages

Abstract

This study conducts a comprehensive review of existing literature on spelling correction models designed for low-resource languages (LRLs). The research examines how these models handle homophone errors, using Khmer as a case study to highlight challenges in LRLs. Out of 174 academic publications, 54 papers were chosen for detailed study. The review covers various spelling correction models, including traditional, deep learning, and large language models. Traditional methods are affordable to implement but struggle to understand word context properly. Deep learning models (DLMs) provide the best balance between cost and effectiveness for correcting Khmer homophones. Large language models (LLMs) offer the highest accuracy but require significant computational power and may favor languages with more available data. The study also provides a three-stage framework for selecting appropriate models. When data is limited, traditional methods work best for homophone correction. When moderate amounts of data are available, DLMs should be the preferred choice. When both sufficient data and computational resources are accessible, LLMs deliver optimal results. These recommendations offer valuable guidance for researchers who develop spelling correction systems for LRLs, particularly when addressing homophone-related challenges.

References

Allamong, M. B., Jeong, J., & Kellstedt, P. M. (2025). Spelling correction with large language models to reduce measurement error in open-ended survey responses. Research & Politics, 12(1), 20531680241311510. https://doi.org/10.1177/20531680241311510

Attia, M., Pecina, P., Samih, Y., Shaalan, K., & Van Genabith, J. (2016). Arabic spelling error detection and correction. Natural Language Engineering, 22(5), 751–773. https://doi.org/10.1017/S1351324915000030

Born, S., May, M., Piau-Toffolon, C., & Iksal, S. (2024). A survey on importance of homophones spelling correction model for Khmer authors. https://doi.org/10.48550/arXiv.2411.10477

Born, S., Valy, D., & Kong, P. (2022). Encoder-decoder language model for Khmer handwritten text recognition in historical documents. 2022 14th International Conference on Software, Knowledge, Information Management and Applications (SKIMA), 234–238. https://doi.org/10.1109/SKIMA57145.2022.10029532

Buoy, R., Iwamura, M., Srun, S., & Kise, K. (2023). Toward a low-resource non-Latin-complete baseline: An exploration of Khmer optical character recognition. IEEE Access, 11, 128044–128060. https://doi.org/10.1109/ACCESS.2023.3332361

Büyük, O., & Arslan, L. M. (2021). Learning from mistakes: Improving spelling correction performance with automatic generation of realistic misspellings. Expert Systems, 38(5), e12692. https://doi.org/10.1111/exsy.12692

Cissé, T. I., & Sadat, F. (2023). Automatic spell checker and correction for under-represented spoken languages: Case study on Wolof. https://doi.org/10.48550/arXiv.2305.12694

Do, D.-T., Nguyen, H. T., Bui, T. N., & Vo, D. H. (2021). VSEC: Transformer-based model for Vietnamese spelling correction. PRICAI 2021: Trends in Artificial Intelligence, 259–272. https://doi.org/10.1007/978-3-030-89363-7_20

Dong, R., Yang, Y., & Jiang, T. (2019). Spelling correction of non-word errors in Uyghur–Chinese machine translation. Information, 10(6), 202. https://doi.org/10.3390/info10060202

Eger, S., vor der Brück, T., & Mehler, A. (2016). A comparison of four character-level string-to-string translation models for (OCR) spelling error correction. Prague Bulletin of Mathematical Linguistics. https://doi.org/10.1515/pralin-2016-0004

Eskander, R., Klavans, J., & Muresan, S. (2019). Unsupervised morphological segmentation for low-resource polysynthetic languages. Proceedings of the 16th Workshop on Computational Research in Phonetics, Phonology, and Morphology, 189–195. https://doi.org/10.18653/v1/W19-4222

Etoori, P., Chinnakotla, M., & Mamidi, R. (2018). Automatic spelling correction for resource-scarce languages using deep learning. Proceedings of ACL 2018, Student Research Workshop, 146–152. https://doi.org/10.18653/v1/P18-3021

Fahrudin, T. M., Sa’diyah, I., Latipah, L., Atha Illah, I. Z., Bey Lirna, C. C., & Acarya, B. S. (2021). KEBI 1.0: Indonesian spelling error detection system for scientific papers using dictionary lookup and Peter Norvig spelling corrector. Lontar Komputer: Jurnal Ilmiah Teknologi Informasi, 12(2), 78. https://doi.org/10.24843/LKJITI.2021.v12.i02.p02

Goslin, K., & Hofmann, M. (2022). English language spelling correction as an information retrieval task using Wikipedia search statistics. Proceedings of the Thirteenth Language Resources and Evaluation Conference (LREC 2022), 458–464. https://aclanthology.org/2022.lrec-1.48/

Gueddah, H., & Lachibi, Y. (2023). Arabic spellchecking: A depth-filtered composition metric to achieve fully automatic correction. International Journal of Electrical and Computer Engineering, 13(5), 5366–5373. https://doi.org/10.11591/ijece.v13i5.pp5366-5373

Gueddah, H., Nejja, M., Iazzi, S., Yousfi, A., & Aouragh, S. L. (2022). Improving spellchecking: An effective ad-hoc probabilistic lexical measure for general typos. Indonesian Journal of Electrical Engineering and Computer Science, 27(1), 521–527. https://doi.org/10.11591/ijeecs.v27.i1.pp521-527

Hasan, S., Heger, C., & Mansour, S. (2015). Spelling correction of user search queries through statistical machine translation. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, 451–460. https://doi.org/10.18653/v1/D15-1051

He, Z., Zhu, Y., Wang, L., & Xu, L. (2023). UMRSpell: Unifying the detection and correction parts of pre-trained models towards Chinese missing, redundant, and spelling correction. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10238–10250. https://doi.org/10.18653/v1/2023.acl-long.570

Hladek, D., Staš, J., & Pleva, M. (2020). Survey of automatic spelling correction. Electronics, 9, 1670. https://doi.org/10.3390/electronics9101670

Hu, Y., Jing, X., Ko, Y., & Rayz, J. T. (2020). Misspelling correction with pre-trained contextual language model. 2020 IEEE 19th International Conference on Cognitive Informatics & Cognitive Computing (ICCI*CC), 144–149. https://doi.org/10.1109/ICCICC50026.2020.9450253

Jiang, L., Shen, X., Zhao, Q., & Yao, J. (2024). MLSL-spell: Chinese spelling check based on multi-label annotation. Applied Sciences, 14(6), 2541. https://doi.org/10.3390/app14062541

Kasmaiee, S., Kasmaiee, S., & Homayounpour, M. (2023). Correcting spelling mistakes in Persian texts with rules and deep learning methods. Scientific Reports, 13(1), 19945. https://doi.org/10.1038/s41598-023-47295-2

Kim, J., Weiss, J. C., & Ravikumar, P. (2022). Context-sensitive spelling correction of clinical text via conditional independence. Conference on Health, Inference, and Learning, 234–247. https://proceedings.mlr.press/v174/kim22b.html

Kusuma, A. T. A., & Ratnasari, C. I. (2023). Comparison of spell correction in Bahasa Indonesia: Peter Norvig, LSTM, and n-gram. Jurnal Informatika dan Komputer, 6(3), 214–220. https://doi.org/10.33387/jiko.v6i3.7072

Kuznetsov, A., & Urdiales, H. (2021). Spelling correction with denoising transformer. https://doi.org/10.48550/arXiv.2105.05977

Lai, K. H., Topaz, M., Goss, F. R., & Zhou, L. (2015). Automated misspelling detection and correction in clinical free-text records. Journal of Biomedical Informatics, 55, 188–195. https://doi.org/10.1016/j.jbi.2015.04.008

Lee, J.-H., Kim, M., & Kwon, H.-C. (2017). Improved statistical language model for context-sensitive spelling error candidates. Journal of Korea Multimedia Society, 20(2), 371–381.

Lee, J.-H., Kim, M., & Kwon, H.-C. (2020). Deep learning-based context-sensitive spelling typing error correction. IEEE Access, 8, 152565–152578.

Lertpiya, A., Chalothorn, T., & Chuangsuwanich, E. (2020). Thai spelling correction and word normalization on social text using a two-stage pipeline with neural contextual attention. IEEE Access, 8, 133403–133419. https://doi.org/10.1109/ACCESS.2020.3010828

Liang, Z., Quan, X., & Wang, Q. (2023). Disentangled phonetic representation for Chinese spelling correction. https://doi.org/10.48550/arXiv.2305.14783

Lin, Y., Zhang, Z., Hu, M., Sun, Y., & Zhang, Y. (2024). Modalities should be appropriately leveraged: Uncertainty guidance for multimodal Chinese spelling correction. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), 11463–11474.

Liu, L., Wu, H., & Zhao, H. (2024). Chinese spelling correction as rephrasing language model. Proceedings of the AAAI Conference on Artificial Intelligence. https://doi.org/10.48550/arXiv.2308.08796

Liu, S., Yang, T., Yue, T., Zhang, F., & Wang, D. (2021). PLOME: Pre-training with misspelled knowledge for Chinese spelling correction. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2991–3000. https://doi.org/10.18653/v1/2021.acl-long.233

Liu, Z., Richardson, C., Hatcher, R., & Prud’hommeaux, E. (2022). Not always about you: Prioritizing community needs when developing endangered language technology. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3933–3944. https://doi.org/10.18653/v1/2022.acl-long.272

Magueresse, A., Carles, V., & Heetderks, E. (2020). Low-resource languages: A review of past work and future challenges. https://doi.org/10.48550/arXiv.2006.07264

Mainsah, B. O., Morton, K. D., Collins, L. M., Sellers, E. W., & Throckmorton, C. S. (2015). Moving away from error-related potentials to achieve spelling correction in P300 spellers. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 23(5), 737–743. https://doi.org/10.1109/TNSRE.2014.2374471

Mammadov, S. (2019). Neural spelling correction for Azerbaijani language. Proceedings of the 2019 IEEE 13th International Conference on Application of Information and Communication Technologies (AICT), 1–5. https://doi.org/10.1109/AICT47866.2019.8981776

Mao, M., Peng, S., Yang, Y., & Park, D.-S. (2022). Bi-directional maximal matching algorithm to segment Khmer words in sentence. Journal of Information Processing Systems, 18(4), 549–561. https://doi.org/10.3745/JIPS.04.0250

Mashod Rana, M., Tipu Sultan, M., Mridha, M. F., Eyaseen Arafat Khan, M., Masud Ahmed, M., & Abdul Hamid, M. (2018). Detection and correction of real-word errors in Bangla language. Proceedings of the 2018 International Conference on Bangla Speech and Language Processing (ICBSLP), 1–4. https://doi.org/10.1109/ICBSLP.2018.8554502

Moslem, Y., Haque, R., & Way, A. (2020). Arabisc: Context-sensitive neural spelling checker. Proceedings of the 6th Workshop on Natural Language Processing Techniques for Educational Applications, 11–19. https://doi.org/10.18653/v1/2020.nlptea-1.2

Mukazhanov, N., Alibiyeva, Z., Yerimbetova, A., Kassymova, A., & Alibiyeva, N. (2023). Development of an augmented Damerau–Levenshtein method for correcting spelling errors in Kazakh texts. Eastern-European Journal of Enterprise Technologies, 5(2), 23–33. https://doi.org/10.15587/1729-4061.2023.289187

Phung, H. T., & Luong, N. V. (2024). Detecting spelling errors in Vietnamese administrative document using large language models. Ho Chi Minh City Open University Journal of Science – Engineering and Technology, 14(1), 31–40. https://doi.org/10.46223/HCMCOUJS.tech.en.14.1.3141.2024

Salhab, M., & Abu-Khzam, F. (2024). AraSpell: A deep learning approach for Arabic spelling correction. https://doi.org/10.48550/arXiv.2405.06981

San, M., Cai, Z., Cai, R., & Dao, J. (2021). Analysis on types of spelling errors in true Tibetan characters. MATEC Web of Conferences, 336, 06019. https://doi.org/10.1051/matecconf/202133606019

Shah, K., & De Melo, G. (2020). Correcting the autocorrect: Context-aware typographical error correction via training data augmentation. https://doi.org/10.48550/arXiv.2005.01158

Shanahan, M. (2024). Talking about large language models. Communications of the ACM, 67(2), 68–79. https://doi.org/10.1145/3624724

Sun, R., Wu, X., & Wu, Y. (2023). An error-guided correction model for Chinese spelling error correction. https://doi.org/10.48550/arXiv.2301.06323

Sung, T., & Hwang, I. (2016). Ternary decomposition and dictionary extension for Khmer word segmentation. Journal of Information Technology Applications and Management, 23(2), 11–28. https://doi.org/10.21219/JITAM.2016.23.2.011

Tran, H., Dinh, C. V., Phan, L., & Nguyen, S. T. (2021). Hierarchical transformer encoders for Vietnamese spelling correction. International Conference on Industrial, Engineering and Other Applications of Applied Intelligent Systems, 547–556. https://doi.org/10.48550/arXiv.2105.13578

Unicode Consortium. (2025). The Unicode standard, version 16.0. Retrieved July 14, 2025, from https://www.unicode.org/charts/PDF/U1780.pdf

Van Nam, T., Thi Hue, N., & Huy Khanh, P. (2017). Building a syllable database to solve the problem of Khmer word segmentation. International Journal on Natural Language Computing, 6(1), 1–12. https://doi.org/10.5121/ijnlc.2017.6101

Zhang, R., Pang, C., Zhang, C., Wang, S., He, Z., Sun, Y., Wu, H., & Wang, H. (2021). Correcting Chinese spelling errors with phonetic pre-training. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 2250–2261. https://doi.org/10.18653/v1/2021.findings-acl.198

Zhang, S., Huang, H., Liu, J., & Li, H. (2020). Spelling error correction with soft-masked BERT. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 882–890. https://doi.org/10.18653/v1/2020.acl-main.82

Zhou, Y., Porwal, U., & Konow, R. (2017). Spelling correction as a foreign language. https://doi.org/10.48550/arXiv.1705.07371

Downloads

Published

2026-01-20

How to Cite

[1]
S. Born, M. May, C. Piau-Toffolon, and S. Iksal, “Model Selection for Homophone Spelling Correction in Low-Resource Languages: A Systematic Review with Khmer Case Study”, AIJR Proc., vol. 8, no. 1, pp. 6–17, Jan. 2026, doi: 10.21467/proceedings.8.1.2.