Model Selection for Homophone Spelling Correction in Low-Resource Languages
A Systematic Review with Khmer Case Study
DOI:
https://doi.org/10.21467/proceedings.8.1.2Keywords:
spelling correction, model selection, low-resource languagesAbstract
This study conducts a comprehensive review of existing literature on spelling correction models designed for low-resource languages (LRLs). The research examines how these models handle homophone errors, using Khmer as a case study to highlight challenges in LRLs. Out of 174 academic publications, 54 papers were chosen for detailed study. The review covers various spelling correction models, including traditional, deep learning, and large language models. Traditional methods are affordable to implement but struggle to understand word context properly. Deep learning models (DLMs) provide the best balance between cost and effectiveness for correcting Khmer homophones. Large language models (LLMs) offer the highest accuracy but require significant computational power and may favor languages with more available data. The study also provides a three-stage framework for selecting appropriate models. When data is limited, traditional methods work best for homophone correction. When moderate amounts of data are available, DLMs should be the preferred choice. When both sufficient data and computational resources are accessible, LLMs deliver optimal results. These recommendations offer valuable guidance for researchers who develop spelling correction systems for LRLs, particularly when addressing homophone-related challenges.
References
Allamong, M. B., Jeong, J., & Kellstedt, P. M. (2025). Spelling correction with large language models to reduce measurement error in open-ended survey responses. Research & Politics, 12(1), 20531680241311510. https://doi.org/10.1177/20531680241311510
Attia, M., Pecina, P., Samih, Y., Shaalan, K., & Van Genabith, J. (2016). Arabic spelling error detection and correction. Natural Language Engineering, 22(5), 751–773. https://doi.org/10.1017/S1351324915000030
Born, S., May, M., Piau-Toffolon, C., & Iksal, S. (2024). A survey on importance of homophones spelling correction model for Khmer authors. https://doi.org/10.48550/arXiv.2411.10477
Born, S., Valy, D., & Kong, P. (2022). Encoder-decoder language model for Khmer handwritten text recognition in historical documents. 2022 14th International Conference on Software, Knowledge, Information Management and Applications (SKIMA), 234–238. https://doi.org/10.1109/SKIMA57145.2022.10029532
Buoy, R., Iwamura, M., Srun, S., & Kise, K. (2023). Toward a low-resource non-Latin-complete baseline: An exploration of Khmer optical character recognition. IEEE Access, 11, 128044–128060. https://doi.org/10.1109/ACCESS.2023.3332361
Büyük, O., & Arslan, L. M. (2021). Learning from mistakes: Improving spelling correction performance with automatic generation of realistic misspellings. Expert Systems, 38(5), e12692. https://doi.org/10.1111/exsy.12692
Cissé, T. I., & Sadat, F. (2023). Automatic spell checker and correction for under-represented spoken languages: Case study on Wolof. https://doi.org/10.48550/arXiv.2305.12694
Do, D.-T., Nguyen, H. T., Bui, T. N., & Vo, D. H. (2021). VSEC: Transformer-based model for Vietnamese spelling correction. PRICAI 2021: Trends in Artificial Intelligence, 259–272. https://doi.org/10.1007/978-3-030-89363-7_20
Dong, R., Yang, Y., & Jiang, T. (2019). Spelling correction of non-word errors in Uyghur–Chinese machine translation. Information, 10(6), 202. https://doi.org/10.3390/info10060202
Eger, S., vor der Brück, T., & Mehler, A. (2016). A comparison of four character-level string-to-string translation models for (OCR) spelling error correction. Prague Bulletin of Mathematical Linguistics. https://doi.org/10.1515/pralin-2016-0004
Eskander, R., Klavans, J., & Muresan, S. (2019). Unsupervised morphological segmentation for low-resource polysynthetic languages. Proceedings of the 16th Workshop on Computational Research in Phonetics, Phonology, and Morphology, 189–195. https://doi.org/10.18653/v1/W19-4222
Etoori, P., Chinnakotla, M., & Mamidi, R. (2018). Automatic spelling correction for resource-scarce languages using deep learning. Proceedings of ACL 2018, Student Research Workshop, 146–152. https://doi.org/10.18653/v1/P18-3021
Fahrudin, T. M., Sa’diyah, I., Latipah, L., Atha Illah, I. Z., Bey Lirna, C. C., & Acarya, B. S. (2021). KEBI 1.0: Indonesian spelling error detection system for scientific papers using dictionary lookup and Peter Norvig spelling corrector. Lontar Komputer: Jurnal Ilmiah Teknologi Informasi, 12(2), 78. https://doi.org/10.24843/LKJITI.2021.v12.i02.p02
Goslin, K., & Hofmann, M. (2022). English language spelling correction as an information retrieval task using Wikipedia search statistics. Proceedings of the Thirteenth Language Resources and Evaluation Conference (LREC 2022), 458–464. https://aclanthology.org/2022.lrec-1.48/
Gueddah, H., & Lachibi, Y. (2023). Arabic spellchecking: A depth-filtered composition metric to achieve fully automatic correction. International Journal of Electrical and Computer Engineering, 13(5), 5366–5373. https://doi.org/10.11591/ijece.v13i5.pp5366-5373
Gueddah, H., Nejja, M., Iazzi, S., Yousfi, A., & Aouragh, S. L. (2022). Improving spellchecking: An effective ad-hoc probabilistic lexical measure for general typos. Indonesian Journal of Electrical Engineering and Computer Science, 27(1), 521–527. https://doi.org/10.11591/ijeecs.v27.i1.pp521-527
Hasan, S., Heger, C., & Mansour, S. (2015). Spelling correction of user search queries through statistical machine translation. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, 451–460. https://doi.org/10.18653/v1/D15-1051
He, Z., Zhu, Y., Wang, L., & Xu, L. (2023). UMRSpell: Unifying the detection and correction parts of pre-trained models towards Chinese missing, redundant, and spelling correction. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10238–10250. https://doi.org/10.18653/v1/2023.acl-long.570
Hladek, D., Staš, J., & Pleva, M. (2020). Survey of automatic spelling correction. Electronics, 9, 1670. https://doi.org/10.3390/electronics9101670
Hu, Y., Jing, X., Ko, Y., & Rayz, J. T. (2020). Misspelling correction with pre-trained contextual language model. 2020 IEEE 19th International Conference on Cognitive Informatics & Cognitive Computing (ICCI*CC), 144–149. https://doi.org/10.1109/ICCICC50026.2020.9450253
Jiang, L., Shen, X., Zhao, Q., & Yao, J. (2024). MLSL-spell: Chinese spelling check based on multi-label annotation. Applied Sciences, 14(6), 2541. https://doi.org/10.3390/app14062541
Kasmaiee, S., Kasmaiee, S., & Homayounpour, M. (2023). Correcting spelling mistakes in Persian texts with rules and deep learning methods. Scientific Reports, 13(1), 19945. https://doi.org/10.1038/s41598-023-47295-2
Kim, J., Weiss, J. C., & Ravikumar, P. (2022). Context-sensitive spelling correction of clinical text via conditional independence. Conference on Health, Inference, and Learning, 234–247. https://proceedings.mlr.press/v174/kim22b.html
Kusuma, A. T. A., & Ratnasari, C. I. (2023). Comparison of spell correction in Bahasa Indonesia: Peter Norvig, LSTM, and n-gram. Jurnal Informatika dan Komputer, 6(3), 214–220. https://doi.org/10.33387/jiko.v6i3.7072
Kuznetsov, A., & Urdiales, H. (2021). Spelling correction with denoising transformer. https://doi.org/10.48550/arXiv.2105.05977
Lai, K. H., Topaz, M., Goss, F. R., & Zhou, L. (2015). Automated misspelling detection and correction in clinical free-text records. Journal of Biomedical Informatics, 55, 188–195. https://doi.org/10.1016/j.jbi.2015.04.008
Lee, J.-H., Kim, M., & Kwon, H.-C. (2017). Improved statistical language model for context-sensitive spelling error candidates. Journal of Korea Multimedia Society, 20(2), 371–381.
Lee, J.-H., Kim, M., & Kwon, H.-C. (2020). Deep learning-based context-sensitive spelling typing error correction. IEEE Access, 8, 152565–152578.
Lertpiya, A., Chalothorn, T., & Chuangsuwanich, E. (2020). Thai spelling correction and word normalization on social text using a two-stage pipeline with neural contextual attention. IEEE Access, 8, 133403–133419. https://doi.org/10.1109/ACCESS.2020.3010828
Liang, Z., Quan, X., & Wang, Q. (2023). Disentangled phonetic representation for Chinese spelling correction. https://doi.org/10.48550/arXiv.2305.14783
Lin, Y., Zhang, Z., Hu, M., Sun, Y., & Zhang, Y. (2024). Modalities should be appropriately leveraged: Uncertainty guidance for multimodal Chinese spelling correction. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), 11463–11474.
Liu, L., Wu, H., & Zhao, H. (2024). Chinese spelling correction as rephrasing language model. Proceedings of the AAAI Conference on Artificial Intelligence. https://doi.org/10.48550/arXiv.2308.08796
Liu, S., Yang, T., Yue, T., Zhang, F., & Wang, D. (2021). PLOME: Pre-training with misspelled knowledge for Chinese spelling correction. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2991–3000. https://doi.org/10.18653/v1/2021.acl-long.233
Liu, Z., Richardson, C., Hatcher, R., & Prud’hommeaux, E. (2022). Not always about you: Prioritizing community needs when developing endangered language technology. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3933–3944. https://doi.org/10.18653/v1/2022.acl-long.272
Magueresse, A., Carles, V., & Heetderks, E. (2020). Low-resource languages: A review of past work and future challenges. https://doi.org/10.48550/arXiv.2006.07264
Mainsah, B. O., Morton, K. D., Collins, L. M., Sellers, E. W., & Throckmorton, C. S. (2015). Moving away from error-related potentials to achieve spelling correction in P300 spellers. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 23(5), 737–743. https://doi.org/10.1109/TNSRE.2014.2374471
Mammadov, S. (2019). Neural spelling correction for Azerbaijani language. Proceedings of the 2019 IEEE 13th International Conference on Application of Information and Communication Technologies (AICT), 1–5. https://doi.org/10.1109/AICT47866.2019.8981776
Mao, M., Peng, S., Yang, Y., & Park, D.-S. (2022). Bi-directional maximal matching algorithm to segment Khmer words in sentence. Journal of Information Processing Systems, 18(4), 549–561. https://doi.org/10.3745/JIPS.04.0250
Mashod Rana, M., Tipu Sultan, M., Mridha, M. F., Eyaseen Arafat Khan, M., Masud Ahmed, M., & Abdul Hamid, M. (2018). Detection and correction of real-word errors in Bangla language. Proceedings of the 2018 International Conference on Bangla Speech and Language Processing (ICBSLP), 1–4. https://doi.org/10.1109/ICBSLP.2018.8554502
Moslem, Y., Haque, R., & Way, A. (2020). Arabisc: Context-sensitive neural spelling checker. Proceedings of the 6th Workshop on Natural Language Processing Techniques for Educational Applications, 11–19. https://doi.org/10.18653/v1/2020.nlptea-1.2
Mukazhanov, N., Alibiyeva, Z., Yerimbetova, A., Kassymova, A., & Alibiyeva, N. (2023). Development of an augmented Damerau–Levenshtein method for correcting spelling errors in Kazakh texts. Eastern-European Journal of Enterprise Technologies, 5(2), 23–33. https://doi.org/10.15587/1729-4061.2023.289187
Phung, H. T., & Luong, N. V. (2024). Detecting spelling errors in Vietnamese administrative document using large language models. Ho Chi Minh City Open University Journal of Science – Engineering and Technology, 14(1), 31–40. https://doi.org/10.46223/HCMCOUJS.tech.en.14.1.3141.2024
Salhab, M., & Abu-Khzam, F. (2024). AraSpell: A deep learning approach for Arabic spelling correction. https://doi.org/10.48550/arXiv.2405.06981
San, M., Cai, Z., Cai, R., & Dao, J. (2021). Analysis on types of spelling errors in true Tibetan characters. MATEC Web of Conferences, 336, 06019. https://doi.org/10.1051/matecconf/202133606019
Shah, K., & De Melo, G. (2020). Correcting the autocorrect: Context-aware typographical error correction via training data augmentation. https://doi.org/10.48550/arXiv.2005.01158
Shanahan, M. (2024). Talking about large language models. Communications of the ACM, 67(2), 68–79. https://doi.org/10.1145/3624724
Sun, R., Wu, X., & Wu, Y. (2023). An error-guided correction model for Chinese spelling error correction. https://doi.org/10.48550/arXiv.2301.06323
Sung, T., & Hwang, I. (2016). Ternary decomposition and dictionary extension for Khmer word segmentation. Journal of Information Technology Applications and Management, 23(2), 11–28. https://doi.org/10.21219/JITAM.2016.23.2.011
Tran, H., Dinh, C. V., Phan, L., & Nguyen, S. T. (2021). Hierarchical transformer encoders for Vietnamese spelling correction. International Conference on Industrial, Engineering and Other Applications of Applied Intelligent Systems, 547–556. https://doi.org/10.48550/arXiv.2105.13578
Unicode Consortium. (2025). The Unicode standard, version 16.0. Retrieved July 14, 2025, from https://www.unicode.org/charts/PDF/U1780.pdf
Van Nam, T., Thi Hue, N., & Huy Khanh, P. (2017). Building a syllable database to solve the problem of Khmer word segmentation. International Journal on Natural Language Computing, 6(1), 1–12. https://doi.org/10.5121/ijnlc.2017.6101
Zhang, R., Pang, C., Zhang, C., Wang, S., He, Z., Sun, Y., Wu, H., & Wang, H. (2021). Correcting Chinese spelling errors with phonetic pre-training. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 2250–2261. https://doi.org/10.18653/v1/2021.findings-acl.198
Zhang, S., Huang, H., Liu, J., & Li, H. (2020). Spelling error correction with soft-masked BERT. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 882–890. https://doi.org/10.18653/v1/2020.acl-main.82
Zhou, Y., Porwal, U., & Konow, R. (2017). Spelling correction as a foreign language. https://doi.org/10.48550/arXiv.1705.07371
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.