Ekstraksi informasi leksikal bahasa Kerinci dari kamus dwibahasa untuk klasifikasi Part-of-Speech otomatis

Authors

  • Agung Kharisma Hidayah Universitas Muhammadiyah Bengkulu
  • Dandi Sunardi Universitas Muhammadiyah Bengkulu
  • Sri Handayani Universitas Muhammadiyah Bengkulu

DOI:

https://doi.org/10.24246/aiti.v23i2.334-347

Keywords:

Kerinci language, Part-of-Speech, POS classification, bilingual dictionary, rule-based NLP

Abstract

This study aims to develop initial linguistic resources for the Kerinci language by extracting lexical information from an Indonesian–Kerinci bilingual dictionary. The dictionary data, originally in digital document format (.docx), was converted into a structured dataset suitable for computational processing. Part-of-Speech (POS) classification was performed automatically using a rule-based approach, referring to the Indonesian equivalents that had been POS-tagged using reference lexicons. The results show that verbs and adjectives dominate the dictionary entries, differing from the common trend in bilingual dictionaries which are usually dominated by nouns. This approach proves effective in low-resource conditions and serves as a preliminary step toward the development of Kerinci NLP. The resulting POS-classified dataset has great potential for training POS taggers, supporting linguistic documentation, and contributing to language preservation through technology

Downloads

Download data is not yet available.

Metrics

Metrics Loading ...

References

L. Y. M. Wirajayadi, N. Suryanirmala, A. Winata, and Z. Haeri, “Cerminan Budaya Dalam Bahasa Daerah: Sebagai Penanda Identitas Diri Masyarakat Sasak,” J. Innov. Res. Knowl., vol. 1, no. 3, pp. 367–372, 2021.

A. Y. Chandra, D. Kurniawan, and R. Musa, “Perancangan Chatbot Menggunakan Dialogflow Natural Language Processing (Studi Kasus: Sistem Pemesanan pada Coffee Shop),” J. Media Inform. Budidarma, vol. 4, no. 1, p. 208, 2020, doi: 10.30865/mib.v4i1.1505.

N. D. Kusumawardani, Y. T. Mursityo, and R. I. Rokhmawati, “Evaluasi Critical Success Factors Pada Implementasi Sistem Informasi Supply Chain Management (ALISTA) Menggunakan Metode Dematel Pada PT. Telkom Akses Malang,” J. Pengemb. Teknol. Inf. dan Ilmu Komput., vol. 3, no. 5, pp. 4316–4326, 2019, [Online]. Available: http://j-ptiik.ub.ac.id

M. Kurniawan, K. Kusrini, and M. R. Arief, “Part of Speech Tagging Pada Teks Bahasa Indonesia dengan BiLSTM + CNN + CRF dan ELMo,” J. Eksplora Inform., vol. 11, no. 1, pp. 29–37, 2022, doi: 10.30864/eksplora.v11i1.506.

D. Winingsih, “Ekstraksi Informasi Metadata Statistik pada Artikel Penelitian Ilmiah menggunakan Algoritma Machine Learning,” Tesis Magister Teknik Elektro Institut Teknologi Bandung, 2023.

J. Awwalu, S. E.-Y. Abdullahi, and A. E. Evwiekpaefe, “Parts of Speech Tagging: a Review of Techniques,” Fudma J. Sci., vol. 4, no. 2, pp. 712–721, 2020, doi: 10.33003/fjs-2020-0402-325.

K. K. Purnamasari and I. S. Suwardi, “Rule-based Part of Speech Tagger for Indonesian Language,” IOP Conf. Ser. Mater. Sci. Eng., vol. 407, no. 1, 2018, doi: 10.1088/1757-899X/407/1/012151.

F. Ramadhanti, Y. Wibisono, and R. A. Sukamto, “Analisis Morfologi untuk Menangani Out-of-Vocabulary Words pada Part-of-Speech Tagger Bahasa Indonesia Menggunakan Hidden Markov Model,” J. Linguist. Komputasional, vol. 2, no. 1, p. 6, 2019, doi: 10.26418/jlk.v2i1.13.

A. Muhammad and N. Widyastuti, “Pengembangan Aplikasi Part-of-Speech Tagger Bahasa Banjar Menggunakan Metode Pengembangan DevOps,” J. Ilm. Ilmu Komput. dan Teknol. Inf., vol. 1, no. 1, 2024.

I. P. P. Ananda, M. A. Bijaksana, and I. Asror, “Pembangunan Synsets untuk WordNet Bahasa Indonesia dengan Metode Komutatif,” e-Proceeding Eng., vol. 5, no. 3, pp. 7845–7866, 2018, [Online]. Available: https://core.ac.uk/download/pdf/299925449.pdf

A. F. Aji et al., “One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia,” Proc. Annu. Meet. Assoc. Comput. Linguist., vol. 1, pp. 7226–7249, 2022, doi: 10.18653/v1/2022.acl-long.500.

M. Alfian, U. L. Yuhana, D. Siahaan, H. Munazharoh, and E. Pardede, “LFF-POS: A linguistic fusion method to handle out-of-vocabulary words in low-resource part-of-speech tagging,” Proc. 2018 Int. Conf. Asian Lang. Process. IALP 2018, vol. 15, pp. 303–307, 2018, doi: 10.1109/IALP.2018.8629236.

Abigail Rai, “Part-of-Speech (POS) Tagging of Low-Resource Language (Limbu) with Deep learning,” Panam. Math. J., vol. 35, no. 1s, pp. 149–157, 2024, doi: 10.52783/pmj.v35.i1s.2297.

A. Mulyanto, Y. A. Nurhuda, and N. Wiyanto, “Penyelesaian Kata Ambigu Pada Proses POS Tagging Menggunakan Algoritma Hidden Markov Model ( HMM ),” Pros. Semin. Nas. Metod. Kuantitatif, no. 978, pp. 347–358, 2017, [Online]. Available: https://jurnal.fmipa.unila.ac.id/snmk/article/view/2102/1540

M. P. Prapasha and H. Thamrin, “Implementasi POS Tagging dan Algoritma ANTLR Parser dalam Memeriksa Struktur Kalimat Bahasa Indonesia,” Publ. Univ. Muhammadiyah Surakarta, vol. 16, no. 2, pp. 39–55, 2024.

A. Himawan, M. R. Maarif, and U. S. Aesyi, “Analisis Hashtag pada Twitter untuk Eksplorasi Pokok Bahasan Terkini Mengenai Business Intelligence,” JISKA (Jurnal Inform. Sunan Kalijaga), vol. 6, no. 2, pp. 106–112, 2021, doi: 10.14421/jiska.2021.6.2.106-112.

W. Hermanto, B. Irawan, and C. Setianingsih, “Klasifikasi Emosi pada Lirik Lagu Menggunakan Algoritma Support Vector Machine dan Optimasi Particle Swarm Optimization,” e-Proceeding Eng., vol. 8, no. 5, pp. 6307–6327, 2021, [Online]. Available: https://lirik.kapanlagi.com/

H. Herpindo, A. Wijayanti, I. Shalima, and R. Ngestrini, “Kategori, fungsi, dan peran sintaksis bahasa Indonesia dengan PoS Tagging berbasis rule dan probability,” KEMBARA J. Sci. Lang. Lit. Teach., vol. 8, no. 1, pp. 51–65, 2022, doi: 10.22219/kembara.v8i1.18602.

I. D. Id, Machine Learning: Teori, Studi Kasus dan Implementasi Menggunakan Python. Pekan Baru: Universitas Riau Press, 2021.

H. Mohamed, N. Omar, and M. J. Ab. Aziz, “Malay Part of Speech Tagger: A Comparative Study on Tagging Tools,” Asia-Pacific J. Inf. Technol. Multimed., vol. 04, no. 01, pp. 11–23, 2015, doi: 10.17576/apjitm-2015-0401-02.

M. Ter Hoeve, D. Grangier, and N. Schluter, “High-Resource Methodological Bias in Low-Resource Investigations,” ArXiv, vol. abs/2211.07534, p., 2022, [Online]. Available: https://consensus.app/papers/highresource-methodological-bias-in-lowresource-hoeve-grangier/d1313c3c202b5621abea76a6fbb95aec/

Published

2026-06-12

How to Cite

[1]
A. K. Hidayah, D. Sunardi, and S. Handayani, “Ekstraksi informasi leksikal bahasa Kerinci dari kamus dwibahasa untuk klasifikasi Part-of-Speech otomatis”, AITI, vol. 23, no. 2, pp. 334–347, Jun. 2026.

Issue

Section

Articles