Penerapan IndoBERT untuk Pencarian Semantik Tafsir Al-Qur'an Berdasarkan Tafsir Al-Misbah

Authors

  • Taslim Sultan UIN Alauddin Makassar
  • Faisal Akib Universitas Islam Negeri Alauddin Makassar
  • Muhammad Hasrul Hasanuddin Universitas Islam Negeri Alauddin Makassar

Keywords:

FAISS, IndoBERT, Semantic Search, Tafsir Al-Misbah, Website

Abstract

Conventional Qur'anic exegesis retrieval systems that rely on keyword matching often return less relevant results because they are unable to capture the semantic meaning and context of users' queries. This study aims to develop a semantic search system for Tafsir Al-Misbah using the IndoBERT model within the CRISP-DM framework. The dataset consists of 6,236 Indonesian-language tafsir summaries processed through data cleaning, case folding, normalization, and tokenization. The IndoBERT model was then fine-tuned to generate semantic representations of user queries and document contents, while FAISS was employed to accelerate retrieval using cosine similarity. System performance was evaluated using Mean Average Precision (mAP), Mean Reciprocal Rank (MRR), and Normalized Discounted Cumulative Gain (nDCG). The evaluation results achieved Top-10 scores of 0.2128 for mAP, 0.2681 for MRR, and 0.2792 for nDCG, while the Top-50 scores reached 0.2301, 0.2803, and 0.3481, respectively. These results indicate adequate retrieval performance in the domain of Qur'anic exegesis and demonstrate that the proposed semantic search approach based on IndoBERT and FAISS is capable of providing contextually relevant retrieval results for users' queries.

Downloads

Download data is not yet available.

References

[1] I. F. Hasanah, U. Hasanah, M. Makhzuniyah, dan A. N. K. B. Zawawi, “Qur’ Anic Learning In The Digital Era : A Study On Digital Applications And Their Impact,” J. PAMATOR, vol. 18, no. 1, hal. 146–159, 2025, doi: https://doi.org/10.21107/pamator.v18i1.28172 Manuscript.

[2] M. Huzaifa, B. Aqil, M. A. Haq, dan W. Zaghouani, Arabic natural language processing for Qur ’ anic research : a systematic review, vol. 56, no. 7. Springer Netherlands, 2023. doi: 10.1007/s10462-022-10313-2.

[3] B. V. Kartika, M. J. Alfredo, dan G. P. Kusuma, “Fine-Tuned IndoBERT Based Model and Data Augmentation for Indonesian Language Paraphrase Identification,” Int. Inf. Eng. Technol. Assoc., vol. 37, no. 3, hal. 733–743, 2023, doi: https://doi.org/10.18280/ria.370322 Received:

[4] L. Trisnawati dkk., “A proposed semantic keywords search engine for Indonesian Qur ’ an translation based on word embedding,” Indones. J. Electr. Eng. Comput. Sci., vol. 35, no. 2, hal. 987–995, 2024, doi: 10.11591/ijeecs.v35.i2.pp987-995.

[5] A. Malik, A. P. Gefadri, E. Sidik, dan A. P. Syadrina, “SoulScripture : Chatbot using Bidirectional Encoder Representations from Transformers as a Medium of Spiritual Guidance,” Khazanah J. Relig. Technol., vol. 2, no. 1, hal. 23–27, 2024, doi: https://doi.org/10.15575/kjrt.v2i1.822 SoulScripture:

[6] A. P. Putra, D. Purnami, S. Putri, A. A. Kt, dan A. Cahyawan, “Scientific Paper Recommendation System : Application of Sentence Transformers and Cosine Similarity Using arXiv Data,” J. Appl. Informatics Comput., vol. 9, no. 4, hal. 1374–1382, 2025, [Daring]. Tersedia pada: http://jurnal.polibatam.ac.id/index.php/JAIC

[7] P. Chapman dkk., “CRISP-DM 1.0: Step-by-Step Data Mining Guide,” 2000.

[8] K. Kowsari, K. J. Meimandi, M. Heidarysafa, dan S. Mendu, “Text Classification Algorithms : A Survey,” MDPI Inf., vol. 10, no. 1, hal. 1–68, 2020, doi: 10.3390/info10040150.

[9] M. F. Ikhsan dan S. ‘Uyun, “Analisis Perbandingan Metode Information Retrieval ( IR ) Pada Implementasi Pencarian Ayat Al-Qur ’ an,” J. SHIFT, vol. 6, no. 1, 2026, doi: https://doi.org/10.24252/shift.v6i1.272.

[10] I. Alpha dan E. Yulianti, “Academic expert finding using BERT pre-trained language model,” Int. J. Adv. Intell. Informatics, vol. 10, no. 2, hal. 280–295, 2024, doi: https://doi.org/10.26555/ijain.v10i2.1497.

[11] M. Pan, J. Wang, J. X. Huang, A. J. Huang, Q. Chen, dan J. Chen, “A probabilistic framework for integrating sentence-level semantics via BERT into pseudo-relevance feedback,” Inf. Process. Manag., vol. 59, no. 1, hal. 102734, 2022, doi: https://doi.org/10.1016/j.ipm.2021.102734.

[12] H. Vo, Q. Nguyen, dan D. T. Tran, “Building an Information Retrieval System for University Documents Based on Generative AI Technologies,” VNUHCM J. Econ. Bus. Law, vol. 10, no. 2, hal. 6508–6514, 2026, doi: https://doi.org/10.32508/vnuhcmjebl.v10i2.1597.

[13] M. Alqarni, “Embedding Search for Quranic Texts based on Large Language Models,” Int. Arab J. Inf. Technol., vol. 21, no. 2, hal. 243–256, 2024, doi: https://doi.org/10.34028/iajit/21/2/7 1.

[14] N. E. Vijayan, “Enhancing Chatbot Response Relevance through Semantic Similarity Measures,” J. Artif. Intell. Cloud Comput., vol. 1, no. 1, hal. 1–5, 2022, doi: doi.org/10.47363/JAICC/2022(1)E182.

[15] T. Formal, B. Piwowarski, C. Lassance, dan S. Clinchant, “SPLADE v2 : Sparse Lexical and Expansion Model for Information Retrieval,” Proc. ACM Conf., vol. 1, no. 1, hal. 1–6, 2021, doi: https://doi.org/10.1145/nnnnnnn.nnnnnnn.

[16] M. Kamil dan D. Cakir, “Advances in Transformer-Based Semantic Search : Techniques , Benchmarks , and Future Directions,” Turk. J. Math. Comput. Sci., vol. 17, no. 1, hal. 145–166, 2025, doi: 10.47000/tjmcs.1633092.

Published

2026-08-18

How to Cite

Sultan, T., Akib , F., & Hasanuddin, M. H. (2026). Penerapan IndoBERT untuk Pencarian Semantik Tafsir Al-Qur’an Berdasarkan Tafsir Al-Misbah. System Information and Computer Technology (SYNCTECH), 2(2), 97–109. Retrieved from https://librarium.id/index.php/synctech/article/view/61

Issue

Section

Articles