Text-to-Text Transfer Transformer Model for SNLI Indonesia

Penulis: Afriyanti, Iis; Yulianti, Evi
Informasi
Jurnal2025 10th International Conference on Informatics and Computing, ICIC 2025
PenerbitInstitute of Electrical and Electronics Engineers Inc.
Halaman -
Tahun Publikasi2025
ISBN979-833157583-0
Jenis SumberScopus
Abstrak
This study investigates the performance of various pre-trained language models on the Natural Language Inference (NLI) task in the Indonesian language using the translated SNLI dataset (SNLI Indonesia), espcecially T5 model. We evaluate IndoBERT + BiLSTM, mT5, XLM-R, and Indo-T5 across multiple dataset sizes (Sk to 7Sk). Indo-T5 consistently achieves the highest accuracy, reaching 80.97% on the 75k dataset, slightly outperforming XLM-R. IndoBERT + BiLSTM and mT5 were excluded from larger datasets due to lower performance. A qualitative analysis reveals that both Indo-T5 and XLM-R struggle with understanding paraphrasing and implicit inference, often misclassifying entailments and contradictions as neutral. Furthermore, translation errors in the dataset introduced semantic mismatches that adversely affected the predictions. The findings highlight the need for improved data quality and suggest that a hybrid ensemble of Indo-TS and XLM-R may further enhance NLI performance in Bahasa Indonesia © 2025 IEEE.
Dokumen & Tautan

© 2025 Universitas Indonesia. Seluruh hak cipta dilindungi.