SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages
Penulis:Â Lovenia, Holy;Â Mahendra, Rahmad;Â Akbar, Salsabil Maulana;Â Miranda, Lester James V.;Â Santoso, Jennifer
Informasi
JurnalEMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
PenerbitAssociation for Computational Linguistics (ACL)
Halaman5155 - 5203
Tahun Publikasi2024
ISBN979-889176164-3
Jenis SumberScopus
Sitasi
Scopus: 10
Abstrak
Southeast Asia (SEA) is a region characterized by rich linguistic diversity and cultural variety, with over 1,300 indigenous languages and a population of 671 million people. However, the performance of contemporary AI models for SEA languages is compromised by a significant lack of representation of texts, images, and auditory datasets from SEA. Evaluating models for SEA languages is challenging due to the scarcity of high-quality datasets, compounded by the predominance of English training data, which raises concerns regarding potential cultural misrepresentation. To address these challenges, we introduce SEACrowd, a collaborative initiative that consolidates a comprehensive resource hub to bridge the resource gap by providing standardized corpora and benchmarks in nearly 1,000 SEA languages across three modalities. We assess the performance of AI models on 36 indigenous languages across 13 tasks included in SEACrowd, offering valuable insights into the current AI landscape in SEA. Furthermore, we propose strategies to facilitate greater AI advancements, maximizing potential utility and resource equity for the future of AI in Southeast Asia. © 2024 Association for Computational Linguistics.
Dokumen & Tautan
