TY - GEN
T1 - Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia
AU - Cahyawijaya, Samuel
AU - Lovenia, Holy
AU - Moniz, Joel Ruben Antony
AU - Wong, Tack Hwa
AU - Farhansyah, Mohammad Rifqi
AU - Maung, Thant Thiri
AU - Hudi, Frederikus
AU - Anugraha, David
AU - Habibi, Muhammad Ravi Shulthan
AU - Qorib, Muhammad Reza
AU - Agarwal, Amit
AU - Imperial, Joseph Marvin
AU - Patel, Hitesh Laxmichand
AU - Feliren, Vicky
AU - Nasution, Bahrul Ilmi
AU - Rufino, Manuel Antonio
AU - Winata, Genta Indra
AU - Rajagede, Rian Adam
AU - Catalan, Carlos Rafael
AU - Imam, Mohamed Fazli
AU - Pattnayak, Priyaranjan
AU - Pranida, Salsabila Zahirah
AU - Pratama, Kevin
AU - Bangera, Yeshil
AU - Na-Thalang, Adisai
AU - Monderin, Patricia Nicole
AU - Song, Yueqi
AU - Simon, Christian
AU - Ng, Lynnette Hui Xian
AU - Sapan, Richardy Lobo
AU - Rafi, Taki Hasan
AU - Wang, Bin
AU - Supryadi,
AU - Veerakanjana, Kanyakorn
AU - Ittichaiwong, Piyalitt
AU - Roque, Matthew Theodore
AU - Vincentio, Karissa
AU - Kreangphet, Takdanai
AU - Artkaew, Phakphum
AU - Palgunadi, Kadek Hendrawan
AU - Yu, Yanzhi
AU - Hastuti, Rochana Prih
AU - Nixon, William
AU - Bangera, Mithil
AU - Lim, Adrian Xuan Wei
AU - Khine, Aye Hninn
AU - Zhafran, Hanif Muhammad
AU - Ferdinan, Teddy
AU - Izzani, Audra Aurora
AU - Singh, Ayushman
AU - Evan,
AU - Krito, Jauza Akbar
AU - Anugraha, Michael
AU - Ilasariya, Fenal Ashokbhai
AU - Li, Haochen
AU - Daniswara, John Amadeo
AU - Tjiaranata, Filbert Aurelian
AU - Yulianrifat, Eryawan Presma
AU - Udomcharoenchaikit, Can
AU - Ansori, Fadil Risdian
AU - Ihsani, Mahardika Krisna
AU - Nguyen, Giang
AU - Barik, Anab Maulana
AU - Velasco, Dan John
AU - Genadi, Rifo Ahmad
AU - Saha, Saptarshi
AU - Wei, Chengwei
AU - Flores, Isaiah
AU - Chen, Kenneth Ko Han
AU - Santos, Anjela Gail
AU - Lim, Wan Shen
AU - Phyo, Kaung Si
AU - Santos, Tim
AU - Dwiastuti, Meisyarah
AU - Luo, Jiayun
AU - Cruz, Jan Christian Blaise
AU - Hee, Ming Shan
AU - Hanif, Ikhlasul Akmal
AU - Alif Al Hakim, M.
AU - Sya'ban, Muhammad Rizky
AU - Kerdthaisong, Kun
AU - Miranda, Lester James V.
AU - Koto, Fajri
AU - Fatyanosa, Tirana Noor
AU - Aji, Alham Fikri
AU - Rosal, Jostin Jerico
AU - Kevin, Jun
AU - Wijaya, Robert
AU - Kampman, Onno P.
AU - Zhang, Ruochen
AU - Karlsson, Börje F.
AU - Limkonchotiwat, Peerat
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Despite Southeast Asia's (SEA) extraordinary linguistic and cultural diversity, the region remains significantly underrepresented in vision-language (VL) research, resulting in AI models that inadequately capture SEA cultural nuances. To fill this gap, we present SEA-VL, an open-source initiative dedicated to developing culturally relevant high-quality datasets for SEA languages. By involving contributors from SEA countries, SEA-VL ensures better cultural relevance and diversity, fostering greater inclusivity of underrepresented languages and cultural depictions in VL research. Our methodology employed three approaches: community-driven crowdsourcing with SEA contributors, automated image crawling, and synthetic image generation. We evaluated each method's effectiveness in capturing cultural relevance. We found that image crawling achieves approximately ∼85% cultural relevance while being more cost- and time-efficient than crowdsourcing, whereas synthetic image generation failed to accurately reflect SEA cultural nuances and contexts. Collectively, we gathered 1.28 million SEA culturally relevant images, more than 50 times larger than other existing datasets. This work bridges the representation gap in SEA, establishes a foundation for developing culturally aware AI systems for this region, and provides a replicable framework for addressing representation gaps in other underrepresented regions.
AB - Despite Southeast Asia's (SEA) extraordinary linguistic and cultural diversity, the region remains significantly underrepresented in vision-language (VL) research, resulting in AI models that inadequately capture SEA cultural nuances. To fill this gap, we present SEA-VL, an open-source initiative dedicated to developing culturally relevant high-quality datasets for SEA languages. By involving contributors from SEA countries, SEA-VL ensures better cultural relevance and diversity, fostering greater inclusivity of underrepresented languages and cultural depictions in VL research. Our methodology employed three approaches: community-driven crowdsourcing with SEA contributors, automated image crawling, and synthetic image generation. We evaluated each method's effectiveness in capturing cultural relevance. We found that image crawling achieves approximately ∼85% cultural relevance while being more cost- and time-efficient than crowdsourcing, whereas synthetic image generation failed to accurately reflect SEA cultural nuances and contexts. Collectively, we gathered 1.28 million SEA culturally relevant images, more than 50 times larger than other existing datasets. This work bridges the representation gap in SEA, establishes a foundation for developing culturally aware AI systems for this region, and provides a replicable framework for addressing representation gaps in other underrepresented regions.
UR - https://www.scopus.com/pages/publications/105021028710
M3 - Conference contribution
AN - SCOPUS:105021028710
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 18685
EP - 18717
BT - Long Papers
A2 - Che, Wanxiang
A2 - Nabende, Joyce
A2 - Shutova, Ekaterina
A2 - Pilehvar, Mohammad Taher
PB - Association for Computational Linguistics (ACL)
T2 - 63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
Y2 - 27 July 2025 through 1 August 2025
ER -