iTANONG-DS: A Collection of Benchmark Datasets for Downstream Natural Language Processing Tasks on Select Philippine Languages
Aunhel John Adoptante
Project Technical Specialist I
Computer Software Division
DOST-ASTI
Benchmark datasets are crucial for evaluating algorithms and models objectively. They provide a standardized basis for comparisons, promote reproducibility, and drive innovation by establishing baselines and encouraging advancements in the field. Limited benchmark datasets exist for various natural language processing tasks in low-resource languages, including most Philippine languages. As part of iTANONG’s 10 billion token dataset initiative, the authors release the first iteration of iTANONG-DS1, a collection of unlabeled and labeled datasets for different NLP tasks such as sentiment analysis, part-of-speech tagging, named entity recognition for Tagalog, and language modeling for Cebuano.
As part of the National Electrical, Electronics and Computer Engineering Conference (NEECECON 2024), this technical session is organized by the UP Electrical and Electronics Engineering Institute with the theme "National Development through Sustainable Industrialization."
NEECECON 2024 is co-located with the Advanced Science, Technology, and Innovation Convention (ASTICON) 2024, held from 18 to 19 July 2024 at the Novotel Manila Araneta City in Quezon City.
ASTICON 2024 showcased DOST-ASTI and UP EEEI's pioneering contributions to the ICT landscape while celebrating the partnerships that drive technological advancement and societal progress in the country.
For more info about the event, visit https://neececon2024.eee.upd.edu.ph.