Pengaruh Augmentasi Teks pada Klasifikasi Berita Menggunakan BERT Embedding dan LSTM
DOI:
https://doi.org/10.24843/Keywords:
Back translation, BERT Embedding, Grid search, Long Short Term Memory, News ClassificationAbstract
According to the 2023 Digital News Report, 84% of Indonesians access or consume news online, reflecting how central digital platforms have become to public information consumption. Furthermore, 2024 data from the Indonesian Press Council recorded 1,015 online news platforms operating in Indonesia. With such a diverse range of providers, it is estimated that thousands or even millions of news articles are published annually, making consistent and efficient content organization increasingly important. News category labels serve as a universal variable that helps both online news platform providers and the general public navigate this volume of information more easily. This study therefore performs automatic news topic classification by combining BERT embedding with Long Short-Term Memory (LSTM), incorporating back-translation as a data augmentation technique to mitigate overfitting. Grid search hyperparameter tuning was employed to determine the optimal configuration, resulting in a dropout rate of 0.2 and a learning rate of 0.0001 for both the non-augmented and 30%-augmented models. The non-augmented model achieved 96% accuracy, while the 30%-augmented model achieved 98% accuracy, establishing it as the best-performing approach in this study.