MUHAMMAD RAMDANY JULIANSYAH, . (2026) Analisis Komparatif Algoritma RNN LSTM dan BILSTM Pada Teks Pendek Berbahasa Indonesia. Sarjana thesis, UNIVERSITAS NEGERI JAKARTA.
|
Text
File Cover.pdf Download (894kB) |
|
|
Text
File BAB 1.pdf Download (406kB) |
|
|
Text
File BAB 2.pdf Restricted to Registered users only Download (742kB) | Request a copy |
|
|
Text
File BAB 3.pdf Restricted to Registered users only Download (516kB) | Request a copy |
|
|
Text
File BAB 4.pdf Restricted to Registered users only Download (1MB) | Request a copy |
|
|
Text
File BAB 5.pdf Restricted to Registered users only Download (389kB) | Request a copy |
|
|
Text
File Daftar Pustaka.pdf Download (367kB) |
|
|
Text
File Lampiran.pdf Restricted to Registered users only Download (3MB) | Request a copy |
Abstract
Teks pendek di media sosial memiliki karakteristik keterbatasan konteks dan penggunaan bahasa tidak baku yang menyulitkan proses klasifikasi sentimen. Hingga saat ini, belum ada penelitian yang secara komprehensif membandingkan performa algoritma Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), dan Bidirectional LSTM (BiLSTM) khusus pada teks pendek berbahasa Indonesia, sehingga tidak tersedia acuan empiris untuk pemilihan model yang optimal. Oleh karena itu, penelitian ini bertujuan untuk menganalisis hasil dan membandingkan performa ketiga model tersebut dalam mengklasifikasikan sentimen terkait judi online pada tahun 2025. Penelitian ini menggunakan data berupa tweet berbahasa Indonesia yang dikumpulkan dari platform X menggunakan kata kunci "judol" atau "judi online". Data tersebut diproses melalui tahapan text preprocessing dan pelabelan berbasis leksikon (InSet) yang menghasilkan 11.664 data bersih, lalu direpresentasikan ke dalam bentuk vektor numerik menggunakan algoritma Word2Vec dengan arsitektur Skip-gram. Evaluasi performa model diukur berdasarkan tingkat akurasi, F1-score, serta analisis efisiensi komputasi yang mencakup waktu pelatihan, utilisasi CPU, dan RAM. Hasil penelitian menunjukkan bahwa model BiLSTM mencapai performa klasifikasi tertinggi dengan akurasi 68,80% dan rata-rata makro F1-Score sebesar 0,68, diikuti secara ketat oleh model LSTM dengan akurasi 67,55%, sedangkan model RNN mengalami kegagalan (underfitting) dengan akurasi 55,98%. Dari segi efisiensi komputasi, terdapat perbandingan terbalik (trade-off) yang sangat signifikan. Meskipun BiLSTM memperoleh akurasi tertinggi, peningkatan performanya tergolong tidak signifikan karena hanya terpaut sekitar 1,25% dibandingkan model LSTM. Di sisi lain, biaya komputasi yang harus dibayar oleh BiLSTM jauh lebih besar, yaitu waktu pelatihan terlama (± 910 detik berbanding ± 330 detik pada LSTM) dan konsumsi memori RAM tertinggi (1.000 - 1.100 MB berbanding 750 - 800 MB pada LSTM). Kesimpulannya, meskipun BiLSTM mampu memproses rentang semantik dari dua arah dengan akurasi terbaik, model LSTM terbukti jauh lebih efisien dan optimal untuk diterapkan pada klasifikasi teks pendek karena memberikan performa yang hampir setara dengan beban komputasi yang jauh lebih ringan. ***** Short texts on social media are characterized by limited context and the use of informal language, which complicates the sentiment classification process. To date, there has been no research that comprehensively compares the performance of Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), and Bidirectional LSTM (BiLSTM) algorithms specifically on Indonesian short texts, resulting in a lack of empirical guidelines for optimal model selection. Therefore, this study aims to analyze the results and compare the performance of these three models in classifying online gambling sentiments in 2025. This research utilized Indonesian tweet data collected from the X platform using the keywords "judol" or "judi online". The data was processed through text preprocessing and lexicon-based labeling (InSet), yielding 11,664 clean data points, and subsequently represented into numeric vectors using the Word2Vec algorithm with a Skip-gram architecture. Model performance evaluation was measured based on accuracy, F1-score, and computational efficiency analysis covering training time, CPU, and RAM utilization. The results showed that the BiLSTM model achieved the highest classification performance with an accuracy of 68.80% and a macro average F1-Score of 0.68, closely followed by the LSTM model with an accuracy of 67.55%, while the RNN model experienced underfitting with an accuracy of 55.98%. In terms of computational efficiency, there is a highly significant trade-off. Although BiLSTM achieved the highest accuracy, the performance gain is relatively insignificant as it only outperforms LSTM by approximately 1.25%. On the other hand, the computational cost demanded by BiLSTM is drastically higher, requiring the longest training time (± 910 seconds compared to ± 330 seconds for LSTM) and the highest memory consumption (1,000 - 1,100 MB compared to 750 - 800 MB for LSTM). In conclusion, while BiLSTM effectively processes bidirectional semantic ranges to yield the highest accuracy, the LSTM model proves to be significantly more efficient and optimal for short text classification, offering nearly identical performance with a considerably lighter computational load.
| Item Type: | Thesis (Sarjana) |
|---|---|
| Additional Information: | 1). Dr. Widodo, S. Kom., M. Kom; 2). Neng Ayu Herawati, S. Pd., M. T. |
| Subjects: | Teknologi dan Ilmu Terapan > Teknik Komputer |
| Divisions: | FT > S1 Pendidikan Teknik Informatika Komputer |
| Depositing User: | Muhammad Ramdany Juliansyah . |
| Date Deposited: | 07 Aug 2026 09:05 |
| Last Modified: | 07 Aug 2026 09:05 |
| URI: | http://repository.unj.ac.id/id/eprint/69166 |
Actions (login required)
![]() |
View Item |
