A Rule-Based Heuristic Approach for Effective Madurese Slang Normalization
Mohammad Nazir Arifin(1*), Aji Prasetya Wibawa(2), Fauzan Prasetyo Eka Putra(3)
(1) Universitas Negeri Malang
(2) Universitas Negeri Malang
(3) Universitas Madura
(*) Corresponding Author
Abstract
Madurese slang on social media presents significant challenges for natural language processing due to its systematic yet unconventional alterations in spelling and pronunciation, for example, “bbh” becoming “pp” or “jh” turning into “cc”. Traditional string-matching techniques struggle with these transformations. This study evaluates four normalization approaches: standard string similarity models such as Jaro–Winkler and Levenshtein, Metaphone, hybrid strategies, and a novel heuristic-based method specifically designed for Madurese. We conducted experiments on a dataset of 105 slang–standard word pairs, assessing each method’s ability to accurately map slang to its standard form, both for top-1 and top-3 candidate matches. Results indicate that heuristic-based models significantly outperform generic methods. Notably, a combination of our heuristic approach with Levenshtein achieved 86.67% accuracy for Top-1 matches, while pairing the heuristic with Jaro–Winkler reached 91.43% accuracy for top-3 matches. These findings underscore the effectiveness of incorporating language-specific rules for normalization in morphophonemically complex, low-resource languages. Our proposed rule-based system offers a practical and adaptable solution for handling Madurese slang, highlighting the superiority of tailored heuristics over generic models in such contexts.
Keywords
Heuristic Normalization; Madurese Language; Low-Resource NLP; Slang Normalization; String Similarity
Full Text:
PDFArticle Metrics
Copyright (c) 2026 IJCCS (Indonesian Journal of Computing and Cybernetics Systems)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
View My Stats1






