Genre-Aware and User Guided Music Generation
Abstract
Music plays a crucial role in human life as a medium for emotional expression and entertainment. However, traditional music composition is time-consuming and requires expert skills, limiting accessibility for non-musicians. Recent advancement in artificial intelligence (AI) such as generative music enables the automatic creation of music, yet many models fail to incorporate explicit user preferences such as genre and composition length. This study proposed, user-guided Transformer XL, a genre-aware and user-guided music generation framework that worked on top of Transformer-XL. The system allowed users to specify desired genre and length, while a built-in genre classifier validated the stylistic accuracy of generated outputs. A genre classification module along with the number of bars was incorporated to ensure the generated outputs reflect the intended musical styles. The model was evaluated through a genre characteristic analysis and expert evaluation, revealing strong consistency in genre-specific features. Extensive experiments using the Lakh MIDI dataset demonstrated strong model performance, achieving an overall accuracy of 78.33%, with notable genre-specific strengths, particularly in jazz with recall 1.00 and classic with recall 0.88. In contrast, subjective evaluation using expert evaluation yielded a more moderate accuracy of 60%, highlighting the model’s ability to capture genre-relevant features even under more stringent, nuanced evaluative standards. Overall, this work demonstrates that integrating user preferences into generative modelling enhances flexibility and musical relevance, offering a robust foundation for future development in adaptive AI-driven composition.
References
G.F. Welch et al., “Editorial: The impact of music on human development and well-being,” Front. Psychol., vol. 11, pp. 1–4, Jun. 2020, doi: 10.3389/fpsyg.2020.01246.
Y. Ren et al., “PopMAG: Pop music accompaniment generation,” in MM '20, Proc. 28th ACM Int. Conf. Multimed., 2020, pp. 1198–1206, doi: 10.1145/3394171.3413721.
S. Sibagariang, “Interpretable machine learning for job placement prediction: A shap-based feature analysis,” J. Nas. Tek. Elekt. Teknol. Inf., vol. 14, no. 3, pp. 190–198, Aug. 2025, doi: 10.22146/jnteti.v14i3.20516.
B.Y. Kasula, “Harnessing machine learning for personalized patient care,” Trans. Latest Trends Artif. Intell., vol. 4, no. 4, pp. 1–9, Nov. 2023.
R. Lumbantoruan, X. Zhou, and Y. Ren, “Declarative user-item profiling based context-aware recommendation,” in 16th Int. Conf. Adv. Data Min. Appl., 2020, pp. 413–427, doi: 10.1007/978-3-030-65390-3_32.
R. Lumbantoruan, X. Zhou, Y. Ren, and Z. Bao, “D-CARS: A declarative context-aware recommender system,” in 2018 IEEE Int. Conf. Data Min. (ICDM), 2018, pp. 1152–1157, doi: 10.1109/Icdm.2018.00151.
Z. Hu, Y. Liu, G. Chen, and Y. Liu, “Can machines generate personalized music? A hybrid favorite-aware method for user preference music transfer,” IEEE Trans. Multimed., vol. 25, pp. 2296–2308, Jan. 2022, doi: 10.1109/TMM.2022.3146002.
C. Hernandez-Olivan and J.R. Beltran, “Music composition with deep learning: A review,” in Advances in Speech and Music Technology, Cham, Switzerland: Springer, 2022, pp. 25–50.
J. Ens and P. Pasquier, “MMM: Exploring conditional multi-track music generation with the transformer,” 2020, arXiv:2008.06048.
S.-L. Wu and Y.-H. Yang, “MuseMorphose: Full-song and fine-grained piano music style transfer with one transformer VAE,” IEEE/ACM Trans. Audio Speech Lang. Process., vol. 31, pp. 1953–1967, May 2023, doi: 10.1109/TASLP.2023.3270726.
L.N. Ferreira and J. Whitehead, “Learning to generate music with sentiment,” 2021, arXiv:2103.06125.
N. Tokui, “Can GAN originate new electronic dance music genres? Generating novel rhythm patterns using GAN with genre ambiguity loss,” 2020, arXiv:2011.13062.
A. Kolokolova et al., “GANs & reels: Creating Irish music using a generative adversarial network,” 2020, arXiv:2010.15772.
T. Greer, X. Shi, B. Ma, and S. Narayanan, “Creating musical features using multi-faceted, multi-task encoders based on transformers,” Sci. Rep., vol. 13, pp. 1–14, Jul. 2023, doi: 10.1038/s41598-023-36714-z.
S. Ji, X. Yang, and J. Luo, “A survey on deep learning for symbolic music generation: Representations, algorithms, evaluations, and challenges,” ACM Comput. Surv., vol. 56, no. 1, pp. 1–39, Aug. 2023, doi: 10.1145/3597493.
C. Hernandez-Olivan, J.A. Puyuelo, and J.R. Beltran, “Subjective evaluation of deep learning models for symbolic music composition,” 2022, arXiv:2203.14641.
L. Wang et al., “A review of intelligent music generation systems,” Neural Comput. Appl., vol. 36, no. 12, pp. 6381–6401, Apr. 2024, doi: 10.1007/s00521-024-09418-2.
O. Green, B. Sturm, G. Born, and M. Wald-Fuhrmann, “A critical survey of research in music genre recognition,” in Proc. 25th Int. Soc. Music Inf. Retr. (ISMIR) Conf., 2024, pp. 745–782, doi: 10.5281/zenodo.14877445.
G. Cideron et al., “Musicrl: Aligning music generation to human preferences,” in ICML'24, Proc. 41st Int. Conf. Mach. Learn., 2024, pp. 8968–8984.
C. Donahue et al., “LakhNES: Improving multi-instrumental music generation with cross-domain pre-training,” 2019, arXiv:1907.04868.
C. Wang, M. Li, and A.J. Smola, “Language models with transformers,” 2019, arXiv:1904.09408.
M.S. Cuthbert et al., “Hidden beyond MIDI’s reach: Feature extraction and machine learning with rich symbolic formats in music21,” in 4th Int. Workshop Mach. Learn. Music, Learn. Music. Struct., 2011. [Online]. Available: https://www.trecento.com/research/Cuthbert_Ariza_Cabal-Ugaz_Hadley_Parikh-Hidden-NIPS2011.pdf
E. Manilow, G. Wichern, P. Seetharaman, and J. Le Roux, “Cutting music source separation some Slakh: A dataset to study the impact of training data quality and quantity,” in 2019 IEEE Workshop Appl. Signal Process. Audio Acoust. (WASPAA), 2019, pp. 45–49, doi: 10.1109/WASPAA.2019.8937170.
A. Agostinelli et al., “Musiclm: Generating music from text,” 2023, arXiv:2301.11325.
© Jurnal Nasional Teknik Elektro dan Teknologi Informasi, under the terms of the Creative Commons Attribution-ShareAlike 4.0 International License.

1.png)

