Abstract
This article introduces a novel algorithm for key detection in musical compositions represented in symbolic form, such as MIDI. The method is based on the analysis of a triple composite signature of fifths (TCSF), which is formed by combining the signatures of fifths obtained individually for the beginning, the end, and the entirety of the analyzed piece. The key is determined by the sign of the angle between the characteristic vector of the TCSF and the major/minor mode axis derived from that signature. To evaluate the algorithm’s effectiveness, experiments were conducted using Chopin’s Preludes, Op. 28, pieces from the Saarland Music Data: MIDI-Audio Piano Music dataset, as well as pieces from the Schubert Winterreise Dataset. As a reference method, a correlation-based algorithm implementing various tonal profiles was used. The proposed algorithm achieved the highest accuracy for Chopin’s Preludes and the Schubert Winterreise Dataset (91.67 % in both cases), while for the Saarland Music Data, its accuracy was lower than that of the correlation-based approaches (72.34 %). The main advantages of the presented algorithm are low computational complexity, stable decision-making when extending the analytical window, as well as incorporation of expert analysis aspects, particularly by focusing on the beginning and end of the composition.
Keywords:
key detection, tonality, music information retrieval, music classificationReferences
- Aarden J. (2003), Dynamic melodic expectancy, Ph.D. Thesis, The Ohio State University.
- Albrecht J., Shanahan D. (2013), The use of large corpora to train a new type of key-finding algorithm: an improved treatment of the minor mode, Music Perception: An Interdisciplinary Journal, 31(1): 59–67, https://doi.org/10.1525/mp.2013.31.1.59
- Bernardes G., Cocharro D., Caetano M., Guedes C., Davies M.E.P. (2016), A multi-level tonal interval space for modelling pitch relatedness and musical dissonance, Journal of New Music Research, 45(4): 281–294, https://doi.org/10.1080/09298215.2016.1182192
- Bellmann (2006), About the determination of key of a musical excerpt, [in:] Computer Music Modeling and Retrieval (CMMR 2005), Kronland-Martinet R., Voinier T., Ystad S. [Eds], 3902: 76–91, Springer, https://doi.org/10.1007/11751069 7.
- Bittner R.M., McFee B., Salamon J., Li P., Bello J.P. (2017), Deep salience representations for F0 estimation in polyphonic music, [in:] Proceedings of the 18th International Society for Music Information Retrieval Conference (ISMIR), pp. 63–70.
- Briot J.-P., Hadjeres G., Pachet F.-D. (2020), Deep Learning Techniques for Music Generation, Springer Cham, https://doi.org/10.1007/978-3-319-70163-9
- Chapin H., Jantzen K., Kelso J.A.S., Steinberg F., Large E. (2020), Dynamic emotional and neural responses to music depend on performance expression and listener experience, PloS ONE, 5(12): e13812, https://doi.org/10.1371/journal.pone.0013812
- Chew E. (2000), Towards a mathematical model of tonality, Ph.D. Thesis, Massachusetts Institute of Technology.
- Chew E. (2008), Out of the grid and into the spiral: geometric interpretations of and comparisons with the spiral-array model, Computing in Musicology, 15: 51–72.
- Cho T. (2013), Improved techniques for automatic chord recognition from music audio signals, Ph.D. Thesis, New York University, New York, USA.
- Dawson M.R.W. (2018), Connectionist Representations of Tonal Music. Discovering Musical Patterns by Interpreting Artificial Neural Networks, AU Press, Athabasca University.
- Deng J., Kwok Y.-K. (2017), Large vocabulary automatic chord estimation using deep neural nets: design framework, system variations and limitations, arXiv, https://doi.org/10.48550/arXiv.1709.07153
- Duan Z., Lu L., Zhang C. (2008), Audio tonality mode classification without tonic annotations, [in:] IEEE International Conference on Multimedia and Expo, pp. 1361–1364.
- Euler L. (1739), An Attempt at a New Theory of Music, Exposed in All Clearness, According to the Most Well-Founded Principles of Harmony [in Latin: Tentamen Novae Theoriae Musicae Ex Certissismis Harmoniae Principiis Dilucide Expositae], Saint Petersburg Academy.
- Grekow J. (2018), Musical performance analysis in terms of emotions it evokes, Journal of Intelligent Information Systems, 51(2): 415–437, https://doi.org/10.1007/s10844-018-0510-y
- Harte C., Sandler M., Gasser M. (2006), Detecting harmonic change in musical audio, [in:] Proceedings of Special 1st ACM Workshop on Audio and Music Computing Multimedia, pp. 21–26, https://doi.org/10.1145/1178723.1178727
- Herremans D., Chew E. (2019), MorpheuS: generating structured music with constrained patterns and tension, IEEE Transactions on Affective Computing, 10(4): 510–523, https://doi.org/10.1109/TAFFC.2017.2737984
- Huang C.-Z.A., Duvenaud D., Gajos K.Z. (2016), ChordRipple: recommending chords to help novice composers go beyond the ordinary, [in:] Proceedings of the 21st International Conference on Intelligent User Interfaces, pp. 241–250, https://doi.org/10.1145/2856767.2856792
- Jiang X., Zhang Y., Lin G., Yu L. (2024), Music emotion recognition based on deep learning: a review, IEEE Access, 12: 157716–157745, https://doi.org/10.1109/ACCESS.2024.3484470
- Ji S., Yang X. (2024), EmoMusicTv: emotion-conditioned symbolic music generation with hierarchical transformer VAE, IEEE Transactions on Multimedia, 26: 1076–1088, https://doi.org/10.1109/TMM.2023.3276177
- Ji S., Yang X., Luo J. (2023), A survey on deep learning for symbolic music generation: representations, algorithms, evaluations, and challenges, ACM Computing Surveys, 56(1): 1–39, https://doi.org/10.1145/3597493
- Juslin P.N., Sloboda J.A. [Eds] (2010), Handbook of Music and Emotion: Theory, Research, Applications, Oxford University Press, https://doi.org/10.1093/acprof:oso/9780199230143.001.0001
- Kania D., Kania P. (2019), A key-finding algorithm based on music signature, Archives of Acoustics, 44(3): 447–457, https://doi.org/10.24425/aoa.2019.129260
- Kania D., Kania P., Łukaszewicz T. (2021a), Trajectory of fifths in music data mining, IEEE Access, 9: 8751–8761, https://doi.org/10.1109/ACCESS.2021.3049266
- Kania M., Kania D. (2022), Trajectory of fifths – a two-dimensional representation of music [in Polish: Trajektoria kwintowa – dwuwymiarowa reprezentacja muzyki], Przegląd Elektrotechniczny, 98(6): 70–73, http://doi.org/10.15199/48.2022.06.13
- Kania M., Łukaszewicz T., Kania D., Mościńska K., Kulisz J. (2022), A comparison of the music key detection approaches utilizing key-profiles with a new method based on the signature of fifths, Applied Sciences, 12(21): 11261, https://doi.org/10.3390/app122111261
- Kania P., Kania D., Łukaszewicz T. (2024), A low complexity key-finding algorithm based on the signature of fifths, Archives of Acoustics, 49(4): 491–505, https://doi.org/10.24425/aoa.2024.148817
- Kania P., Kania D., Łukaszewicz T. (2021b), A hardware-oriented algorithm for real-time music key signature recognition, Applied Sciences, Computing and Artificial Intelligence, 11(18): 8753, https://doi.org/10.3390/app11188753
- Korzeniowski F., Widmer G. (2017), End-to-end musical key estimation using a convolutional neural network, [in:] Proceedings of the 25th European Signal Processing Conference (EUSIPCO), pp. 966–970.
- Kostrzewa D., Mazur W., Brzeski R. (2022), Wide ensembles of neural networks in music genre classification, [in:] Computational Science – ICCS 2022. ICCS 2022. Lecture Notes in Computer Science, 13351: 64–71, https://doi.org/10.1007/978-3-031-08754-7_9
- Krumhansl L. (1990), Cognitive Foundations of Musical Pitch, New York: Oxford University Press, pp. 77–110, https://doi.org/10.1093/acprof:oso/9780195148367.001.0001
- Longuet-Higgins C. (1962a), Letter to a musical friend, The Music Review, 23: 244–238.
- Longuet-Higgins C. (1962b), Second letter to a musical friend, The Music Review, 23: 271–280.
- Łukaszewicz T., Kania D. (2022), A music classification approach based on the trajectory of fifths, IEEE Access, 10: 73494–73502, https://doi.org/10.1109/ACCESS.2022.3190016
- Łukaszewicz T., Kania D. (2025a), Trajectory of fifths in tonal mode detection, IEEE Access, 13: 113763–113772, https://doi.org/10.1109/ACCESS.2025.3584400
- Łukaszewicz T., Kania D. (2025b), Automatic key detection in audio recordings using the signature of fifths, IEEE Access, 13: 160149–160157, https://doi.org/10.1109/ACCESS.2025.3609536
- Łukaszewicz T., Kania D. (2025c), Trajectory of fifths based on chroma subbands extraction – a new approach to music representation, analysis, and classification, IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(3): 2157–2169, https://doi.org/10.1109/TPAMI.2024.3519420
- Mauch (2010), Automatic chord transcription from audio using computational models of musical context, Ph.D. Thesis, Queen Mary University of London.
- McFee B., Salamon J., Bello J.P. (2018), Adaptive pooling operators for weakly labeled sound event detection, IEEE Transactions on Audio, Speech, and Language Processing, 26(11): 2180–2193, https://doi.org/10.1109/TASLP.2018.2858559
- Muller M., Konz V., Arifi-Muller V., Bogler W., Zeitleri J. (2024), Saarland music data: MIDI-audio piano music, Zenodo, https://doi.org/10.5281/zenodo.13753319
- Pauwels J., O’Hanlon K., Gomez E., Sandler M.B. (2019), 20 years of automatic chord recognition from audio, [in:] 20th International Society for Music Information Retrieval Conference, pp. 54–63.
- Roig C., Tardon L.J., Barbancho I., Barbancho A.M. (2014), Automatic melody composition based on a probabilistic model of music style and harmonic rules, Knowledge-Based Systems, 71: 419–434, https://doi.org/10.1016/j.knosys.2014.08.018
- Rosner A., Kostek B. (2018), Automatic music genre classification based on musical instrument track separation, Journal of Intelligent Information Systems, 50(2): 363–384, https://doi.org/10.1007/s10844-017-0464-5
- Sabathé R., Coutinho E., Schuller B. (2017), Deep recurrent music writer: memory-enhanced variational
autoencoder-based musical score composition and an objective measure, [in:] 2017 International Joint Conference on Neural Networks (IJCNN), pp. 3467–3474, https://doi.org/10.1109/IJCNN.2017.7966292 - Schedl M., Gomez E., Urbano J. (2014), Music information retrieval: recent developments and applications, Foundations and Trends in Information Retrieval, 8(2–3): 127–261, https://doi.org/10.1561/1500000042
- Shepard N. (1982), Geometrical approximations to the structure of musical pitch, Psychological Review, 89(4): 305–333, https://doi.org/10.1037/0033-295X.89.4.305
- Sturm L. (2013), Classification accuracy is not enough: on the evaluation of music genre recognition systems, Journal of Intelligent Information Systems, 41(3): 371–406, https://doi.org/10.1007/s10844-013-0250-y
- Temperley D., Marvin E.W. (2008), Pitch-class distribution and the identification of key, Music Perception, 25(3): 193–212, https://doi.org/10.1525/mp.2008.25.3.193
- Weiß C., Brand F., Muller M. (2019), Mid-level chord transition features for musical style analysis, [in:] ICASSP 2019 – 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 341–345, https://doi.org/10.1109/ICASSP.2019.8682293
- Weiß C. et al. (2020), Schubert Winterreise dataset, Zenodo, https://doi.org/10.5281/zenodo.4122060
- Wu X. et al. (2023), MelodyGLM: multi-task pre-training for symbolic melody generation, arXiv, https://doi.org/10.48550/arXiv.2309.10738
- Yang S., Reed C.N., Chew E., Barthet M. (2023), Examining emotion perception agreement in live music performance, IEEE Transactions on Affective Computing, 14(2): 1442–1460, https://doi.org/10.1109/TAFFC.2021.3093787
- Zhou X., Lerch A. (2015), Chord detection using deep learning, [in:] Proceedings of the 16th International Society for Music Information Retrieval Conference, ISMIR 2015, pp. 52–58.

