Audio Strips Network (ASNet) and Amalgamated Audio Features (A2F): A Synergistic Approach for Audio Source Separation

Downloads

Authors

Abstract

Audio source separation refers to the procedure of decomposing a mixed audio signal into its constituent components. This technique enables numerous applications, including creative music production, educational tools, karaoke, transcription, and music analysis. Despite the recent success of deep learning-based source separation techniques, these techniques often do not perform very accurately and do not provide high-quality separation of sources when many contain complex combinations in their mixtures. Source separation techniques generally rely on temporal or spectral features for analysis, which does not fully capture the complex dynamics of audio signals. To address these limitations, this paper proposes the amalgamated audio features (A2F), a hybrid representation combining temporal and spectral features. Then, the audio strip network (ASNet) is proposed, a novel framework designed to achieve clean and precise separation of individual audio sources with enhanced performance. ASNet utilizes A2F to separate sources more effectively. The model is trained and evaluated on the MUSDB, DSD100, and MUSDB18-HQ datasets, a benchmark for music source separation, and its standard measures, such as the signal-to-distortion ratio (SDR) and signal-to-interference ratio (SIR), are used to examine performance. ASNet achieves enhanced separation performance with SDR values of 12.63 for drums, 11.42 for vocal, 12.01 for bass, and 11.14 for other, and SIR values of 9.57 for drums, 9.61 for vocal, 9.66 for bass, and 9.67 for other. This advancement benefits musicians through high-quality remixing and creativity while aiding researchers in improving deep learning and hybrid audio processing models.

Keywords:

feature extraction, amalgamated audio features (A2F), source extraction, audio strip network (ASNet)

References


  1. Chen J., Vekkot S., Shukla P. (2024), Music source separation based on a lightweight deep learning framework (DTTNET: DUAL-PATH TFC-TDF UNET), [in:] ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 656–660, https://doi.org/10.1109/ICASSP48485.2024.10448020

  2. Chuang S.Y., Wang H.M., Tsao Y. (2022), Improved lite audio-visual speech enhancement, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 1345–1359, https://doi.org/10.1109/TASLP.2022.3153265

  3. Defossez A. (2021), Hybrid spectrogram and waveform source separation, arXiv, https://doi.org/10.48550/arXiv.2111.03600

  4. Deng B., Yang H., Kim N.Y. (2024), A denoising autoencoder based on U-Net and bidirectional long short-term memory for multi-level random telegraph signal analysis, Engineering Applications of Artificial Intelligence, 135: 108685, https://doi.org/10.1016/j.engappai.2024.108685

  5. Hamza A. et al. (2022), Deepfake audio detection via MFCC features using machine learning, IEEE Access, 10: 134018–134028, https://doi.org/10.1109/ACCESS.2022.3231480

  6. Huber D.M., Caballero E., Runstein R. (2023), Modern Recording Techniques: A Practical Guide to Modern Music Production, Focal Press.

  7. Issa R.J., Al-Irhaym Y.F. (2021), Audio source separation using supervised deep neural network, Journal of Physics: Conference Series, 1879(2): 022077, https://doi.org/10.1088/1742-6596/1879/2/022077

  8. Kong Q., Cao Y., Liu H., Choi K., Wang Y. (2021), Decoupling magnitude and phase estimation with deep ResUNet for music source separation, arXiv, https://doi.org/10.48550/arXiv.2109.05418

  9. Luo Y., Yu J. (2023), Music source separation with band-split RNN, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31: 1893–1901, https://doi.org/10.1109/TASLP.2023.3271145

  10. O’Connell D.C., Kowal S. (2009), Transcription systems for spoken discourse, [in:] The Pragmatics of Interaction, pp. 240–254, John Benjamins Publishing Company.

  11. Rafii Z., Liutkus A., Stoter F.R., Mimilakis S.I., Bittner R. (2019), MUSDB18-HQ-an uncompressed version of MUSDB18, Zenodo, https://doi.org/10.5281/zenodo.3338372

  12. Rajaby E., Sayedi S.M. (2022), A structured review of sparse fast Fourier transform algorithms, Digital Signal Processing, 123: 103403, https://doi.org/10.1016/j.dsp.2022.103403

  13. Renukadevi P., John S., Shivani N. (2024), Forensic science: AI-powered image and audio analysis, [in:] 2024 5th International Conference on Smart Electronics and Communication (ICOSEC), pp. 1519–1525, https://doi.org/10.1109/ICOSEC61587.2024.10722068

  14. Rezaul K.M. et al. (2024), Enhancing audio classification through MFCC feature extraction and data augmentation with CNN and RNN models, International Journal of Advanced Computer Science and Applications, 15(7): 37–53, https://doi.org/10.14569/IJACSA.2024.0150704

  15. Sanders M.E., Kant E., Smit A.L., Stegeman I. (2021), The effect of hearing aids on cognitive function: a systematic review, PLoS One, 16(12): e0261207, https://doi.org/10.1371/journal.pone.0261207

  16. Sun C. et al. (2021), A convolutional recurrent neural network with attention framework for speech separation in monaural recordings, Scientific Reports, 11(1): 1434, https://doi.org/10.1038/s41598-020-80713-3

  17. Takahashi N., Mitsufuji Y. (2020), D3Net: densely connected multidilated DenseNet for music source separation, arXiv, https://doi.org/10.48550/arXiv.2010.01733

  18. Tong W. et al. (2024), SCNet: sparse compression network for music source separation, [in:] ICASSP 2024 – 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1276–1280, https://doi.org/10.1109/ICASSP48485.2024.10446651

  19. Wang J.-C., Lu W.-T., Won M. (2023), Mel-band roformer for music source separation, arXiv, https://doi.org/10.48550/arXiv.2310.01809

  20. Wu X., Dang B., Zhang T., Wu X., Yang Y. (2024), Spatiotemporal audio feature extraction with dynamic memristor-based time-surface neurons, Science Advances, 10(14): eadl2767, https://doi.org/10.1126/sciadv.adl2767

  21. Zhang Y. et al. (2025), Comparison of performance for cochlear-implant listeners using audio processing strategies based on short-time fast Fourier transform or spectral feature extraction, Ear and Hearing, 46(1): 163–183, https://doi.org/10.1097/AUD.0000000000001565

  22. Zolzer U. (2022), Digital Audio Signal Processing, John Wiley & Sons.