Audio Strips Network (ASNet) and Amalgamated Audio Features (A2F): A Synergistic Approach for Audio Source Separation
Abstract
Audio source separation refers to the procedure of decomposing a mixed audio signal into its constituent components. This technique enables numerous applications, including creative music production, educational tools, karaoke, transcription, and music analysis. Despite the recent success of deep learning-based source separation techniques, these techniques often do not perform very accurately and do not provide high-quality separation of sources when many contain complex combinations in their mixtures. Source separation techniques generally rely on temporal or spectral features for analysis, which does not fully capture the complex dynamics of audio signals. To address these limitations, this paper proposes the amalgamated audio features (A2F), a hybrid representation combining temporal and spectral features. Then, the audio strip network (ASNet) is proposed, a novel framework designed to achieve clean and precise separation of individual audio sources with enhanced performance. ASNet utilizes A2F to separate sources more effectively. The model is trained and evaluated on the MUSDB, DSD100, and MUSDB18-HQ datasets, a benchmark for music source separation, and its standard measures, such as the signal-to-distortion ratio (SDR) and signal-to-interference ratio (SIR), are used to examine performance. ASNet achieves enhanced separation performance with SDR values of 12.63 for drums, 11.42 for vocal, 12.01 for bass, and 11.14 for other, and SIR values of 9.57 for drums, 9.61 for vocal, 9.66 for bass, and 9.67 for other. This advancement benefits musicians through high-quality remixing and creativity while aiding researchers in improving deep learning and hybrid audio processing models.
Keywords:
feature extraction, amalgamated audio features (A2F), source extraction, audio strip network (ASNet)References
- Chen J., Vekkot S., Shukla P. (2024), Music source separation based on a lightweight deep learning framework (DTTNET: DUAL-PATH TFC-TDF UNET), [in:] ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 656–660, https://doi.org/10.1109/ICASSP48485.2024.10448020
- Chuang S.Y., Wang H.M., Tsao Y. (2022), Improved lite audio-visual speech enhancement, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 1345–1359, https://doi.org/10.1109/TASLP.2022.3153265
- Defossez A. (2021), Hybrid spectrogram and waveform source separation, arXiv, https://doi.org/10.48550/arXiv.2111.03600
- Deng B., Yang H., Kim N.Y. (2024), A denoising autoencoder based on U-Net and bidirectional long short-term memory for multi-level random telegraph signal analysis, Engineering Applications of Artificial Intelligence, 135: 108685, https://doi.org/10.1016/j.engappai.2024.108685
- Hamza A. et al. (2022), Deepfake audio detection via MFCC features using machine learning, IEEE Access, 10: 134018–134028, https://doi.org/10.1109/ACCESS.2022.3231480
- Huber D.M., Caballero E., Runstein R. (2023), Modern Recording Techniques: A Practical Guide to Modern Music Production, Focal Press.
- Issa R.J., Al-Irhaym Y.F. (2021), Audio source separation using supervised deep neural network, Journal of Physics: Conference Series, 1879(2): 022077, https://doi.org/10.1088/1742-6596/1879/2/022077
- Kong Q., Cao Y., Liu H., Choi K., Wang Y. (2021), Decoupling magnitude and phase estimation with deep ResUNet for music source separation, arXiv, https://doi.org/10.48550/arXiv.2109.05418
- Luo Y., Yu J. (2023), Music source separation with band-split RNN, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31: 1893–1901, https://doi.org/10.1109/TASLP.2023.3271145
- O’Connell D.C., Kowal S. (2009), Transcription systems for spoken discourse, [in:] The Pragmatics of Interaction, pp. 240–254, John Benjamins Publishing Company.
- Rafii Z., Liutkus A., Stoter F.R., Mimilakis S.I., Bittner R. (2019), MUSDB18-HQ-an uncompressed version of MUSDB18, Zenodo, https://doi.org/10.5281/zenodo.3338372
- Rajaby E., Sayedi S.M. (2022), A structured review of sparse fast Fourier transform algorithms, Digital Signal Processing, 123: 103403, https://doi.org/10.1016/j.dsp.2022.103403
- Renukadevi P., John S., Shivani N. (2024), Forensic science: AI-powered image and audio analysis, [in:] 2024 5th International Conference on Smart Electronics and Communication (ICOSEC), pp. 1519–1525, https://doi.org/10.1109/ICOSEC61587.2024.10722068
- Rezaul K.M. et al. (2024), Enhancing audio classification through MFCC feature extraction and data augmentation with CNN and RNN models, International Journal of Advanced Computer Science and Applications, 15(7): 37–53, https://doi.org/10.14569/IJACSA.2024.0150704
- Sanders M.E., Kant E., Smit A.L., Stegeman I. (2021), The effect of hearing aids on cognitive function: a systematic review, PLoS One, 16(12): e0261207, https://doi.org/10.1371/journal.pone.0261207
- Sun C. et al. (2021), A convolutional recurrent neural network with attention framework for speech separation in monaural recordings, Scientific Reports, 11(1): 1434, https://doi.org/10.1038/s41598-020-80713-3
- Takahashi N., Mitsufuji Y. (2020), D3Net: densely connected multidilated DenseNet for music source separation, arXiv, https://doi.org/10.48550/arXiv.2010.01733
- Tong W. et al. (2024), SCNet: sparse compression network for music source separation, [in:] ICASSP 2024 – 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1276–1280, https://doi.org/10.1109/ICASSP48485.2024.10446651
- Wang J.-C., Lu W.-T., Won M. (2023), Mel-band roformer for music source separation, arXiv, https://doi.org/10.48550/arXiv.2310.01809
- Wu X., Dang B., Zhang T., Wu X., Yang Y. (2024), Spatiotemporal audio feature extraction with dynamic memristor-based time-surface neurons, Science Advances, 10(14): eadl2767, https://doi.org/10.1126/sciadv.adl2767
- Zhang Y. et al. (2025), Comparison of performance for cochlear-implant listeners using audio processing strategies based on short-time fast Fourier transform or spectral feature extraction, Ear and Hearing, 46(1): 163–183, https://doi.org/10.1097/AUD.0000000000001565
- Zolzer U. (2022), Digital Audio Signal Processing, John Wiley & Sons.

