Abstract
Repetition is a core principle in music. This is especially true for popular songs, generally marked by a noticeable repeating musical structure, over which the singer performs varying lyrics. On this basis, we propose a simple method for separating music and voice, by extraction of the repeating musical structure. First, the period of the repeating structure is found. Then, the spectrogram is segmented at period boundaries and the segments are averaged to create a repeating segment model. Finally, each time-frequency bin in a segment is compared to the model, and the mixture is partitioned using binary time-frequency masking by labeling bins similar to the model as the repeating background. Evaluation on a dataset of 1,000 song clips showed that this method can improve on the performance of an existing music/voice separation method without requiring particular features or complex frameworks.
Original language | English (US) |
---|---|
Title of host publication | 2011 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2011 - Proceedings |
Pages | 221-224 |
Number of pages | 4 |
DOIs | |
State | Published - Aug 18 2011 |
Event | 36th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2011 - Prague, Czech Republic Duration: May 22 2011 → May 27 2011 |
Other
Other | 36th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2011 |
---|---|
Country | Czech Republic |
City | Prague |
Period | 5/22/11 → 5/27/11 |
Keywords
- Binary Time-Frequency Masking
- Music/Voice Separation
- Repeating Pattern
ASJC Scopus subject areas
- Software
- Signal Processing
- Electrical and Electronic Engineering