Hello,
I have been trying to use SpatialNet for single-channel speech enhancement, but I have encountered an issue in our lab experiments. When the trained model is exposed to relatively strong noise, the enhanced speech still contains quite noticeable residual noise (the model was trained using an MSE loss on the complex spectrum), as illustrated in the example below (fileid_6.wav in DNS Challenge no-reverb test set).
I find this somewhat puzzling, because with models I have trained in the past, when facing low signal-to-noise ratio conditions, the typical behavior is that the speech tends to be poorly preserved, rather than exhibiting obvious residual noise. I am therefore wondering whether the authors have encountered similar issues when using SpatialNet for single-channel speech enhancement, or whether there is any relevant experience or insight they could share.
Thank you very much!
fileid_6_enh.wav
fileid_6_noisy.wav
Hello,
I have been trying to use SpatialNet for single-channel speech enhancement, but I have encountered an issue in our lab experiments. When the trained model is exposed to relatively strong noise, the enhanced speech still contains quite noticeable residual noise (the model was trained using an MSE loss on the complex spectrum), as illustrated in the example below (fileid_6.wav in DNS Challenge no-reverb test set).
I find this somewhat puzzling, because with models I have trained in the past, when facing low signal-to-noise ratio conditions, the typical behavior is that the speech tends to be poorly preserved, rather than exhibiting obvious residual noise. I am therefore wondering whether the authors have encountered similar issues when using SpatialNet for single-channel speech enhancement, or whether there is any relevant experience or insight they could share.
Thank you very much!
fileid_6_enh.wav
fileid_6_noisy.wav