Hi, I'm wondering if the data is normalized to equal power or the
model receive input of different power during the training? I check the code and don't find anything like that and the paper doesn't mention this, too.
That's to say, if I train the model with negative sample from different speaker, than there is probability for the model to make use of speaking volume of the speaker the distinguish positive sample from negative sample?
Hi, I'm wondering if the data is normalized to equal power or the
model receive input of different power during the training? I check the code and don't find anything like that and the paper doesn't mention this, too.
That's to say, if I train the model with negative sample from different speaker, than there is probability for the model to make use of speaking volume of the speaker the distinguish positive sample from negative sample?