Hi,
Amazing work! I'm just curious about the pretraining dataset used for experiments on Kinetics-50C. Is the original CAV-MAE model completely fine-tuned on Kinetics-50 or are the CAV-MAE weights, initialized from VGGSound, fixed with "only" the classifiers being fine-tuned?
I don't understand this statement in the Appendix (for both datasets) - "During the fine-tuning phase, we maintain the visual and audio encoders of the pre-trained model and add one randomly initialized classification head upon them." What are the pre-trained model weights here?
Hi,
Amazing work! I'm just curious about the pretraining dataset used for experiments on Kinetics-50C. Is the original CAV-MAE model completely fine-tuned on Kinetics-50 or are the CAV-MAE weights, initialized from VGGSound, fixed with "only" the classifiers being fine-tuned?
I don't understand this statement in the Appendix (for both datasets) - "During the fine-tuning phase, we maintain the visual and audio encoders of the pre-trained model and add one randomly initialized classification head upon them." What are the pre-trained model weights here?