Skip to content

Pretraining data of released source models. #10

Description

@sarthaxxxxx

Hi,

Amazing work! I'm just curious about the pretraining dataset used for experiments on Kinetics-50C. Is the original CAV-MAE model completely fine-tuned on Kinetics-50 or are the CAV-MAE weights, initialized from VGGSound, fixed with "only" the classifiers being fine-tuned?

I don't understand this statement in the Appendix (for both datasets) - "During the fine-tuning phase, we maintain the visual and audio encoders of the pre-trained model and add one randomly initialized classification head upon them." What are the pre-trained model weights here?

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions