AI model DenseAV learns by comparing pairs of audio and visual signals and determines what data is important. To learn these patterns it uses audio-video contrastive learning to associate a particular sound with the observable world. This mode of learning means the visual side of the model can’t gain any insights from the audio side (and vice-versa) forcing the algorithm to recognize objects in a meaningful way. Will AI become smart the geniuses of this world?

