AudioSet

Introduced by Jort F. Gemmeke et al. in Audio Set: An ontology and human-labeled dataset for audio events

Audioset is an audio event dataset, which consists of over 2M human-annotated 10-second video clips. These clips are collected from YouTube, therefore many of which are in poor-quality and contain multiple sound-sources. A hierarchical ontology of 632 event classes is employed to annotate these data, which means that the same sound could be annotated as different labels. For example, the sound of barking is annotated as Animal, Pets, and Dog. All the videos are split into Evaluation/Balanced-Train/Unbalanced-Train set.

Source: Curriculum Audiovisual Learning

Homepage

Benchmarks

Add a new result Link an existing benchmark

Task	Dataset Variant	Best Model
Audio Classification	AudioSet	OmniVec
Audio Tagging	AudioSet	CAV-MAE
Zero-shot Audio Classification	AudioSet	LanguageBind
Audio Source Separation	AudioSet	ST-SED-SEP
Multi-modal Classification	AudioSet	CAV-MAE
Target Sound Extraction	AudioSet	CLAPSep