Acoustically Grounded Cost Learning for Open-Vocabulary Audio-Visual Semantic Segmentation

Authors: Tianrui Hui, shaofei-huangShaofei Huang, Qisong Han, yaxiong-wangYaxiong Wang, Lechao Cheng, Zhedong Zheng, zhun-zhongZhun Zhong, Richang Hong, Meng Wang

Published in ACM International Conference on Multimedia (ACM MM), 2026

Recommended citation: Tianrui Hui, Shaofei Huang, Qisong Han, Yaxiong Wang, Lechao Cheng, Zhedong Zheng, Zhun Zhong, Richang Hong, Meng Wang, "Acoustically Grounded Cost Learning for Open-Vocabulary Audio-Visual Semantic Segmentation." ACM MM, 2026.

Abstract: Open-Vocabulary Audio-Visual Semantic Segmentation (OV-AVSS) aims to perform pixel-level segmentation of sound-emitting objects from an open set of categories. The previous method relies on a class-agnostic foreground definition, which groups semantically diverse objects into a heterogeneous positive set, causing the model to learn unstable sounding patterns and produce unreliable proposals. To address this, we reformulate the objective to be category-specific and propose a novel Acoustically Grounded Cost Learning (AGCL) framework to transform the static, audio-agnostic visual-text priors into dynamic, audio-grounded cost representations. For intra-category soundingness discovery, we devise Audio-Modulated Cost Generation (AMCG) and Audio-Guided Temporal Aggregation (AGTA) modules to enable both frame-level sounding region highlighting and video-level temporal refinement with a low-intrusive audio injection mechanism. For inter-category distractor discrimination, we introduce a Synergistic Distractor Mining (SDM) strategy, which selectively penalizes acoustically and semantically confusing negative categories to learn more discriminative decision boundaries. Extensive experiments on the AVSBench-OV dataset demonstrate that our method significantly outperforms previous state-of-the-art approaches, particularly on unseen categories. Code will be released to facilitate further research.

@inproceedings{hui2026acoustically,
author = "Hui, Tianrui and Huang, Shaofei and Han, Qisong and Wang, Yaxiong and Cheng, Lechao and Zheng, Zhedong and Zhong, Zhun and Hong, Richang and Wang, Meng",
title = "Acoustically Grounded Cost Learning for Open-Vocabulary Audio-Visual Semantic Segmentation",
abstract = "Open-Vocabulary Audio-Visual Semantic Segmentation (OV-AVSS) aims to perform pixel-level segmentation of sound-emitting objects from an open set of categories. The previous method relies on a class-agnostic foreground definition, which groups semantically diverse objects into a heterogeneous positive set, causing the model to learn unstable sounding patterns and produce unreliable proposals. To address this, we reformulate the objective to be category-specific and propose a novel Acoustically Grounded Cost Learning (AGCL) framework to transform the static, audio-agnostic visual-text priors into dynamic, audio-grounded cost representations. For intra-category soundingness discovery, we devise Audio-Modulated Cost Generation (AMCG) and Audio-Guided Temporal Aggregation (AGTA) modules to enable both frame-level sounding region highlighting and video-level temporal refinement with a low-intrusive audio injection mechanism. For inter-category distractor discrimination, we introduce a Synergistic Distractor Mining (SDM) strategy, which selectively penalizes acoustically and semantically confusing negative categories to learn more discriminative decision boundaries. Extensive experiments on the AVSBench-OV dataset demonstrate that our method significantly outperforms previous state-of-the-art approaches, particularly on unseen categories. Code will be released to facilitate further research.",
booktitle = "ACM MM",
year = "2026" }