Causal Intervention with LMMs: Generative Text Facilitates Fine-Grained Pedestrian Anomaly Behavior Retrieval
Authors: Weifeng Xu,
Shaofei Huang,
Yaxiong Wang, Zhedong Zheng
Published in Findings of EMNLP, 2026
Recommended citation: Weifeng Xu, Shaofei Huang, Yaxiong Wang, Zhedong Zheng, "Causal Intervention with LMMs: Generative Text Facilitates Fine-Grained Pedestrian Anomaly Behavior Retrieval." Findings of EMNLP, 2026.
Abstract: Pedestrian anomaly behavior retrieval requires disambiguating subtle anomalous actions from routine ones using natural language. Existing cross-modal alignment methods often struggle with confounding biases: visual entanglement (dominant environments overshadowing target actions) and textual biases (limited stylistic diversity). These spurious correlations lead to erroneous matches among visually similar but semantically distinct samples. To address this, we reformulate the task from a causal intervention perspective. We propose a dual generative framework leveraging Large Multimodal Models (LMMs) to break environment-action and style-text confounders. Specifically, we synthesize causally meaningful hard negatives to decouple actions from environments, and augment positive queries with stylistically diverse rephrasings to eliminate narrative dependency. Extensive experiments on the Pedestrian Anomaly Behavior (PAB) benchmark yield consistent improvements of +1.71% in Recall@1 and +1.01% in mAP. Furthermore, zero-shot evaluation on the out-of-distribution UCC dataset demonstrates markedly enhanced generalization, achieving +3.23% and +2.46% gains in Recall@1 and mAP, respectively.
@inproceedings{xu2026causal,
author = "Xu, Weifeng and Huang, Shaofei and Wang, Yaxiong and Zheng, Zhedong",
title = "Causal Intervention with {LMMs}: Generative Text Facilitates Fine-Grained Pedestrian Anomaly Behavior Retrieval",
abstract = "Pedestrian anomaly behavior retrieval requires disambiguating subtle anomalous actions from routine ones using natural language. Existing cross-modal alignment methods often struggle with confounding biases: visual entanglement (dominant environments overshadowing target actions) and textual biases (limited stylistic diversity). These spurious correlations lead to erroneous matches among visually similar but semantically distinct samples. To address this, we reformulate the task from a causal intervention perspective. We propose a dual generative framework leveraging Large Multimodal Models (LMMs) to break environment-action and style-text confounders. Specifically, we synthesize causally meaningful hard negatives to decouple actions from environments, and augment positive queries with stylistically diverse rephrasings to eliminate narrative dependency. Extensive experiments on the Pedestrian Anomaly Behavior (PAB) benchmark yield consistent improvements of +1.71\% in Recall@1 and +1.01\% in mAP. Furthermore, zero-shot evaluation on the out-of-distribution UCC dataset demonstrates markedly enhanced generalization, achieving +3.23\% and +2.46\% gains in Recall@1 and mAP, respectively.",
booktitle = "Findings of EMNLP",
year = "2026" }