Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts

arXiv Dataset License

Accepted to EMNLP 2026 Findings 🎉

Tri-PvP overview

Omni-modal LLMs (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexplored. Existing benchmarks conflate two forms of evidence within a modality: perceptual signals (a photo or recording of a dog) and propositional signals (the declarative claim "this is a dog"). Thus, any measured modality bias is confounded with evidence-form bias. Tri-PvP is an 8,000-sample tri-modal conflict benchmark that varies the two axes independently, letting modality bias be attributed cleanly.

Key Findings

  • Visual bias dominates. BIAS_IMAGE is the top single-modality label in 18 of 20 model × condition cells, often above 60%. BIAS_AUDIO is consistently the smallest, typically under 10%.
  • Evidence form modulates but never reverses the preference. Models prefer perceptual evidence in vision but propositional evidence in audio. This may be related to encoder pretraining: vision encoders target perceptual content, while audio encoders are usually ASR-initialized and recover linguistic content.
  • Fully propositional inputs raise unbiased responses. NO_BIAS peaks under Prop-image + Prop-audio, reaching 60.1% on Qwen2.5.
  • Bias is already encoded in the representations. Layer-wise linear probes decode image bias well above chance within the first few layers, before any token is generated. Contrastive decoding, used as a diagnostic, reduces image bias at inference time with no parameter updates and preserves general competence on OmniBench (37.4% vs. 38.4%), with part of the displaced mass appearing as text bias.

Benchmark

Each sample is a tuple of image, audio, text, and question. The three modality labels are mutually distinct, so the modalities always disagree, and none is more reliable by construction. Image and audio are either perceptual or propositional, while text is always propositional.

Axis Values
Domains animal, emotion, environment, music
Evidence conditions Perc-I × Perc-A, Perc-I × Prop-A, Prop-I × Perc-A, Prop-I × Prop-A
Size 2,000 per domain, 8,000 total

Responses are scored by an LLM judge into eight mutually exclusive labels: BIAS_IMAGE, BIAS_AUDIO, BIAS_TEXT, BIAS_IMAGE_AUDIO, BIAS_IMAGE_TEXT, BIAS_AUDIO_TEXT, HALLUCINATION, and NO_BIAS.

Five OLLMs are evaluated: Qwen2.5-Omni-7B, MiniCPM-o 4.5, Qwen3-Omni-30B-A3B-Thinking, Gemma 4 E4B, and Gemini 3 Flash.

Run Inference

# Gemini 3 Flash
python inference/inference.py --model gemini-3.0-flash --hf-repo ModaSense/animal --output results/

# Qwen2.5-Omni
python inference/inference.py --model qwen2.5-omni-7b --backend vllm --hf-repo ModaSense/animal --output results/

# Qwen3-Omni
python inference/inference_Qwen3Omni.py --model_name qwen3-omni-thinking --image_type perceptual --audio_type propositional --output_dir raw_output/

# Gemma 4 E4B
python inference/inference_gemma4.py --model gemma4-e4b-it --dataset ModaSense/animal --split test --output results/

# MiniCPM-o 4.5
python inference/inference_minicpm.py --model minicpm-o-4_5 --dataset ModaSense/animal --split test --output results/

Run LLM Judge

Remember to set the OpenAI API key in your environment variables before running the judge script.

python evaluation/llm_judge.py --input_file <your_inference_results.json>

Analysis

# Layer-wise linear probing
python analysis/gemma4_probing.py --image_type perceptual --audio_type propositional --judge_dir judge_results
python analysis/qwen2.5_probing.py --image_type perceptual --audio_type propositional

# Contrastive decoding mitigation
python analysis/contrastive_decoding.py --image_type propositional --audio_type propositional --alpha 1.5 --beta_filter 0.1

License

Code is released for non-commercial research and educational use. The benchmark compilation and generated data are CC BY-NC 4.0; perceptual source files remain governed by their original licenses, including the ImageNet Terms of Access and ESC-50's CC BY-NC 3.0.

Citation

@misc{piao2026tripvpexposingmodalitybias,
      title={Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts}, 
      author={Yen-Ting Piao and Shu-Yun Chen and Chin-Hui Chu and Chun-Wei Chen and Shih-Yun Shan Kuan and Hung-yi Lee and Yun-Nung Chen},
      year={2026},
      eprint={2609.06011},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2609.06011}, 
}

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages