ECCV 2026 Workshop

2nd Workshop on
Benchmarking Evidence‑Aligned
Multimodal Reasoning

Moving beyond accuracy-only evaluation — verifying that multimodal predictions are supported by correct perceptual signals.

Half-Day Workshop — 4 Hours
14:00 – 18:00
September 8, 2026
Malmömässan — Room C1

Workshop Overview

Multimodal foundation models now achieve strong performance on audio-visual QA, vision–language, and video understanding benchmarks, yet accuracy can hide shortcut reasoning. Models may answer correctly by relying on linguistic priors or spurious correlations rather than attending to the visual regions, frames, or audio events that provide the true evidence.

Most existing benchmarks score only the final answer, offering limited insight into whether a model actually "saw" or "heard" what it needed. BEAM 2 centers on evidence-aligned multimodal reasoning: datasets, protocols, and metrics that jointly evaluate (i) answer correctness and (ii) perceptual grounding of the reasoning process.

We emphasize evaluation designs with verifiable anchors — bounding boxes, representative frames, temporal segments, and audio-event timestamps — and model outputs that include structured evidence references and short grounded rationales.

4 Invited Talks
4h Half-Day Program

Evidence Alignment

Verifying that predictions are supported by the correct perceptual signals — not linguistic shortcuts.

Perceptual Grounding

Linking multimodal reasoning to concrete audio-visual cues with interpretable anchors.

Composite Metrics

Jointly scoring answer correctness and grounding quality beyond accuracy-only reporting.

Community Standards

Establishing standardized, interpretable metrics for the next generation of trustworthy multimodal systems.

Topics of Interest

BEAM 2 welcomes contributions spanning methodology, empirical analysis, and systems-level advances in evidence-aligned multimodal evaluation.

01

Reasoning-Aware Benchmarks

Multimodal benchmarks for video, audio, vision–language, and audio-visual QA that go beyond final-answer scoring.

02

Perceptual Grounding & Localization

Spatial, temporal, and audio-event grounding and evidence localization for multimodal reasoning.

03

Explanation Faithfulness

Evaluation of explanation faithfulness and evidence alignment beyond plausibility checks.

04

Composite Metrics

Metrics that jointly score answer correctness and grounding quality (composite or multi-objective).

05

Counterfactual & Perturbation Evaluations

Muting audio, masking regions, or swapping objects/events to probe genuine model understanding.

06

Human Evaluation & Auditing

Human-in-the-loop protocols and auditing tools for interpretable multimodal reasoning and agentic deployment.

07

Bias, Robustness & Fairness

Robustness and fairness in multimodal reasoning across accents, dialects, demographics, and domain shift.

08

Benchmark Governance

Documentation, licensing, privacy-preserving releases, and responsible leaderboards.

Workshop Schedule

A half-day (4-hour) afternoon program with invited talks, oral presentations, and a poster session. Held in Room C1 at Malmömässan.

14:00 – 14:05
Opening

Opening Remarks

From accuracy to evidence alignment

14:05 – 14:35
Invited Talk 1

Dacheng Tao

Nanyang Technological University

14:35 – 15:05
Oral Session 1

3 Contributed Papers

10 minutes per paper, including questions and transitions

  • Order Effects: Do Multimodal LLMs Ground Their Answers in Option Content or in Position? Sergei Olegovich Kurashkin, Vadim Tynchenko, Aleksei Borodulin, Vladimir Nelyub, Vladislav Kukartsev
  • Chartography: Benchmarking Grounded Multimodal Reasoning over Professional Charts Sushant Mehta
  • Warping Earth Observations for Better Ice Labelling in the Marginal Ice Zone Tom Kelly, Martin Samuel James Rogers
15:05 – 15:40
Break

Poster Session & Coffee

Posters from all 6 accepted papers

15:40 – 16:10
Invited Talk 2

Giorgos Tolias

Czech Technical University in Prague

16:10 – 16:40
Oral Session 2

3 Contributed Papers

10 minutes per paper, including questions and transitions

  • Grounded on Nothing: The Video-Blind Floor of Grounded VideoQA Metrics Krishna Harish
  • Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails Suyoung Lee, Myungsub Choi
  • TAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film Kaishuu Shinozaki-Conefrey, Olivier Pascaud, Robin Courant, Xi Wang, Dimitris Samaras, Vicky Kalogeiton
16:40 – 17:10
Invited Talk 3

Xiangyu Yue

The Chinese University of Hong Kong

17:10 – 17:20
Invited Paper from the Main Conference

HippoCamp

17:20 – 17:40
Invited Talk 4

Swaroop Mishra

Google DeepMind

17:40 – 18:00
Closing

Closing Remarks

Invited Speakers

Dacheng Tao

Dacheng Tao

Distinguished University Professor

Nanyang Technological University

Applies statistics and mathematics to AI and data science with 200+ publications and multiple best paper and test-of-time awards. Fellow of the Australian Academy of Science, AAAS, ACM, and IEEE.

Giorgos Tolias

Giorgos Tolias

Associate Professor

Czech Technical University in Prague

Leads a research team within the Visual Recognition Group at CTU in Prague. His research focuses on computer vision, with a particular emphasis on visual representation learning, large-scale image retrieval, and open-vocabulary segmentation.

Xiangyu Yue

Xiangyu Yue

Assistant Professor

The Chinese University of Hong Kong

Works on multi-modal AI, embodied AI, generative models, and agentic AI within the Multimedia Laboratory (MMLab). Received his Ph.D. from UC Berkeley at Berkeley AI Research, and has been recognized with the Lotfi A. Zadeh Award and an AI 2000 Most Influential Scholar honorable mention in robotics.

Swaroop Mishra

Swaroop Mishra

Research Scientist

Google DeepMind

Advances the reasoning capabilities of Gemini and contributed to the team's historic IMO silver-medal effort. A pioneer in instruction-tuning and self-improving AI systems, his foundational contributions to large language models have earned him major industry honors, including the AI2 Lasting Impact Paper Award.

Organizing Committee

A team spanning Carnegie Mellon University, Amazon AGI, Google, and Adobe, combining academic and industrial expertise in computer vision, multimodal learning, and large-scale evaluation.

Carnegie Mellon University

Laszlo A. Jeni

Laszlo A. Jeni

Assistant Professor

Carnegie Mellon University

Joel Julin

Joel Julin

PhD Student

Carnegie Mellon University

Souraja Kundu

Souraja Kundu

PhD Student

Carnegie Mellon University

Liza Dahiya

Liza Dahiya

MS Student

Carnegie Mellon University

Ananya Bal

Ananya Bal

PhD Student

Carnegie Mellon University

Amazon AGI

Louise Xie

Louise Xie

Applied Scientist

Amazon AGI

Davide Modolo

Davide Modolo

Senior Manager

Amazon AGI

Google

Xu Zhang

Xu Zhang

Staff Research Scientist

Google DeepMind

Hao Yang

Hao Yang

Senior Staff Engineer

Google

Ashwin Swaminathan

Ashwin Swaminathan

Director

Google

Adobe

Jingru Yi

Jingru Yi

Senior Applied Scientist

Adobe Firefly

Submit Your Work

We welcome original papers, position statements, and benchmark reports aligned with the workshop themes. Accepted original submissions will be archival.

Paper Format

Submissions should follow ECCV formatting guidelines and be between 4 and 14 pages excluding references.

Important Dates

  • Submission Opens May 31, 2026
  • Paper Deadline Jul 31, 2026
  • Notification Aug 7, 2026
  • Camera Ready Aug 13, 2026
  • Workshop Date September 8, 2026

Contact

Questions about submissions or the workshop program:

laszlojeni@cmu.edu