A benchmark for selective history use

When History Helps and Hurts:Selective History Use across Multimodal Turns

History can supply what a model needs—and distract it from what matters.
Can multimodal assistants tell the difference?

Shuoyang Sun1*Kerui Gu2*Hao Fang3Shaoli Huang2†Bin Chen1✉

1 Harbin Institute of Technology, Shenzhen   ·   2 AgiBot
3 Tsinghua Shenzhen International Graduate School, Tsinghua University

* Equal contribution   † Project lead   ✉ Corresponding author

Full dataset coming soon.

7,000base tasks · 3,500 matched pairs
4 conditionstwo roles of conversational history
Vision + Audiotwo evidence modalities
13 modelsOmni · LVLM · LALM

01 / Selective history use

Keep what matters.
Leave interference behind.

LEFT · TASK-CARRYING HISTORYAn outdated answer persists

The earlier clothing-color question must be applied to the new clip. The model repeats “White,” although the correct answer is “Black.” It should retain the historical question but update its answer using current evidence.

RIGHT · EVIDENCE-CARRYING HISTORYCurrent evidence takes over

The current question targets the first clip, whose answer is “Fish.” The model instead answers “Dog” from the newer clip. The question is explicit; the challenge is selecting historical evidence despite a competing current observation.

Selective history use means retaining the relevant question or evidence, while resisting information that conflicts with the intended target.

Paper examples of task-carrying and evidence-carrying history failures

To assess selective history use, ReTurn varies historical agreement on the task side and current-media competition on the evidence side.

Task-carrying history

Historical question → Current evidence

Reconfirm Reference

Apply the historical question to current media; the answer stays the same.

Reground Conflict

Apply the historical question to current media; the earlier answer is outdated.

Evidence-carrying history

Current question → Historical evidence

Retrieve Reference

Answer about historical media, with weaker competition from current media.

Rebind Conflict

Answer about historical media, despite a conflicting answer in current media.

Within each pair, the target question, media, and answer remain fixed. Each task also has a direct Single-turn counterpart.

02 / Results

Answerable in isolation.
Harder in conversation.

Median OpenQA accuracy

93.7%72.3%Single-turnMulti-turn

Across 13 models · Short + Long pooled · Test

Single-turn vs. Multi-turn

OpenQA · Long and Short shown separately

OpenQA Single-turn and Multi-turn accuracy by model, with Long conversations on the left and Short conversations on the right
Test OpenQA results. Blue: Single-turn. Coral: Multi-turn. Gaps show Single-turn minus Multi-turn accuracy (pp), matching Δconv in the table. OMNI models use visual and audio evidence; LVLMs use visual evidence; LALMs use audio evidence.

Detailed scores

Model Family Long ↑ Δconv ↓ Short ↑ Δconv ↓
Qwen3.8-Omni-Flash OMNI 83.2 11.8 91.1 3.4
Doubao-Seed-2.0-Lite OMNI 63.0 33.6 89.8 6.6
Qwen3-Omni OMNI 62.9 31.7 91.6 3.0
MiMo-V2.5 OMNI 47.3 44.7 80.2 11.2
MiniCPM-o 4.5 OMNI 42.6 50.5 78.7 14.4
Nemotron-3-Nano-Omni OMNI 40.5 50.8 85.9 5.2
Qwen3.7-Plus LVLM 81.4 12.6 92.6 0.9
MiniCPM-V 4.5 LVLM 64.8 24.3 82.3 6.8
Qwen3-VL LVLM 56.7 34.5 87.9 3.1
InternVL3.5 LVLM 31.8 53.4 56.4 28.7
StepAudio 3 LALM 85.2 13.7 80.0 18.7
Audio Flamingo 3 LALM 54.2 41.7 72.6 23.2
MiMo-Audio LALM 37.4 58.2 62.7 32.2

Multi-turn accuracy (%), with conversational degradation Δconv = Accsingle − Accmulti in pp. Short / Long contain 2 / 5 media-bearing user turns, including the final turn. Rows retain the same order across interfaces, sorted by Long OpenQA accuracy within each family. Bold marks each family’s best Long score for the selected interface. Results use valid Single-turn/Multi-turn pairs.

03 / Construction & evaluation

From verified seeds
to matched conversations.

Each model generates its own historical replies. Paired Single-turn and Multi-turn evaluation measures conversational degradation across Short and Long horizons, using both OpenQA and MCQ.

ReTurn construction pipeline: verified seed questions, four paired history conditions, and matched Single-turn and Multi-turn evaluation
7,000 tasks · 3,500 matched pairs · 3,500 visual and 3,500 audio tasks. Dev: 1,500 training + 500 validation tasks. Test: 5,000 tasks.
04 / What the benchmark reveals

Accessing history is only
part of the challenge.

Matched conditions and behavioral probes distinguish access to relevant history from its application and the selection of the intended evidence.

Two history-use ladders across OpenQA and MCQ
Mean accuracy across 13 model configurations on complete valid operation pairs, pooling Short and Long. Bars overlay accuracies from a shared zero baseline; they do not add. Each model uses its supported modalities. Plotted values ↗

Two controlled contrasts

History access, then conflict resistance.

Task side. Single-turn → Reconfirm introduces a question carried by history. Reconfirm → Reground changes whether the historical answer remains valid, exposing interference from an outdated answer.

Evidence side. Single-turn → Retrieve requires using historical media. Retrieve → Rebind strengthens competition from current media, testing whether the model selects the intended source.

These contrasts motivate two diagnostic questions: can a model access the relevant history, and can it bind the request to the right turn?

01Recall ≠ application

Qwen3.8-Omni-Flash recalls the historical question in 99.8% of separate Reground probes, yet achieves 84.1% task accuracy. Restating the question raises accuracy to 96.4%.

02Competing evidence redirects answers

For Qwen3-Omni, matching the replacement-media answer rises from 0.1% to 36.5% after a Rebind media swap, while target accuracy barely changes.

03Adaptation only partly closes the gap

Qwen3-Omni SFT improves Multi-turn accuracy by 3.42 pp on OpenQA and 2.36 pp on MCQ. Long Rebind remains difficult.

05 / Resources

Resources

Read the paper, explore the dataset, or evaluate a model with ReTurn.