NeurIPS 2026

What Should a Streaming Video Model Remember?

SelectStream — a selective latent-memory framework that keeps the current observation visible to a frozen VLM and exposes history only through a compact, query-conditioned evidence budget.

Haonan Ge1,2 Yiwei Wang2 Hang Wu2 Yujun Cai3,†

1University of California, Santa Barbara 2University of California, Merced 3The University of Queensland

†Corresponding author

Motivation: sparse surprise peaks over a long stream, a query-conditioned memory graph, and a radar chart comparing SelectStream with HERMES, StreamForest and Flash-VStream.
Figure 1. Relevant events may appear sparsely over a long stream, while later queries need evidence beyond the recent context. In this counting example, inserted movie clips form sparse surprise peaks. SelectStream strengthens long-range memory and efficiency without sacrificing current-scene perception.

Abstract

Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation budgets. Existing methods address this by adding memory banks, retrieval modules, or visual token compression to preserve long-range history. However, strong recent-window baselines show that indiscriminate history injection can dilute current-scene perception, suggesting that the key challenge is not whether to use memory, but how to allocate it selectively. We formulate this as budgeted online latent evidence allocation and propose SelectStream, a selective latent-memory framework that keeps the current observation directly visible to a frozen VLM while exposing historical information only through a compact, query-conditioned evidence budget. Three coordinated mechanisms govern when to write, what to preserve, and how to retrieve: surprise-driven adaptive windowing, priority-preserving consolidation, and query-conditioned graph reasoning over a fixed-capacity latent memory graph. Retrieved evidence is calibrated and injected as latent tokens for answer generation, without replaying frames or growing the context with stream length. Experimental results show that SelectStream achieves strong online streaming performance and preserves general video understanding, reaching 82.67% on StreamingBench, 67.03% on OVO-Bench, and 74.4% average accuracy on offline video benchmarks, while outperforming strong recent-window baselines and prior streaming memory methods.

82.67
StreamingBench
67.03
OVO-Bench overall
74.4
Offline average (VideoMME / MLVU / MVBench)

Method

SelectStream overview: surprise-driven adaptive windowing, latent visual memory with a dynamic memory graph, graph attention reasoning, and latent evidence injection into a frozen MLLM.
Figure 2. SelectStream writes projected VLM visual embeddings into a budgeted latent memory graph, retrieves a query-conditioned evidence subgraph, and injects calibrated latent evidence tokens into a frozen MLLM for answer generation.
When to write

Surprise-driven Adaptive Windowing

Surprise combines attention shift and feature change. A segment closes on a surprise spike, accumulated surprise energy, or the maximum length — coarse entries for stable intervals, finer ones near transitions.

What to preserve

Latent Visual Memory

Segments are encoded and written into a fixed-capacity memory graph with gated updates. When the budget N is exceeded, priority-preserving consolidation merges the pair with the smallest graph-aware penalty, protecting surprising, frequently read, and recent evidence.

How to read

Graph Attention Reasoning

The query scores active nodes, routes through temporal and semantic neighborhoods within a subgraph budget B, and refines them with relational graph attention before selecting the top-M nodes.

How much to expose

Latent Evidence Injection

Only M calibrated evidence tokens join the prompt and current observation. There is no frame replay or text summary, and the context does not grow with stream length.

Results

Model#FramesStreamingBenchOVO RT Avg.OVO BT Avg.OVO RT/BT Avg.OVO Overall
Offline video LLMs
Qwen2.5-VL-7B1 fps73.3159.9044.7052.28–
LLaVA-OneVision-7B3271.1264.0043.7053.8552.74
Online / streaming video LLMs
Flash-VStream-7B1 fps23.2328.4027.4027.9033.61
StreamForest-7B1 fps77.2661.2052.0056.60–
Streamo-7B2 fps–67.4449.1858.3157.86
HERMES-7B1 fps79.4469.0049.4059.20–
ThinkStream-7B1 fps75.0069.1260.6864.90–
Qwen2.5-VL-7B + 4f (recent window)1 fps78.4778.4051.9065.13–
Qwen3-VL-8B + 4f (recent window)1 fps80.5981.4054.0067.70–
SelectStream-Qwen2.5-VL-7B1 fps81.4280.8561.0570.9565.71
SelectStream-Qwen3-VL-8B1 fps82.6782.7662.2072.4867.03

Online streaming benchmarks (subset of Table 1). RT: Real-Time Visual Perception; BT: Backward Tracing. The largest gains appear on Backward Tracing, which directly tests the use of prior visual context, while current-scene perception is retained.

Model#FramesVideoMMEMLVUMVBenchAvg.
LLaVA-Video-7B6463.370.858.664.2
StreamForest-7B1 fps61.470.070.267.2
Qwen2.5-VL-7Bmax 76865.170.269.668.3
Qwen3-VL-8B2 fps, max 204871.478.168.772.7
SelectStream-Qwen2.5-VL-7B1 fps, max 102467.873.070.470.4
SelectStream-Qwen3-VL-8B1 fps, max 102473.280.069.974.4

Offline video generalization (subset of Table 2), evaluated with a causal final-query protocol. SelectStream preserves the backbone's general video understanding while adding long-range evidence.

Variant (SelectStream-Qwen2.5-VL-7B)StreamingBenchOVO-BenchMLVU
Memory allocation
Fixed segments w/o SAW79.8664.2171.2
w/o gated writing80.2164.4871.7
FIFO consolidation79.4263.6170.4
Similarity-only merging80.0364.0371.0
Evidence readout
Top-k retrieval w/o GAR79.7163.9271.1
Fixed-hop expansion80.0964.2871.6
w/o evidence calibration78.9663.0870.2
w/o Lret80.2764.4271.8
w/o Lspar80.7664.8272.3
Full SelectStream81.4265.7173.0

Ablations (Tables 3 and 4) under the same node and evidence budgets. Every write, consolidation and readout component contributes; removing evidence calibration causes the largest drop.

Budgets and Efficiency

Budget sensitivity: accuracy, Recall@M and latency as memory capacity N, subgraph budget B and evidence budget M vary.
Figure 3. Sensitivity to memory capacity N, subgraph budget B and evidence budget M. Larger N improves retention, B mainly improves retrieval recall before saturating, and M exposes more evidence at a higher time to first token.
Streaming latency (TTFT) and peak GPU memory versus processed frames for Flash-VStream, StreamForest, TimeChat-Online, HERMES and SelectStream.
Figure 4. Efficiency scaling with stream length. Query latency and peak GPU memory stay nearly flat, since the stream is compressed into at most N active nodes and each query materializes only a bounded subgraph and M evidence tokens.

BibTeX

@article{ge2026should,
  title={What Should a Streaming Video Model Remember?},
  author={Ge, Haonan and Wang, Yiwei and Wu, Hang and Cai, Yujun},
  journal={arXiv preprint arXiv:2606.16353},
  year={2026}
}