NTH

Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization

AuthorsHyangsuk Min, Hwanjun Song

August 17, 2026 2 min read
Watch on YouTube
The one-line take

ReMEMBER helps dialogue summarizers retrieve the missing past context needed to understand and summarize the present conversation.

Key results

900
Benchmark instances

Window-level instances drawn from LoCoMo, RealTalk21, and EverMemBench.

160K
Maximum history length

Longest dialogue histories evaluated.

0.6984
ReMEMBER memory recall

Overall fraction of annotated contextual gaps whose evidence appears in memory.

0.157
Recall gain over Hybrid

Absolute memory-recall improvement over the strongest retrieval baseline.

0.17
Maximum gap-resolution gain

Largest improvement over Hybrid in reflecting recovered evidence in summaries.

4
Memory construction runtime

Approximate seconds per construction run on an NVIDIA H200.

What the paper found

This paper defines streaming dialogue summarization as producing a self-contained summary of a recent 1,024-token window using selective memory from an unbounded history. Its central finding is that memory quality depends less on how much history is retrieved than on whether it recovers evidence missing from the window, such as antecedents, attributes, and causal relations. The authors introduce a benchmark of 900 window-level instances drawn from LoCoMo, RealTalk21, and EverMemBench, with histories reaching 160K tokens, and separate memory recall from summary gap-resolution completeness. ReMEMBER detects unresolved dependencies with Qwen3.5-4B, forms targeted evidence queries, combines BM25 with Qwen3-Embedding-0.6B through reciprocal rank fusion, refines 128-token retrieval chunks to relevant turns, and allocates them round-robin under a fixed budget. It reaches 0.6984 memory recall, improving over the Hybrid baseline by 0.157, while gap-resolution completeness improves by up to 0.17. The gains hold across Qwen3.5-4B, Qwen3.5-9B, and DeepMind’s Gemma-4-E2B-it, showing that the approach is not tied to one summarizer. ReMEMBER takes approximately 4 seconds to build memory on an NVIDIA H200, trading modest latency for substantially denser, more usable historical evidence than recency, compression, or ordinary retrieval.

Original abstract

Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context to be interpreted on its own. We formalize this setting as streaming dialogue summarization, where a system must summarize a current window using selective memory from an unbounded history under a fixed budget. We show that the central challenge is not how much history is accessed, but whether memory recovers the evidence that the current window presupposes. We construct a benchmark and evaluation protocol that separately assesses whether memory contains gap-resolving evidence and whether the generated summary reflects it. We propose ReMEMBER, a missing-evidence memory framework that conditions retrieval on unresolved window dependencies and refines retrieved chunks into evidence-dense memory under a fixed budget. Experiments on dialogues with histories up to 160K tokens show that ReMEMBER improves memory recall and gap-resolution completeness over memory construction baselines under the same budget.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →