Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
AuthorsHyangsuk Min, Hwanjun Song
Resources
ReMEMBER helps dialogue summarizers retrieve the missing past context needed to understand and summarize the present conversation.
Key results
Window-level instances drawn from LoCoMo, RealTalk21, and EverMemBench.
Longest dialogue histories evaluated.
Overall fraction of annotated contextual gaps whose evidence appears in memory.
Absolute memory-recall improvement over the strongest retrieval baseline.
Largest improvement over Hybrid in reflecting recovered evidence in summaries.
Approximate seconds per construction run on an NVIDIA H200.
What the paper found
This paper defines streaming dialogue summarization as producing a self-contained summary of a recent 1,024-token window using selective memory from an unbounded history. Its central finding is that memory quality depends less on how much history is retrieved than on whether it recovers evidence missing from the window, such as antecedents, attributes, and causal relations. The authors introduce a benchmark of 900 window-level instances drawn from LoCoMo, RealTalk21, and EverMemBench, with histories reaching 160K tokens, and separate memory recall from summary gap-resolution completeness. ReMEMBER detects unresolved dependencies with Qwen3.5-4B, forms targeted evidence queries, combines BM25 with Qwen3-Embedding-0.6B through reciprocal rank fusion, refines 128-token retrieval chunks to relevant turns, and allocates them round-robin under a fixed budget. It reaches 0.6984 memory recall, improving over the Hybrid baseline by 0.157, while gap-resolution completeness improves by up to 0.17. The gains hold across Qwen3.5-4B, Qwen3.5-9B, and DeepMind’s Gemma-4-E2B-it, showing that the approach is not tied to one summarizer. ReMEMBER takes approximately 4 seconds to build memory on an NVIDIA H200, trading modest latency for substantially denser, more usable historical evidence than recency, compression, or ordinary retrieval.
Original abstract
Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context to be interpreted on its own. We formalize this setting as streaming dialogue summarization, where a system must summarize a current window using selective memory from an unbounded history under a fixed budget. We show that the central challenge is not how much history is accessed, but whether memory recovers the evidence that the current window presupposes. We construct a benchmark and evaluation protocol that separately assesses whether memory contains gap-resolving evidence and whether the generated summary reflects it. We propose ReMEMBER, a missing-evidence memory framework that conditions retrieval on unresolved window dependencies and refines retrieved chunks into evidence-dense memory under a fixed budget. Experiments on dialogues with histories up to 160K tokens show that ReMEMBER improves memory recall and gap-resolution completeness over memory construction baselines under the same budget.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.