Periscope: Extending Frozen Language Models Beyond Their Context Window
AuthorsMohamed Eltahir, Anas Obayd, Raed Rashid, Abdulrahman Alghamdi, Abdulrahman Mousa, Abdallah Ahmed, Tanveer Hussain, Naeemullah Khan
AffiliationsKing Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia · Department of Computer Science, Edge Hill University, Ormskirk, England (a) (b) · Oct 2026 65 window read (262k) 262k Periscope (select)
Resources
Periscope lets frozen language models search million-token documents on a single GPU by replacing one enormous read with a smart grid of small probes.
Key results
Qwen3.5-4B matched its best window read while reading only 9k tokens.
Tokens read in Periscope’s second-stage LongBench v2 selection.
Periscope’s selected read exceeded the best window read by 5 points.
Periscope’s best score across the six evaluated retrieval methods.
Key-value cache required for a 4.5M-token Qwen3.5-27B read.
Peak cache for Periscope probes on the same 4.5M-token context.
What the paper found
Periscope is a training-free inference method that lets frozen language models process texts beyond their context window without changing weights or passing hidden state between calls. It divides a document into chunks, arranges them on a square grid, and issues 2K one-token probes: K local spans of consecutive chunks and K strided spans sampling the entire document. Each probe returns answer log-odds, and combining the strongest local and strided evidence produces an evidence map that can answer a finite-choice question directly, rank documents, locate supporting chunks, or select a compact second-stage read. Because each probe is about the square root of the document length, attention cost grows approximately as s^1.5 rather than s^2, while peak key-value memory remains bounded by one probe. On LongBench v2, Qwen3.5-4B achieved 41.9 accuracy using a selected read of only 9k tokens, matching its best window read across windows from 32k to 1M tokens. On InfiniteBench En.MC, the same model reached 82.1, exceeding the best window read by 5 points. On BRIGHT, Qwen3-4B achieved .493 NDCG@10, the best among six methods, showing that model-based probing can outperform embedding retrieval on reasoning-intensive relevance. The approach also works with Google DeepMind’s Gemma-4 and Qwen3.5-27B: a 27B model processed a 4.5M-token context with 3.1GB of probe cache, versus 296GB for a single-pass read. Its limitation is finite answer spaces and weaker performance when evidence must be jointly synthesized across multiple documents.
Original abstract
A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which passage is the evidence. Periscope, a training-free inference method, arranges the $N$ chunks of a text on a $K{\times}K$ grid with $K{=}\lceil\sqrt{N}\rceil$ and asks a frozen model the same question about $K$ local spans of consecutive chunks and $K$ strided spans that sample the whole text, reading the log-odds of every answer at one token. Each answer takes its best local and strided score, and scoring every chunk by its two spans gives an evidence map at no further cost, whose peak is the chunk behind the answer. Every probe is about $\sqrt{sc}$ tokens for a text of $s$ tokens and chunk size $c$, so a window of $W$ tokens reaches $W^{2}/c$ tokens at $s^{1.5}$ cost. The map replaces the long read. On LongBench v2, reading only the $K$ chunks the map ranks highest, 9k tokens, matches the same model's best window read across windows from 32k to 1M tokens, and on InfiniteBench, where the median context is 150k tokens, it leads the best window read by 5 points. The same map ranks BRIGHT's long-document corpora with the best NDCG@10 of six methods. Each call caches only one probe, so a 27B model reads 4.5M-token contexts on one 80GB GPU, where a single pass would need 296GB of cache. A long read then needs a GPU that holds the model, not one that holds the text.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.