NTH

Periscope: Extending Frozen Language Models Beyond Their Context Window

AuthorsMohamed Eltahir, Anas Obayd, Raed Rashid, Abdulrahman Alghamdi, Abdulrahman Mousa, Abdallah Ahmed, Tanveer Hussain, Naeemullah Khan

AffiliationsKing Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia · Department of Computer Science, Edge Hill University, Ormskirk, England (a) (b) · Oct 2026 65 window read (262k) 262k Periscope (select)

October 7, 2026 3 min read
Watch on YouTube
The one-line take

Periscope lets frozen language models search million-token documents on a single GPU by replacing one enormous read with a smart grid of small probes.

Key results

41.9
LongBench v2 selected-read accuracy

Qwen3.5-4B matched its best window read while reading only 9k tokens.

9k
Selected-read budget

Tokens read in Periscope’s second-stage LongBench v2 selection.

82.1
InfiniteBench En.MC accuracy

Periscope’s selected read exceeded the best window read by 5 points.

.493
BRIGHT NDCG@10

Periscope’s best score across the six evaluated retrieval methods.

296GB
Single-pass KV cache

Key-value cache required for a 4.5M-token Qwen3.5-27B read.

3.1GB
Probe KV cache

Peak cache for Periscope probes on the same 4.5M-token context.

What the paper found

Periscope is a training-free inference method that lets frozen language models process texts beyond their context window without changing weights or passing hidden state between calls. It divides a document into chunks, arranges them on a square grid, and issues 2K one-token probes: K local spans of consecutive chunks and K strided spans sampling the entire document. Each probe returns answer log-odds, and combining the strongest local and strided evidence produces an evidence map that can answer a finite-choice question directly, rank documents, locate supporting chunks, or select a compact second-stage read. Because each probe is about the square root of the document length, attention cost grows approximately as s^1.5 rather than s^2, while peak key-value memory remains bounded by one probe. On LongBench v2, Qwen3.5-4B achieved 41.9 accuracy using a selected read of only 9k tokens, matching its best window read across windows from 32k to 1M tokens. On InfiniteBench En.MC, the same model reached 82.1, exceeding the best window read by 5 points. On BRIGHT, Qwen3-4B achieved .493 NDCG@10, the best among six methods, showing that model-based probing can outperform embedding retrieval on reasoning-intensive relevance. The approach also works with Google DeepMind’s Gemma-4 and Qwen3.5-27B: a 27B model processed a 4.5M-token context with 3.1GB of probe cache, versus 296GB for a single-pass read. Its limitation is finite answer spaces and weaker performance when evidence must be jointly synthesized across multiple documents.

Original abstract

A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which passage is the evidence. Periscope, a training-free inference method, arranges the $N$ chunks of a text on a $K{\times}K$ grid with $K{=}\lceil\sqrt{N}\rceil$ and asks a frozen model the same question about $K$ local spans of consecutive chunks and $K$ strided spans that sample the whole text, reading the log-odds of every answer at one token. Each answer takes its best local and strided score, and scoring every chunk by its two spans gives an evidence map at no further cost, whose peak is the chunk behind the answer. Every probe is about $\sqrt{sc}$ tokens for a text of $s$ tokens and chunk size $c$, so a window of $W$ tokens reaches $W^{2}/c$ tokens at $s^{1.5}$ cost. The map replaces the long read. On LongBench v2, reading only the $K$ chunks the map ranks highest, 9k tokens, matches the same model's best window read across windows from 32k to 1M tokens, and on InfiniteBench, where the median context is 150k tokens, it leads the best window read by 5 points. The same map ranks BRIGHT's long-document corpora with the best NDCG@10 of six methods. Each call caches only one probe, so a 27B model reads 4.5M-token contexts on one 80GB GPU, where a single pass would need 296GB of cache. A long read then needs a GPU that holds the model, not one that holds the text.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis