NTH

When Can You Correct Distribution Drift in Temporal Graph Generation? A Sharpening--Drift Tension and an Impossibility for Observation-Based Correction

AuthorsTianpeng Li, Xuan Guo, Wenjun Wang, Wang Zhang, Pengfei Jiao

July 29, 2026 3 min read
Watch on YouTube
The one-line take

Temporal graph generators cannot reliably correct deployment drift from observations alone, because unseen structural changes create an irreducible error floor.

Key results

0.605
Sharpening-drift power-law exponent

Magnitude of the empirical exponent linking deployment loss to in-period loss

0.9977
Power-law fit R2

Fit quality for the sharpening-drift trade-off

50
Sampling-budget range

Upper end of the 1-to-50-step sweep

34.3
Maximum drift floor ratio

Largest drift-period versus in-period joint-error floor ratio

60%
Oracle error reduction

Marginal-error reduction from knowing current target marginals

5.7%
Observation-based gain recovered

Fraction of the oracle gain recovered using past observations

What the paper found

Researchers at Tianjin University and Hangzhou Dianzi University analyze why temporal graph generators trained on one period degrade on the next. For masked flow matching, they prove an exact decomposition of loss into irreducible conditional entropy plus a divergence from the deployment distribution, without assuming independent graph slots. Along a source-training sharpening path, the divergence increases specifically for structures rare in training but common later, with the empirical trade-off following a power law of exponent −0.605 and R2 = 0.9977. Across seven conditions spanning Enron, MOOC, Reddit, and Wikipedia, changing sampling from 1 to 50 steps alters drift-period marginal error by at most 6.0%, while the converged joint-error floor is 2.2× to 34.3× higher than in-period performance, showing that drift changes the destination rather than sampling speed. The central impossibility theorem states that any correction based only on past unlabeled observations retains at least the target statistic’s conditional variance. An oracle knowing current marginals removes 60% of marginal error, but the best observation-based correction recovers only 5.7% of that gain. Because measured drift is trendless and mean-reverting, extrapolation is worse than persistence, with error ratios of 1.27× to 1.51×. DiGress shows the same directional degradation, though its weaker sharpening does not reproduce the power law. The paper therefore directs future remedies toward exogenous side information or deliberately less-confident generators, rather than more test-time adaptation.

Original abstract

Generative models of temporal graphs are trained on one stretch of an evolving network and deployed on the next, and they degrade badly in the gap. We show this degradation is derivable, general, and not fixable from observations. The masked flow-matching loss decomposes exactly, with no independence assumption, into an irreducible entropy plus a divergence whose derivative along the training path is positive precisely for structures rare during training and common at deployment, diverging as their training probability goes to zero. Empirically the trade-off is a power law with exponent $-0.605$ ($R^2=0.9977$), and drift raises the sampler's error floor without changing how many steps reach it: across seven well-powered conditions the drift-period marginal error varies by at most $6\%$ over a $50\times$ range of sampling budgets, while the floor sits $2.2\times$ to $34.3\times$ above the in-period floor. Because the deployment period is observed, correction looks like a matter of measurement. It is not. We prove that any corrector measurable with respect to past observations leaves at least the conditional variance of the statistic it tracks, and that trend extrapolation beats trusting the last observation only when $μ^2>v(1-2ρ)$. Both premises are measurable and both go the wrong way: the drift is trendless and mean-reverting, with a one-step innovation as large as the drift itself. An oracle removes $60\%$ of the error, the best observation-based corrector recovers $5.7\%$ of that, and extrapolation is strictly worse than doing nothing clever.

Read the original paper

More in Graph Learning

Browse all 32 papers →
02Graph Learning

GraphWrit3R: End-to-End 3D Scene Graph Writing

Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel

GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.

Read analysis