When the Model Retires: An Empirical Study of LLM Migration in Open-Source Applications
AuthorsHyungjin Lukas Kim
AffiliationsDepartment of Future and Convergence Business Administration, Myongji University, Seoul, Republic of Korea
Resources
Most open-source LLM applications migrate only after a provider shuts down their model, revealing a major and measurable dependency risk.
Key results
Commits mined across the open-source ecosystem.
Non-fork repositories included in the study.
Estimated share of genuine migrations committed after shutdown.
Event-level reactive share for Anthropic retirement events.
Share of migrating applications editing source code rather than only configuration or documentation.
Median added lines for migrations involving fine-tuned applications.
What the paper found
This study examines what happens when commercial LLM endpoints disappear, mining 22,555 GitHub commits across 17,703 non-fork repositories and matching migrations to official retirement events from OpenAI, Anthropic, and Google. The central finding is that 82% of genuine migrations occurred after shutdown, meaning applications often failed before developers repaired them. Notice policy strongly predicted preparedness: event-level reactive migration reached 89% for Anthropic’s short Claude retirement notices but only 13% for OpenAI’s one-year Assistants API notice; statistically, each e-fold increase in notice length reduced post-shutdown odds by roughly three quarters. Model identifiers were hard-coded in 94% of migrating applications, while abstraction layers such as LiteLLM, OpenRouter, LangChain, and provider gateways were uncommon and did not make migrations more timely. Migration effort depended sharply on architecture: prompt-only applications required a median of 6 added lines, compared with 693 for fine-tuned systems, where replacing the base model can invalidate the tuned checkpoint. Only 8% of migrations switched providers, so most applications moved from retiring GPT or Claude versions to successors within the same ecosystem. The study also finds weak operational learning: repositories affected by earlier retirements were not more proactive later, and only a small minority added evaluation or regression testing. The results frame model retirement as a lifecycle dependency problem requiring machine-readable deprecation notices, CI checks, externalized model identifiers, and fail-loudly error handling.
Original abstract
Applications built on commercial large language model (LLM) APIs depend on model versions that providers retire on their own schedule, with notice periods ranging from one year to two weeks. We ask what actually happens to applications when a model is retired. We mine GitHub for commits that migrate away from officially deprecated models and endpoints of OpenAI, Anthropic, and Google, matching each commit to the provider's published announcement and shutdown dates. From 22,555 commits in 17,703 non-fork repositories (2024-2026), 5,139 are matched to an official event; two independent coders validated a stratified sample of 300 (kappa = 0.89-0.95), and we reweight all estimates by their labels. We find that an estimated 82% (95% CI 79-84) of migrations away from retired models were committed after the shutdown date - after the application had started failing - regardless of repository popularity, prior retirement experience, or the presence of a provider-abstraction layer. The share tracks the provider's notice policy: 89% for Anthropic's 60-114-day notices versus 13% for OpenAI's one-year Assistants API notice, and each e-fold increase in notice length reduces the odds of post-shutdown migration by about three quarters. Model identifiers are hard-coded in 94% of migrating applications, migration effort scales from a median of 6 added lines for prompt-only applications to nearly 700 for fine-tuned ones, and only 8% of migrations switch provider. We release the dataset and pipeline and discuss implications for deprecation policy, dependency-risk assessment of LLM products, and tooling.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.