End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery
AuthorsAkira Ito, Takayuki Miura, Yosuke Todo
Resources
A new query-efficient technique makes it possible to steal the parameters of small black-box neural networks using only their predicted labels.
Key results
Intersection spaces used per evaluated neuron.
Correctly recovered out of 510 evaluated signs in the MNIST-trained 784-128(5)-10 model.
Label agreement between extracted and victim models.
Label agreement between extracted and victim models.
Label agreement between extracted and victim models.
Label agreement between extracted and victim models.
What the paper found
This paper presents the first fully black-box, end-to-end extraction of trained deep ReLU multilayer perceptrons using only hard labels, meaning the attacker sees class predictions but not logits or gradients. Its central innovation is the cosine sign-recovery method, especially signature-weighted cosine, which reuses intersection spaces already collected during signature recovery, applies pseudoinverse projection and whitening against competing neuron signatures, and requires no dedicated sign-recovery queries. With 50 intersection spaces per neuron, it correctly recovered 506 of 510 evaluated signs in an MNIST-trained 784-128(5)-10 model, while the prior boundary-walking approach produced far fewer correct signs at the same sample count. The complete pipeline addresses deeper-layer failures through decision-boundary validation, active-weighted cosine correction, projection-based sign recovery, adaptive searches for missing coordinates, and cross-layer extraction of unreachable weights. On trained MNIST and Fashion-MNIST models with 784 inputs, 4 or 6 hidden layers of width 16, and 10 outputs, it recovered every active neuron and achieved label agreements of 98.477% and 99.948% for MNIST, and 100.000% and 99.744% for Fashion-MNIST, evaluated on 100,000 standard-Gaussian inputs. The work targets model confidentiality risks relevant to deployed systems such as OpenAI services, although its experiments use compact ReLU classifiers; OpenAI Codex (GPT-5.6) was used only for implementation assistance, debugging, and code review.
Original abstract
The importance of deep neural networks (DNNs) is widely recognized, and the parameters obtained through training are regarded as valuable assets. Recently, attacks that extract these parameters using only oracle queries to a DNN have been actively studied at IACR conferences. The hard-label setting is the most challenging setting for model extraction, where an adversary can observe only the final output label, such as "dog" or "cat." At Eurocrypt 2025, Carlini et al. proposed polynomial-time hard-label extraction of ReLU-based MLPs. However, one step of this attack process, i.e., sign recovery, requires a large number of queries and substantial computation. Implementing this step in a black-box setting remains difficult. Consequently, a fully black-box end-to-end demonstration on trained deep ReLU MLPs has remained a challenge. In this paper, we propose a new sign-recovery algorithm based on a completely different principle from the existing method. Our method requires no dedicated queries for sign recovery. In our experiments, it achieves higher sign-recovery accuracy than the existing method. Consequently, it enables efficient sign recovery even for trained models. With our sign-recovery algorithm, all steps of hard-label model extraction can be implemented in a black-box setting. By combining these implementations, we demonstrate end-to-end model extraction from models trained on MNIST and Fashion-MNIST, with width 16 and 4 or 6 hidden layers, achieving over 98% label agreement.
Read the original paperMore in Neural Networks
Browse all 22 papers →Retrieving Individual Stems from Music Mixtures with Slot Embeddings
David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein
Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.
The Linear Representation Hypothesis Needs a Group Action
Louie Hong Yao, Yuhao Li, Shengchao Liu
This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.
HypLTSF: A Hyperbolic Geometric View of Multi-Scale Hierarchies for Long-Term Time Series Forecasting
Namwoo Kim, Hyungryul Baik, Yoonjin Yoon
HypLTSF maps multi-scale time-series patterns into hyperbolic space so that fine-to-coarse temporal hierarchies become explicit and useful for long-term forecasting.