NTH

End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery

AuthorsAkira Ito, Takayuki Miura, Yosuke Todo

September 29, 2026 2 min read
Watch on YouTube
The one-line take

A new query-efficient technique makes it possible to steal the parameters of small black-box neural networks using only their predicted labels.

Key results

50
Sign-recovery samples

Intersection spaces used per evaluated neuron.

506
Correct signs

Correctly recovered out of 510 evaluated signs in the MNIST-trained 784-128(5)-10 model.

98.477%
MNIST agreement, 4 hidden layers

Label agreement between extracted and victim models.

99.948%
MNIST agreement, 6 hidden layers

Label agreement between extracted and victim models.

100.000%
Fashion-MNIST agreement, 4 hidden layers

Label agreement between extracted and victim models.

99.744%
Fashion-MNIST agreement, 6 hidden layers

Label agreement between extracted and victim models.

What the paper found

This paper presents the first fully black-box, end-to-end extraction of trained deep ReLU multilayer perceptrons using only hard labels, meaning the attacker sees class predictions but not logits or gradients. Its central innovation is the cosine sign-recovery method, especially signature-weighted cosine, which reuses intersection spaces already collected during signature recovery, applies pseudoinverse projection and whitening against competing neuron signatures, and requires no dedicated sign-recovery queries. With 50 intersection spaces per neuron, it correctly recovered 506 of 510 evaluated signs in an MNIST-trained 784-128(5)-10 model, while the prior boundary-walking approach produced far fewer correct signs at the same sample count. The complete pipeline addresses deeper-layer failures through decision-boundary validation, active-weighted cosine correction, projection-based sign recovery, adaptive searches for missing coordinates, and cross-layer extraction of unreachable weights. On trained MNIST and Fashion-MNIST models with 784 inputs, 4 or 6 hidden layers of width 16, and 10 outputs, it recovered every active neuron and achieved label agreements of 98.477% and 99.948% for MNIST, and 100.000% and 99.744% for Fashion-MNIST, evaluated on 100,000 standard-Gaussian inputs. The work targets model confidentiality risks relevant to deployed systems such as OpenAI services, although its experiments use compact ReLU classifiers; OpenAI Codex (GPT-5.6) was used only for implementation assistance, debugging, and code review.

Original abstract

The importance of deep neural networks (DNNs) is widely recognized, and the parameters obtained through training are regarded as valuable assets. Recently, attacks that extract these parameters using only oracle queries to a DNN have been actively studied at IACR conferences. The hard-label setting is the most challenging setting for model extraction, where an adversary can observe only the final output label, such as "dog" or "cat." At Eurocrypt 2025, Carlini et al. proposed polynomial-time hard-label extraction of ReLU-based MLPs. However, one step of this attack process, i.e., sign recovery, requires a large number of queries and substantial computation. Implementing this step in a black-box setting remains difficult. Consequently, a fully black-box end-to-end demonstration on trained deep ReLU MLPs has remained a challenge. In this paper, we propose a new sign-recovery algorithm based on a completely different principle from the existing method. Our method requires no dedicated queries for sign recovery. In our experiments, it achieves higher sign-recovery accuracy than the existing method. Consequently, it enables efficient sign recovery even for trained models. With our sign-recovery algorithm, all steps of hard-label model extraction can be implemented in a black-box setting. By combining these implementations, we demonstrate end-to-end model extraction from models trained on MNIST and Fashion-MNIST, with width 16 and 4 or 6 hidden layers, achieving over 98% label agreement.

Read the original paper

More in Neural Networks

Browse all 22 papers →
01Neural Network

Retrieving Individual Stems from Music Mixtures with Slot Embeddings

David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein

Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.

Read analysis
02Neural Network

The Linear Representation Hypothesis Needs a Group Action

Louie Hong Yao, Yuhao Li, Shengchao Liu

This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.

Read analysis