Consistency-Preserving Logit Shaping for Robust Model Stealing Defense with Applications in Wireless Spectrum Security

Abstract

Model stealing (model extraction) threatens ML-as-a-service APIs by enabling adversaries to reconstruct proprietary models from queried probability outputs, undermining both intellectual property and privacy. We seek a defense that reduces information leakage without harming utility. We introduce Consistency-Preserving Logit Shaping (CPLS), a simple, margin-adaptive perturbation applied to logits that provably preserves the top-1 label while reducing mutual information in released soft labels. CPLS adds deterministic, input-keyed, class-orthogonal noise bounded by a fraction of the decision margin, yielding closed-form argmax invariance and resistance to expectation-over-transformation averaging. CPLS has a wide range of applications, and achieves up to 13.1% relative mutual information reduction with zero accuracy loss and zero flip rate, and degrades surrogate training on both MNIST and CIFAR-10 datasets under standard knowledge distillation protocols. We further evaluate CPLS on a wireless spectrum sensing dataset, demonstrating generalization to non-image domains. These results suggest CPLS offers a practical, theory-backed mechanism for curbing extraction while retaining model utility.

Publication
The ACM Workshop on Wireless Security and Machine Learning (WiSec-WiseML) 2026
Click the Cite button above to demo the feature to enable visitors to import publication metadata into their reference management software.
Create your slides in Markdown - click the Slides button to check out the example.

Add the publication’s full text or supplementary notes here. You can use rich formatting such as including code, math, and images.