Toward Parking Spot Occupancy Recognition: A Self-Supervised Approach
AuthorsLuan Marko Kujavski, Rayson Laroca, Paulo Lisboa de Almeida
This paper shows that self-supervised learning can remove the need for target-lot labels while still achieving very high parking-spot occupancy accuracy, making deployment cheaper and more scalable.
Key results
Average accuracy across the three-dataset leave-one-out evaluation
Average accuracy after switching to the Specialized Model after day 7
Average accuracy of the supervised transfer-learning baseline
Average accuracy of the self-supervised transfer-learning baseline
Approximate GPU hours needed to train a Strong General Model or Specialized Model
Seconds per parking-space image on a Raspberry Pi 5
What the paper found
Toward Parking Spot Occupancy Recognition: A Self-Supervised Approach proposes a label-efficient pipeline for parking-space empty/occupied recognition built on SimCLR with a ResNet-50 encoder, where the model is first self-supervised on ImageNet, then self-supervised again on unlabeled parking-lot imagery, and finally supervised fine-tuned only on labeled source-domain parking data. Evaluated in a leave-one-out cross-environment protocol on PKLot, CNRPark-EXT, and PLds, the method introduces two deployment modes: a Strong General Model for the first 7 days and a Specialized Model trained from unlabeled target-lot data collected during those days. The Strong General Model reaches 97.2% average accuracy, while the two-stage deployment improves that to 97.8%, outperforming both supervised and self-supervised transfer-learning baselines at 97.1% and 97.0%. On labeled-target adaptation, the Specialized Model reaches 98.8% with 1,000 target labels, exceeding the 97% reference point reported by prior deployment-strategy work. The approach is computationally practical for edge use: training takes about 25 GPU hours, and inference on a Raspberry Pi 5 takes 0.2282 seconds per parking-space image, with 0.2140 seconds spent in inference and the remainder in preprocessing.
Original abstract
As urban areas expand, automatic monitoring of parking lots becomes essential for efficient and sustainable cities. This work proposes a self-supervised approach for parking spot occupancy recognition that requires no labeled samples from the target parking lot. Building upon a self-supervised transfer learning fine-tuning protocol, the proposed training strategy consists of two self-supervised stages: first on unlabeled generic data and then on unlabeled target-specific data, followed by supervised fine-tuning using only generic parking lot labels. We adopt SimCLR with a ResNet-50 encoder and evaluate the method under a leave-one-out cross-environment protocol on three public datasets: PKLot, CNRPark-EXT, and PLds. We also introduce a two-stage deployment strategy in which a Strong General Model is initially deployed, followed by a Specialized Model that incorporates unlabeled images collected during the first N days of deployment in a self-supervised manner. Experimental results show that the Strong General Model alone outperforms supervised and self-supervised baselines, achieving an average accuracy of 97.2%, which further improves to 97.8% with the proposed two-stage strategy. These results demonstrate that self-supervised learning enables a scalable and labelefficient solution for real-world parking occupancy monitoring. Our trained models and source code are publicly available at https://github.com/LoanMaikon/Parking-Spot-Occupancy-Recognition.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.