Ünsal Öztürk, Sébastien Marcel
Face recognition models place each face as a point on a unit sphere, pulling images of the same person together and pushing different people apart. Here is the subtle leak: the training objective only tightens clusters for identities the model actually saw, so people who were in the training set can end up with differently shaped clusters than those who were not. That difference is a privacy signal about who was used to train the model.
The authors measure it across 180 models built in a factorial design over backbone size, loss, training length, and number of identities, using four cluster-geometry statistics. The biggest factor by far is how many identities were in training, and oddly the signal shrinks as you add more. Cross-domain test sets can make the leak look bigger than it really is.
This reflects the abstract, so see the paper for the exact statistics and results.
Face recognition models represent each face as an embedding vector on the unit hypersphere by clustering embeddings of the same identity while pushing different identities apart through angular-margin losses. Because these losses act only on training identities, non-member identities may form clusters with different geometric properties. In this paper, we quantify the magnitude of this difference and what training-time factors control it. We compute four statistics based on…
JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models
arXiv (cs.CL) · July 17, 2026Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D
arXiv (cs.AI) · July 17, 2026LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization
arXiv (cs.AI) · July 17, 2026Spatial Normalization for Cross-Domain Retinal Layer Segmentation in Optical Coherence Tomography
Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.