Ünsal Öztürk, Sébastien Marcel
Face recognition models place each face as a point on a unit sphere, pulling images of the same person together and pushing different people apart. Here is the subtle leak: the training objective only tightens clusters for identities the model actually saw, so people who were in the training set can end up with differently shaped clusters than those who were not. That difference is a privacy signal about who was used to train the model.
The authors measure it across 180 models built in a factorial design over backbone size, loss, training length, and number of identities, using four cluster-geometry statistics. The biggest factor by far is how many identities were in training, and oddly the signal shrinks as you add more. Cross-domain test sets can make the leak look bigger than it really is.
This reflects the abstract, so see the paper for the exact statistics and results.
Face recognition models represent each face as an embedding vector on the unit hypersphere by clustering embeddings of the same identity while pushing different identities apart through angular-margin losses. Because these losses act only on training identities, non-member identities may form clusters with different geometric properties. In this paper, we quantify the magnitude of this difference and what training-time factors control it. We compute four statistics based on…
CanonicalPhys: Pose-Robust Remote Photoplethysmography via Canonical-Space Priors
arXiv (cs.AI) · July 17, 2026Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI
arXiv (cs.AI) · July 17, 2026A Formally Grounded ODRL Evaluator: Implementation and Comparison
arXiv (cs.LG) · July 17, 2026DebrisTracer: Reliable Tracking in Hypervelocity Impact Fast Imaging