Invisible Shortcuts: Why Vision Encoders Know Your Camera
Abstract
Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-background or texture correlations. We identify a different source of shortcut learning: invisible metadata traces embedded at the pixel level, for metadata such as image processing and photo acquisition. We hypothesize that large-scale semantic supervision, whether through categorical labels (ImageNet) or billion-scale captions (LAION), naturally induces metadata-semantics correlations during pretraining, leading models to convert low-level signals into predictive features. By introducing controlled metadata-semantics correlations, we show that stronger ones produce systematically higher sensitivity to metadata traces and larger performance degradation under metadata distribution shifts. We further explore mitigation strategies applied during and after pretraining that reduce sensitivity not only to targeted metadata but also to unseen ones, without sacrificing performance on downstream tasks. Metadata sensitivity also has a positive side: it partly explains the strong generated-image detection ability of some encoders, while its mitigation can improve out-of-distribution generalization. Code: https://github.com/ryan-caesar-ramos/visual-encoder-traces
Community
Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-background or texture correlations. We identify a different source of shortcut learning: invisible metadata traces embedded at the pixel level, for metadata such as image processing and photo acquisition. We hypothesize that large-scale semantic supervision, whether through categorical labels (ImageNet) or billion-scale captions (LAION), naturally induces metadata-semantics correlations during pretraining, leading models to convert low-level signals into predictive features. By introducing controlled metadata-semantics correlations, we show that stronger ones produce systematically higher sensitivity to metadata traces and larger performance degradation under metadata distribution shifts. We further explore mitigation strategies applied during and after pretraining that reduce sensitivity not only to targeted metadata but also to unseen ones, without sacrificing performance on downstream tasks. Metadata sensitivity also has a positive side: it partly explains the strong generated-image detection ability of some encoders, while its mitigation can improve out-of-distribution generalization.
Does this mean every model I've trained on scraped web images has been quietly memorizing camera fingerprints, not just semantics? If so, the fix isn't just better augmentation — it's asking whether the shortcut survives when you strip EXIF and re-encode. I'd love to see how much of the effect persists through JPEG recompression and resizing, because that's what actually happens to images before they hit most training pipelines. If it survives that, this is a much bigger deal than a dataset hygiene footnote.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Unleashing the Potential of Vision-Language Models for Generalizable AI-Generated Image Detection (2026)
- Domain Generalization via Text-Anchored Information Bottleneck (2026)
- Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval (2026)
- Dissect and Prune: Enhancing Robustness in AI-Generated Image Detection (2026)
- Spatially Localized Image Degradation Embeddings for Image Quality Assessment (2026)
- Representation and Reference Selection in Training-Free Synthetic Image Attribution (2026)
- TaskTok: Delving into Task Tokens for Task-driven Image Restoration (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper