Local Ai
Frozen DINO Localizes Image Edits Without a Localizer
arXiv:2608.18968v1 Announce Type: new Abstract: Localized image edits can change a photograph's meaning while leaving most of it authentic, so forensic analysis must identify where an edit occurred. W
arXiv:2608.18968v1 Announce Type: new Abstract: Localized image edits can change a photograph's meaning while leaving most of it authentic, so forensic analysis must identify where an edit occurred. We show that patch-level perturbation responses from frozen DINO encoders are themselves localization maps. Training-free Localization of AI-image Edits from patch-token Drift (TRAIL) applies one global Haar perturbation and maps cosine drift between corresponding patch tokens. On 80 source-disjoint CocoGlide test images, TRAIL reaches .903 patch AUROC versus .912 for the mask-supervised Detective SAM; fixed-threshold Dice is .619 versus .709, while an oracle threshold raises TRAIL to .790. Transferred unchanged to Poisson image interpolation, TRAIL reaches .855 AUROC versus .864, showing that the cue persists without a generator. Across sixteen DINO encoders, the best block lies at normalized depth .80-.94. Global context matters: AUROC falls from .903 globally to .857 for local-in-canvas perturbations and .735 for independently encoded crops. Frozen DINO patch tokens therefore contain a strong late-layer localization signal whose visibility depends on the perturbation and preserved context. Code: https://github.com/VishalJ99/trail-image-edit-localization.
Related
- EditSleuth: A Dataset of Grounded Reasoning Chains for Image-Edit Forensics
- Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback
- Synthetic Aperture Radar Image Change Detection Based on Global Dynamic Context-Aware Network
Source: arXiv cs.CV | 2026-08-20