The rapid advancement of generative models poses significant challenges to reliable forgery detection, as existing methods often lack generalization by overfitting to a single type of artifact. We observe that generative processes introduce clues at two distinct levels: low-level forensic traces from up-sampling and high-level semantic features from the generative model. To exploit this observation, we propose a novel forensic–semantic framework that analyzes forgery evidence across these different levels. At the forensic level, the framework introduces a residual analysis method that isolates up-sampling artifacts. By computing the difference between an image and its re-sampled version, this approach effectively removes content-related information to focus on critical forgery traces. Simultaneously, at the semantic level, it employs a vision-language model fine-tuned with low-rank adaptation (LoRA) to capture the inherent semantic features. Evaluated on 30 sub-datasets, our method outperforms representative approaches by an average of 2.1% in accuracy (ACC) and 1.8% in average precision (AP), demonstrating strong generalization.

All-around forgery clues for generalizable AI-generated image detection / Zhou X., Fei J., Dai Y., Yu P., Xie J., Liu X., Xia Z., Piva A.. - In: PATTERN RECOGNITION. - ISSN 0031-3203. - ELETTRONICO. - 181:(2027), pp. 114661.0-114661.0. [10.1016/j.patcog.2026.114661]

All-around forgery clues for generalizable AI-generated image detection

Fei J.;Piva A.
2027

Abstract

The rapid advancement of generative models poses significant challenges to reliable forgery detection, as existing methods often lack generalization by overfitting to a single type of artifact. We observe that generative processes introduce clues at two distinct levels: low-level forensic traces from up-sampling and high-level semantic features from the generative model. To exploit this observation, we propose a novel forensic–semantic framework that analyzes forgery evidence across these different levels. At the forensic level, the framework introduces a residual analysis method that isolates up-sampling artifacts. By computing the difference between an image and its re-sampled version, this approach effectively removes content-related information to focus on critical forgery traces. Simultaneously, at the semantic level, it employs a vision-language model fine-tuned with low-rank adaptation (LoRA) to capture the inherent semantic features. Evaluated on 30 sub-datasets, our method outperforms representative approaches by an average of 2.1% in accuracy (ACC) and 1.8% in average precision (AP), demonstrating strong generalization.
2027
181
0
0
Zhou X.; Fei J.; Dai Y.; Yu P.; Xie J.; Liu X.; Xia Z.; Piva A.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in FLORE sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificatore per citare o creare un link a questa risorsa: https://hdl.handle.net/2158/1486672
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact