ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation
| Source: arXiv AI
Tags: AI image detection, ChatGPT, synthetic media, deepfake detection, provenance
Researchers collected 3,478 images from 2,440 posts within 51 hours of the ChatGPT Images 2.5 launch and tested 6 AI image detectors — finding detection rates ranging from 3.7% to 56.4%, far below the 83-100% recall these detectors claim on benchmark datasets.
Details
When image tools change their underlying generators while keeping the same product name, attribution of AI-generated images becomes ambiguous. This paper documents the ChatGPT Images 2.5 launch window to study this problem at scale. The frozen dataset contains 3,478 images from 2,440 posts across 8 sources, all recorded within the first 51.1 hours after the product announcement. Attribution information comes from caption claims and host records — not independently verified generator identity. The content profile skews toward fantasy scenes, partly because NightCafe (one source) contributes 39% of images but 77% of CLIP-assigned fantasy content. Six AI image detectors were evaluated at thresholds calibrated to a 5% flag rate on reference photographs. Detection rates ranged from 3.7% to 56.4% on the ChatGPT Images 2.5 collection — compared to 83-100% recall these same detectors achieve on GenImage benchmark. Artwork false-positive rates ranged from 1.5% to 96.5%, showing that a higher detection rate on one collection does not imply lower false positives on another. The paper provides a methodology reference for building real-world AI image datasets and shows that detection benchmarks substantially overstate practical detection capabilities during model transitions.