Viktoriia Vovchenko, Bielefeld University of Applied Sciences and Arts, Germany
Vincenzo Barberi, Bielefeld University of Applied Sciences and Arts, Germany
Sergej Schultenkämper, Bielefeld University of Applied Sciences and Arts, Germany
Frederik Simon Bäumer, Bielefeld University of Applied Sciences and Arts, Germany
Instruction-guided image editing models have made photorealistic facial manipulation accessible through natural language prompts, yet existing facial forensics datasets do not cover this manipulation paradigm. We present ADRIAN InstructFace-Edit (AIFE), a dataset of 36,000 edited facial image pairs at 1024×1024 resolution generated by four state-of-the-art instruction-guided image editing models (FLUX.1-Kontext-dev, FLUX.2-dev, Qwen-Image-Edit-2511, and Step1X-Edit-v1.2) across six semantically distinct edit types, of which 18,768 are annotated as high quality. The full dataset, including all pairs with quality labels, is made publicly available. To ensure data quality, we design a multi-stage filtering pipeline that integrates structural, semantic, and multimodal large language model–based quality checks, with thresholds derived empirically from a human-annotated subset of 1,500-images. Because open-weight vision-language models perform poorly on VIEScore evaluation out of the box, we fine-tune a LoRA adapter for Qwen3-VL-32B-Instruct that improves within ±1 SC accuracy from near-random to 73.2%. We will publicly release the dataset, generation pipeline, and LoRA adapter to support the development of detectors robust to modern instruction-guided facial manipulation. The code is available at: https://github.com/vika-v-v/faceguard-dataset. The dataset is available at: https://hsbi.sciebo.de/s/pb6ZB4ZeCtmJxGm (password: adrian).