Search

The Phish, The Spam, and The Valid: Generating Feature-Rich Emails for Benchmarking LLMs

Contribution type: article

Title: The Phish, The Spam, and The Valid: Generating Feature-Rich Emails for Benchmarking LLMs

Authors:

Rebeka Toth, University of Oslo, Norway
Nils Gruschka, University of Oslo, Norway
Tamas Bisztray, University of Oslo, Norway

Keywords: phishing, spam, email security, large language models, email detection, emotion analysis, dataset

Abstract:

In this paper, we introduce a metadata-enriched generation framework (PhishFuzzer) that seeds real emails into Large Language Models (LLMs) to produce 23,100 diverse, structurally consistent email variants across controlled entity and length dimensions. Unlike prior corpora, our dataset features strict three-class labels (Phishing, Spam, Valid), provides full URLs and attachment metadata, and annotates each email with attacker intent. Using this dataset, we benchmark two state-of-the-art LLMs (Qwen-2.5-72B and Gemini-3.1-Pro) under both Basic (body, subject) and Full (+URL, sender, attachment) settings. Using formal confidence metrics (Task Success Rate and Confidence Index), we analyze model reliability, robustness to linguistic fuzzing, and the impact of structural metadata on detection accuracy. Our fully open-source framework and dataset provide a rigorous foundation for evaluating next-generation email security systems. To support open science, we make the PhishFuzzer Dataset, the generation scripts, and prompts available on GitHub: https://github.com/DataPhish/PhishFuzzer

Publication Date: June 7, 2026

Presented during:

Dates: June 7, 2026 to June 11, 2026

Location: Porto, Portugal

Venue:

Hotel Novotel Porto Gaia

Rua Martir Sao Sebastiao, Afurada,
4400-499 Vila Nova de Gaia

Hotel website

Copyright (c) DTR Society, 2026

Contact Us

Proposals

The submission system is currently being prepared. We invite you to subscribe to the newsletter so you will be the first to know when it goes online.