How to Spot Fake Email Threads with AI Summarizer (Avoid Hidden Text)
Artificial Intelligence Forcepoint X-Labs has released findings demonstrating that indirect prompt injection attacks on AI summarization tools can succeed without relying on hidden text or explicit instructions.
Fake Thread Manipulates AI Summarizer Without Hidden Text Security
Artificial Intelligence Forcepoint X-Labs has released findings demonstrating that indirect prompt injection attacks on AI summarization tools can succeed without relying on hidden text or explicit instructions. The research, led by senior researcher Ben Gibney, builds on a previous August 2026 proof of concept that used hidden HTML content to alter AI-generated summaries. This latest study reveals that malicious content can influence summaries simply by appearing as part of a thread’s natural structure.
A Forged Thread Altered the Summarizer’s Output
Researchers conducted experiments using six test emails designed with three presentation techniques: standard view, 30 lines of blank padding, and hidden styling. Each method was tested both with and without direct instructions to the AI. The trials used an unsecured Outlook-based pipeline powered by Claude Haiku 4.5 at a temperature setting of zero, with each sample processed 10 times across 60 trials.
Researchers Conducted Experiments
The original message included details about a quarterly supplier review scheduled for August 24, 2026, and an outstanding invoice of €8,650. A second, fabricated header block inserted into the thread modified these details to September 3, 2026, and €46,200. All 60 trials produced summaries reflecting the altered information.
The Attack’s Success
In the plain-view, no-instruction test, the forged message consistently generated false details across all 10 runs. The attack succeeded because the malicious content appeared as a legitimate message within the thread, avoiding detection by keyword or pattern-based filters. Gibney noted in the analysis that “the instruction in these samples is conveyed through the structure of two messages in a thread rather than explicit wording.” This contrasts with the earlier proof of concept, which combined hidden HTML with direct commands.
The new findings show that concealment is not necessary for fabricated data to appear in summaries.
Direct Instructions Influenced Summary Content
The study also found that explicit user instructions altered which information the AI retained. Samples with direct commands consistently excluded genuine details, while those without instructions allowed some true data to remain, though often relegated to a “Note” section. An unexpected outcome occurred when 30 blank lines were added to the test emails, causing the summarizer to discard half of the pre-existing content.
Broader Implications and Limitations
The research highlights broader concerns about how content is perceived by users, security tools, and AI systems. Microsoft recently reported attacks using invisible Unicode characters to bypass phishing filters. However, the study does not claim universal security failures. The experiments involved a single client, one model, synthetic data, and an intentionally unguarded pipeline. Forcepoint emphasized that results may vary with different models, temperature settings, or guardrails.
The findings underscore the limitations of hidden-text detection and instruction scanning as standalone defenses. Malicious content presented as normal thread elements can still manipulate AI-generated summaries without overt prompts.
