Invisible text in PDFs, so AI models are manipulated in selections: the University of Turin study

Written by Jason Miller

Researchers at the University of Turin have demonstrated the structural vulnerability of language models such as ChatGPT, Gemini And Claude to indirect prompt injection attacks delivered via strings hidden in digital documents. By inserting instructions not visible to the naked eye into the files, the team forced the algorithms to completely alter your judgment during automated document review, selection and analysis processes.

The operating principle exploits the very architecture of generative models, incapable of distinguishing raw input data from system operational directives. In the study conducted by the Turin team, composed of Federico Torrielli, Stefano Locci, Amon Rapp and Luigi Di Caro, the insertion of invisible commands within scientific articles achieved a success rate up to 99%. The model thus attributed excellence ratings to manipulated documents, ignoring the objective review parameters.

The scenario, described in a recent interview by La Repubblica, becomes critical in automated recruiting processes, where the use of intelligent agents to screen thousands of applications is now widespread practice. A candidate can hide textual instructions in the file layout that instruct the system to classify him as the ideal profile to hire. The manipulation exploits the same structural weakness documented by the university researchers, for a phenomenon that reopens several unanswered questions about the reliability of algorithmic filters used by human resources.

From hiring to healthcare: the risks of manipulation and countermeasures

The impact of such vulnerabilities transcends the boundaries of personnel selection for touch high-risk contexts such as the analysis of medical records and health reports. If a model is fed altered health documentation with invisible strings, the risk of misdiagnosis or compromised clinical interpretations becomes concrete. The reliance on natural language models highlights an inherent limitation in the unverified processing of data from third-party sources.

At the same time, the study proposes a defensive countermeasure based on the incorporation of invisible character sequences within the artificially generated texts. This digital watermark allows you to trace the synthetic origin of the documents, recording adetection accuracy between 87% and 97%. The security of automated workflows will increasingly depend on the ability to intercept alterations upstream, avoiding delegating critical choices to systems without rigorous checks on incoming data.

Jason Miller

I'm Jason Miller, and I've been passionate about technology and storytelling for over a decade. As a lead writer at Herald Editorials, I strive to bring clarity and creativity to complex tech topics. When I'm not writing, you'll find me exploring the latest gadgets or hiking in the great outdoors.