Advancing text adversarial example generation using large language models

dc.contributor.authorMadrueño, Natalia
dc.contributor.authorFernández-Isabel, Alberto
dc.contributor.authorR. Fernández, Rubén
dc.contributor.authorMartín de Diego, Isaac
dc.date.accessioned2026-02-05T15:47:25Z
dc.date.issued2025-08-30
dc.description.abstractRecent advances in Natural Language Processing (NLP) are highly based on black-box score-based models that provide only final predictions along with their score. This opacity impedes the comprehension of their internal decision-making processes, complicating the identification of potential weaknesses. A powerful strategy for analyzing model vulnerabilities is the generation of text adversarial examples. These attacks introduce subtle text perturbations that cause victim models to make incorrect predictions while preserving the original semantic meaning. This paper presents a novel method for generating text adversarial examples through Large Language Models (LLMs). The proposed method uses the outstanding text generation capabilities of LLMs to modify the original input text at multiple granularities: character-, word-, and sentence-level. First, sentence-level perturbations are introduced by generating paraphrases with an LLM instruction prompt. Next, further character- and word-level perturbations are introduced to words that most affect predictions using another set of LLM instruction prompts. In particular, vulnerable words are perturbed by replacing them with their synonyms or misspelled variants, or by inserting additional neutral words adjacent to them. Experiments were conducted to assess the proposal’s viability on two sentiment classification tasks: sentence-level reviews and full-length reviews. The proposal demonstrates an advantage over many well-known approaches based on LLMs. It preserves the original semantics to a similar extent, while increasing the deception of victim models by 29–85 % over the best-analyzed state-of-the-art methods.
dc.description.sponsorshipThis research has been supported by grants from the Spanish Ministry of Science and Innovation, under the Knowledge Generation Projects program: XMIDAS (Ref: PID2021-122640OB-100), and the Public-Private Collaboration program: DICYME (Ref: CPP2021-009025).
dc.identifier.citationNatalia Madrueño, Alberto Fernández-Isabel, Rubén R. Fernández, Isaac Martín de Diego, Advancing text adversarial example generation using large language models, Knowledge-Based Systems, Volume 329, Part B, 2025, 114361, ISSN 0950-7051, https://doi.org/10.1016/j.knosys.2025.114361. (https://www.sciencedirect.com/science/article/pii/S0950705125014005)
dc.identifier.doihttps://doi.org/10.1016/j.knosys.2025.114361
dc.identifier.issn0950-7051 (print)
dc.identifier.issn1872-7409 (online)
dc.identifier.publicationfirstpage1
dc.identifier.publicationlastpage13
dc.identifier.publicationtitleKnowledge-Based Systems
dc.identifier.publicationvolume329
dc.identifier.urihttps://hdl.handle.net/10115/160977
dc.language.isoen
dc.publisherElsevier
dc.rightsAttribution 4.0 Internationalen
dc.rights.accessRightsinfo:eu-repo/semantics/openAccess
dc.rights.urihttp://creativecommons.org/licenses/by/4.0/
dc.subjectAdversarial attack
dc.subjectText adversarial example
dc.subjectLarge language model
dc.subjectNatural language processing
dc.subjectText classification
dc.titleAdvancing text adversarial example generation using large language models
dc.typeArticle
dc.type.hasVersionhttp://purl.org/coar/version/c_970fb48d4fbd8a85

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Name:
Advancing Text Adversarial Example Generation Using Large Language Models.pdf
Size:
2.21 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Name:
license.txt
Size:
2.96 KB
Format:
Item-specific license agreed upon to submission
Description: