Skip to main navigation Skip to search Skip to main content

Detecting Sentiment Steering Attacks on RAG-enabled Large Language Models

  • Alan Mohan
  • , Shalaka S. Mahadik
  • , Mithun Mukherjee
  • , Pranav M. Pawar
  • , Raja Muthalagu
  • , Jaime Lloret
  • Birla Institute of Technology and Science Pilani
  • Polytechnic University of Valencia

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Retrieval Augmented Generation (RAG) is a technique that enhances the accuracy and reliability of Large Language Models (LLMs), enabling them to answer questions about data they weren't trained on by fetching relevant documents and adding them as context to the prompts sent to an LLM. As companies are incorporating RAG in their LLM applications at a large scale, the technique opens a new attack surface, and sentiment-steering data poisoning attacks are becoming one of the most critical vulnerabilities in RAG, where an attacker can inject biased data to influence the sentiment of an LLM towards an open-ended topic. However, there is a lack of research detecting sentiment-steering attacks. This work lays a foundation towards this by developing a novel detection model to identify poisoned biased passages before they can make their way into databases. From the results, we observe that the detection model achieves a high recall score of 97.75%.

Original languageEnglish
Title of host publicationICC 2026 - IEEE International Conference on Communications, Proceedings
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798319542090
DOIs
StatePublished - 2026
Externally publishedYes
Event2026 IEEE International Conference on Communications, ICC 2026 - Glasgow, United Kingdom
Duration: 24 May 202628 May 2026

Publication series

NameIEEE International Conference on Communications
ISSN (Print)1550-3607

Conference

Conference2026 IEEE International Conference on Communications, ICC 2026
Country/TerritoryUnited Kingdom
CityGlasgow
Period24/05/2628/05/26

Bibliographical note

Publisher Copyright:
© 2026 IEEE.

Keywords

  • Data Poisoning
  • Generative-AI
  • LLM Security
  • Large Language Models
  • RAG Security
  • Retrieval-Augmented Generation (RAG)

Fingerprint

Dive into the research topics of 'Detecting Sentiment Steering Attacks on RAG-enabled Large Language Models'. Together they form a unique fingerprint.

Cite this