Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

8 Commits
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents



Comparison of the traditional task execution paradigm (top) and our IPIGUARD (bottom)

📢 News

  • [2025.09.15] IPIGuard is selected for Oral presentation at EMNLP 2025
  • [2025.08.21]🎉 Our paper "IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents" has been accepted to EMNLP 2025 Main Conference!

🔧 Installation

We recommend using Python ≥3.10.

# git clone
git clone https://github.com/Greysahy/ipiguard.git
cd ipiguard

# create conda environment
conda create -n ipiguard python=3.10
conda activate ipiguard

# install agentdojo
cd agentdojo
pip install -e .

🚀 Quick Start

1. Set Your OpenAI API Key in eval.sh

export OPENAI_API_KEY='<your_api_key>'
export OPENAI_BASE_URL='<your_base_url>'

2. Run Evaluation

We provide a ready-to-use shell script for running evaluations with specific agent, attack, and defense configurations. You can modify the arguments to evaluate different settings.

To run the script:

bash eval.sh
Argument Description
OPENAI_API_KEY Your OpenAI API key, required to access the model.
OPENAI_BASE_URL Base URL of the OpenAI API. Keep default unless using a proxy or private deployment.
agent_model The agent model used for evaluation (e.g., gpt-4o-mini-2024-07-18).
attack_name The adversarial attack type to simulate (e.g., important_instructions).
defense_name The defense strategy applied during evaluation (e.g., ipiguard).
suite_name The task suite/domain being evaluated (e.g., travel, workspace, slack, banking).
mode Running mode: benign → Standard tasks without attacks; attack → Adversarial tasks with injected attacks
output_dir Directory to store evaluation results.
--uid / --iid User task ID and injection attack ID, used for resuming evaluation after interruption.

📊 Results

Attack Workspace Slack Travel Banking Overall
ASR↓ / UA↑ ASR↓ / UA↑ ASR↓ / UA↑ ASR↓ / UA↑ ASR↓ / UA↑
Ignore Previous 0.00 / 68.33 0.00 / 59.05 0.00 / 62.86 2.78 / 49.31 0.64 / 61.21
InjectAgent 0.42 / 67.92 0.95 / 63.81 0.00 / 65.00 0.00 / 47.92 0.32 / 61.84
Tool Knowledge 0.00 / 69.58 1.90 / 59.05 0.00 / 59.29 2.78 / 47.92 0.95 / 60.57
Important Instr. 0.83 / 65.00 0.00 / 49.52 0.00 / 57.14 1.39 / 49.31 0.64 / 57.07
Average 0.31 / 67.71 0.71 / 57.86 0.00 / 61.07 1.74 / 48.44 0.69 / 58.77

More results and ablations are available in the paper.


📌 Citation

If you use this code or find our work helpful, please cite:

@misc{an2025ipiguardnoveltooldependency,
      title={IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents}, 
      author={Hengyu An and Jinghuai Zhang and Tianyu Du and Chunyi Zhou and Qingming Li and Tao Lin and Shouling Ji},
      year={2025},
      eprint={2508.15310},
      archivePrefix={arXiv},
      primaryClass={cs.CR},
      url={https://arxiv.org/abs/2508.15310}, 
}

🏷️ License

Apache License 2.0 - See LICENSE for details.

About

[EMNLP 2025 Oral] IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents

Resources

Stars

22 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages