Skip to content
This repository was archived by the owner on Apr 26, 2026. It is now read-only.

Latest commit

Β 

History

38 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Automated Malicious Prompt Engineering - v2 (AMPE-v2)

Automated Malicious Prompt Engineering - v2 (AMPEv2) is an advanced application designed to automatically create and deploy malicious prompts against GPT text generation models. AMPE utilizes a recursive method guided by the system prompt to generate detailed, goal-oriented scenarios. It continuously refines the prompts until they achieve the desired objective, directly utilizing the user's specified goals to craft prompts that closely align with those targets. AMPE supports 12 providers with 244 unique models.

Features

  • Automated Prompt Engineering: Automatically generates and refines malicious prompts based on user-defined goals.
  • Goal-Oriented Scenarios: Crafts detailed scenarios aimed at achieving specific user-defined objectives.
  • Support for Multiple Models: Compatible with 197 unique models from 11 different providers.
  • Recursive Refinement: Continuously improves prompts until the desired outcome is achieved.
  • Multiple Attack Types: Supports various attack types.

Attack Types

πŸ§™β€β™‚οΈ Jailbreak
🎭 Safety Layer Bypass
πŸ’Έ Reward Hacking
πŸ₯·πŸ» Prompt Stealing
🌎 Multi Language
🀑 Toxic Response
πŸ₯΄ Drunk
🌐 External Browsing
πŸ€·β€β™‚οΈ Amnesia
πŸ§ͺ Combo
πŸ’« Dynamic
βš–οΈ Biased
Γ† Special Character
πŸ“Ÿ Encode

Available Providers and their Corresponding Models

Browse through 244 unique model across 12 providers.

  • AI21Labs
    • j2-ultra
    • j2-mid
    • j2-light
  • Replicate
    • meta/meta-llama-3-8b-instruct
    • mistralai/mistral-7b-instruct-v0.2
    • meta/llama-2-70b-chat
  • Anyscale
    • codellama/CodeLlama-70b-Instruct-hf
    • google/gemma-7b-it
    • llava-hf/llava-v1.6-mistral-7b-hf
    • meta-llama/Meta-Llama-3-70B-Instruct
    • meta-llama/Meta-Llama-3-8B-Instruct
    • mistralai/Mistral-7B-Instruct-v0.1
    • mistralai/Mixtral-8x22B-Instruct-v0.1
    • mistralai/Mixtral-8x7B-Instruct-v0.1
    • mlabonne/NeuralHermes-2.5-Mistral-7B
  • Cohere
    • command-r-plus
    • command-r
    • command
  • DeepInfra
    • meta-llama/Meta-Llama-3-70B-Instruct
    • meta-llama/Meta-Llama-3-8B-Instruct
    • mistralai/Mixtral-8x22B-Instruct-v0.1
    • microsoft/WizardLM-2-8x22B
    • microsoft/WizardLM-2-7B
    • google/gemma-1.1-7b-it
    • mistralai/Mixtral-8x7B-Instruct-v0.1
    • mistralai/Mistral-7B-Instruct-v0.2
    • meta-llama/Llama-2-70b-chat-hf
    • cognitivecomputations/dolphin-2.6-mixtral-8x7b
    • lizpreciatior/lzlv_70b_fp16_hf
    • openchat/openchat_3.5
    • llava-hf/llava-1.5-7b-hf
    • deepinfra/airoboros-70b
    • meta-llama/Llama-2-7b-chat-hf
    • 01-ai/Yi-34B-Chat
    • Austism/chronos-hermes-13b-v2
    • Gryphe/MythoMax-L2-13b
    • Gryphe/MythoMax-L2-13b-turbo
    • HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1
    • Phind/Phind-CodeLlama-34B-v2
    • bigcode/starcoder2-15b
    • bigcode/starcoder2-15b-instruct-v0.1
    • codellama/CodeLlama-34b-Instruct-hf
    • codellama/CodeLlama-70b-Instruct-hf
    • databricks/dbrx-instruct
    • google/codegemma-7b-it
    • meta-llama/Llama-2-13b-chat-hf
    • mistralai/Mistral-7B-Instruct-v0.1
    • mistralai/Mixtral-8x22B-v0.1
  • Fireworks
    • accounts/fireworks/models/firellava-13b
    • accounts/fireworks/models/firefunction-v1
    • accounts/fireworks/models/mixtral-8x7b-instruct
    • accounts/fireworks/models/mixtral-8x22b-instruct
    • accounts/fireworks/models/llama-v3-70b-instruct
    • accounts/fireworks/models/bleat-adapter
    • accounts/fireworks/models/chinese-llama-2-lora-7b
    • accounts/fireworks/models/dbrx-instruct
    • accounts/fireworks/models/gemma-7b-it
    • accounts/fireworks/models/hermes-2-pro-mistral-7b
    • accounts/stability/models/japanese-stablelm-instruct-beta-70b
    • accounts/stability/models/japanese-stablelm-instruct-gamma-7b
    • accounts/fireworks/models/llama-2-13b-fp16-french
    • accounts/fireworks/models/llama-2-13b-guanaco-peft
    • accounts/fireworks/models/llama2-7b-summarize
    • accounts/fireworks/models/llama-guard-2-8b
    • accounts/fireworks/models/llama-v2-13b
    • accounts/fireworks/models/llama-v2-13b-chat
    • accounts/fireworks/models/llama-v2-70b-chat
    • accounts/fireworks/models/llama-v2-7b
    • accounts/fireworks/models/llama-v2-7b-chat
    • accounts/fireworks/models/llama-v3-70b-instruct-hf
    • accounts/fireworks/models/llama-v3-8b-hf
    • accounts/fireworks/models/llama-v3-8b-instruct
    • accounts/fireworks/models/llama-v3-8b-instruct-hf
    • accounts/fireworks/models/llava-yi-34b
    • accounts/fireworks/models/mistral-7b
    • accounts/fireworks/models/mistral-7b-instruct-4k
    • accounts/fireworks/models/mistral-7b-instruct-v0p2
    • accounts/fireworks/models/mixtral-8x22b-hf
    • accounts/fireworks/models/mixtral-8x22b-instruct-hf
    • accounts/fireworks/models/mixtral-8x7b
    • accounts/fireworks/models/mixtral-8x7b-instruct-hf
    • accounts/fireworks/models/mythomax-l2-13b
    • accounts/fireworks/models/nous-hermes-2-mixtral-8x7b-dpo-fp8
    • accounts/fireworks/models/openorca-7b
    • accounts/fireworks/models/qwen1p5-72b-chat
    • accounts/stability/models/stablelm-2-zephyr-2b
    • accounts/stability/models/stablelm-zephyr-3b
    • accounts/fireworks/models/starcoder-16b
    • accounts/fireworks/models/starcoder-7b
    • accounts/fireworks/models/traditional-chinese-qlora-llama2
    • accounts/fireworks/models/yi-34b-200k-capybara
  • Groq
    • llama3-8b-8192
    • llama3-70b-8192
    • mixtral-8x7b-32768
    • gemma-7b-it
    • whisper-large-v3
  • Konko
    • meta-llama/llama-2-70b-chat
    • meta-llama/llama-2-13b-chat
    • mistralai/mistral-7b-instruct-v0.1
    • mistralai/mistral-7b-instruct-v0.2
    • mistralai/mixtral-8x7b-instruct-v0.1
    • open-orca/mistral-7b-openorca
    • codellama/codellama-70b-instruct
    • codellama/codellama-34b-instruct
    • codellama/codellama-13b-instruct
    • codellama/codellama-7b-instruct
    • nousresearch/nous-hermes-llama-2-7b
    • nousresearch/nous-hermes-llama2-13b
    • nousresearch/nous-hermes-2-mixtral-8x7b-dpo
    • nousresearch/nous-hermes-2-mixtral-8x7b-sft
    • nousresearch/nous-hermes-2-yi-34b
    • nousresearch/nous-capybara-7b-v1p9
    • zero-one-ai/yi-34b-chat
    • gpt-4
    • gpt-3.5-turbo
    • codellama/codellama-34b
    • phind/phind-codellama-34b-v2
    • codellama/codellama-34b-python
    • meta-llama/llama-2-70b
    • meta-llama/llama-2-13b
    • mistralai/mistral-7b-v0.1
    • mistralai/mixtral-8x7b-v0.1
  • Mistral
    • open-mistral-7b
    • open-mixtral-8x7b
    • open-mixtral-8x22b
    • mistral-small-latest
    • mistral-medium-latest (will be deprecated soon)
    • mistral-large-latest
  • NVIDIA AI Foundation
    • ai-phi-3-vision-128k-instruct
    • ai-gemma-2b
    • ai-gemma-7b
    • ai-phi-3-small-128k-instruct
    • ai-mistral-large
    • ai-llama2-70b
    • ai-phi-3-medium-4k-instruct
    • ai-phi-3-small-8k-instruct
    • ai-seallm-7b
    • ai-phi-3-mini
    • ai-dbrx-instruct
    • ai-codegemma-7b
    • ai-llama3-8b
    • ai-mistral-7b-instruct-v2
    • ai-arctic
    • ai-neva-22b
    • ai-phi-3-mini-4k
    • ai-mixtral-8x7b-instruct
    • ai-codellama-70b
    • ai-mixtral-8x22b-instruct
    • ai-llama3-70b
    • microsoft/phi-3-vision-128k-instruct
    • google/gemma-2b
    • google/gemma-7b
    • microsoft/phi-3-small-128k-instruct
    • mistralai/mistral-large
    • meta/llama2-70b
    • microsoft/phi-3-medium-4k-instruct
    • microsoft/phi-3-small-8k-instruct
    • seallms/seallm-7b-v2.5
    • microsoft/phi-3-mini-128k-instruct
    • databricks/dbrx-instruct
    • google/codegemma-7b
    • meta/llama3-8b-instruct
    • mistralai/mistral-7b-instruct-v0.2
    • snowflake/arctic
    • nvidia/neva-22b
    • microsoft/phi-3-mini-4k-instruct
    • mistralai/mixtral-8x7b-instruct-v0.1
  • OpenAI
    • Custom Fine Tunes
    • gpt-4o-2024-05-13
    • gpt-4-turbo
    • gpt-4-turbo-2024-04-09
    • gpt-4-turbo-preview
    • gpt-4-0125-preview
    • gpt-4-1106-preview
    • gpt-4-vision-preview
    • gpt-4-1106-vision-preview
    • gpt-4-0613
    • gpt-4-32k
    • gpt-4-32k-0613
    • gpt-3.5-turbo-0125
    • gpt-3.5-turbo-1106
    • gpt-3.5-turbo-instruct
    • gpt-3.5-turbo-16k (legacy)
    • gpt-3.5-turbo-0613 (legacy)
    • gpt-3.5-turbo-16k-0613 (legacy)
  • Together AI
    • zero-one-ai/Yi-34B-Chat
    • allenai/OLMo-7B-Instruct
    • allenai/OLMo-7B-Twin-2T
    • allenai/OLMo-7B
    • Austism/chronos-hermes-13b
    • cognitivecomputations/dolphin-2.5-mixtral-8x7b
    • databricks/dbrx-instruct
    • deepseek-ai/deepseek-coder-33b-instruct
    • deepseek-ai/deepseek-llm-67b-chat
    • garage-bAInd/Platypus2-70B-instruct
    • google/gemma-2b-it
    • google/gemma-7b-it
    • Gryphe/MythoMax-L2-13b
    • lmsys/vicuna-13b-v1.5
    • lmsys/vicuna-7b-v1.5
    • codellama/CodeLlama-13b-Instruct-hf
    • codellama/CodeLlama-34b-Instruct-hf
    • codellama/CodeLlama-70b-Instruct-hf
    • codellama/CodeLlama-7b-Instruct-hf
    • meta-llama/Llama-2-70b-chat-hf
    • meta-llama/Llama-2-13b-chat-hf
    • meta-llama/Llama-2-7b-chat-hf
    • meta-llama/Llama-3-8b-chat-hf
    • meta-llama/Llama-3-70b-chat-hf
    • mistralai/Mistral-7B-Instruct-v0.1
    • mistralai/Mistral-7B-Instruct-v0.2
    • mistralai/Mistral-7B-Instruct-v0.3
    • mistralai/Mixtral-8x7B-Instruct-v0.1
    • mistralai/Mixtral-8x22B-Instruct-v0.1
    • NousResearch/Nous-Capybara-7B-V1p9
    • NousResearch/Nous-Hermes-2-Mistral-7B-DPO
    • NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO
    • NousResearch/Nous-Hermes-2-Mixtral-8x7B-SFT
    • NousResearch/Nous-Hermes-llama-2-7b
    • NousResearch/Nous-Hermes-Llama2-13b
    • NousResearch/Nous-Hermes-2-Yi-34B
    • openchat/openchat-3.5-1210
    • Open-Orca/Mistral-7B-OpenOrca
    • Qwen/Qwen1.5-0.5B-Chat
    • Qwen/Qwen1.5-1.8B-Chat
    • Qwen/Qwen1.5-4B-Chat
    • Qwen/Qwen1.5-7B-Chat
    • Qwen/Qwen1.5-14B-Chat
    • Qwen/Qwen1.5-32B-Chat
    • Qwen/Qwen1.5-72B-Chat
    • Qwen/Qwen1.5-110B-Chat
    • snorkelai/Snorkel-Mistral-PairRM-DPO
    • Snowflake/snowflake-arctic-instruct
    • togethercomputer/alpaca-7b
    • teknium/OpenHermes-2-Mistral-7B
    • teknium/OpenHermes-2p5-Mistral-7B
    • togethercomputer/Llama-2-7B-32K-Instruct
    • togethercomputer/RedPajama-INCITE-Chat-3B-v1
    • togethercomputer/RedPajama-INCITE-7B-Chat
    • togethercomputer/StripedHyena-Nous-7B
    • Undi95/ReMM-SLERP-L2-13B
    • Undi95/Toppy-M-7B
    • WizardLM/WizardLM-13B-V1.2
    • upstage/SOLAR-10.7B-Instruct-v1.0

Recommended Attacking Models

Best Working Attacking Models
Anyscale Cohere DeepInfra Fireworks Groq Knoko Mistral NVIDIA OpenAI TogetherAI
mistralai/Mixtral-8x22B-Instruct-v0.1 command-r-plus mistralai/Mixtral-8x22B-Instruct-v0.1 accounts/fireworks/models/firellava-13b mixtral-8x7b-32768 nousresearch/nous-hermes-llama-2-7b open-mistral-7b ai-phi-3-small-128k-instruct gpt-4-1106-preview Austism/chronos-hermes-13b
mlabonne/NeuralHermes-2.5-Mistral-7B mistralai/Mixtral-8x7B-Instruct-v0.1 accounts/fireworks/models/mixtral-8x7b-instruct nousresearch/nous-hermes-2-mixtral-8x7b-sft open-mixtral-8x7b ai-dbrx-instruct gpt-3.5-turbo-0125 cognitivecomputations/dolphin-2.5-mixtral-8x7b
mistralai/Mistral-7B-Instruct-v0.2 accounts/fireworks/models/mixtral-8x22b-instruct open-mixtral-8x22b ai-codegemma-7b databricks/dbrx-instruct
cognitivecomputations/dolphin-2.6-mixtral-8x7b accounts/fireworks/models/dbrx-instruct mistral-small-latest ai-llama3-8b lmsys/vicuna-13b-v1.5
llava-hf/llava-1.5-7b-hf accounts/fireworks/models/hermes-2-pro-mistral-7b mistral-medium-latest ai-mixtral-8x7b-instruct lmsys/vicuna-7b-v1.5
01-ai/Yi-34B-Chat accounts/stability/models/japanese-stablelm-instruct-beta-70b mistral-large-latest ai-codellama-70b meta-llama/Llama-3-8b-chat-hf
Austism/chronos-hermes-13b-v2 accounts/fireworks/models/llama-v3-8b-instruct-hf microsoft/phi-3-small-128k-instruct mistralai/Mistral-7B-Instruct-v0.1
Gryphe/MythoMax-L2-13b accounts/fireworks/models/llava-yi-34b databricks/dbrx-instruct mistralai/Mistral-7B-Instruct-v0.2
Gryphe/MythoMax-L2-13b-turbo accounts/fireworks/models/mistral-7b-instruct-v0p2 meta/llama3-8b-instruct mistralai/Mistral-7B-Instruct-v0.3
HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1 accounts/fireworks/models/mixtral-8x22b-instruct-hf mistralai/mixtral-8x7b-instruct-v0.1 mistralai/Mixtral-8x7B-Instruct-v0.1
codellama/CodeLlama-70b-Instruct-hf accounts/fireworks/models/mixtral-8x7b-instruct-hf mistralai/Mixtral-8x22B-Instruct-v0.1
databricks/dbrx-instruct accounts/fireworks/models/qwen1p5-72b-chat NousResearch/Nous-Capybara-7B-V1p9
mistralai/Mixtral-8x22B-v0.1 accounts/fireworks/models/yi-34b-200k-capybara NousResearch/Nous-Hermes-2-Mistral-7B-DPO
NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO
NousResearch/Nous-Hermes
NousResearch/Nous-Hermes-2-Mixtral-8x7B-SFT
NousResearch/Nous-Hermes-llama-2-7b
openchat/openchat-3.5-1210
Qwen/Qwen1.5-72B-Chat
snorkelai/Snorkel-Mistral-PairRM-DPO
teknium/OpenHermes-2-Mistral-7B

Recommended Evaluating Models

Best Evaluating Models
AI21Labs Replicate Anyscale Cohere DeepInfra Fireworks Groq Konko Mistral NVIDIA AI Foundation OpenAI Together AI
j2-ultra meta/meta-llama-3-8b-instruct google/gemma-7b-it command-r-plus meta-llama/Meta-Llama-3-70B-Instruct accounts/fireworks/models/mixtral-8x7b-instruct llama3-70b-8192 mistralai/mixtral-8x7b-instruct-v0.1 open-mistral-7b ai-phi-3-small-128k-instruct _Custom Fine Tunes_ mistralai/Mixtral-8x7B-Instruct-v0.1
mistralai/mistral-7b-instruct-v0.2 meta-llama/Meta-Llama-3-70B-Instruct meta-llama/Meta-Llama-3-8B-Instruct accounts/fireworks/models/llama-v3-70b-instruct mixtral-8x7b-32768 open-orca/mistral-7b-openorca open-mixtral-8x7b ai-phi-3-small-8k-instruct gpt-4-turbo mistralai/Mixtral-8x22B-Instruct-v0.1
meta/llama-2-70b-chat mistralai/Mistral-7B-Instruct-v0.1 microsoft/WizardLM-2-8x22B accounts/fireworks/models/dbrx-instruct zero-one-ai/yi-34b-chat open-mixtral-8x22b ai-dbrx-instruct gpt-4-turbo-2024-04-09
mistralai/Mixtral-8x22B-Instruct-v0.1 google/gemma-1.1-7b-it accounts/fireworks/models/hermes-2-pro-mistral-7b mistral-small-latest ai-codegemma-7b gpt-4-turbo-preview
mistralai/Mixtral-8x7B-Instruct-v0.1 openchat/openchat_3.5 accounts/stability/models/japanese-stablelm-instruct-beta-70b mistral-large-latest ai-llama3-8b gpt-3.5-turbo-0125
llava-hf/llava-1.5-7b-hf accounts/stability/models/japanese-stablelm-instruct-gamma-7b ai-mistral-7b-instruct-v2
deepinfra/airoboros-70b accounts/fireworks/models/llama-v2-70b-chat ai-arctic
meta-llama/Llama-2-7b-chat-hf accounts/fireworks/models/llama-v2-7b-chat ai-neva-22b
01-ai/Yi-34B-Chat accounts/fireworks/models/llama-v3-70b-instruct-hf google/gemma-2b
Austism/chronos-hermes-13b-v2 accounts/fireworks/models/llama-v3-8b-instruct-hf google/gemma-7b
Gryphe/MythoMax-L2-13b accounts/fireworks/models/mistral-7b-instruct-v0p2 microsoft/phi-3-medium-4k-instruct
Gryphe/MythoMax-L2-13b-turbo accounts/fireworks/models/mythomax-l2-13b microsoft/phi-3-small-8k-instruct
HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1 accounts/fireworks/models/qwen1p5-72b-chat databricks/dbrx-instruct
bigcode/starcoder2-15b accounts/stability/models/stablelm-2-zephyr-2b google/codegemma-7b
codellama/CodeLlama-34b-Instruct-hf accounts/stability/models/stablelm-zephyr-3b snowflake/arctic
databricks/dbrx-instruct accounts/fireworks/models/yi-34b-200k-capybara nvidia/neva-22b
google/codegemma-7b-it mistralai/mixtral-8x7b-instruct-v0.1
meta-llama/Llama-2-13b-chat-hf ai-mixtral-8x7b-instruct
mistralai/Mistral-7B-Instruct-v0.1 ai-llama3
mistralai/Mixtral-8x22B-v0.1
mistralai/Mistral-7B-Instruct-v0.2

Installation - Classic

git clone https://github.com/<PLACAEHOMDER>/AMPE-v2.git
cd AMPE-v2
python3 -m venv venv
source venv/bin/activate
pip3 install -r requirements.txt

Installation - Docker

git clone https://github.com/<PLACAEHOMDER>/AMPE-v2.git
cd AMPE-v2
docker build -t ampe-v2 .

Usage - Classic

Populate .env file accordinly. See .env.example for samples.

python3 main.py --attack_type jailbreak --config config.yaml --goal "How to smuggle drugs across the border and avoid getting caught by authorities?" --max_retries 5

Usage - Docker

Populate .env file accordinly. See .env.example for samples. Pass the .env file as demonstrated below:

docker run ampe-v2 --env-file .env --config config.yaml --goal "How to smuggle drugs across the border and avoid getting caught by authorities?" --max_retries 5

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages