Author: Sunil Gentyala, IEEE Senior Member | Lead Cybersecurity and AI Security Consultant, HCLTech Contact: sunil.gentyala@ieee.org Website: sunilgentyala.github.io/gsh-framework License: Apache 2.0
Most enterprise security stacks were not built for the threat surface that agentic AI creates. Endpoint agents cannot see what an LLM gateway is doing. SIEMs have no baselines for multi-agent tool call chains. The GSH Framework closes that gap.
GSH is an open-source research artifact for autonomous agentic AI threat hunting. It provides structured detection playbooks, behavioral baselining logic, and a policy-driven enforcement engine (Sovereign Sentinel) designed for the cognitive cyber domain: the operational layer where large language models, autonomous agents, and multi-agent pipelines interact with enterprise infrastructure.
All detection signals are mapped to MITRE ATLAS and NIST CSF 2.0, giving practitioners framework-aligned coverage they can operationalize immediately.
The hunt playbooks, detection logic, thresholds, and policy schema are complete and documented.
Hunt-001 through Hunt-004 (scripts/gsh-sentinel-deploy.py, scripts/gsh-probe-eval.py) implement the full baselining, drift-scoring, and ZTLV enforcement logic end-to-end, but ship with a synthetic telemetry generator (clearly marked SIMULATION MODE in the script output and # Replace this block in source) so you can see the detection logic run without a live environment first. Wiring --target to a real LLM gateway event stream is the integration step you complete before using this for actual enforcement.
Hunt-005 (adapters/mcp_proxy.py, scripts/gsh-mcp-proxy.py, scripts/gsh-baseline.py) is different: it is a real MCP JSON-RPC stdio proxy that intercepts actual tool definitions and tool calls between a real MCP host and a real MCP server - approval-time schema hashing, drift detection, semantic poisoning scans, and per-call enforcement (permit/alert/block) all run against live traffic, not synthetic data. A captured baseline is never auto-trusted: it starts as UNVERIFIED and only becomes a trusted comparison point through the gsh-baseline.py capture -> review -> approve -> verify workflow; --mode aggressive refuses to even launch the wrapped server without an approved baseline. See tests/test_mcp_proxy.py and tests/test_gsh_baseline.py for subprocess-driven end-to-end tests of both CLIs. Known gaps: canary/response-asymmetry comparison and tool-return-value scanning are not implemented yet (see playbooks/hunt-005-mcp-tool-poisoning.md section 5.2 for details), and only the stdio transport is supported (not streamable HTTP/SSE MCP servers).
SIEM output (adapters/splunk_hec.py, adapters/elastic_bulk.py, adapters/windows_eventlog.py) is also real: set siem_output: splunk, siem_output: elastic, or siem_output: windows_eventlog in your policy YAML (see configs/sentinel-policy-default.yaml) and both gsh-sentinel-deploy.py and gsh-mcp-proxy.py will send findings there (Splunk HEC / Elasticsearch _bulk over real HTTP, or a registered source in the local Windows Application Event Log). A failed or unconfigured send always falls back to local file output - a finding is never silently dropped. The Windows Event Log adapter is Windows-only and requires pywin32; on any other platform (or without pywin32) it logs a warning and falls back like any other unconfigured destination. See tests/test_siem_adapters.py and tests/test_windows_eventlog.py (the latter includes a test that writes a real event and reads it back, not just a mocked one).
LangChain telemetry (adapters/langchain_callback.py) is a fourth real integration: GSHCallbackHandler attaches to any LangChain Runnable/agent via config={"callbacks": [handler]} and evaluates real tool-call rate, token velocity, unauthorized-tool invocations, and suspicious call parameters against Hunt-001/Hunt-004 thresholds - no synthetic data. Important limitation: LangChain callback handlers are notification hooks, not gates - by default LangChain swallows exceptions raised inside a callback rather than stopping the tool call, so this adapter can only alert, never block. Every finding it emits is explicitly marked enforcement_mode: "alert_only" and action_taken: "ALERTED", regardless of policy mode. It also has no visibility into DNS queries (Hunt-002). See tests/test_langchain_callback.py, tested against langchain-core 1.4.x.
See open issues for remaining work (SARIF reporting, the Hunt-006 playbook, and the Docker Compose demo).
Full release notes (including known limitations at each release) are on the Releases page. Summary:
| Version | Highlights |
|---|---|
| v1.5.0 | Real Windows Application Event Log output adapter (adapters/windows_eventlog.py); optional and Windows-only, safe no-op elsewhere |
| v1.4.0 | Real LangChain callback adapter (adapters/langchain_callback.py) for Hunt-001/Hunt-004 telemetry - alert-only by design, since LangChain callbacks cannot block a tool call |
| v1.3.0 | Real Splunk HEC and Elastic bulk SIEM output adapters, wired into both the Sentinel and the MCP proxy via a shared dispatcher; a failed/unconfigured SIEM send now always falls back to local file output |
| v1.2.0 | Real MCP JSON-RPC stdio proxy for Hunt-005 (adapters/mcp_proxy.py) - schema-hash drift detection, semantic poisoning scan, and real per-call enforcement against live MCP traffic, not simulated |
| v1.1.0 | Hunt-004 (rogue agent) completed; Hunt-005 (MCP supply chain / tool poisoning) added as a playbook; project website launched |
| v1.0.0-beta | Initial public release: Hunt-001 through Hunt-003 playbooks, Sentinel reference scripts (synthetic telemetry), default policy schema |
| Component | Description |
|---|---|
| Sovereign Sentinel | Policy-driven behavioral enforcement agent deployed alongside LLM gateways |
| Hunt Playbooks | Structured threat detection playbooks for high-severity agentic AI threats |
| DDI-AI Fusion | DNS/DHCP/IPAM telemetry layer with AI-agent-aware baselining |
| Zero-Trust Logic Validation (ZTLV) Gate | Per-invocation tool call authorization engine |
| Behavioral Baseline Engine | Continuous model output drift detection and probe evaluation pipeline |
| Playbook | Threat Class | Severity | Status |
|---|---|---|---|
| Hunt-001 | Agentic Loop / Resource Exhaustion | High | Active |
| Hunt-002 | DDI Covert Channel / C2 via DNS | Critical | Active |
| Hunt-003 | ML Model Poisoning / Behavioral Drift | Critical | Active |
| Hunt-004 | Rogue Agent / Unauthorized Tool Use | Critical | Active |
| Hunt-005 | MCP Supply Chain / Tool Poisoning | Critical | Active |
git clone https://github.com/sunilgentyala/gsh-framework.git
cd gsh-framework
pip install -r requirements.txtAlternative: install as a package. The repo is also pip-installable from source (not yet published to PyPI, so install from the checkout, not pip install gsh-framework):
pip install . # adapters + all CLIs, PyYAML only (no SIEM/LangChain/Windows extras)
pip install ".[splunk]" # + Splunk/Elastic HTTP output (adapters/splunk_hec.py, elastic_bulk.py)
pip install ".[langchain]" # + LangChain callback adapter
pip install ".[windows]" # + Windows Event Log adapter (Windows only)
pip install ".[llm]" # + OpenAI-compatible client for gsh-probe-eval.py
pip install ".[dev]" # + pytest, ruff, mypy, blackThis installs gsh-sentinel-deploy, gsh-mcp-proxy, gsh-baseline, gsh-probe-eval, and gsh-ddi-log-parser as commands (equivalent to python scripts/<name>.py), and makes adapters importable without manually adjusting sys.path. Extras can be combined, e.g. pip install ".[splunk,langchain]".
cat configs/sentinel-policy-default.yamlEdit it to set your organization name, SIEM output destination, and egress allowlist before deploying.
Start in passive mode to build a 7-day behavioral baseline, then move to standard enforcement:
python scripts/gsh-sentinel-deploy.py \
--target "llm-gateway-01" \
--mode passive \
--policy configs/sentinel-policy-default.yaml \
--baseline-window 7dAs shipped, this generates synthetic telemetry (SIMULATION MODE, logged at startup) so you can watch the baselining and scoring logic run immediately. Replace the telemetry-generation block noted in the script (real LLM gateway/API metrics or LangChain callbacks) to run it against live traffic.
Unlike step 3, this runs against real MCP traffic. A captured baseline is never auto-trusted - capture it, review it, then approve it:
python scripts/gsh-baseline.py capture \
--server-id "corp-tools-mcp-01" \
--server-cmd "npx -y @modelcontextprotocol/server-filesystem /srv/data"
python scripts/gsh-baseline.py review --baseline baselines/mcp/corp-tools-mcp-01.json
python scripts/gsh-baseline.py approve \
--baseline baselines/mcp/corp-tools-mcp-01.json --reviewer "your-name-or-email"Then configure your MCP host to launch the proxy instead of the real server directly:
python scripts/gsh-mcp-proxy.py \
--server-cmd "npx -y @modelcontextprotocol/server-filesystem /srv/data" \
--server-id "corp-tools-mcp-01" \
--mode standard \
--baseline baselines/mcp/corp-tools-mcp-01.jsonThe proxy will alert on (or, in --mode aggressive, block) definition drift, poisoned tool descriptions, invisible Unicode content, and unauthorized tool calls. In --mode aggressive, the proxy refuses to launch the wrapped server at all unless the baseline above has been approved - see playbooks/hunt-005-mcp-tool-poisoning.md section 5.1 for why, and section 5.2 for the full detection logic.
pip install langchain-corefrom adapters.langchain_callback import GSHCallbackHandler
handler = GSHCallbackHandler(
target="my-langchain-agent",
allowlist=["web_search", "calculator"], # unlisted tools trigger an immediate alert
)
# Attach to any LLM, tool, or chain via the standard LangChain callbacks config:
llm.invoke(prompt, config={"callbacks": [handler]})
my_tool.invoke(args, config={"callbacks": [handler]})
handler.flush() # evaluate any partial window at the end of a runThis is alert-only, not enforcement - see adapters/langchain_callback.py's module docstring for why LangChain callback handlers cannot reliably block a tool call.
Each playbook is a self-contained Markdown document with detection logic, data sources, MITRE ATLAS mapping, triage decision tree, and response actions. Start with Hunt-001 for loop detection:
cat playbooks/hunt-001-agentic-loop-detection.mdA companion research paper covering the full technical rationale, design decisions, and threat model is in preparation and not yet submitted. Per publication policy, the manuscript is not included in this repository. For research inquiries, contact sunil.gentyala@ieee.org.
gsh-framework/
├── README.md
├── SECURITY.md
├── LICENSE
├── CITATION.cff
├── CONTRIBUTING.md
├── requirements.txt
├── pyproject.toml # pip-installable package + console-script CLIs, extras
├── .github/
│ └── workflows/
│ └── ci.yml # pytest + ruff + mypy across Python 3.10-3.13
├── adapters/
│ ├── mcp_proxy.py # Real MCP JSON-RPC proxy (Hunt-005) + baseline approval governance
│ ├── langchain_callback.py # Real LangChain telemetry, alert-only (Hunt-001/004)
│ ├── splunk_hec.py # Real Splunk HTTP Event Collector output
│ ├── elastic_bulk.py # Real Elasticsearch/OpenSearch _bulk output
│ ├── windows_eventlog.py # Real Windows Application Event Log output
│ └── siem_dispatch.py # Shared dispatcher used by all three SIEM adapters
├── configs/
│ └── sentinel-policy-default.yaml
├── docs/
│ └── index.html # Project website (GitHub Pages)
├── playbooks/
│ ├── hunt-001-agentic-loop-detection.md
│ ├── hunt-002-ddi-tunneling-anomaly.md
│ ├── hunt-003-model-poisoning-baseline.md
│ ├── hunt-004-rogue-agent-detection.md
│ └── hunt-005-mcp-tool-poisoning.md
├── probes/
│ └── standardized-probe-set-v1.json
├── scripts/
│ ├── _cli_shims.py # console-script entry points (pyproject.toml) for the scripts below
│ ├── ddi-log-parser-ai.py
│ ├── gsh-baseline.py # capture/review/approve/verify CLI for MCP baseline governance
│ ├── gsh-mcp-proxy.py # CLI for adapters/mcp_proxy.py
│ ├── gsh-probe-eval.py
│ └── gsh-sentinel-deploy.py
├── tests/
│ ├── test_ddi_log_parser.py
│ ├── test_gsh_baseline.py
│ ├── test_mcp_proxy.py
│ ├── test_siem_adapters.py
│ ├── test_windows_eventlog.py
│ ├── test_langchain_callback.py
│ └── fixtures/
│ ├── mock_mcp_server.py # Minimal MCP stdio server for testing
│ └── mock_http_sink.py # Minimal HTTP server for testing SIEM adapters
├── baselines/
└── reports/
| Threat | MITRE ATLAS | MITRE ATT&CK | NIST CSF 2.0 |
|---|---|---|---|
| Agentic Loop / Resource Exhaustion | AML.T0048, AML.T0040 | DE.AE-02, DE.CM-01, RS.MI-01 | |
| DDI Covert Channel Exfiltration | AML.T0048, AML.T0051 | T1071.004, T1048, T1568 | DE.CM-01, DE.AE-04, PR.DS-01 |
| ML Model Poisoning / Behavioral Drift | AML.T0020, AML.T0043, AML.T0044 | ID.RA-01, DE.AE-02, DE.CM-06 | |
| Rogue Agent / Unauthorized Tool Use | AML.T0051, AML.T0053, AML.T0054 | PR.PS-04, DE.CM-01, RS.AN-03 | |
| MCP Supply Chain / Tool Poisoning | AML.T0010, AML.T0051, AML.T0053 | T1195 | ID.SC-04, PR.PS-04, DE.CM-06 |
Security practitioners, AI safety researchers, and detection engineers are welcome. Read CONTRIBUTING.md before opening a Pull Request.
High-priority contributions include: additional hunt playbooks, refined detection thresholds, and integration adapters for LangChain, AutoGen, CrewAI, and MCP host platforms.
If you use the GSH Framework in your research, please cite:
@misc{gentyala2026gsh,
author = {Gentyala, Sunil},
title = {The Governed Security Hunting (GSH): An Autonomous Agentic Framework
for Defending the Cognitive Cyber Domain},
year = {2026},
howpublished = {Open Source Research Artifact, GitHub},
url = {https://github.com/sunilgentyala/gsh-framework}
}To report a vulnerability in the GSH Framework itself, use GitHub's private vulnerability reporting or email sunil.gentyala@ieee.org with the subject [GSH Security Vulnerability] - [brief description]. Do not open a public GitHub Issue. See SECURITY.md for the full policy, supported versions, and response timeline.
- ContextGuard: Zero-trust middleware for Model Context Protocol (MCP) server security. Precision 100%, Recall 96.7%, F1 98.3% at 1.005ms latency.
- ARGUS: LLM application security scanner.
- IEEE Senior Member Profile: ORCID 0009-0005-2642-3479