AI Security Logging and Observability
What to log for AI security events and what to leave out: request IDs, categories, policy versions and tool calls, with raw prompts kept separate.

When an AI feature misbehaves, the first question is "what exactly happened?" If your logs contain full prompts and responses, you can answer it, but you have also created a store of everything your users and systems ever typed, including passwords, personal data and confidential documents. If your logs contain nothing useful, you cannot answer at all. The goal is a record that explains decisions without becoming a second data leak.
Log the decision, not the content
Prefer structured metadata. For every AI request, one event per boundary can carry:
- A request ID shared across the input check, model call, output check and tool calls.
- The application, workflow and model involved.
- The policy version and the thresholds in force.
- The category and score for each check, and the action taken (allow, warn, block, review).
- For agent steps: the tool, target and argument summary, plus the outcome.
- Timestamps and latency for each stage.
An example event:
{
"request_id": "req_8f21",
"workflow": "support-copilot",
"stage": "post_llm",
"model": "provider/model-name",
"policy_version": "2026-09-01",
"category": "sensitive_data_exposure",
"score": 0.87,
"action": "block",
"latency_ms": 212
}
Nothing in this record contains a customer message, yet it lets you count blocks by category, spot a policy change that shifted the rate and correlate the event with the surrounding application logs.
Treat raw content separately
Sometimes you do need the text, for example to reproduce a false positive or investigate an incident. Keep that in a separate store with:
- Shorter retention than security metadata.
- Access controls limited to the people who investigate.
- Redaction of secrets and personal data before storage where possible.
- A documented reason for collecting it, tied to your privacy commitments.
Raw content can contain the very secrets the control exists to protect. Do not dump full prompts into every centralized log. That is the most common mistake, and it turns the logging system into an attractive target. See preventing credential leakage for how secrets end up in text they should not be in.
Correlate across layers
An investigation crosses layers: user request, retrieval, model, scanner, tool, downstream service. A stable request ID lets you rebuild the chain. Include the same ID in your web logs, scanner calls and tool audit records. For agents, log every tool call, including reads, because leaks often happen through "harmless" reads.
What to alert on
- Repeated blocked attempts from one user or source, which may indicate probing.
- Sudden shifts in category rates after a model, prompt or policy change.
- Unusual outbound destinations or a burst of tool calls.
- Any request that bypassed a check, as described in fail open or fail closed.
- Scanner errors and timeouts.
Use logs to improve the controls
Blocked events feed the regression suite. A false positive becomes a legitimate test case, and a confirmed attack becomes an attack case, as covered in the prompt-injection testing checklist. Record the policy version with every decision so you can compare the effect of a threshold change, which is the raw material for tuning latency and false positives.
Common mistakes
- Dumping full prompts into every log stream.
- Recording a block with no category or policy version.
- Ignoring tool-call outcomes when investigating an agent incident.
Sources and further reading
Frequently asked questions
What should I log for AI security events?
Request ID, workflow, model, policy version, risk category and score, the action taken, tool and target for agent steps, and timings. Keep raw prompts and outputs out of general logs.
Is it safe to log full prompts?
Usually not by default. Prompts and responses can hold credentials, personal data and confidential text. If you need them, store them separately with short retention, access controls and redaction.
How do I connect logs across services?
Generate a request ID at the entry point and pass it through every call: the scanner, the model, retrieval and each tool. Include it in every log line so the chain can be rebuilt.
Which events should trigger an alert?
Repeated blocked attempts from one source, category-rate jumps after a change, unexpected outbound destinations, bypassed checks and scanner errors. Tie each alert to a named owner and a first step for the response, as described in LLM data leakage incident response.

