Evaluation & Ops

AI Security Logging and Observability

What to log for AI security events and what to leave out: request IDs, categories, policy versions and tool calls, with raw prompts kept separate.

4 min read
AI Security Logging and Observability

When an AI feature misbehaves, the first question is "what exactly happened?" If your logs contain full prompts and responses, you can answer it, but you have also created a store of everything your users and systems ever typed, including passwords, personal data and confidential documents. If your logs contain nothing useful, you cannot answer at all. The goal is a record that explains decisions without becoming a second data leak.

Log the decision, not the content

Prefer structured metadata. For every AI request, one event per boundary can carry:

  • A request ID shared across the input check, model call, output check and tool calls.
  • The application, workflow and model involved.
  • The policy version and the thresholds in force.
  • The category and score for each check, and the action taken (allow, warn, block, review).
  • For agent steps: the tool, target and argument summary, plus the outcome.
  • Timestamps and latency for each stage.

An example event:

{
  "request_id": "req_8f21",
  "workflow": "support-copilot",
  "stage": "post_llm",
  "model": "provider/model-name",
  "policy_version": "2026-09-01",
  "category": "sensitive_data_exposure",
  "score": 0.87,
  "action": "block",
  "latency_ms": 212
}

Nothing in this record contains a customer message, yet it lets you count blocks by category, spot a policy change that shifted the rate and correlate the event with the surrounding application logs.

Treat raw content separately

Sometimes you do need the text, for example to reproduce a false positive or investigate an incident. Keep that in a separate store with:

  • Shorter retention than security metadata.
  • Access controls limited to the people who investigate.
  • Redaction of secrets and personal data before storage where possible.
  • A documented reason for collecting it, tied to your privacy commitments.

Raw content can contain the very secrets the control exists to protect. Do not dump full prompts into every centralized log. That is the most common mistake, and it turns the logging system into an attractive target. See preventing credential leakage for how secrets end up in text they should not be in.

Correlate across layers

An investigation crosses layers: user request, retrieval, model, scanner, tool, downstream service. A stable request ID lets you rebuild the chain. Include the same ID in your web logs, scanner calls and tool audit records. For agents, log every tool call, including reads, because leaks often happen through "harmless" reads.

What to alert on

  • Repeated blocked attempts from one user or source, which may indicate probing.
  • Sudden shifts in category rates after a model, prompt or policy change.
  • Unusual outbound destinations or a burst of tool calls.
  • Any request that bypassed a check, as described in fail open or fail closed.
  • Scanner errors and timeouts.
AI security events joined by one request ID A request ID is created at entry. The input check, output check and tool calls each write a metadata event. Raw prompts and outputs go to a separate store with short retention and restricted access. Request ID created at entry Passed to every call Input check event Category, score, action, policy version Output check event Category, score, action, latency Tool call event Tool, target, arguments summary, outcome Raw content, if needed Separate store, short retention, restricted access
Each AI request produces one event per boundary, joined by a request ID. Raw content is stored elsewhere under stricter rules.

Use logs to improve the controls

Blocked events feed the regression suite. A false positive becomes a legitimate test case, and a confirmed attack becomes an attack case, as covered in the prompt-injection testing checklist. Record the policy version with every decision so you can compare the effect of a threshold change, which is the raw material for tuning latency and false positives.

Common mistakes

  • Dumping full prompts into every log stream.
  • Recording a block with no category or policy version.
  • Ignoring tool-call outcomes when investigating an agent incident.

Sources and further reading

Frequently asked questions

What should I log for AI security events?

Request ID, workflow, model, policy version, risk category and score, the action taken, tool and target for agent steps, and timings. Keep raw prompts and outputs out of general logs.

Is it safe to log full prompts?

Usually not by default. Prompts and responses can hold credentials, personal data and confidential text. If you need them, store them separately with short retention, access controls and redaction.

How do I connect logs across services?

Generate a request ID at the entry point and pass it through every call: the scanner, the model, retrieval and each tool. Include it in every log line so the chain can be rebuilt.

Which events should trigger an alert?

Repeated blocked attempts from one source, category-rate jumps after a change, unexpected outbound destinations, bypassed checks and scanner errors. Tie each alert to a named owner and a first step for the response, as described in LLM data leakage incident response.

Keep reading