Systems
Security
Data
Python

PLEXData Engineering for people and agents that act on systems

Permissions, Limits, Evidence & eXecution.

No. 012 Latest feature

Wednesday, 30 September 2026

plexdata.online

001

Published
Updated 29 Sep 2026
Author
Dorian Sotpyrc
Reading
10 minutes · 7 sections
Fields
  • AI Security
  • Identity
  • Tooling
  • Observability

Limits · AI Security

The Fortified AI Stack: How Teams Are Locking Down ML Workflows in 2026

The model is no longer the security boundary. Identity, data, tools, runtime, network and evidence all need their own gate — and their own owner.

References
In this article · 7 sections
  1. 01The model is not the boundary
  2. 02Identity before intelligence
  3. 03Secrets and the data path
  4. 04Tools need external policy
  5. 05Contain execution and egress
  6. 06Monitoring needs evidence
  7. 07A practical build order

The old mental model for AI security was a protected model endpoint: authenticate the caller, encrypt the traffic and keep the API key out of Git. That is no longer enough once the model can search internal data, call tools, write files or trigger another service.

The security boundary has moved outward. A production workflow now includes identity, context, retrieval, tools, runtimes, networks and logs. A weak assumption in any one of those layers can become authority for the whole chain.

That is the useful meaning of a “fortified AI stack”. It is not six products around a model. It is six independent decisions about what the workflow is allowed to know and do.

§ 01The model is not the boundary

NSA's May 2026 guidance on Model Context Protocol makes the shift explicit. Authentication, authorization and input validation remain necessary, but agentic systems add dynamic tool invocation, implicit trust relationships and context sharing. NSA's conclusion is the important part: the environment has to be treated as a continuum, because a bad assumption in one stage can propagate into the next.1

OWASP describes the same problem from the application side. Its Excessive Agency risk is usually caused by too much functionality, too much permission or too much autonomy.2

So the first design question is not “how do we stop the model hallucinating?” It is “what happens if the model makes the wrong decision?” A fortified stack assumes that bad outputs, malicious inputs and compromised context are possible, then limits the blast radius.

Table 1 — Six boundaries, six owners
BoundaryQuestionOwnerEvidence
IdentityWho or what is acting?IAM / platformToken subject, scope, expiry
Secrets & dataWhat can it read?Data / securityClassification, retrieval decision
PolicyWhat may it do?App / securityAllow, deny, approval
IsolationWhere can code execute?PlatformSandbox, filesystem, egress
MonitoringWhat is happening now?Ops / SOCRuntime events, alerts
AuditWhat can we prove later?Risk / securityImmutable trace, retention

§ 02Identity before intelligence

Every agent, workflow and tool call should have an identity that can be limited independently. Reusing one broad service account for every model task is convenient, but it destroys the boundary between “the model needs this file” and “the service account can read the whole drive”.

The practical pattern is boring on purpose: short-lived credentials, task-specific scopes, separate identities for different trust levels and no permission inherited merely because a connector happens to expose it.

OWASP's examples of excessive agency include tools that expose functions the task does not need and downstream identities that have write or delete permission when read-only access would have been enough.2

§ 03Secrets and the data path

Secrets belong in a secret store or credential broker, not in the system prompt. OWASP's 2025 guidance is blunt: the system prompt should not be treated as a secret or as a security control, and credentials or connection strings should not be placed there.3

The same discipline applies to retrieval. A model should not receive a document simply because the retrieval system can find it. Classification, user entitlement and task context should be checked before content enters the model context.

That gives you a clean separation:

  • Identity says who is asking.
  • Data policy says whether this identity may receive this object for this task.
  • The model only sees the content after both checks pass.

§ 04Tools need policy outside the model

Tool use is where a chat system becomes an operational system. Reading a calendar is different from sending an email. Listing files is different from deleting them. A single “tools enabled” switch is too coarse once the workflow can change state.

Use a small policy layer between the model and each tool. The model can request an action. The policy layer decides whether the action is allowed, whether it needs human approval and what arguments are acceptable.

Table 2 — Class the action before the tool runs
ActionDefaultExample
ReadAllow + logSearch approved documentation
CreatePolicy checkCreate a draft ticket
External sendApprovalSend email or publish content
Delete / privilegeDeny by defaultDelete file, change access

§ 05Contain execution and egress

If the workflow can run code, treat the runtime as untrusted work. Use a short-lived sandbox with the smallest filesystem view possible. Mount only the data the task needs. Keep privileged sockets and host credentials out of reach.

Then control egress. A workload that can reach any address on the internet can turn a prompt-injection failure into data exfiltration. A gateway or allowlist gives you one place to say which destinations exist for this task and to record what left.

NSA and partner guidance on deploying AI systems securely has long framed AI security as protection of the model, data and surrounding infrastructure, not just the inference endpoint.4

§ 06Monitoring is useful only if it leaves evidence

NIST's Generative AI Profile treats monitoring and testing as lifecycle work, including post-deployment monitoring, incident handling and provenance. NIST updated the publication in April 2026, but the operational point has not changed: deployment is not the end of evaluation.5

For a production AI workflow, one trace should let you reconstruct:

  • which user or workload initiated the task;
  • which model and policy version ran;
  • which documents entered context;
  • which tools were requested and allowed;
  • which network destinations were contacted;
  • which approval or denial changed the path.

Logging everything is not the answer. Log the decisions that define authority. Protect those logs from casual modification and give every event a trace ID that follows the workflow across services.

§ 07A practical build order

Do not begin by buying six security products. Begin by making the boundaries explicit, then fill the gaps in an order that reduces authority fastest.

Deployment gates

Seven checks before production

§References

  1. 1NSA — Model Context Protocol: Security Design Considerations for AI-Driven Automationnsa.gov
  2. 2OWASP GenAI Security Project — LLM06:2025 Excessive Agencyowasp.org
  3. 3OWASP GenAI Security Project — LLM07:2025 System Prompt Leakageowasp.org
  4. 4NSA / partners — Deploying AI Systems Securelynsa.gov
  5. 5NIST AI 600-1 — Generative Artificial Intelligence Profilenist.gov