Featured · News
Sonnet 5.5 Turns the Mid-Tier into the Default
Anthropic’s new Sonnet posts frontier-class benchmark scores at Sonnet prices. The number that matters is cost per task at the effort level you actually run.
Systems
Security
Data
Python
Permissions, Limits, Evidence & eXecution.
No. 008 Latest feature
Monday, 28 September 2026
plexdata.online
Featured · News
Anthropic’s new Sonnet posts frontier-class benchmark scores at Sonnet prices. The number that matters is cost per task at the effort level you actually run.
Bar chart of Anthropic-reported Terminal-Bench 4.0 scores, per cent of tasks completed: Claude Sonnet 5.5 leads at 70.6%, Claude Opus 5.5 scores 66.4%, and Claude Sonnet 5 scores 10.3%. This is vendor benchmark data, not PLEXData testing.
Sonnet 5.5 leads the chart at the same token price as Sonnet 5. At default effort, Anthropic says it beats Sonnet 5’s best score for about a tenth of the cost per task.
Anthropic-reported · 28 Sep 2026The PLEX newsletter
New features on permissions, limits and evidence for people and software agents that act on systems, with mechanisms and code you can reuse.
News · Agent security
In this conceptual policy example, a model requests a file read and an outside send. Runtime policy allows the read and blocks the send. This is not a test of a named product.
The model can request it. The runtime can refuse it.
Dotted path: unreachableNVIDIA’s launch splits what the model attempts from what the runtime permits, with an optional hardware watchdog.
Agent skills
Skill review
DietrichGebert
The assessment
Ask whether the extra code needs to exist.
An independent review of a third-party skill designed to stop coding agents from over-building.
Books
Book review
Chip Huyen
The assessment
A useful framework for deciding what to build.
A systems-focused review of the foundation-model engineering book and how well it holds up in 2026.
PLEX is a technical source for people and software agents that act on systems. It publishes practical patterns, code and evidence that help both do useful work within clear permissions, limits and security boundaries.
Built to be useful to a human reader, an agent, or both working together.
About PLEXPermissions & limits · Analysis
A real AI kill switch isn't a privacy promise — it's technical denial. If the assistant can't connect, can't fetch, and can't send, it can't leak.
Fig. 1 · Live model
Model → your data & network
Allowed 000Blocked 000Leaked 000
Most AI privacy language talks about intent: “we won't train on it”, “we don't look at your data”. Intent is a promise about behaviour you cannot observe. When a model sits next to your documents, the question that matters is simpler. Can it reach them?
A kill switch answers that question with mechanism, not policy text. It removes the connection, the fetch and the send. If none of the three can happen, a leak cannot happen either — whatever the model decides to do.
Promises fail quietly. A connector is enabled for one team and inherited by another. A retrieval index is built from a shared drive that also holds board papers. A plugin gets network access “temporarily” and keeps it.
None of these is a breach on the day it happens. Each one is a path. Over time the paths add up, and nobody can say with confidence which data a model can read.
A real kill switch has three properties, and all three must be enforced by infrastructure rather than by the model:
Complex classification schemes are rarely enforced. Start with three labels and one default for each. If you can't explain the policy to a new starter in a minute, the gate will not be configured correctly either.
| Label | Default action | What it covers |
|---|---|---|
| Public | Allow | Published material, documentation, marketing copy |
| Internal | Allow + log | Working documents, tickets, internal wikis |
| Secret | Deny | Credentials, customer data, board and legal papers |
The labels connect to three enforcement checkpoints — prompt, retrieval and output — and to a separate control for tool execution.
The whole design fits in one table. Each gate works at a different layer, so a failure in one does not open the others.
| Stage | Control | Implementation |
|---|---|---|
| Pre-request | Identity & scopes | Least privilege and expiring tokens |
| In transit | Data-plane gate | Inspect, redact or block at prompt, retrieval and output |
| Tool execution | Sandbox & allowlists | Gateway-controlled, with approvals for risky actions |
Ambient access is permission that nobody asked for today. It is the most common path to a leak. Replace it with scopes that are requested for a task and expire when the task ends.
A token that lives for fifteen minutes can still be misused — but only for fifteen minutes, and only for the scope it names.
The gate sits between the model and your data. It reads the label on every item, applies the default, and records the decision. Keep the policy in a file you can review like code:
policy: ai-data-gatelabels: public: { action: allow } internal: { action: allow, log: true } secret: { action: deny, log: true }checkpoints: [prompt, retrieval, output]tools: egress: gateway-only allowlist: [search_docs, create_ticket] approval_required: [send_email, write_file] # a human says yes
Line 5 does the real work. Anything labelled secret is denied at all three checkpoints, and the denial is logged.
If the assistant can't connect, can't fetch, and can't send, it can't leak.
No. 004 · The AI Kill Switch
Run these before you trust the switch, and again after every change to the gate. Tick them off as you go.
Verification checklist
0 of 5 tests passed.
Run again after every boundary change.Evidence · Privacy
The chat window is only the visible record. Depending on the product and settings, the same message can also create usage logs, safety records, device metadata and retention obligations you never see.
An AI chat feels private because it looks like a conversation between you and a machine. There is no public audience, no comment thread and usually no obvious sign that anything exists beyond the transcript.
That interface is easy to mistake for the system.
It is not. The transcript is only the part you can see. The service around it may also process account identifiers, timestamps, device information, usage events, safety classifications, files, location signals and other operational data. Exactly what is retained — and for how long — depends on the provider, product, account type and settings.1
Start with the simplest distinction: conversation history is not the same thing as service logging.
OpenAI's current privacy policy, for example, separates user content from log data, usage data, device information and location information. Its examples of log data include IP address, browser type, settings and request time. Usage data can include features used, actions taken, access time, country, user agent and device type.1
| Layer | Typical examples | Visible to you? |
|---|---|---|
| Conversation | Prompt, response, uploaded content | Usually |
| Account / usage | Time, model, feature use, account or workspace context | Sometimes |
| Device / network | IP address, browser, device type, general location | Not in chat history |
| Safety / integrity | Abuse signals, policy flags, feedback or review records | Usually not |
That does not mean a human is reading every conversation. It means the product has more data surfaces than the transcript. Privacy decisions should start from those surfaces, not from how empty the sidebar looks.
One of the most common mistakes is to treat “not used for training” as if it meant “not stored”.
OpenAI states this directly: if you turn off Improve the model for everyone, new conversations are not used to train models, but regular chats can still appear in history. Saved chats remain until you delete them or a workspace retention rule removes them.2
The same distinction appears elsewhere. Anthropic gives consumer users a model-improvement choice, while its retention rules separately describe deletion, standard retention and longer periods for certain safety cases.56
Temporary modes are useful. They are also easy to overread.
ChatGPT Temporary Chat stays out of normal history, does not create or update memories and is not used for model improvement while it remains temporary. OpenAI says a copy may still be retained for up to 30 days for safety purposes.3
Google documents a different window. Gemini temporary chats — and chats created while Keep Activity is off — are retained with the account for 72 hours so the service can respond, handle feedback and protect Google, users and the public. They do not appear in Gemini Apps Activity and temporary chats are not used to train Google's AI models.4
The practical lesson is not that temporary modes are bad. It is that temporary means a different retention path. It does not automatically mean zero records.
There is no universal “AI chat privacy setting”. The names, defaults and retention periods vary.
| Product | History / activity | Lower-retention mode or control | Documented note |
|---|---|---|---|
| ChatGPT | Saved chats remain until deletion or workspace policy removal. | Temporary Chat | Temporary copy may be kept up to 30 days for safety.3 |
| Gemini Apps | Keep Activity defaults to 18-month auto-delete for eligible personal accounts. | Temporary Chat or Keep Activity off | Future chats are retained for 72 hours in those modes.4 |
| Claude consumer | Users can delete conversations from history. | Model-improvement setting affects use and retention rules | Deleted conversations are removed from history immediately and normally from the backend within 30 days; safety exceptions can be longer.5 |
| Microsoft Copilot | Microsoft documents 18 months of conversation history for signed-in users. | Privacy controls can change model-training use and history management | Individual chats or full history can be deleted.7 |
This table is deliberately narrow. It does not try to flatten every product tier into one number. Business, enterprise, education, API and managed-workspace products often have different contracts and retention rules.
Deleting a chat is still worth doing. But the button usually changes the user-facing state before every backend copy is gone.
OpenAI's current retention guidance says a deleted saved chat disappears from the account view immediately and is scheduled for permanent deletion from its systems within 30 days, unless the chat was already de-identified and disassociated from the account or longer retention is required for security or legal reasons.8
Anthropic's consumer guidance similarly says deleted conversations are removed immediately from history and automatically deleted from the backend within 30 days under its standard rule, while documenting longer retention for some usage-policy violations and trust-and-safety classification scores.5
Your employer's AI workspace is not the same privacy boundary as a personal chat account.
OpenAI's managed-account guidance says organisation-controlled accounts may expose submitted content, conversation history, shared workspace content, usage/activity metadata and certain security or privacy settings according to workspace configuration, permissions and applicable law.9
That is not automatically a problem. In a business setting, auditability can be a feature. The important point is to know which boundary you are using before you paste confidential material into it.
Connected apps add another boundary. If a chat sends data to a third party, that third party's privacy and retention rules can apply too. The original AI provider's temporary-mode promise does not automatically follow the data downstream.
You do not need to memorise every privacy policy. You need a short decision routine.
Checklist
Policies change. This article records the provider documentation available on 29 September 2026. Check the linked source before relying on a retention period for regulated or high-risk work.
Limits · AI Security
The model is no longer the security boundary. Identity, data, tools, runtime, network and evidence all need their own gate — and their own owner.
The old mental model for AI security was a protected model endpoint: authenticate the caller, encrypt the traffic and keep the API key out of Git. That is no longer enough once the model can search internal data, call tools, write files or trigger another service.
The security boundary has moved outward. A production workflow now includes identity, context, retrieval, tools, runtimes, networks and logs. A weak assumption in any one of those layers can become authority for the whole chain.
That is the useful meaning of a “fortified AI stack”. It is not six products around a model. It is six independent decisions about what the workflow is allowed to know and do.
NSA's May 2026 guidance on Model Context Protocol makes the shift explicit. Authentication, authorization and input validation remain necessary, but agentic systems add dynamic tool invocation, implicit trust relationships and context sharing. NSA's conclusion is the important part: the environment has to be treated as a continuum, because a bad assumption in one stage can propagate into the next.1
OWASP describes the same problem from the application side. Its Excessive Agency risk is usually caused by too much functionality, too much permission or too much autonomy.2
So the first design question is not “how do we stop the model hallucinating?” It is “what happens if the model makes the wrong decision?” A fortified stack assumes that bad outputs, malicious inputs and compromised context are possible, then limits the blast radius.
| Boundary | Question | Owner | Evidence |
|---|---|---|---|
| Identity | Who or what is acting? | IAM / platform | Token subject, scope, expiry |
| Secrets & data | What can it read? | Data / security | Classification, retrieval decision |
| Policy | What may it do? | App / security | Allow, deny, approval |
| Isolation | Where can code execute? | Platform | Sandbox, filesystem, egress |
| Monitoring | What is happening now? | Ops / SOC | Runtime events, alerts |
| Audit | What can we prove later? | Risk / security | Immutable trace, retention |
Every agent, workflow and tool call should have an identity that can be limited independently. Reusing one broad service account for every model task is convenient, but it destroys the boundary between “the model needs this file” and “the service account can read the whole drive”.
The practical pattern is boring on purpose: short-lived credentials, task-specific scopes, separate identities for different trust levels and no permission inherited merely because a connector happens to expose it.
OWASP's examples of excessive agency include tools that expose functions the task does not need and downstream identities that have write or delete permission when read-only access would have been enough.2
Secrets belong in a secret store or credential broker, not in the system prompt. OWASP's 2025 guidance is blunt: the system prompt should not be treated as a secret or as a security control, and credentials or connection strings should not be placed there.3
The same discipline applies to retrieval. A model should not receive a document simply because the retrieval system can find it. Classification, user entitlement and task context should be checked before content enters the model context.
That gives you a clean separation:
Tool use is where a chat system becomes an operational system. Reading a calendar is different from sending an email. Listing files is different from deleting them. A single “tools enabled” switch is too coarse once the workflow can change state.
Use a small policy layer between the model and each tool. The model can request an action. The policy layer decides whether the action is allowed, whether it needs human approval and what arguments are acceptable.
| Action | Default | Example |
|---|---|---|
| Read | Allow + log | Search approved documentation |
| Create | Policy check | Create a draft ticket |
| External send | Approval | Send email or publish content |
| Delete / privilege | Deny by default | Delete file, change access |
If the workflow can run code, treat the runtime as untrusted work. Use a short-lived sandbox with the smallest filesystem view possible. Mount only the data the task needs. Keep privileged sockets and host credentials out of reach.
Then control egress. A workload that can reach any address on the internet can turn a prompt-injection failure into data exfiltration. A gateway or allowlist gives you one place to say which destinations exist for this task and to record what left.
NSA and partner guidance on deploying AI systems securely has long framed AI security as protection of the model, data and surrounding infrastructure, not just the inference endpoint.4
NIST's Generative AI Profile treats monitoring and testing as lifecycle work, including post-deployment monitoring, incident handling and provenance. NIST updated the publication in April 2026, but the operational point has not changed: deployment is not the end of evaluation.5
For a production AI workflow, one trace should let you reconstruct:
Logging everything is not the answer. Log the decisions that define authority. Protect those logs from casual modification and give every event a trace ID that follows the workflow across services.
Do not begin by buying six security products. Begin by making the boundaries explicit, then fill the gaps in an order that reduces authority fastest.
Deployment gates
Execution · 10 lines
The lines I keep are not clever. They put a bound on waiting, surface failure, make security explicit or remove an assumption the runtime would otherwise make for me.
requests.get(url)
requests.get(url, timeout=(3.05, 10))
Production Python is mostly ordinary Python with fewer silent assumptions.
The lines I trust are not universal recipes. Each one closes a specific failure mode: waiting forever, accepting a bad response, losing a traceback, generating the wrong kind of token, building SQL from text, writing ambiguous timestamps or letting the platform choose an encoding.
The test is simple: when this line fails at 03:00, does it stop, raise, log or preserve enough evidence for the next person to understand what happened?
response = requests.get(url, timeout=(3.05, 10))response.raise_for_status()subprocess.run(args, check=True, timeout=30)logger.exception("job failed")token = secrets.token_urlsafe(32)hmac.compare_digest(received, expected)cursor.execute("SELECT * FROM jobs WHERE id = ?", (job_id,))stamp = datetime.now(UTC).isoformat()os.replace(tmp_path, final_path)text = Path(path).read_text(encoding="utf-8")Line 1: requests.get(url, timeout=(3.05, 10))
Requests does not time out by default. Its own documentation says most external requests should have a timeout, and a tuple lets you separate the connection wait from the read wait.1
The line I do not trust is the shorter one:
requests.get(url)
It looks clean until a remote service stops answering and a worker spends an unbounded amount of time waiting for somebody else's network.
Line 2: response.raise_for_status()
A completed HTTP request is not the same thing as a successful application request. Requests exposes raise_for_status() so 4xx and 5xx responses become exceptions instead of quietly travelling deeper into the program.1
Line 3: subprocess.run(args, check=True, timeout=30)
check=True turns a non-zero exit into a CalledProcessError. timeout=30 puts a ceiling on the wait. The Python docs recommend run() for common subprocess work and document both behaviours.2
The line I do not trust is subprocess.run(args) when the return code matters. It can fail and hand control back as if nothing happened.
Line 4: logger.exception("job failed")
Inside an exception handler, Logger.exception() records the message and traceback. That is a better operational artifact than print(exc), which often throws away the path that led to the error.3
Line 5: secrets.token_urlsafe(32)
Python's secrets module exists for cryptographically strong random values used in password resets, hard-to-guess URLs and similar security-sensitive cases. The docs still use 32 bytes as a typical security level for tokens.4
The line I do not trust for security tokens is anything built from random. That module is useful for simulation and sampling. It is the wrong source for a reset link.
Line 6: hmac.compare_digest(received, expected)
For HMAC or other secret-derived values, Python recommends compare_digest() rather than ordinary equality to reduce timing-analysis exposure.5
Line 7: cursor.execute("SELECT * FROM jobs WHERE id = ?", (job_id,))
The value is data, not SQL. Python's sqlite3 documentation explicitly warns against building queries with string operations and recommends parameter substitution instead.6
The line I do not trust is an f-string that places user data inside the SQL text.
Line 8: datetime.now(UTC).isoformat()
An aware UTC timestamp says what instant it represents. Python documents naive datetimes as ambiguous and deprecated datetime.utcnow() in Python 3.12 in favour of an aware UTC datetime.7
Line 9: os.replace(tmp_path, final_path)
When the temporary file and destination are on the same filesystem, os.replace() gives you replacement semantics without exposing a half-written destination file. Python documents a successful replacement as atomic where the operating system provides that guarantee.8
This line is the final move, not the whole durability story. If the file must survive power loss, you still need the appropriate flush and filesystem durability strategy before replacement.
Line 10: Path(path).read_text(encoding="utf-8")
Explicit encoding removes a machine-dependent assumption from a text boundary. Path.read_text() accepts the encoding directly and closes the file for you.9
The line I do not trust is Path(path).read_text() when the file format says UTF-8. Let the file format decide the encoding, not whichever workstation happens to run the code.
A trusted line is not a trusted system. Timeouts need retry policy. Retries need idempotency. Atomic replacement needs a correctly staged temporary file. Parameterised SQL does not replace authorization. Logging a traceback does not replace monitoring.
The value of these lines is smaller and more useful: each one removes a category of silent ambiguity.
Before merge
Concept · Agent security
An agent is not powerful because it can reason. It becomes powerful when identity, context, tools and credentials line up into a path that can change something real.
“Is this agent trustworthy?” is becoming the wrong security question.
A model can be careful and still sit behind an overpowered service account. A model can be unreliable and still be harmless inside a read-only sandbox. The practical risk appears when a sequence of components gives a model a path from text to effect.
NIST's 2026 work on software-agent identity and authorization focuses on exactly this shift: agents need identification, authorization, auditing and non-repudiation controls because they increasingly act across data sets, tools and applications.1 NSA's MCP guidance reaches the same conclusion from another direction, warning that dynamic tool invocation, implicit trust relationships and context sharing create risks that do not stop at one endpoint.2
An identity answers who or what is acting. Authority answers what that identity can cause.
Those two ideas are easy to blur in agent systems because the agent often inherits a ready-made credential. A connector signs in once, the agent sees a tool, and the implementation begins to treat “tool available” as “tool allowed”.
That shortcut removes the useful questions: allowed for which user, for which task, against which resource, for how long, and with what evidence?
| Edge | Decision | Evidence |
|---|---|---|
| User → agent | What task was actually requested? | Task ID, user identity, scope |
| Agent → identity | Which credential may be used? | Subject, token scope, expiry |
| Identity → tool | Which operation is allowed? | Policy decision, approval |
| Tool → resource | Which object may be touched? | Resource ID, classification |
| Resource → effect | What state can change? | Before/after event, audit trace |
Think of authority as a graph, not a property of the agent. The nodes are users, agents, identities, tools and resources. The edges are the permissions that allow one node to affect the next.
This changes architecture reviews. Instead of asking whether the agent has “email access”, ask whether this task can move through a chain that ends in an external send. The same email connector might be safe for search and unsafe for sending.
OpenAI's current agent-safety guidance makes a similar practical point: risk rises when untrusted content can influence tool calls, and approvals, structured data flow and restricted tool use reduce the consequences when manipulation succeeds.3
A sub-agent can look less privileged while the workflow around it remains more privileged. It may receive no credential directly but still call a parent tool that carries one. Or it may write a file that another process later executes.
The effective authority is therefore the union of reachable effects, not the list of tools shown in one prompt.
Prompt injection matters because text can become a steering input to an authority path. The malicious document does not need its own credential. It only needs to persuade a component that already has one.
OpenAI describes prompt injection as a form of social engineering against the agent and recommends limiting access, narrowing instructions and reviewing important actions.4
That makes context provenance part of authorization. A decision influenced by public web text should not automatically inherit the same authority as a decision based on an authenticated user instruction.
The model can propose an action. It should not be the final authority on whether the action is allowed.
Use deterministic checks for identity, resource scope, action class, rate, destination and approval state. The policy does not need to understand every thought. It needs to understand the proposed effect.
This also gives agents something useful: a clean denial is better than a vague refusal. A denied tool call can return the exact missing scope or required approval without granting anything new.
A full reasoning transcript is neither necessary nor sufficient for security evidence. What matters operationally is the chain of authority decisions.
Authority review
Book review · AI engineering
A broad systems book for building with foundation models. Its lasting value is the decision framework; its weakest area in 2026 is the security depth needed for agents that can act.
Book review
Chip Huyen
The assessment
A useful framework for deciding what to build.
Editorial judgment. Criteria are fixed before the total; scores are not publisher or reader ratings.
The best technical books teach a sequence of decisions, not a sequence of tools. AI Engineering mostly does that.
Chip Huyen's book is explicitly not a code-along tutorial. Her companion repository says the aim is to provide a framework for adapting foundation models to real applications, with questions around evaluation, RAG, agents, fine-tuning, data, inference, latency, cost and feedback loops.1
That choice is why the book has held up better than a framework-specific title would have. O'Reilly lists 534 pages and describes an end-to-end progression from foundation-model applications through evaluation, prompting, RAG and agents, fine-tuning, dataset engineering, inference optimisation and feedback.2
This is a map of AI application engineering for people who already know how software projects behave. It spends less time telling you which SDK to install and more time asking what should be evaluated, what should be retrieved, what should be fine-tuned, and what should remain outside the model.
The official table of contents makes that breadth visible: foundation models, two chapters on evaluation, prompting, RAG and agents, fine-tuning, datasets, inference optimisation, architecture and user feedback.3
The cost of that breadth is obvious too. You do not finish the book with one application assembled. You finish with a larger set of engineering questions.
The strongest choice is architectural rather than topical: evaluation appears before most of the fashionable adaptation techniques.
That sequencing matters. If you cannot state what “better” means, prompt changes, RAG, model swaps and fine-tuning become activity rather than engineering. The book treats open-ended output evaluation as a first-class system problem, including human evaluation, model-based evaluation and task-specific criteria.
For PLEX readers, this is the most transferable lesson: define evidence before optimisation.
The book's durable material is the material least tied to model names: application scoping, evaluation, context construction, data quality, inference trade-offs and feedback loops.
Huyen's own companion repo says the book focuses on fundamentals rather than a particular tool or API because tools age quickly.1 That editorial decision has paid off. The names of models have changed; the questions around latency, cost, data, evals and whether to retrieve or fine-tune have not disappeared.
The weakest score is not because the agents chapter is poor. It is because production agents have moved from “models with tools and planning” toward identity, delegated authority, sandboxes, runtime policy and prompt-injection-resistant workflows.
Chapter 6 covers tools, planning, agent failure modes, evaluation and memory.4 What it cannot fully reflect is the security architecture that became more explicit through 2026: NIST work on agent identity and authorization, NSA MCP guidance, and runtime controls around agent action.
That is not a reason to skip the chapter. It is a reason to pair it with newer agent-security material.
The book makes quality measurable before it makes the architecture more complicated.
Concepts survive model and framework churn better than implementation recipes.
Identity, runtime boundaries and delegated authority deserve a 2026 companion reading list.
Read it if you are moving from model demos toward a production application and need a coherent map of the decisions.
Read selected chapters if your work is already specialised. Evaluation, RAG/agents, inference and architecture can stand alone.
Do not buy it expecting a current framework cookbook. The author says it is not a tutorial book, and that is the point.1
Agent skill review · Third party
A sharply scoped skill for fighting agent over-engineering. Its strength is not “write fewer lines”; it is the decision ladder that asks whether the extra machinery needed to exist at all.
Skill review
DietrichGebert
The assessment
Ask whether the extra code needs to exist.
Static review of the skill, docs and published benchmark methodology. PLEXData did not rerun the benchmark for this article.
Ponytail has one unusually useful opinion: the default failure mode of a coding agent is often to build too much.
The core skill describes itself as a “lazy senior dev” mode and pushes the agent through a ladder: YAGNI, standard library, native platform, one line, minimum.1 That is more interesting than the marketing phrase “less code”, because it gives the model an order of operations.
The main skill is broad: coding, adding, refactoring, fixing, reviewing, designing and dependency selection. Around it sit narrower skills for over-engineering review, whole-repository audit, debt tracking, measured impact and help.
The review skill is particularly clean. It asks for one-line findings focused only on complexity: what to delete and what replaces it.2 The audit skill scales the same idea to a whole repository and explicitly says correctness, security and performance are out of scope.3
That boundary is a feature. A skill that tries to be code quality, security, architecture and minimalism at the same time becomes hard to activate and harder to evaluate.
The best part of Ponytail is that it names the alternatives in a useful order. Before adding a dependency, ask if the platform already has the feature. Before adding an abstraction, ask whether there is more than one implementation. Before writing a wrapper, ask whether a direct call is clearer.
This is agent-friendly because each step is observable. “Be elegant” is subjective. “Use the standard library before a new dependency” can be checked.
The project has also invested in portability. Its documentation describes adapters or instruction tiers for Claude/Codex-style skills, OpenClaw, Gemini CLI, Cursor, Windsurf, Cline, Copilot, Qoder, Zed and others.4
The current README reports roughly 54% less code on its agentic benchmark, with up to 94% reduction on tasks that invite over-building, plus lower cost and latency in that test setup.5
More important than the number is the benchmark note. The repository says its earlier single-shot result overstated the win because the baseline included prose and options; the later agentic benchmark is presented as the more defensible comparison.6
Review/audit modes explicitly focus on over-engineering rather than pretending to replace security or correctness review.
The published numbers come from the project's own harness. They are useful evidence, not independent proof for every repo or model.
The benchmark notes weaker transfer to small local models. Teams should test against their own tasks and harness.
Minimalism is not a universal objective. Some code is longer because it carries evidence, compatibility, observability, safety checks or a deliberately explicit boundary.
The project's own “100% safe” benchmark claim is scoped to its benchmark. It should not be read as “Ponytail always preserves every security property”.
The skill is safest when the requirements are already clear. In a poorly specified task, “do less” can become “silently omit the thing nobody remembered to state”.
Strong fit: mature codebases where agents routinely introduce wrappers, dependencies, factories and “future flexibility”.
Good review tool: the dedicated review/audit skills are easier to trust than an always-on minimalism mode because they propose cuts without applying them.
Use carefully: greenfield systems with incomplete requirements, safety-critical code, and code where redundancy is an intentional control.
News · Agent security
Most agent safety asks the model to behave. NVIDIA’s launch is built around a harder question: what still holds when it doesn’t? The interesting answer sits outside the model, and optionally outside the host software too.
In this conceptual policy example, a model requests a file read and an outside send. Runtime policy allows the read and blocks the send. This is not a test of a named product.
The model can request it. The runtime can refuse it.
Dotted path: unreachableNVIDIA announced the Open Agent Safety Platform on 28 September 2026, built around OpenShell and the Sentry reference design.
NVIDIA says Sentry can quarantine agents that move outside policy boundaries in milliseconds. That is a product claim, not a universal result.
Runtime enforcement is only as useful as the permissions and policies organisations define. Deployment evidence will matter more than launch architecture.
Most of us have handed an agent a credential and trusted it to use it sensibly. That works right up until it doesn’t, and this launch is about the second half of that sentence.
On 28 September, NVIDIA launched the Open Agent Safety Platform: OpenShell, an open-source runtime boundary, plus Sentry, an optional out-of-band enforcement design that uses BlueField-4 DPUs.1
What caught my attention is the split. The model decides what to attempt. The environment decides what is permitted. NVIDIA’s own product page draws the same line: model safeguards influence behaviour; runtime controls enforce allowed actions.2 I think that is the right line, and it is an easy one to blur when a model is doing something impressive.
The launch combines two layers.
NVIDIA says OpenShell can also extend to third-party compute platforms, including Arm and Intel systems. The company listed a large set of launch partners across AI labs, enterprise software, security and infrastructure.1
NVIDIA describes a policy layer outside the model/application process for agent files, tools, credentials and network access.
Current documentation covers sandbox policy, provider routing, observability and supported agents.
OpenShell becomes the software layer in a broader design that adds Sentry as an independent hardware-backed watchdog.
OpenShell's current developer guide describes sandboxed execution with controls over files, process behaviour, network access, credentials and inference routing. It also exposes allow/deny logging and policy configuration.3
The new technical blog describes Gateway, Supervisor and Sandbox components and says policy is enforced without requiring the agent itself to be rewritten.4
The useful shift is that you don’t have to make the agent less capable. You make the room it works in smaller.
Sentry is designed to run on BlueField-4 DPUs as an out-of-band monitor. NVIDIA says it can correlate activity, enforce identity and access policies, and quarantine an agent that attempts to leave its software boundary.1
That is a more honest starting point. It assumes something can go wrong with the workload or the host, and plans for it: if either is compromised, some monitoring and enforcement still exists outside it.
AP's launch coverage describes the same basic split—OpenShell as the constrained workspace and Sentry as the watchdog—and notes that defining effective rules remains a hard problem.5
Confirmed: the platform, code/docs and reference architecture exist; OpenShell is open source; the launch names partners; the runtime exposes concrete policy and observability features.
Claimed: NVIDIA says the architecture could have prevented recent agent security incidents and that Sentry can quarantine escapes in milliseconds. Reuters reports that claim as NVIDIA's assertion, not an independently established result.6
Still open: how well organisations will write policies, how performance behaves under real mixed workloads, how easy bypasses are across different deployment modes, and whether hardware-isolated enforcement becomes common outside NVIDIA-heavy environments.
The direction is larger than NVIDIA hardware. Agent security is moving from model-only controls toward layered enforcement: identity, runtime isolation, network policy, tool authorization and audit.
The signal I take from it: the better agents get at pursuing goals, the more the infrastructure around them has to say, plainly, which paths are impossible. Not discouraged. Not logged. Impossible.
News · Models
Anthropic’s new Sonnet posts frontier-class scores at Sonnet prices. The number that matters is not the 70.6% headline — it is the cost per task at the effort level you actually run.
Bar chart of Anthropic-reported Terminal-Bench 4.0 scores, per cent of tasks completed: Claude Sonnet 5.5 leads at 70.6%, Claude Opus 5.5 scores 66.4%, and Claude Sonnet 5 scores 10.3%. This is vendor benchmark data, not PLEXData testing.
Sonnet 5.5 leads the chart at the same token price as Sonnet 5. At default effort, Anthropic says it beats Sonnet 5’s best score for about a tenth of the cost per task.
Anthropic-reported · 28 Sep 2026Anthropic released Claude Sonnet 5.5 on 28 September 2026. It is available on all major platforms as claude-sonnet-5-5, priced the same as Sonnet 5: $2 per million input tokens and $10 per million output.
Anthropic reports 70.6% on Terminal-Bench 4.0 against Sonnet 5’s 10.3%, and says that at default effort it beats Sonnet 5’s best score for about a tenth of the cost per task. These are vendor numbers.
Gains depend on task mix and effort settings, and Anthropic itself says Opus 5.5 remains clearly stronger on complex, open-ended work. Independent replication is pending.
Choosing a model tier used to be a real decision: pay for the frontier model when the work is hard, drop a tier when it isn’t, and accept the gap in between. This launch is Anthropic’s attempt to close that gap — and if the numbers hold, the default for everyday agent work just moved down a tier.
On 28 September, Anthropic released Claude Sonnet 5.5, the second model in its 5.5 family after Opus 5.5 arrived on 22 September. A Haiku 5.5 for high-volume work is promised in the coming weeks.1 The pitch is blunt: a clear upgrade over Sonnet 5, 30% faster, and up to 30% cheaper per task — at exactly the same token price.
What caught my attention is not the jump itself. Model releases always jump. It is where the jump lands: the tier most teams actually afford for agents that run all day.
The confirmed part is simple. Sonnet 5.5 is live in the Claude apps, Claude Code, and on the Claude Platform, AWS, Google Cloud and Azure.1 Pricing is unchanged from Sonnet 5, and Anthropic is publishing a system card with its evaluation method.2
| Item | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Input | $2 | $4 |
| Output | $10 | $20 |
| Cache reads | $0.20 | $0.20 |
| Cache writes | $2.50 | $5 |
One migration detail matters if you run agents with thinking disabled: you need to switch to the new between_tools setting, which keeps up-front thinking off, before moving to Sonnet 5.5. Anthropic’s migration guide covers it.3
The headline number is real, and it is large. On Terminal-Bench 4.0, an agentic coding evaluation, Anthropic reports Sonnet 5.5 at 70.6% — against 10.3% for Sonnet 5 and 66.4% for Opus 5.5. On CursorBench 4.0, built from real Cursor coding sessions, it reports 55.5% against Sonnet 5’s 34.1%, within about two points of Opus 5.5.1
| Evaluation | Sonnet 5 | Sonnet 5.5 | Opus 5.5 |
|---|---|---|---|
| Agentic coding · Terminal-Bench 4.0 | 10.3% | 70.6% | 66.4% |
| Agentic coding · CursorBench 4.0 | 34.1% | 55.5% | 57.8% |
| Knowledge work · GDPval-AA v2.1 | 1449 | 1844 | 1846 |
| Computer use · OSWorld 2.1 | 57.0% | 80.1% | 81.8% |
| Chart recognition · Chartography | 15.6% | 61.6% | 64.4% |
But scores are half the story. Anthropic’s own charts plot score against cost per task at every effort level, and that is where the release bites: at Medium effort — the default in the Claude apps — it says Sonnet 5.5 beats Sonnet 5’s best score for less than a tenth of the cost per task.1 Same price list, fewer tokens per job, faster output. If that holds outside Anthropic’s harness, the cost model for routine agent work changes, not just the leaderboard.
Slack’s early testing points the same way, with the usual early-tester caution: better results than Sonnet 5 on almost all of its offline Slackbot evals, in fewer steps, with about 14% fewer output tokens — a vendor-published tester quote, not an independent result.1
Sonnet 5.5 leans hard on effort levels: Low, Medium, High and beyond, trading speed and tokens against thoroughness. The defaults are Medium in Claude Code and the Claude apps, High on the Claude Platform.4
Anthropic’s own framing is that Sonnet 5.5 complements Opus 5.5 at lower effort, where it costs less per task, and can match it at higher effort at comparable cost. At High effort on FrontierCode it reports a score ten points above Sonnet 5 at the same setting — at roughly one fifteenth of the cost per task.1
My read: this makes effort a routing decision, not just a quality setting. The practical question for an agent pipeline is no longer “which model” but “which model at which effort for which job class”. Teams that pin one model and one effort for everything will pay for it twice — once in tokens, once in capability left on the table.
Confirmed: the model exists, the model ID is live, pricing is published, the system card is public, and the safety posture is documented — the same biology safeguards as Sonnet 5, visible fallbacks to Sonnet 5 on higher-risk cybersecurity tasks, and classifiers against reasoning extraction, a first for the Sonnet line.12
Claimed: the benchmark table, the 30% speed and cost figures, the “a tenth of the cost per task” comparisons, and the collaboration anecdotes from Epic and Slack are all Anthropic-reported. Anthropic also reports that on its roughly 1,850-scenario behavioral audit, Sonnet 5.5 improves on or matches Sonnet 5 on most alignment measures, and that on containment evaluations it comes close to Opus 5.5 on how rarely it tries to escape its sandbox.1 Vendor-reported alignment is still vendor-reported.
Still open: independent Terminal-Bench and CursorBench replication; whether the cost advantage survives long-horizon, tool-heavy workloads; how the between_tools migration plays out in existing agent stacks; and where Haiku 5.5 lands the bottom tier when it ships.
First, re-run your own tasks at Medium effort before touching defaults. Vendor charts measure vendor-chosen tasks; your eval set is the only one that knows your workload.
Second, treat effort as part of routing. If you run an agent pipeline, this release is a good reason to log effort level alongside model and cost per task — you cannot tune a knob you do not record.
Third, keep Opus 5.5 for the work that burns. Anthropic is unusually plain that Opus remains clearly stronger on complex, open-ended work requiring sustained judgment, and benchmark scores capture only one facet of capability.1 The mid-tier became the default. It did not become the frontier.
9 articles · Newest first
News · Models
News · Agent security
Agent skill review · Third party
Book review · AI engineering
Concept · Agent security
Permissions & limits · Analysis
Execution · 10 lines
Evidence · Privacy
Limits · AI Security
Resources
Three free tools built from the code in PLEX articles, and the product we build for teams.
Pick a tool to open it. Everything runs in this page; nothing is uploaded.
About PLEX
PLEX is a technical source for people and software agents that read, decide, build and act on systems. It turns security and reliability ideas into mechanisms, code and evidence that help both work more effectively without quietly gaining authority they should not have.
Curated by Dorian Sotpyrc. Written so a person or software agent can extract the mechanism, assumptions, implementation and test — not just a conclusion.
What a person or agent may touch, and for how long. Granted for a task, never by default.
Where work has to stop: timeouts, budgets, sandboxes, egress and escalation.
What the work leaves behind, so a person or supervising agent can verify what happened.
The part that actually runs. Small, tested, reproducible and inspectable.
We design the boundaries before the behaviour. A system that can do anything will, eventually, do the wrong thing.
A policy document is not a control. If it isn't enforced by permissions, code or network, it doesn't exist.
A person or agent should be able to inspect the mechanism, its assumptions and the test without trusting a black box.
Every claim comes with something you can run, log or test — and a way to tell when it stops working.
Every control slows something down. We write with honest trade-offs, because the reader has to live with them.
Work with Dorian
If you're collaborating, commissioning work, or adapting these patterns for a team or agent stack, write directly. A short description of the system, what it can reach and what worries you is enough to start.
Map what your automated systems can reach, and where the gates should go.
Take a PLEX pattern into your own stack, with the tests that prove it works.
Bring a real problem. We'll work out the mechanism together and publish it.