More of your visitors are now agents sent by people, and they arrive without the context a person picks up in seconds. Making a site readable for them starts with one plain file.
Minimal animated illustration. A small robot agent rides a surfboard over rolling waves. Ahead of it, a coral flag reading AGENT.md flies from a marker float, with a dotted line of sight from the agent to the flag. Buoys labelled evidence, reports and opinion float past on the waves. Along the bottom, five stages light up in turn: discover, orient, navigate, verify, use.
Agent Surfing
Give an agent readable water and a map, and it can pick its own line.
Line chart of US money growth and consumer price inflation, quarter-end readings from 2019 to mid-2026. Money growth (M2) rises from 6.7% in December 2019 to 24.5% in December 2020, falls below zero in late 2022, reaches minus 3.9% in early 2023, and recovers to 5.3% by June 2026. Consumer price inflation peaks around 9.0% in June 2022, about eighteen months later, and is 3.5% in June 2026. Official data, with year-on-year rates calculated by PLEXData.
Fig. 1 · US money growth and pricesOfficial data
Money growth peaked in 2020–21 and went negative in 2023. Prices peaked in mid-2022. By June 2026 M2 was growing 5.3% a year and CPI inflation was 3.5%.
US money growth peaked first and prices followed. The bill landed on the people with the least room to absorb it.
30 Sep 2026Opinion
Opinion · Models
Bar chart of AutomationBench 1.0.6 correct completion rates at maximum effort, with cost per task. GPT-6.1 Sol completes 36.1% of workflows at $0.30 per task. Claude Opus 5.5 with default fallbacks completes 42.5% at $1.44. GPT-6 Sol completes 32.0% at $0.34. Vendor-reported figures from OpenAI's release chart, checked 30 September 2026. Business workflows, not a coding benchmark. No independent PLEXData testing.
Fig. 1 · AutomationBench 1.0.6Vendor-reported
Opus completes more workflows. GPT-6.1 Sol costs less per task. Completion and cost are separate measures here.
Max effort · Opus includes default fallbacks · checked 30 Sep 2026
A failed integration job, an Opus handoff that worked, and what the new Sol has to prove.
30 Sep 2026Opinion
News · Models
Bar chart of Anthropic-reported Terminal-Bench 4.0 scores, per cent of tasks completed: Claude Sonnet 5.5 leads at 70.6%, Claude Opus 5.5 scores 66.4%, and Claude Sonnet 5 scores 10.3%. This is vendor benchmark data, not PLEXData testing.
Fig. 1 · Terminal-Bench 4.0Vendor data
Sonnet 5.5 leads the chart at the same token price as Sonnet 5. At default effort, Anthropic says it beats Sonnet 5’s best score for about a tenth of the cost per task.
PLEX is a technical source for people and software agents that act on systems. It publishes practical patterns, code and evidence that help both do useful work within clear permissions, limits and security boundaries.
Built to be useful to a human reader, an agent, or both working together.
Most AI privacy language talks about intent: “we won't train on it”, “we don't look at your data”. Intent is a promise about behaviour you cannot observe. When a model sits next to your documents, the question that matters is simpler. Can it reach them?
A kill switch answers that question with mechanism, not policy text. It removes the connection, the fetch and the send. If none of the three can happen, a leak cannot happen either — whatever the model decides to do.
§ 01Why “privacy promises” fail in practice
Promises fail quietly. A connector is enabled for one team and inherited by another. A retrieval index is built from a shared drive that also holds board papers. A plugin gets network access “temporarily” and keeps it.
None of these is a breach on the day it happens. Each one is a path. Over time the paths add up, and nobody can say with confidence which data a model can read.
§ 02What a real kill switch is
A real kill switch has three properties, and all three must be enforced by infrastructure rather than by the model:
It denies by default. Access exists only where you granted it, for as long as you granted it.
It acts before content moves. The check happens at the boundary, not in a report the next morning.
It leaves evidence. Every denial is logged with who, what and when.
§ 03Start with one policy you can explain
Complex classification schemes are rarely enforced. Start with three labels and one default for each. If you can't explain the policy to a new starter in a minute, the gate will not be configured correctly either.
Table 1 — Three labels, three defaults
Label
Default action
What it covers
Public
Allow
Published material, documentation, marketing copy
Internal
Allow + log
Working documents, tickets, internal wikis
Secret
Deny
Credentials, customer data, board and legal papers
The labels connect to three enforcement checkpoints — prompt, retrieval and output — and to a separate control for tool execution.
§ 04The kill switch in one page
The whole design fits in one table. Each gate works at a different layer, so a failure in one does not open the others.
Table 2 — Where each gate sits
Stage
Control
Implementation
Pre-request
Identity & scopes
Least privilege and expiring tokens
In transit
Data-plane gate
Inspect, redact or block at prompt, retrieval and output
Tool execution
Sandbox & allowlists
Gateway-controlled, with approvals for risky actions
§ 05Pattern 1: Kill ambient access with scopes that expire
Ambient access is permission that nobody asked for today. It is the most common path to a leak. Replace it with scopes that are requested for a task and expire when the task ends.
A token that lives for fifteen minutes can still be misused — but only for fifteen minutes, and only for the scope it names.
§ 06Pattern 2: Put a policy gate in the data plane
The gate sits between the model and your data. It reads the label on every item, applies the default, and records the decision. Keep the policy in a file you can review like code:
Your AI Chat History Isn't Private — Here's What's Actually Logged
The chat window is only the visible record. Depending on the product and settings, the same message can also create usage logs, safety records, device metadata and retention obligations you never see.
An AI chat feels private because it looks like a conversation between you and a machine. There is no public audience, no comment thread and usually no obvious sign that anything exists beyond the transcript.
That interface is easy to mistake for the system.
It is not. The transcript is only the part you can see. The service around it may also process account identifiers, timestamps, device information, usage events, safety classifications, files, location signals and other operational data. Exactly what is retained — and for how long — depends on the provider, product, account type and settings.1
§ 01The chat is not the log
Start with the simplest distinction: conversation history is not the same thing as service logging.
OpenAI's current privacy policy, for example, separates user content from log data, usage data, device information and location information. Its examples of log data include IP address, browser type, settings and request time. Usage data can include features used, actions taken, access time, country, user agent and device type.1
Table 1 — One chat can create more than one record
Layer
Typical examples
Visible to you?
Conversation
Prompt, response, uploaded content
Usually
Account / usage
Time, model, feature use, account or workspace context
Sometimes
Device / network
IP address, browser, device type, general location
Not in chat history
Safety / integrity
Abuse signals, policy flags, feedback or review records
Usually not
That does not mean a human is reading every conversation. It means the product has more data surfaces than the transcript. Privacy decisions should start from those surfaces, not from how empty the sidebar looks.
§ 02Training is a separate question
One of the most common mistakes is to treat “not used for training” as if it meant “not stored”.
OpenAI states this directly: if you turn off Improve the model for everyone, new conversations are not used to train models, but regular chats can still appear in history. Saved chats remain until you delete them or a workspace retention rule removes them.2
The same distinction appears elsewhere. Anthropic gives consumer users a model-improvement choice, while its retention rules separately describe deletion, standard retention and longer periods for certain safety cases.56
§ 03Temporary is not zero retention
Temporary modes are useful. They are also easy to overread.
ChatGPT Temporary Chat stays out of normal history, does not create or update memories and is not used for model improvement while it remains temporary. OpenAI says a copy may still be retained for up to 30 days for safety purposes.3
Google documents a different window. Gemini temporary chats — and chats created while Keep Activity is off — are retained with the account for 72 hours so the service can respond, handle feedback and protect Google, users and the public. They do not appear in Gemini Apps Activity and temporary chats are not used to train Google's AI models.4
The practical lesson is not that temporary modes are bad. It is that temporary means a different retention path. It does not automatically mean zero records.
§ 04Providers do this differently
There is no universal “AI chat privacy setting”. The names, defaults and retention periods vary.
This table is deliberately narrow. It does not try to flatten every product tier into one number. Business, enterprise, education, API and managed-workspace products often have different contracts and retention rules.
§ 05Delete does not mean instant erasure
Deleting a chat is still worth doing. But the button usually changes the user-facing state before every backend copy is gone.
01DELETEremoved from your account view
→
02DELETION QUEUEprovider removes backend copies
→
03EXCEPTIONSlegal, security or de-identified data may follow different rules
Fig. 2Example pattern, not a universal timer. OpenAI says deleted saved chats are scheduled for permanent deletion within 30 days, subject to stated legal, security and de-identification exceptions.
OpenAI's current retention guidance says a deleted saved chat disappears from the account view immediately and is scheduled for permanent deletion from its systems within 30 days, unless the chat was already de-identified and disassociated from the account or longer retention is required for security or legal reasons.8
Anthropic's consumer guidance similarly says deleted conversations are removed immediately from history and automatically deleted from the backend within 30 days under its standard rule, while documenting longer retention for some usage-policy violations and trust-and-safety classification scores.5
§ 06Work accounts change the boundary
Your employer's AI workspace is not the same privacy boundary as a personal chat account.
OpenAI's managed-account guidance says organisation-controlled accounts may expose submitted content, conversation history, shared workspace content, usage/activity metadata and certain security or privacy settings according to workspace configuration, permissions and applicable law.9
That is not automatically a problem. In a business setting, auditability can be a feature. The important point is to know which boundary you are using before you paste confidential material into it.
Connected apps add another boundary. If a chat sends data to a third party, that third party's privacy and retention rules can apply too. The original AI provider's temporary-mode promise does not automatically follow the data downstream.
§ 07What to do before you paste something sensitive
You do not need to memorise every privacy policy. You need a short decision routine.
6Anthropic — Consumer Terms and Privacy Policy updateanthropic.com
7Microsoft Support — Copilot privacy controls / conversation historysupport.microsoft.com
8OpenAI Help — Chat and file retention in ChatGPThelp.openai.com
9OpenAI Help — Data access for your managed ChatGPT accounthelp.openai.com
Policies change. This article records the provider documentation available on 29 September 2026. Check the linked source before relying on a retention period for regulated or high-risk work.
001
Published
Updated 29 Sep 2026
Author
Dorian Sotpyrc
Reading
10 minutes · 7 sections
Fields
AI Security
Identity
Tooling
Observability
Limits · AI Security
The Fortified AI Stack: How Teams Are Locking Down ML Workflows in 2026
The model is no longer the security boundary. Identity, data, tools, runtime, network and evidence all need their own gate — and their own owner.
The old mental model for AI security was a protected model endpoint: authenticate the caller, encrypt the traffic and keep the API key out of Git. That is no longer enough once the model can search internal data, call tools, write files or trigger another service.
The security boundary has moved outward. A production workflow now includes identity, context, retrieval, tools, runtimes, networks and logs. A weak assumption in any one of those layers can become authority for the whole chain.
That is the useful meaning of a “fortified AI stack”. It is not six products around a model. It is six independent decisions about what the workflow is allowed to know and do.
§ 01The model is not the boundary
NSA's May 2026 guidance on Model Context Protocol makes the shift explicit. Authentication, authorization and input validation remain necessary, but agentic systems add dynamic tool invocation, implicit trust relationships and context sharing. NSA's conclusion is the important part: the environment has to be treated as a continuum, because a bad assumption in one stage can propagate into the next.1
OWASP describes the same problem from the application side. Its Excessive Agency risk is usually caused by too much functionality, too much permission or too much autonomy.2
So the first design question is not “how do we stop the model hallucinating?” It is “what happens if the model makes the wrong decision?” A fortified stack assumes that bad outputs, malicious inputs and compromised context are possible, then limits the blast radius.
Table 1 — Six boundaries, six owners
Boundary
Question
Owner
Evidence
Identity
Who or what is acting?
IAM / platform
Token subject, scope, expiry
Secrets & data
What can it read?
Data / security
Classification, retrieval decision
Policy
What may it do?
App / security
Allow, deny, approval
Isolation
Where can code execute?
Platform
Sandbox, filesystem, egress
Monitoring
What is happening now?
Ops / SOC
Runtime events, alerts
Audit
What can we prove later?
Risk / security
Immutable trace, retention
§ 02Identity before intelligence
Every agent, workflow and tool call should have an identity that can be limited independently. Reusing one broad service account for every model task is convenient, but it destroys the boundary between “the model needs this file” and “the service account can read the whole drive”.
The practical pattern is boring on purpose: short-lived credentials, task-specific scopes, separate identities for different trust levels and no permission inherited merely because a connector happens to expose it.
OWASP's examples of excessive agency include tools that expose functions the task does not need and downstream identities that have write or delete permission when read-only access would have been enough.2
§ 03Secrets and the data path
Secrets belong in a secret store or credential broker, not in the system prompt. OWASP's 2025 guidance is blunt: the system prompt should not be treated as a secret or as a security control, and credentials or connection strings should not be placed there.3
The same discipline applies to retrieval. A model should not receive a document simply because the retrieval system can find it. Classification, user entitlement and task context should be checked before content enters the model context.
That gives you a clean separation:
Identity says who is asking.
Data policy says whether this identity may receive this object for this task.
The model only sees the content after both checks pass.
§ 04Tools need policy outside the model
Tool use is where a chat system becomes an operational system. Reading a calendar is different from sending an email. Listing files is different from deleting them. A single “tools enabled” switch is too coarse once the workflow can change state.
Use a small policy layer between the model and each tool. The model can request an action. The policy layer decides whether the action is allowed, whether it needs human approval and what arguments are acceptable.
Table 2 — Class the action before the tool runs
Action
Default
Example
Read
Allow + log
Search approved documentation
Create
Policy check
Create a draft ticket
External send
Approval
Send email or publish content
Delete / privilege
Deny by default
Delete file, change access
§ 05Contain execution and egress
If the workflow can run code, treat the runtime as untrusted work. Use a short-lived sandbox with the smallest filesystem view possible. Mount only the data the task needs. Keep privileged sockets and host credentials out of reach.
Then control egress. A workload that can reach any address on the internet can turn a prompt-injection failure into data exfiltration. A gateway or allowlist gives you one place to say which destinations exist for this task and to record what left.
NSA and partner guidance on deploying AI systems securely has long framed AI security as protection of the model, data and surrounding infrastructure, not just the inference endpoint.4
§ 06Monitoring is useful only if it leaves evidence
NIST's Generative AI Profile treats monitoring and testing as lifecycle work, including post-deployment monitoring, incident handling and provenance. NIST updated the publication in April 2026, but the operational point has not changed: deployment is not the end of evaluation.5
For a production AI workflow, one trace should let you reconstruct:
which user or workload initiated the task;
which model and policy version ran;
which documents entered context;
which tools were requested and allowed;
which network destinations were contacted;
which approval or denial changed the path.
Logging everything is not the answer. Log the decisions that define authority. Protect those logs from casual modification and give every event a trace ID that follows the workflow across services.
§ 07A practical build order
Do not begin by buying six security products. Begin by making the boundaries explicit, then fill the gaps in an order that reduces authority fastest.
10 Python Lines I Trust in Production (And the Ones I Don't)
The lines I keep are not clever. They put a bound on waiting, surface failure, make security explicit or remove an assumption the runtime would otherwise make for me.
Production Python is mostly ordinary Python with fewer silent assumptions.
The lines I trust are not universal recipes. Each one closes a specific failure mode: waiting forever, accepting a bad response, losing a traceback, generating the wrong kind of token, building SQL from text, writing ambiguous timestamps or letting the platform choose an encoding.
The test is simple: when this line fails at 03:00, does it stop, raise, log or preserve enough evidence for the next person to understand what happened?
07cursor.execute("SELECT * FROM jobs WHERE id = ?", (job_id,))
08stamp = datetime.now(UTC).isoformat()
09os.replace(tmp_path, final_path)
10text = Path(path).read_text(encoding="utf-8")
§ 01Bound the wait
Line 1:requests.get(url, timeout=(3.05, 10))
Requests does not time out by default. Its own documentation says most external requests should have a timeout, and a tuple lets you separate the connection wait from the read wait.1
The line I do not trust is the shorter one:
requests.get(url)
It looks clean until a remote service stops answering and a worker spends an unbounded amount of time waiting for somebody else's network.
Line 2:response.raise_for_status()
A completed HTTP request is not the same thing as a successful application request. Requests exposes raise_for_status() so 4xx and 5xx responses become exceptions instead of quietly travelling deeper into the program.1
§ 02Make failure loud
Line 3:subprocess.run(args, check=True, timeout=30)
check=True turns a non-zero exit into a CalledProcessError. timeout=30 puts a ceiling on the wait. The Python docs recommend run() for common subprocess work and document both behaviours.2
The line I do not trust is subprocess.run(args) when the return code matters. It can fail and hand control back as if nothing happened.
Line 4:logger.exception("job failed")
Inside an exception handler, Logger.exception() records the message and traceback. That is a better operational artifact than print(exc), which often throws away the path that led to the error.3
§ 03Treat secrets as secrets
Line 5:secrets.token_urlsafe(32)
Python's secrets module exists for cryptographically strong random values used in password resets, hard-to-guess URLs and similar security-sensitive cases. The docs still use 32 bytes as a typical security level for tokens.4
The line I do not trust for security tokens is anything built from random. That module is useful for simulation and sampling. It is the wrong source for a reset link.
Line 6:hmac.compare_digest(received, expected)
For HMAC or other secret-derived values, Python recommends compare_digest() rather than ordinary equality to reduce timing-analysis exposure.5
§ 04Keep data boundaries explicit
Line 7:cursor.execute("SELECT * FROM jobs WHERE id = ?", (job_id,))
The value is data, not SQL. Python's sqlite3 documentation explicitly warns against building queries with string operations and recommends parameter substitution instead.6
The line I do not trust is an f-string that places user data inside the SQL text.
Line 8:datetime.now(UTC).isoformat()
An aware UTC timestamp says what instant it represents. Python documents naive datetimes as ambiguous and deprecated datetime.utcnow() in Python 3.12 in favour of an aware UTC datetime.7
§ 05Files need boring guarantees
Line 9:os.replace(tmp_path, final_path)
When the temporary file and destination are on the same filesystem, os.replace() gives you replacement semantics without exposing a half-written destination file. Python documents a successful replacement as atomic where the operating system provides that guarantee.8
This line is the final move, not the whole durability story. If the file must survive power loss, you still need the appropriate flush and filesystem durability strategy before replacement.
Line 10:Path(path).read_text(encoding="utf-8")
Explicit encoding removes a machine-dependent assumption from a text boundary. Path.read_text() accepts the encoding directly and closes the file for you.9
The line I do not trust is Path(path).read_text() when the file format says UTF-8. Let the file format decide the encoding, not whichever workstation happens to run the code.
§ 06What these lines cannot do
A trusted line is not a trusted system. Timeouts need retry policy. Retries need idempotency. Atomic replacement needs a correctly staged temporary file. Parameterised SQL does not replace authorization. Logging a traceback does not replace monitoring.
The value of these lines is smaller and more useful: each one removes a category of silent ambiguity.
Stop Trusting the Agent: Map the Authority Path Instead
An agent is not powerful because it can reason. It becomes powerful when identity, context, tools and credentials line up into a path that can change something real.
“Is this agent trustworthy?” is becoming the wrong security question.
A model can be careful and still sit behind an overpowered service account. A model can be unreliable and still be harmless inside a read-only sandbox. The practical risk appears when a sequence of components gives a model a path from text to effect.
NIST's 2026 work on software-agent identity and authorization focuses on exactly this shift: agents need identification, authorization, auditing and non-repudiation controls because they increasingly act across data sets, tools and applications.1 NSA's MCP guidance reaches the same conclusion from another direction, warning that dynamic tool invocation, implicit trust relationships and context sharing create risks that do not stop at one endpoint.2
§ 01Identity is not authority
An identity answers who or what is acting. Authority answers what that identity can cause.
Those two ideas are easy to blur in agent systems because the agent often inherits a ready-made credential. A connector signs in once, the agent sees a tool, and the implementation begins to treat “tool available” as “tool allowed”.
That shortcut removes the useful questions: allowed for which user, for which task, against which resource, for how long, and with what evidence?
Table 1 — An authority path is a set of edges
Edge
Decision
Evidence
User → agent
What task was actually requested?
Task ID, user identity, scope
Agent → identity
Which credential may be used?
Subject, token scope, expiry
Identity → tool
Which operation is allowed?
Policy decision, approval
Tool → resource
Which object may be touched?
Resource ID, classification
Resource → effect
What state can change?
Before/after event, audit trace
§ 02The authority path
Think of authority as a graph, not a property of the agent. The nodes are users, agents, identities, tools and resources. The edges are the permissions that allow one node to affect the next.
This changes architecture reviews. Instead of asking whether the agent has “email access”, ask whether this task can move through a chain that ends in an external send. The same email connector might be safe for search and unsafe for sending.
OpenAI's current agent-safety guidance makes a similar practical point: risk rises when untrusted content can influence tool calls, and approvals, structured data flow and restricted tool use reduce the consequences when manipulation succeeds.3
§ 03Delegation compounds faster than it looks
A sub-agent can look less privileged while the workflow around it remains more privileged. It may receive no credential directly but still call a parent tool that carries one. Or it may write a file that another process later executes.
The effective authority is therefore the union of reachable effects, not the list of tools shown in one prompt.
§ 04Context can steer power without owning it
Prompt injection matters because text can become a steering input to an authority path. The malicious document does not need its own credential. It only needs to persuade a component that already has one.
OpenAI describes prompt injection as a form of social engineering against the agent and recommends limiting access, narrowing instructions and reviewing important actions.4
That makes context provenance part of authorization. A decision influenced by public web text should not automatically inherit the same authority as a decision based on an authenticated user instruction.
§ 05Put the final policy decision outside the model
The model can propose an action. It should not be the final authority on whether the action is allowed.
Use deterministic checks for identity, resource scope, action class, rate, destination and approval state. The policy does not need to understand every thought. It needs to understand the proposed effect.
This also gives agents something useful: a clean denial is better than a vague refusal. A denied tool call can return the exact missing scope or required approval without granting anything new.
§ 06Audit the edges, not the monologue
A full reasoning transcript is neither necessary nor sufficient for security evidence. What matters operationally is the chain of authority decisions.
AI Engineering by Chip Huyen — Strongest When It Refuses the Shortcut
A broad systems book for building with foundation models. Its lasting value is the decision framework; its weakest area in 2026 is the security depth needed for agents that can act.
Editorial judgment. Criteria are fixed before the total; scores are not publisher or reader ratings.
Systems framing19/20
Evaluation20/20
Production practicality18/20
Agent/security depth14/20
Shelf life17/20
The best technical books teach a sequence of decisions, not a sequence of tools. AI Engineering mostly does that.
Chip Huyen's book is explicitly not a code-along tutorial. Her companion repository says the aim is to provide a framework for adapting foundation models to real applications, with questions around evaluation, RAG, agents, fine-tuning, data, inference, latency, cost and feedback loops.1
That choice is why the book has held up better than a framework-specific title would have. O'Reilly lists 534 pages and describes an end-to-end progression from foundation-model applications through evaluation, prompting, RAG and agents, fine-tuning, dataset engineering, inference optimisation and feedback.2
§ 01What the book is
This is a map of AI application engineering for people who already know how software projects behave. It spends less time telling you which SDK to install and more time asking what should be evaluated, what should be retrieved, what should be fine-tuned, and what should remain outside the model.
The official table of contents makes that breadth visible: foundation models, two chapters on evaluation, prompting, RAG and agents, fine-tuning, datasets, inference optimisation, architecture and user feedback.3
The cost of that breadth is obvious too. You do not finish the book with one application assembled. You finish with a larger set of engineering questions.
§ 02The real centre of the book is evaluation
The strongest choice is architectural rather than topical: evaluation appears before most of the fashionable adaptation techniques.
That sequencing matters. If you cannot state what “better” means, prompt changes, RAG, model swaps and fine-tuning become activity rather than engineering. The book treats open-ended output evaluation as a first-class system problem, including human evaluation, model-based evaluation and task-specific criteria.
For PLEX readers, this is the most transferable lesson: define evidence before optimisation.
§ 03What ages well
The book's durable material is the material least tied to model names: application scoping, evaluation, context construction, data quality, inference trade-offs and feedback loops.
Huyen's own companion repo says the book focuses on fundamentals rather than a particular tool or API because tools age quickly.1 That editorial decision has paid off. The names of models have changed; the questions around latency, cost, data, evals and whether to retrieve or fine-tune have not disappeared.
§ 04Where 2026 has moved on
The weakest score is not because the agents chapter is poor. It is because production agents have moved from “models with tools and planning” toward identity, delegated authority, sandboxes, runtime policy and prompt-injection-resistant workflows.
Chapter 6 covers tools, planning, agent failure modes, evaluation and memory.4 What it cannot fully reflect is the security architecture that became more explicit through 2026: NIST work on agent identity and authorization, NSA MCP guidance, and runtime controls around agent action.
That is not a reason to skip the chapter. It is a reason to pair it with newer agent-security material.
Strong
Evaluation discipline
The book makes quality measurable before it makes the architecture more complicated.
Strong
Tool-agnostic framing
Concepts survive model and framework churn better than implementation recipes.
Pair with newer work
Agent security
Identity, runtime boundaries and delegated authority deserve a 2026 companion reading list.
§ 05Who should read it
Read it if you are moving from model demos toward a production application and need a coherent map of the decisions.
Read selected chapters if your work is already specialised. Evaluation, RAG/agents, inference and architecture can stand alone.
Do not buy it expecting a current framework cookbook. The author says it is not a tutorial book, and that is the point.1
Ponytail: The Best Code May Be the Code Your Agent Never Writes
A sharply scoped skill for fighting agent over-engineering. Its strength is not “write fewer lines”; it is the decision ladder that asks whether the extra machinery needed to exist at all.
Static review of the skill, docs and published benchmark methodology. PLEXData did not rerun the benchmark for this article.
Activation clarity18/20
Scope discipline19/20
Portability18/20
Evidence16/20
Safety & limits15/20
Ponytail has one unusually useful opinion: the default failure mode of a coding agent is often to build too much.
The core skill describes itself as a “lazy senior dev” mode and pushes the agent through a ladder: YAGNI, standard library, native platform, one line, minimum.1 That is more interesting than the marketing phrase “less code”, because it gives the model an order of operations.
§ 01What Ponytail does
The main skill is broad: coding, adding, refactoring, fixing, reviewing, designing and dependency selection. Around it sit narrower skills for over-engineering review, whole-repository audit, debt tracking, measured impact and help.
The review skill is particularly clean. It asks for one-line findings focused only on complexity: what to delete and what replaces it.2 The audit skill scales the same idea to a whole repository and explicitly says correctness, security and performance are out of scope.3
That boundary is a feature. A skill that tries to be code quality, security, architecture and minimalism at the same time becomes hard to activate and harder to evaluate.
§ 02Why the scope works
The best part of Ponytail is that it names the alternatives in a useful order. Before adding a dependency, ask if the platform already has the feature. Before adding an abstraction, ask whether there is more than one implementation. Before writing a wrapper, ask whether a direct call is clearer.
This is agent-friendly because each step is observable. “Be elegant” is subjective. “Use the standard library before a new dependency” can be checked.
The project has also invested in portability. Its documentation describes adapters or instruction tiers for Claude/Codex-style skills, OpenClaw, Gemini CLI, Cursor, Windsurf, Cline, Copilot, Qoder, Zed and others.4
§ 03The benchmark is useful because the project corrected its own claim
The current README reports roughly 54% less code on its agentic benchmark, with up to 94% reduction on tasks that invite over-building, plus lower cost and latency in that test setup.5
More important than the number is the benchmark note. The repository says its earlier single-shot result overstated the win because the baseline included prose and options; the later agentic benchmark is presented as the more defensible comparison.6
Source-backed
Clear skill boundary
Review/audit modes explicitly focus on over-engineering rather than pretending to replace security or correctness review.
Project benchmark
Large code reductions
The published numbers come from the project's own harness. They are useful evidence, not independent proof for every repo or model.
Needs local test
Transfer to your stack
The benchmark notes weaker transfer to small local models. Teams should test against their own tasks and harness.
§ 04Where it can go wrong
Minimalism is not a universal objective. Some code is longer because it carries evidence, compatibility, observability, safety checks or a deliberately explicit boundary.
The project's own “100% safe” benchmark claim is scoped to its benchmark. It should not be read as “Ponytail always preserves every security property”.
The skill is safest when the requirements are already clear. In a poorly specified task, “do less” can become “silently omit the thing nobody remembered to state”.
§ 05Who should use it
Strong fit: mature codebases where agents routinely introduce wrappers, dependencies, factories and “future flexibility”.
Good review tool: the dedicated review/audit skills are easier to trust than an always-on minimalism mode because they propose cuts without applying them.
Use carefully: greenfield systems with incomplete requirements, safety-critical code, and code where redundancy is an intentional control.
NVIDIA Moves Agent Safety Below the Model With OpenShell and Sentry
Most agent safety asks the model to behave. NVIDIA’s launch is built around a harder question: what still holds when it doesn’t? The interesting answer sits outside the model, and optionally outside the host software too.
In this conceptual policy example, a model requests a file read and an outside send. Runtime policy allows the read and blocks the send. This is not a test of a named product.
Fig. 1 · Outside the modelConcept model
The model can request it. The runtime can refuse it.
NVIDIA announced the Open Agent Safety Platform on 28 September 2026, built around OpenShell and the Sentry reference design.
Vendor claim
Could stop escapes
NVIDIA says Sentry can quarantine agents that move outside policy boundaries in milliseconds. That is a product claim, not a universal result.
Open question
Real-world policy quality
Runtime enforcement is only as useful as the permissions and policies organisations define. Deployment evidence will matter more than launch architecture.
Most of us have handed an agent a credential and trusted it to use it sensibly. That works right up until it doesn’t, and this launch is about the second half of that sentence.
On 28 September, NVIDIA launched the Open Agent Safety Platform: OpenShell, an open-source runtime boundary, plus Sentry, an optional out-of-band enforcement design that uses BlueField-4 DPUs.1
What caught my attention is the split. The model decides what to attempt. The environment decides what is permitted. NVIDIA’s own product page draws the same line: model safeguards influence behaviour; runtime controls enforce allowed actions.2 I think that is the right line, and it is an easy one to blur when a model is doing something impressive.
§ 01Two layers, one idea
The launch combines two layers.
OpenShell — an open-source runtime that places agents in sandboxed environments governed by policy.
Sentry — a reference system design that runs an independent watchdog on BlueField-4 DPUs and can enforce policy outside the agent workload.
NVIDIA says OpenShell can also extend to third-party compute platforms, including Arm and Intel systems. The company listed a large set of launch partners across AI labs, enterprise software, security and infrastructure.1
23 Mar 2026
OpenShell appears as a secure agent runtime
NVIDIA describes a policy layer outside the model/application process for agent files, tools, credentials and network access.
17 Sep 2026
OpenShell docs show a mature runtime surface
Current documentation covers sandbox policy, provider routing, observability and supported agents.
28 Sep 2026
Open Agent Safety Platform launches
OpenShell becomes the software layer in a broader design that adds Sentry as an independent hardware-backed watchdog.
§ 02OpenShell lets a capable agent work in a smaller room
OpenShell's current developer guide describes sandboxed execution with controls over files, process behaviour, network access, credentials and inference routing. It also exposes allow/deny logging and policy configuration.3
The new technical blog describes Gateway, Supervisor and Sandbox components and says policy is enforced without requiring the agent itself to be rewritten.4
The useful shift is that you don’t have to make the agent less capable. You make the room it works in smaller.
§ 03Sentry starts from the assumption that the host might fail
Sentry is designed to run on BlueField-4 DPUs as an out-of-band monitor. NVIDIA says it can correlate activity, enforce identity and access policies, and quarantine an agent that attempts to leave its software boundary.1
That is a more honest starting point. It assumes something can go wrong with the workload or the host, and plans for it: if either is compromised, some monitoring and enforcement still exists outside it.
AP's launch coverage describes the same basic split—OpenShell as the constrained workspace and Sentry as the watchdog—and notes that defining effective rules remains a hard problem.5
§ 04What shipped, what is claimed, and what nobody knows yet
Confirmed: the platform, code/docs and reference architecture exist; OpenShell is open source; the launch names partners; the runtime exposes concrete policy and observability features.
Claimed: NVIDIA says the architecture could have prevented recent agent security incidents and that Sentry can quarantine escapes in milliseconds. Reuters reports that claim as NVIDIA's assertion, not an independently established result.6
Still open: how well organisations will write policies, how performance behaves under real mixed workloads, how easy bypasses are across different deployment modes, and whether hardware-isolated enforcement becomes common outside NVIDIA-heavy environments.
§ 05The part worth keeping, even if you never buy the hardware
The direction is larger than NVIDIA hardware. Agent security is moving from model-only controls toward layered enforcement: identity, runtime isolation, network policy, tool authorization and audit.
The signal I take from it: the better agents get at pursuing goals, the more the infrastructure around them has to say, plainly, which paths are impossible. Not discouraged. Not logged. Impossible.
Anthropic’s new Sonnet posts frontier-class scores at Sonnet prices. The number that matters is not the 70.6% headline — it is the cost per task at the effort level you actually run.
Bar chart of Anthropic-reported Terminal-Bench 4.0 scores, per cent of tasks completed: Claude Sonnet 5.5 leads at 70.6%, Claude Opus 5.5 scores 66.4%, and Claude Sonnet 5 scores 10.3%. This is vendor benchmark data, not PLEXData testing.
Fig. 1 · Terminal-Bench 4.0Vendor data
Sonnet 5.5 leads the chart at the same token price as Sonnet 5. At default effort, Anthropic says it beats Sonnet 5’s best score for about a tenth of the cost per task.
Anthropic released Claude Sonnet 5.5 on 28 September 2026. It is available on all major platforms as claude-sonnet-5-5, priced the same as Sonnet 5: $2 per million input tokens and $10 per million output.
Vendor claim
Frontier class, a tenth of the cost
Anthropic reports 70.6% on Terminal-Bench 4.0 against Sonnet 5’s 10.3%, and says that at default effort it beats Sonnet 5’s best score for about a tenth of the cost per task. These are vendor numbers.
Open question
Your workload, your effort level
Gains depend on task mix and effort settings, and Anthropic itself says Opus 5.5 remains clearly stronger on complex, open-ended work. Independent replication is pending.
Choosing a model tier used to be a real decision: pay for the frontier model when the work is hard, drop a tier when it isn’t, and accept the gap in between. This launch is Anthropic’s attempt to close that gap — and if the numbers hold, the default for everyday agent work just moved down a tier.
On 28 September, Anthropic released Claude Sonnet 5.5, the second model in its 5.5 family after Opus 5.5 arrived on 22 September. A Haiku 5.5 for high-volume work is promised in the coming weeks.1 The pitch is blunt: a clear upgrade over Sonnet 5, 30% faster, and up to 30% cheaper per task — at exactly the same token price.
What caught my attention is not the jump itself. Model releases always jump. It is where the jump lands: the tier most teams actually afford for agents that run all day.
§ 01What changed on 28 September
The confirmed part is simple. Sonnet 5.5 is live in the Claude apps, Claude Code, and on the Claude Platform, AWS, Google Cloud and Azure.1 Pricing is unchanged from Sonnet 5, and Anthropic is publishing a system card with its evaluation method.2
API pricing, per million tokens
Item
Sonnet 5.5
Opus 5.5
Input
$2
$4
Output
$10
$20
Cache reads
$0.20
$0.20
Cache writes
$2.50
$5
One migration detail matters if you run agents with thinking disabled: you need to switch to the new between_tools setting, which keeps up-front thinking off, before moving to Sonnet 5.5. Anthropic’s migration guide covers it.3
§ 02The number that matters is cost per task, not the headline score
The headline number is real, and it is large. On Terminal-Bench 4.0, an agentic coding evaluation, Anthropic reports Sonnet 5.5 at 70.6% — against 10.3% for Sonnet 5 and 66.4% for Opus 5.5. On CursorBench 4.0, built from real Cursor coding sessions, it reports 55.5% against Sonnet 5’s 34.1%, within about two points of Opus 5.5.1
Anthropic-reported scores, 28 September 2026 announcement
Evaluation
Sonnet 5
Sonnet 5.5
Opus 5.5
Agentic coding · Terminal-Bench 4.0
10.3%
70.6%
66.4%
Agentic coding · CursorBench 4.0
34.1%
55.5%
57.8%
Knowledge work · GDPval-AA v2.1
1449
1844
1846
Computer use · OSWorld 2.1
57.0%
80.1%
81.8%
Chart recognition · Chartography
15.6%
61.6%
64.4%
But scores are half the story. Anthropic’s own charts plot score against cost per task at every effort level, and that is where the release bites: at Medium effort — the default in the Claude apps — it says Sonnet 5.5 beats Sonnet 5’s best score for less than a tenth of the cost per task.1 Same price list, fewer tokens per job, faster output. If that holds outside Anthropic’s harness, the cost model for routine agent work changes, not just the leaderboard.
Slack’s early testing points the same way, with the usual early-tester caution: better results than Sonnet 5 on almost all of its offline Slackbot evals, in fewer steps, with about 14% fewer output tokens — a vendor-published tester quote, not an independent result.1
§ 03Effort is now the knob, and the default moved
Sonnet 5.5 leans hard on effort levels: Low, Medium, High and beyond, trading speed and tokens against thoroughness. The defaults are Medium in Claude Code and the Claude apps, High on the Claude Platform.4
Anthropic’s own framing is that Sonnet 5.5 complements Opus 5.5 at lower effort, where it costs less per task, and can match it at higher effort at comparable cost. At High effort on FrontierCode it reports a score ten points above Sonnet 5 at the same setting — at roughly one fifteenth of the cost per task.1
My read: this makes effort a routing decision, not just a quality setting. The practical question for an agent pipeline is no longer “which model” but “which model at which effort for which job class”. Teams that pin one model and one effort for everything will pay for it twice — once in tokens, once in capability left on the table.
§ 04What shipped, what is claimed, and what nobody knows yet
Confirmed: the model exists, the model ID is live, pricing is published, the system card is public, and the safety posture is documented — the same biology safeguards as Sonnet 5, visible fallbacks to Sonnet 5 on higher-risk cybersecurity tasks, and classifiers against reasoning extraction, a first for the Sonnet line.12
Claimed: the benchmark table, the 30% speed and cost figures, the “a tenth of the cost per task” comparisons, and the collaboration anecdotes from Epic and Slack are all Anthropic-reported. Anthropic also reports that on its roughly 1,850-scenario behavioral audit, Sonnet 5.5 improves on or matches Sonnet 5 on most alignment measures, and that on containment evaluations it comes close to Opus 5.5 on how rarely it tries to escape its sandbox.1 Vendor-reported alignment is still vendor-reported.
Still open: independent Terminal-Bench and CursorBench replication; whether the cost advantage survives long-horizon, tool-heavy workloads; how the between_tools migration plays out in existing agent stacks; and where Haiku 5.5 lands the bottom tier when it ships.
§ 05What I would do before switching anything
First, re-run your own tasks at Medium effort before touching defaults. Vendor charts measure vendor-chosen tasks; your eval set is the only one that knows your workload.
Second, treat effort as part of routing. If you run an agent pipeline, this release is a good reason to log effort level alongside model and cost per task — you cannot tune a knob you do not record.
Third, keep Opus 5.5 for the work that burns. Anthropic is unusually plain that Opus remains clearly stronger on complex, open-ended work requiring sustained judgment, and benchmark scores capture only one facet of capability.1 The mid-tier became the default. It did not become the frontier.
Personal account · benchmark figures vendor-reported
Opinion · Models
GPT-6.1 Sol Has to Earn Back My Trust
Hours of unsuccessful revisions with GPT-6 Sol, an Astra escalation that exhausted my usage, and an Opus 5.5 handoff that finally worked. OpenAI’s new release arrives with a better value proposition, and some trust to recover.
Bar chart of AutomationBench 1.0.6 correct completion rates at maximum effort, with cost per task. GPT-6.1 Sol completes 36.1% of workflows at $0.30 per task. Claude Opus 5.5 with default fallbacks completes 42.5% at $1.44. GPT-6 Sol completes 32.0% at $0.34. Vendor-reported figures from OpenAI's release chart, checked 30 September 2026. Business workflows, not a coding benchmark. No independent PLEXData testing.
Fig. 1 · AutomationBench 1.0.6Vendor-reported
Opus completes more workflows. GPT-6.1 Sol costs less per task. Completion and cost are separate measures here.
Max effort · Opus includes default fallbacks · checked 30 Sep 2026
GPT-6 Sol did not fix the integration issue, Astra used up my allowance, and Claude Opus 5.5 delivered the functionality with roughly 80% less code. That is my account of one job, not a controlled test.
Vendor claim
The new Sol is cheaper per task
OpenAI’s chart has GPT-6.1 Sol at 36.1% correct completion and $0.30 per task on AutomationBench. Opus 5.5 with fallbacks completes more, at $1.44.
Open question
Does it fix that job?
I have not used GPT-6.1 Sol on this incident. The release earns a retest, and only finished work earns the trust.
Last week, I spent hours trying to fix an issue in a large codebase connecting several systems. GPT-6 Sol kept circling the problem, giving me iterations of the code that achieved nothing. The root cause remained unresolved, and the functionality I needed still did not work.
I eventually handed the job to GPT-6 Astra. My available usage was exhausted almost immediately, before it had completed the work. I then moved it to Claude Opus 5.5.
Opus reduced and simplified the bloated codebase by roughly 80%, delivered the desired functionality, and integrated it correctly. After the previous attempts, that was a substantial difference.
The experience left me disappointed with OpenAI. It also changed what I wanted to hear about its next model. A claim about greater intelligence matters much more when it translates into a finished job.
GPT-6.1 Sol now arrives with reported improvements in capability and cost efficiency. I hope those improvements reach the kind of work that went badly last week. That remains an open question: I have not used it on this incident.
§ 01The code changed. The problem remained.
The frustrating part was the repeated lack of progress. Another explanation, another revision, another attempt, and the same underlying issue.
There is a cost to that beyond the model’s usage allowance. You have to follow the changes, keep track of what has been attempted, and decide whether the next explanation deserves more of your time. A model that produces a plausible patch can keep a session moving long after it has stopped being useful.
Opus’s smaller solution made an impression because it also worked. The approximate 80% reduction is my account of that job, rather than an independently measured benchmark. The models received sequential handoffs, so this was not a controlled comparison from identical starting conditions.
That limits what I can claim about the models in general. It does not change which one completed the work I needed.
§ 02Other users report similar frustration
Public complaints after the move from GPT-5.6 Sol to GPT-6 Sol describe problems that sound familiar. One Reddit user reported instruction-following failures and a costly return to Astra. Another described a maintenance check in which GPT-6 Sol relied on stale package information while its predecessor checked more thoroughly.12
These are self-selected accounts. They show that similar frustration exists, but they do not establish how common it is or whether the same cause explains my experience.
A more detailed report in OpenAI’s Codex issue tracker compared three pairs of tasks at maximum reasoning effort. Its author recorded context and completion omissions, but also a serious database defect found by GPT-6 Sol and missed by GPT-5.6 Sol. One pair had a contamination caveat.3
The picture is mixed. There is evidence worth taking seriously, including evidence against a blanket claim that GPT-6 Sol is worse. My disappointment is specific: on a difficult integration job, repeated attempts consumed time without delivering the required result.
§ 03Completed work is the more useful price comparison
The metric closest to that concern is cost per successful task: total spending across attempts divided by the number that finish correctly. Arize and Fireworks have published a benchmark using that approach, although its model set does not cover this three-way comparison.6
AutomationBench provides a useful comparison here. It tests business workflows across applications and grades the final system state. A task passes only when every required assertion holds. An agent saying it has finished does not count as success.5
OpenAI’s GPT-6.1 Sol release chart reports the following AutomationBench 1.0.6 results. All three rows use the named maximum effort setting; that does not imply identical compute budgets across providers. The Opus configuration includes default fallback models.45
AutomationBench 1.0.6, maximum effort, vendor-reported
Model / configuration
Correctly completed tasks
Reported cost per task, USD
Approx. cost per correct completion, USD*
GPT-6.1 Sol · Max
36.10%
$0.2989
$0.83
Claude Opus 5.5 · default fallbacks · Max
42.47%
$1.4400
$3.39
GPT-6 Sol · Max
31.96%
$0.3406
$1.07
*The final column is an editorial calculation from the published figures: reported cost per task ÷ success rate expressed as a fraction. For example, $0.2989 ÷ 0.361 = approximately $0.83. It treats the reported cost as the average across attempted tasks. It is not a separately measured result, a guarantee that retries will solve a particular failure, or a forecast of subscription usage.
In these results, Opus completes the largest share of tasks. The new Sol has the lowest reported cost per task and the lowest calculated spend per correct completion. Those are different advantages, and both matter.
AutomationBench concerns business workflow execution, rather than debugging my codebase. OpenAI’s published DeepSWE coding comparison does not include Opus 5.5, so it cannot supply the requested three-way coding chart.4
The fallback qualification matters too. Anthropic separately reports an Opus AutomationBench result without fallbacks; that is a different configuration and should not be mixed into this table.7
These are published evaluation results, not testing performed by PLEXData. API costs also cannot tell us how quickly my subscription allowance would have been consumed. The time spent supervising unsuccessful work is another cost this table does not capture.
§ 04What I would like to see from the new Sol
I would like to see GPT-6.1 Sol hold together the details of a connected system well enough to recognise where the failure begins. In last week’s job, another version of the code was of little value while the underlying problem remained.
I would like fewer confident explanations that lead back to the same fault. When the model has not established a cause, I would rather that uncertainty be clear before more time and capacity disappear into another round of changes.
The Opus handoff also raised my expectations about simplicity. It showed that, in this case, the desired functionality could be delivered with substantially less code. I would like the next Sol to recognise when complexity is getting in the way, while preserving the behaviour the system actually needs.
Above all, I would like enough usable capacity to finish difficult work. A model’s capability and the amount of it available to a customer belong in the same conversation. The Astra escalation was a reminder that access to a more capable model has limited practical value when the allowance is gone before the result arrives.
§ 05The release earns a retest; the work earns the trust
The published numbers give GPT-6.1 Sol a stronger case than its predecessor. They do not erase last week’s experience.
Opus 5.5 completed that job. GPT-6 Sol did not, and the move to Astra exhausted my available usage before completion. That is one personal experience, but it is the experience against which I will read the new claims.
I want OpenAI to have addressed it. A model that finishes more of the work, wastes less of the user’s time, and leaves a simpler working system would be welcome.
The release earns a retest. The work earns the trust.
Sources checked 30 September 2026, Australia/Brisbane. Public accounts and published evaluations were reviewed; no model runs were performed for this article.
4OpenAI: Introducing GPT-6.1 Sol — AutomationBench scores and costs extracted from the embedded chart data; vendor-reported. Opus result also corroborated by Zapier’s leaderboard.openai.com
7Anthropic: Claude Opus 5.5 — AutomationBench reporting without fallback models; explains why configurations must remain distinct.anthropic.com
011
Published
Author
Dorian Sotpyrc
Reading
7 minutes · Opinion
Status
Opinion · US data to June 2026
Opinion · Money & Banking
The Great Print: What the Data Shows About Money Printing in the West
US broad money grew 40% in two years. Consumer inflation peaked sixteen months after money growth did. I think the surge belongs in the inflation story alongside supply shocks, and the cost mattered most to households with little room to absorb it.
Minimal animated illustration. A small printing press feeds banknotes onto a growing stack labelled Money. A price tag on a pole to the right starts low and rises later, after the stack has grown.
Illustration · Money first, prices laterEditorial
Money surged in 2020–21; consumer inflation peaked in June 2022. The sequence is clear. How much money caused the rise needs more evidence.
US M2 rose from $15.35 trillion in December 2019 to $21.50 trillion in December 2021, a rise of 40%. Growth peaked at 26.8% a year in February 2021.1
My view
Money and supply shocks both mattered
Consumer inflation peaked sixteen months after money growth did. Energy, food and reopening also mattered. I think the honest reading keeps money and supply shocks in view together.
Open question
How much was money?
The timing alone cannot tell us how much inflation came from monetary demand, reopening or supply shocks. Establishing that split needs more than two lines on a chart.
In 2020 and 2021, the central banks of the West did something they had never done at this scale in peacetime. By the end of it, roughly 29 cents of every dollar of US broad money had been created in the previous two years. The phrase people reached for was “money printing”, and it has been argued over ever since. I think the phrase is a bad description and a fair warning.
This piece is the 2026 edition of an article first published on 2 December 2025. I have kept its structure, updated the data to June 2026, and asked a harder question of it: how much weight the money numbers deserve now that the dust has settled.
§ 01The story people tell about “money printing”
The popular version is short. Governments and central banks created huge amounts of money, prices rose, and everyone got poorer. Versions of it circulated with claims that between 20 and 40 per cent of all dollars in existence were created in a couple of years.
Check the arithmetic and the claim is closer to true than its critics allow. US M2 was $15.35 trillion at the end of 2019 and $21.50 trillion at the end of 2021.1 That is a 40% rise in two years, and it means about 29% of the end-2021 stock was new. The claim only goes wrong when it treats all of that money as cash from a press. Most of it was bank deposits.
§ 02What economists mean by money, and by “printing”
Base money, broad money and deposits. Base money is currency in circulation and the reserve balances commercial banks hold at the central bank. Broad money measures currency held by the public, deposits and other liquid balances: M2 in the United States, M3 in the euro area. Bank reserves are not part of those broad-money measures. When a bank makes a loan, it creates a deposit. Broad money can grow without a single new note.3
Three engines at once. In 2020 three things ran together. Central banks bought bonds through quantitative easing, which added reserves; purchases from non-bank sellers could also add bank deposits. Governments ran large deficits paid for by selling bonds, and the cheques landed in bank accounts. Banks kept lending.43 No single one of them is “the printing”, and that matters for what came next.
§ 03The Covid money surge in the US, Europe, the UK, Australia and Canada
In the United States the jump was unmistakable. The year-on-year growth rate of M2 was 6.7% in December 2019, reached 22.8% by June 2020 and 24.5% by December 2020, and peaked at 26.8% in February 2021.1
Large bond-purchase programmes were part of the response in all five economies. The figures below show the scale reported by each central bank, with the programme and date stated alongside each amount.
Selected bond-purchase figures from the pandemic response; currencies, programmes and reporting dates differ
Economy
Reported amount
Programme and scope
United States
US$4.4 trillion
Fed bond acquisitions since February 2020, reported in January 2022.5
Euro area
About €1.7 trillion
Net PEPP purchases from March 2020 through March 2022; excludes the separate APP.13
United Kingdom
£895 billion
Total QE bond purchases, including rounds before the pandemic; £875 billion government and £20 billion corporate bonds.7
Australia
A$280.7 billion
Bond Purchase Program, November 2020 to February 2022; excludes earlier market-function and yield-target purchases.6
Canada
More than C$180 billion
Government of Canada Bond Purchase Program, from its March 2020 launch to the December 2020 report.8
These are programme snapshots, not a ranking of comparable totals. The UK figure includes older QE; the other rows cover different periods and assets. None is a measure of broad-money growth. The distinction matters: bond buying, fiscal spending and bank lending work through different channels, so a larger purchase programme need not produce a larger rise in deposits.34
§ 04From the money surge to the inflation that followed
The monthly series puts the US M2 growth peak at 26.8% in February 2021. Consumer price inflation, calculated from the seasonally adjusted CPI index, peaked at 9.0% in June 2022, sixteen months later.12 The chart uses quarter-end readings, which put its highest sampled M2 growth point in December 2020; that is not the monthly peak. By the end of 2021 consumer prices were 8.6% above their end-2019 level, and by June 2026 they were 28.6% above it.2
Line chart of US money growth and consumer price inflation, quarter-end readings from 2019 to mid-2026. Money growth (M2) rises from 6.7% in December 2019 to 24.5% in December 2020, falls below zero in late 2022, reaches minus 3.9% in early 2023, and recovers to 5.3% by June 2026. Consumer price inflation peaks around 9.0% in June 2022, sixteen months after the monthly money-growth peak, and is 3.5% in June 2026. Official data, with year-on-year rates calculated by PLEXData.
Fig. 1 · US money growth and pricesOfficial data
Money growth peaked in 2020–21 and went negative in 2023. Prices peaked in mid-2022. By June 2026 M2 was growing 5.3% a year and CPI inflation was 3.5%.
Quarter-end readings · FRED, seasonally adjusted
The timing is consistent with a delayed monetary effect, but it does not establish one. Energy and food prices also surged, and reopening demand ran into constrained supply. OECD-area headline inflation peaked at 10.7% in October 2022.9 The OECD’s household study found that energy prices drove much of the purchasing-power loss in countries including Denmark, Italy and the United Kingdom.10
My position, with US data through June 2026, is that both camps were partly right, and the argument between them was less useful than it looked. The money surge made the economy easier to push prices up in. The supply shocks lit the match. How much each economy paid depended on how large both were.
§ 05How high inflation hit younger households
Inflation is not one rate. Renters and first-time buyers meet it through rent, mortgage payments, childcare, food and energy. Owners with fixed-rate mortgages and assets that rose in value met it very differently.
The housing evidence gives this argument firmer ground. The RBA reported in March 2023 that around half of Australian renter-household heads were aged 25 to 44. Renters also tended to have lower incomes, less wealth and smaller savings buffers than owner-occupiers.11 In Great Britain, an ONS survey covering 8 February to 1 May 2023 found that 43% of renters had difficulty affording their rent, compared with 28% of mortgage holders reporting difficulty with mortgage payments.12 These are specific household findings, not a claim that every young person paid the same cost.
An OECD study of purchasing-power losses between August 2021 and August 2022 found that inflation weighed more heavily on low-income than high-income households in every country it examined. It also found substantial exposure among rural households; it did not establish a universal ranking by age.10 That matches how I think about it. The same shock that raised the price of a flat did nothing to the wages of someone trying to rent one, and people who owned assets had a cushion that people who owned nothing did not.
§ 06After the Great Print: QT, higher rates and the long tail
The reversal came quickly. Central banks moved from emergency easing to quantitative tightening, and raised policy rates. US M2 fell from its pre-contraction high of $21.79 trillion in March 2022 and was down 3.9% on a year earlier by March 2023, the sharpest contraction of the period in the data I pulled.1 It did not stay negative. US M2 growth was 4.0% in December 2025, and 5.3% by June 2026.1
By June 2026, broad money was 6.1% above its March 2022 high, and consumer inflation was 3.5%.12 Money growth had slowed markedly from the pandemic surge. A slower inflation rate has not reset the price level.
The US consumer price index was still 28.6% above its December 2019 level in June 2026.2 Whether a household could absorb that increase depended on what happened to its income, savings and housing costs. Falling M2 did not take prices back to 2019.
§ 07What I take from it
First, the surge was as large as people said, and the plain-English version of the story is mostly fair. Second, money growth explained the timing better than it explained the size. Supply shocks and the way policy was run explain a lot of the rest. Third, the bill did not land evenly. It landed on the people with the least room to absorb it.
The fix I would argue for is unglamorous. Watch broad money as a signal that something is moving, not as a verdict. Keep the fiscal and monetary decisions in view together, because they came as a pair. And count the cost by who paid it, not by the average.
Agent Surfing: Build a Website an AI Agent Can Actually Navigate
More of your visitors are now agents sent by people, and they arrive without the context a person picks up in seconds. Making a site readable for them starts with one plain file and a few old habits done properly.
Minimal animated illustration. A small robot agent rides a surfboard over rolling waves. Ahead of it, a coral flag reading AGENT.md flies from a marker float, with a dotted line of sight from the agent to the flag. Buoys labelled evidence, reports and opinion float past on the waves. Along the bottom, five stages light up in turn: discover, orient, navigate, verify, use.
Agent Surfing
Give an agent readable water and a map, and it can pick its own line.
Most of the care that goes into a website assumes a person on the other end. Someone who glances at the navigation, reads a headline, scrolls a little and knows within seconds whether they are in the right place. The agent that person sends gets none of that for free. It comes in through whatever link it found, often deep inside the site, carrying a broad instruction and no feel for the place.
That visitor is now common enough to design for. I have started calling the design problem Agent Surfing: making a site easy for an agent to discover, get its bearings in, move around, check, and bring something useful back from.
§ 01Being found is the easy part
SEO still matters. It is how a machine finds your page at all. But it deals with the moment before arrival, and the questions an agent has afterwards are different. What is this site? Is this page the real version or a copy? Who wrote it, and when? Is it a measurement or somebody’s view? Where would I go to check it?
A person answers most of those without noticing. An agent has to work them out from markup, link text and whatever the page says about itself. Every guess costs it steps, and some guesses end with it quoting the wrong page back to the person who sent it.
So the line I keep coming back to is this: SEO helps a machine find your page; Agent Surfing helps an agent understand where it has landed and where it should explore next.
What I am not claiming matters as much. Agent products differ, and they change month to month. I have not measured how any one of them moves through a site. This is about what any careful agent needs, whichever product it is.
§ 02Reading the water
The surfing picture is deliberate. A good surfer does not fight the sea. They read it, pick a line and commit, and they can only do that because the water gives them something to read. A website is either readable water or chop.
When I break down what an agent has to do on a site it has never seen, I get five moves.
Agent Surfing: five moves
Move
What the agent needs to know
What the site can give it
Discover
Is there a map, and where is it?
A downloadable AGENT.md at a stable, linked address
Orient
What is this site, and what does it hold?
A plain description, the main sections, who runs it
Navigate
Where next?
Stable URLs, semantic HTML, link text that says where it goes
Verify
Can I trust this, and where did it come from?
Named authors, dates, canonical sources, evidence kept apart from opinion
Use
Can I take this back to my user?
Downloads it can open, and statements it can cite cleanly
Most sites manage some of these by accident. The work is doing all five on purpose, and the first is the cheapest to fix.
§ 03Leave a map at the door
Start with one file. A short, plain AGENT.md at a stable address, linked from the footer so it can be reached from any page. Think of it as the note you would leave for a capable stranger who has to find their way around without you: what the site is and who runs it, where the authoritative material lives, how the sections relate, and where to look for the questions people usually bring.
AGENT.mdListing 1 · Illustrative example
# Example Field Notes<!-- Illustrative example. Names and paths are placeholders. -->What this is: an independent publication on home-battery systems.Run by: Jane Example. Contact: /aboutLast updated:2026-09-30## Authoritative sources/data/ original measurements, CSV, with method notes/reports/ our reporting, dated, each linking its data/opinion/ our views, labelled as opinion## How the sections relateOpinion cites reports. Reports cite data. Start at data to verify.## To investigate furtherPricing question: /reports/tag/pricing, then /data/prices.csvWho wrote this: /about and the byline on each page
17 lines · plain MarkdownIt describes and points. It does not command.
Keep it short enough to read in one pass, and keep it true. A map of last year’s site is worse than no map, because the agent has no reason to doubt it. A community proposal called llms.txt is aimed at a similar problem.5 The name matters less than having one clear file and linking to it.
PLEX keeps its own at plexdata.online/AGENT.md. Every article here also has an Agent MD button that exports the piece as clean Markdown, with its title, author, date and sources, and each of those exports points back to the site map. An agent that starts from a single article can still find its way to the rest of the site.
§ 04Make the ground match the map
A map only helps if the ground matches it, and most of the ground is ordinary web practice that has quietly become more important.
Addresses come first. An agent that saves a link, or cites one to its user, is trusting that address to mean the same thing next week. I am learning this on this site. It is part-way through a migration, and one of the old article addresses still brings in steady traffic. If that address simply broke, every link, bookmark and index entry pointing at it would break with it. Move a page if you have to, but redirect it.
Then structure. Real headings, lists, tables and links carry meaning without anyone seeing the page.6 A layout built from anonymous boxes looks fine to a person and says nothing to a machine. Titles, descriptions and structured data do the same job one level up: they tell an agent what a page is before it has read it.3
Then provenance. Put a name and a date on every page that makes a claim, and say when it changed. Where the same material lives at several addresses, declare which one is canonical, so the agent is not left choosing between near-copies.4 If the evidence is a dataset, publish the dataset with a name and a description, not just a picture of the chart.
The habit I care about most is keeping evidence, reporting and opinion visibly apart. PLEX labels each piece as news, concept, opinion or review, and inside many articles it separates what is confirmed from what a vendor claims and what is still open. That labelling was written for human readers. It turns out to be exactly what an agent needs to report a finding honestly, as a measurement or as somebody’s view.
The older plumbing still counts. A robots.txt file and a sitemap are still how many crawlers find their way in at all.12
§ 05Don’t write to the agent
There are two tempting wrong turns.
The first is a second website for machines: a stripped-down copy that drifts away from the real one until you are maintaining two sources of truth and hoping they agree. Make the one site legible instead.
The second is filling pages with instructions aimed at AI systems. Hidden or pushy text telling an agent what to do looks exactly like the prompt-injection attacks agent builders defend against, and a well-built agent should ignore it. It also runs against the argument in Stop Trusting the Agent: an agent’s authority should come from its user and its permissions, not from whatever it happened to read along the way. An AGENT.md should describe and point. It should never give orders.
§ 06Try it on your own site
You do not need a lab for a first look. Give an agent you already use a broad job about your own site, something like “find out what this site says about pricing and come back with sources”. Then read what it did, not only what it said. Did it find the map, or wander? Did it describe the site the way you would? Did it keep what you measured apart from what you think? Would you be happy to see its links and dates quoted back to you?
Each fumble points at something specific to fix. I have not run this across agent products for this piece, so there are no results here, only what I would look for.
If I had one afternoon, I would write the AGENT.md, put a name and a date on every page missing them, and make sure nobody, human or agent, has to guess whether a page is evidence or opinion. None of that is new work. The visitor is new.
PLEX is a technical source for people and software agents that read, decide, build and act on systems. It turns security and reliability ideas into mechanisms, code and evidence that help both work more effectively without quietly gaining authority they should not have.
Curated by Dorian Sotpyrc. Written so a person or software agent can extract the mechanism, assumptions, implementation and test — not just a conclusion.
P
Permissions
What a person or agent may touch, and for how long. Granted for a task, never by default.
L
Limits
Where work has to stop: timeouts, budgets, sandboxes, egress and escalation.
E
Evidence
What the work leaves behind, so a person or supervising agent can verify what happened.
X
eXecution
The part that actually runs. Small, tested, reproducible and inspectable.
What we believe
Constrain first, automate second.
We design the boundaries before the behaviour. A system that can do anything will, eventually, do the wrong thing.
Mechanism over promise.
A policy document is not a control. If it isn't enforced by permissions, code or network, it doesn't exist.
Small enough to verify.
A person or agent should be able to inspect the mechanism, its assumptions and the test without trusting a black box.
Show the evidence.
Every claim comes with something you can run, log or test — and a way to tell when it stops working.
Say what it costs.
Every control slows something down. We write with honest trade-offs, because the reader has to live with them.
Work with Dorian
Bring PLEX a system that needs boundaries.
If you're collaborating, commissioning work, or adapting these patterns for a team or agent stack, write directly. A short description of the system, what it can reach and what worries you is enough to start.