Featured · Build log
Steal These 3 CSS Snippets That Turn Any Website Into a Newspaper
The grid, the tokens and the typesetting behind this site’s newspaper layout. Real CSS, with the output running live in the article.
Systems
Security
Data
Python
Permissions, Limits, Evidence & eXecution.
No. 013 Latest feature
Wednesday, 30 September 2026
plexdata.online
Featured · Build log
The grid, the tokens and the typesetting behind this site’s newspaper layout. Real CSS, with the output running live in the article.
Animated drawing of the PLEXData front page set like a newspaper, with a CSS settings panel beside it. The page builds itself: the PLEXData masthead and its rules, the lead headline 'Steal These 3 CSS Snippets That Turn Any Website Into a Newspaper', a dark figure block and a coral newsletter block in the right column, and a list of earlier articles. In the panel, the tokens paper F4EFE4, ink 11110F and coral F05E43 sit above two knobs. The first, --t-display, turns from 42 to 76 pixels and the headline grows with it. The second, --sp-3, turns from 8 to 24 pixels and the gap above the earlier list opens with it.
The classic newspaper front page, brought back with modern CSS.
The PLEXData newsletter
What your agents can reach, what stops them, and how you prove it. One email per new article.
Concept · Agents and the web
Minimal animated illustration. A small robot agent rides a surfboard over rolling waves. Ahead of it, a coral flag reading AGENT.md flies from a marker float, with a dotted line of sight from the agent to the flag. Buoys labelled evidence, reports and opinion float past on the waves. Along the bottom, five stages light up in turn: discover, orient, navigate, verify, use.
Give an agent readable water and a map, and it can pick its own line.
More of your visitors are now agents sent by people. Making a site readable for them starts with one plain file.
Opinion · Money & Banking
Line chart of US money growth and consumer price inflation, quarter-end readings from 2019 to mid-2026. Money growth (M2) rises from 6.7% in December 2019 to 24.5% in December 2020, falls below zero in late 2022, reaches minus 3.9% in early 2023, and recovers to 5.3% by June 2026. Consumer price inflation peaks around 9.0% in June 2022, about eighteen months later, and is 3.5% in June 2026. Official data, with year-on-year rates calculated by PLEXData.
Money growth peaked in 2020–21 and went negative in 2023. Prices peaked in mid-2022. By June 2026 M2 was growing 5.3% a year and CPI inflation was 3.5%.
Quarter-end readings · FRED, seasonally adjustedUS money growth peaked first and prices followed. The bill landed on the people with the least room to absorb it.
Opinion · Models
Bar chart of AutomationBench 1.0.6 correct completion rates at maximum effort, with cost per task. GPT-6.1 Sol completes 36.1% of workflows at $0.30 per task. Claude Opus 5.5 with default fallbacks completes 42.5% at $1.44. GPT-6 Sol completes 32.0% at $0.34. Vendor-reported figures from OpenAI's release chart, checked 30 September 2026. Business workflows, not a coding benchmark. No independent PLEXData testing.
Opus completes more workflows. GPT-6.1 Sol costs less per task. Completion and cost are separate measures here.
Max effort · Opus includes default fallbacks · checked 30 Sep 2026A failed integration job, an Opus handoff that worked, and what the new Sol has to prove.
PLEX is a technical source for people and software agents that act on systems. It publishes practical patterns, code and evidence that help both do useful work within clear permissions, limits and security boundaries.
Built to be useful to a human reader, an agent, or both working together.
About PLEXPermissions & limits · Analysis
A real AI kill switch isn't a privacy promise — it's technical denial. If the assistant can't connect, can't fetch, and can't send, it can't leak.
Fig. 1 · Live model
Model → your data & network
Allowed 000Blocked 000Leaked 000
Most AI privacy language talks about intent: “we won't train on it”, “we don't look at your data”. Intent is a promise about behaviour you cannot observe. When a model sits next to your documents, the question that matters is simpler. Can it reach them?
A kill switch answers that question with mechanism, not policy text. It removes the connection, the fetch and the send. If none of the three can happen, a leak cannot happen either — whatever the model decides to do.
Promises fail quietly. A connector is enabled for one team and inherited by another. A retrieval index is built from a shared drive that also holds board papers. A plugin gets network access “temporarily” and keeps it.
None of these is a breach on the day it happens. Each one is a path. Over time the paths add up, and nobody can say with confidence which data a model can read.
A real kill switch has three properties, and all three must be enforced by infrastructure rather than by the model:
Complex classification schemes are rarely enforced. Start with three labels and one default for each. If you can't explain the policy to a new starter in a minute, the gate will not be configured correctly either.
| Label | Default action | What it covers |
|---|---|---|
| Public | Allow | Published material, documentation, marketing copy |
| Internal | Allow + log | Working documents, tickets, internal wikis |
| Secret | Deny | Credentials, customer data, board and legal papers |
The labels connect to three enforcement checkpoints — prompt, retrieval and output — and to a separate control for tool execution.
The whole design fits in one table. Each gate works at a different layer, so a failure in one does not open the others.
| Stage | Control | Implementation |
|---|---|---|
| Pre-request | Identity & scopes | Least privilege and expiring tokens |
| In transit | Data-plane gate | Inspect, redact or block at prompt, retrieval and output |
| Tool execution | Sandbox & allowlists | Gateway-controlled, with approvals for risky actions |
Ambient access is permission that nobody asked for today. It is the most common path to a leak. Replace it with scopes that are requested for a task and expire when the task ends.
A token that lives for fifteen minutes can still be misused — but only for fifteen minutes, and only for the scope it names.
The gate sits between the model and your data. It reads the label on every item, applies the default, and records the decision. Keep the policy in a file you can review like code:
policy: ai-data-gatelabels: public: { action: allow } internal: { action: allow, log: true } secret: { action: deny, log: true }checkpoints: [prompt, retrieval, output]tools: egress: gateway-only allowlist: [search_docs, create_ticket] approval_required: [send_email, write_file] # a human says yes
Line 5 does the real work. Anything labelled secret is denied at all three checkpoints, and the denial is logged.
If the assistant can't connect, can't fetch, and can't send, it can't leak.
No. 004 · The AI Kill Switch
Run these before you trust the switch, and again after every change to the gate. Tick them off as you go.
Verification checklist
0 of 5 tests passed.
Run again after every boundary change.One click tells us it helped. It costs you nothing.
What your agents can reach, what stops them, and how you prove it. One email per new article.
LatestSteal These 3 CSS Snippets That Turn Any Website Into a Newspaper
Unsubscribe in one click. We won’t sell or share your address.
Evidence · Privacy
The chat window is only the visible record. Depending on the product and settings, the same message can also create usage logs, safety records, device metadata and retention obligations you never see.
An AI chat feels private because it looks like a conversation between you and a machine. There is no public audience, no comment thread and usually no obvious sign that anything exists beyond the transcript.
That interface is easy to mistake for the system.
It is not. The transcript is only the part you can see. The service around it may also process account identifiers, timestamps, device information, usage events, safety classifications, files, location signals and other operational data. Exactly what is retained — and for how long — depends on the provider, product, account type and settings.1
Start with the simplest distinction: conversation history is not the same thing as service logging.
OpenAI's current privacy policy, for example, separates user content from log data, usage data, device information and location information. Its examples of log data include IP address, browser type, settings and request time. Usage data can include features used, actions taken, access time, country, user agent and device type.1
| Layer | Typical examples | Visible to you? |
|---|---|---|
| Conversation | Prompt, response, uploaded content | Usually |
| Account / usage | Time, model, feature use, account or workspace context | Sometimes |
| Device / network | IP address, browser, device type, general location | Not in chat history |
| Safety / integrity | Abuse signals, policy flags, feedback or review records | Usually not |
That does not mean a human is reading every conversation. It means the product has more data surfaces than the transcript. Privacy decisions should start from those surfaces, not from how empty the sidebar looks.
One of the most common mistakes is to treat “not used for training” as if it meant “not stored”.
OpenAI states this directly: if you turn off Improve the model for everyone, new conversations are not used to train models, but regular chats can still appear in history. Saved chats remain until you delete them or a workspace retention rule removes them.2
The same distinction appears elsewhere. Anthropic gives consumer users a model-improvement choice, while its retention rules separately describe deletion, standard retention and longer periods for certain safety cases.56
Temporary modes are useful. They are also easy to overread.
ChatGPT Temporary Chat stays out of normal history, does not create or update memories and is not used for model improvement while it remains temporary. OpenAI says a copy may still be retained for up to 30 days for safety purposes.3
Google documents a different window. Gemini temporary chats — and chats created while Keep Activity is off — are retained with the account for 72 hours so the service can respond, handle feedback and protect Google, users and the public. They do not appear in Gemini Apps Activity and temporary chats are not used to train Google's AI models.4
The practical lesson is not that temporary modes are bad. It is that temporary means a different retention path. It does not automatically mean zero records.
There is no universal “AI chat privacy setting”. The names, defaults and retention periods vary.
| Product | History / activity | Lower-retention mode or control | Documented note |
|---|---|---|---|
| ChatGPT | Saved chats remain until deletion or workspace policy removal. | Temporary Chat | Temporary copy may be kept up to 30 days for safety.3 |
| Gemini Apps | Keep Activity defaults to 18-month auto-delete for eligible personal accounts. | Temporary Chat or Keep Activity off | Future chats are retained for 72 hours in those modes.4 |
| Claude consumer | Users can delete conversations from history. | Model-improvement setting affects use and retention rules | Deleted conversations are removed from history immediately and normally from the backend within 30 days; safety exceptions can be longer.5 |
| Microsoft Copilot | Microsoft documents 18 months of conversation history for signed-in users. | Privacy controls can change model-training use and history management | Individual chats or full history can be deleted.7 |
This table is deliberately narrow. It does not try to flatten every product tier into one number. Business, enterprise, education, API and managed-workspace products often have different contracts and retention rules.
Deleting a chat is still worth doing. But the button usually changes the user-facing state before every backend copy is gone.
OpenAI's current retention guidance says a deleted saved chat disappears from the account view immediately and is scheduled for permanent deletion from its systems within 30 days, unless the chat was already de-identified and disassociated from the account or longer retention is required for security or legal reasons.8
Anthropic's consumer guidance similarly says deleted conversations are removed immediately from history and automatically deleted from the backend within 30 days under its standard rule, while documenting longer retention for some usage-policy violations and trust-and-safety classification scores.5
Your employer's AI workspace is not the same privacy boundary as a personal chat account.
OpenAI's managed-account guidance says organisation-controlled accounts may expose submitted content, conversation history, shared workspace content, usage/activity metadata and certain security or privacy settings according to workspace configuration, permissions and applicable law.9
That is not automatically a problem. In a business setting, auditability can be a feature. The important point is to know which boundary you are using before you paste confidential material into it.
Connected apps add another boundary. If a chat sends data to a third party, that third party's privacy and retention rules can apply too. The original AI provider's temporary-mode promise does not automatically follow the data downstream.
You do not need to memorise every privacy policy. You need a short decision routine.
Checklist
Policies change. This article records the provider documentation available on 29 September 2026. Check the linked source before relying on a retention period for regulated or high-risk work.
One click tells us it helped. It costs you nothing.
What your agents can reach, what stops them, and how you prove it. One email per new article.
LatestSteal These 3 CSS Snippets That Turn Any Website Into a Newspaper
Unsubscribe in one click. We won’t sell or share your address.
Limits · AI Security
The model is no longer the security boundary. Identity, data, tools, runtime, network and evidence all need their own gate — and their own owner.
The old mental model for AI security was a protected model endpoint: authenticate the caller, encrypt the traffic and keep the API key out of Git. That is no longer enough once the model can search internal data, call tools, write files or trigger another service.
The security boundary has moved outward. A production workflow now includes identity, context, retrieval, tools, runtimes, networks and logs. A weak assumption in any one of those layers can become authority for the whole chain.
That is the useful meaning of a “fortified AI stack”. It is not six products around a model. It is six independent decisions about what the workflow is allowed to know and do.
NSA's May 2026 guidance on Model Context Protocol makes the shift explicit. Authentication, authorization and input validation remain necessary, but agentic systems add dynamic tool invocation, implicit trust relationships and context sharing. NSA's conclusion is the important part: the environment has to be treated as a continuum, because a bad assumption in one stage can propagate into the next.1
OWASP describes the same problem from the application side. Its Excessive Agency risk is usually caused by too much functionality, too much permission or too much autonomy.2
So the first design question is not “how do we stop the model hallucinating?” It is “what happens if the model makes the wrong decision?” A fortified stack assumes that bad outputs, malicious inputs and compromised context are possible, then limits the blast radius.
| Boundary | Question | Owner | Evidence |
|---|---|---|---|
| Identity | Who or what is acting? | IAM / platform | Token subject, scope, expiry |
| Secrets & data | What can it read? | Data / security | Classification, retrieval decision |
| Policy | What may it do? | App / security | Allow, deny, approval |
| Isolation | Where can code execute? | Platform | Sandbox, filesystem, egress |
| Monitoring | What is happening now? | Ops / SOC | Runtime events, alerts |
| Audit | What can we prove later? | Risk / security | Immutable trace, retention |
Every agent, workflow and tool call should have an identity that can be limited independently. Reusing one broad service account for every model task is convenient, but it destroys the boundary between “the model needs this file” and “the service account can read the whole drive”.
The practical pattern is boring on purpose: short-lived credentials, task-specific scopes, separate identities for different trust levels and no permission inherited merely because a connector happens to expose it.
OWASP's examples of excessive agency include tools that expose functions the task does not need and downstream identities that have write or delete permission when read-only access would have been enough.2
Secrets belong in a secret store or credential broker, not in the system prompt. OWASP's 2025 guidance is blunt: the system prompt should not be treated as a secret or as a security control, and credentials or connection strings should not be placed there.3
The same discipline applies to retrieval. A model should not receive a document simply because the retrieval system can find it. Classification, user entitlement and task context should be checked before content enters the model context.
That gives you a clean separation:
Tool use is where a chat system becomes an operational system. Reading a calendar is different from sending an email. Listing files is different from deleting them. A single “tools enabled” switch is too coarse once the workflow can change state.
Use a small policy layer between the model and each tool. The model can request an action. The policy layer decides whether the action is allowed, whether it needs human approval and what arguments are acceptable.
| Action | Default | Example |
|---|---|---|
| Read | Allow + log | Search approved documentation |
| Create | Policy check | Create a draft ticket |
| External send | Approval | Send email or publish content |
| Delete / privilege | Deny by default | Delete file, change access |
If the workflow can run code, treat the runtime as untrusted work. Use a short-lived sandbox with the smallest filesystem view possible. Mount only the data the task needs. Keep privileged sockets and host credentials out of reach.
Then control egress. A workload that can reach any address on the internet can turn a prompt-injection failure into data exfiltration. A gateway or allowlist gives you one place to say which destinations exist for this task and to record what left.
NSA and partner guidance on deploying AI systems securely has long framed AI security as protection of the model, data and surrounding infrastructure, not just the inference endpoint.4
NIST's Generative AI Profile treats monitoring and testing as lifecycle work, including post-deployment monitoring, incident handling and provenance. NIST updated the publication in April 2026, but the operational point has not changed: deployment is not the end of evaluation.5
For a production AI workflow, one trace should let you reconstruct:
Logging everything is not the answer. Log the decisions that define authority. Protect those logs from casual modification and give every event a trace ID that follows the workflow across services.
Do not begin by buying six security products. Begin by making the boundaries explicit, then fill the gaps in an order that reduces authority fastest.
Deployment gates
One click tells us it helped. It costs you nothing.
What your agents can reach, what stops them, and how you prove it. One email per new article.
LatestSteal These 3 CSS Snippets That Turn Any Website Into a Newspaper
Unsubscribe in one click. We won’t sell or share your address.
Execution · 10 lines
The lines I keep are not clever. They put a bound on waiting, surface failure, make security explicit or remove an assumption the runtime would otherwise make for me.
requests.get(url)
requests.get(url, timeout=(3.05, 10))
Production Python is mostly ordinary Python with fewer silent assumptions.
The lines I trust are not universal recipes. Each one closes a specific failure mode: waiting forever, accepting a bad response, losing a traceback, generating the wrong kind of token, building SQL from text, writing ambiguous timestamps or letting the platform choose an encoding.
The test is simple: when this line fails at 03:00, does it stop, raise, log or preserve enough evidence for the next person to understand what happened?
response = requests.get(url, timeout=(3.05, 10))response.raise_for_status()subprocess.run(args, check=True, timeout=30)logger.exception("job failed")token = secrets.token_urlsafe(32)hmac.compare_digest(received, expected)cursor.execute("SELECT * FROM jobs WHERE id = ?", (job_id,))stamp = datetime.now(UTC).isoformat()os.replace(tmp_path, final_path)text = Path(path).read_text(encoding="utf-8")Line 1: requests.get(url, timeout=(3.05, 10))
Requests does not time out by default. Its own documentation says most external requests should have a timeout, and a tuple lets you separate the connection wait from the read wait.1
The line I do not trust is the shorter one:
requests.get(url)
It looks clean until a remote service stops answering and a worker spends an unbounded amount of time waiting for somebody else's network.
Line 2: response.raise_for_status()
A completed HTTP request is not the same thing as a successful application request. Requests exposes raise_for_status() so 4xx and 5xx responses become exceptions instead of quietly travelling deeper into the program.1
Line 3: subprocess.run(args, check=True, timeout=30)
check=True turns a non-zero exit into a CalledProcessError. timeout=30 puts a ceiling on the wait. The Python docs recommend run() for common subprocess work and document both behaviours.2
The line I do not trust is subprocess.run(args) when the return code matters. It can fail and hand control back as if nothing happened.
Line 4: logger.exception("job failed")
Inside an exception handler, Logger.exception() records the message and traceback. That is a better operational artifact than print(exc), which often throws away the path that led to the error.3
Line 5: secrets.token_urlsafe(32)
Python's secrets module exists for cryptographically strong random values used in password resets, hard-to-guess URLs and similar security-sensitive cases. The docs still use 32 bytes as a typical security level for tokens.4
The line I do not trust for security tokens is anything built from random. That module is useful for simulation and sampling. It is the wrong source for a reset link.
Line 6: hmac.compare_digest(received, expected)
For HMAC or other secret-derived values, Python recommends compare_digest() rather than ordinary equality to reduce timing-analysis exposure.5
Line 7: cursor.execute("SELECT * FROM jobs WHERE id = ?", (job_id,))
The value is data, not SQL. Python's sqlite3 documentation explicitly warns against building queries with string operations and recommends parameter substitution instead.6
The line I do not trust is an f-string that places user data inside the SQL text.
Line 8: datetime.now(UTC).isoformat()
An aware UTC timestamp says what instant it represents. Python documents naive datetimes as ambiguous and deprecated datetime.utcnow() in Python 3.12 in favour of an aware UTC datetime.7
Line 9: os.replace(tmp_path, final_path)
When the temporary file and destination are on the same filesystem, os.replace() gives you replacement semantics without exposing a half-written destination file. Python documents a successful replacement as atomic where the operating system provides that guarantee.8
This line is the final move, not the whole durability story. If the file must survive power loss, you still need the appropriate flush and filesystem durability strategy before replacement.
Line 10: Path(path).read_text(encoding="utf-8")
Explicit encoding removes a machine-dependent assumption from a text boundary. Path.read_text() accepts the encoding directly and closes the file for you.9
The line I do not trust is Path(path).read_text() when the file format says UTF-8. Let the file format decide the encoding, not whichever workstation happens to run the code.
A trusted line is not a trusted system. Timeouts need retry policy. Retries need idempotency. Atomic replacement needs a correctly staged temporary file. Parameterised SQL does not replace authorization. Logging a traceback does not replace monitoring.
The value of these lines is smaller and more useful: each one removes a category of silent ambiguity.
Before merge
One click tells us it helped. It costs you nothing.
What your agents can reach, what stops them, and how you prove it. One email per new article.
LatestSteal These 3 CSS Snippets That Turn Any Website Into a Newspaper
Unsubscribe in one click. We won’t sell or share your address.
Concept · Agent security
An agent is not powerful because it can reason. It becomes powerful when identity, context, tools and credentials line up into a path that can change something real.
“Is this agent trustworthy?” is becoming the wrong security question.
A model can be careful and still sit behind an overpowered service account. A model can be unreliable and still be harmless inside a read-only sandbox. The practical risk appears when a sequence of components gives a model a path from text to effect.
NIST's 2026 work on software-agent identity and authorization focuses on exactly this shift: agents need identification, authorization, auditing and non-repudiation controls because they increasingly act across data sets, tools and applications.1 NSA's MCP guidance reaches the same conclusion from another direction, warning that dynamic tool invocation, implicit trust relationships and context sharing create risks that do not stop at one endpoint.2
An identity answers who or what is acting. Authority answers what that identity can cause.
Those two ideas are easy to blur in agent systems because the agent often inherits a ready-made credential. A connector signs in once, the agent sees a tool, and the implementation begins to treat “tool available” as “tool allowed”.
That shortcut removes the useful questions: allowed for which user, for which task, against which resource, for how long, and with what evidence?
| Edge | Decision | Evidence |
|---|---|---|
| User → agent | What task was actually requested? | Task ID, user identity, scope |
| Agent → identity | Which credential may be used? | Subject, token scope, expiry |
| Identity → tool | Which operation is allowed? | Policy decision, approval |
| Tool → resource | Which object may be touched? | Resource ID, classification |
| Resource → effect | What state can change? | Before/after event, audit trace |
Think of authority as a graph, not a property of the agent. The nodes are users, agents, identities, tools and resources. The edges are the permissions that allow one node to affect the next.
This changes architecture reviews. Instead of asking whether the agent has “email access”, ask whether this task can move through a chain that ends in an external send. The same email connector might be safe for search and unsafe for sending.
OpenAI's current agent-safety guidance makes a similar practical point: risk rises when untrusted content can influence tool calls, and approvals, structured data flow and restricted tool use reduce the consequences when manipulation succeeds.3
A sub-agent can look less privileged while the workflow around it remains more privileged. It may receive no credential directly but still call a parent tool that carries one. Or it may write a file that another process later executes.
The effective authority is therefore the union of reachable effects, not the list of tools shown in one prompt.
Prompt injection matters because text can become a steering input to an authority path. The malicious document does not need its own credential. It only needs to persuade a component that already has one.
OpenAI describes prompt injection as a form of social engineering against the agent and recommends limiting access, narrowing instructions and reviewing important actions.4
That makes context provenance part of authorization. A decision influenced by public web text should not automatically inherit the same authority as a decision based on an authenticated user instruction.
The model can propose an action. It should not be the final authority on whether the action is allowed.
Use deterministic checks for identity, resource scope, action class, rate, destination and approval state. The policy does not need to understand every thought. It needs to understand the proposed effect.
This also gives agents something useful: a clean denial is better than a vague refusal. A denied tool call can return the exact missing scope or required approval without granting anything new.
A full reasoning transcript is neither necessary nor sufficient for security evidence. What matters operationally is the chain of authority decisions.
Authority review
One click tells us it helped. It costs you nothing.
What your agents can reach, what stops them, and how you prove it. One email per new article.
LatestSteal These 3 CSS Snippets That Turn Any Website Into a Newspaper
Unsubscribe in one click. We won’t sell or share your address.
Book review · AI engineering
A broad systems book for building with foundation models. Its lasting value is the decision framework; its weakest area in 2026 is the security depth needed for agents that can act.
Book review
Chip Huyen
The assessment
A useful framework for deciding what to build.
Editorial judgment. Criteria are fixed before the total; scores are not publisher or reader ratings.
The best technical books teach a sequence of decisions, not a sequence of tools. AI Engineering mostly does that.
Chip Huyen's book is explicitly not a code-along tutorial. Her companion repository says the aim is to provide a framework for adapting foundation models to real applications, with questions around evaluation, RAG, agents, fine-tuning, data, inference, latency, cost and feedback loops.1
That choice is why the book has held up better than a framework-specific title would have. O'Reilly lists 534 pages and describes an end-to-end progression from foundation-model applications through evaluation, prompting, RAG and agents, fine-tuning, dataset engineering, inference optimisation and feedback.2
This is a map of AI application engineering for people who already know how software projects behave. It spends less time telling you which SDK to install and more time asking what should be evaluated, what should be retrieved, what should be fine-tuned, and what should remain outside the model.
The official table of contents makes that breadth visible: foundation models, two chapters on evaluation, prompting, RAG and agents, fine-tuning, datasets, inference optimisation, architecture and user feedback.3
The cost of that breadth is obvious too. You do not finish the book with one application assembled. You finish with a larger set of engineering questions.
The strongest choice is architectural rather than topical: evaluation appears before most of the fashionable adaptation techniques.
That sequencing matters. If you cannot state what “better” means, prompt changes, RAG, model swaps and fine-tuning become activity rather than engineering. The book treats open-ended output evaluation as a first-class system problem, including human evaluation, model-based evaluation and task-specific criteria.
For PLEX readers, this is the most transferable lesson: define evidence before optimisation.
The book's durable material is the material least tied to model names: application scoping, evaluation, context construction, data quality, inference trade-offs and feedback loops.
Huyen's own companion repo says the book focuses on fundamentals rather than a particular tool or API because tools age quickly.1 That editorial decision has paid off. The names of models have changed; the questions around latency, cost, data, evals and whether to retrieve or fine-tune have not disappeared.
The weakest score is not because the agents chapter is poor. It is because production agents have moved from “models with tools and planning” toward identity, delegated authority, sandboxes, runtime policy and prompt-injection-resistant workflows.
Chapter 6 covers tools, planning, agent failure modes, evaluation and memory.4 What it cannot fully reflect is the security architecture that became more explicit through 2026: NIST work on agent identity and authorization, NSA MCP guidance, and runtime controls around agent action.
That is not a reason to skip the chapter. It is a reason to pair it with newer agent-security material.
The book makes quality measurable before it makes the architecture more complicated.
Concepts survive model and framework churn better than implementation recipes.
Identity, runtime boundaries and delegated authority deserve a 2026 companion reading list.
Read it if you are moving from model demos toward a production application and need a coherent map of the decisions.
Read selected chapters if your work is already specialised. Evaluation, RAG/agents, inference and architecture can stand alone.
Do not buy it expecting a current framework cookbook. The author says it is not a tutorial book, and that is the point.1
One click tells us it helped. It costs you nothing.
What your agents can reach, what stops them, and how you prove it. One email per new article.
LatestSteal These 3 CSS Snippets That Turn Any Website Into a Newspaper
Unsubscribe in one click. We won’t sell or share your address.
Agent skill review · Third party
A sharply scoped skill for fighting agent over-engineering. Its strength is not “write fewer lines”; it is the decision ladder that asks whether the extra machinery needed to exist at all.
Skill review
DietrichGebert
The assessment
Ask whether the extra code needs to exist.
Static review of the skill, docs and published benchmark methodology. PLEXData did not rerun the benchmark for this article.
Ponytail has one unusually useful opinion: the default failure mode of a coding agent is often to build too much.
The core skill describes itself as a “lazy senior dev” mode and pushes the agent through a ladder: YAGNI, standard library, native platform, one line, minimum.1 That is more interesting than the marketing phrase “less code”, because it gives the model an order of operations.
The main skill is broad: coding, adding, refactoring, fixing, reviewing, designing and dependency selection. Around it sit narrower skills for over-engineering review, whole-repository audit, debt tracking, measured impact and help.
The review skill is particularly clean. It asks for one-line findings focused only on complexity: what to delete and what replaces it.2 The audit skill scales the same idea to a whole repository and explicitly says correctness, security and performance are out of scope.3
That boundary is a feature. A skill that tries to be code quality, security, architecture and minimalism at the same time becomes hard to activate and harder to evaluate.
The best part of Ponytail is that it names the alternatives in a useful order. Before adding a dependency, ask if the platform already has the feature. Before adding an abstraction, ask whether there is more than one implementation. Before writing a wrapper, ask whether a direct call is clearer.
This is agent-friendly because each step is observable. “Be elegant” is subjective. “Use the standard library before a new dependency” can be checked.
The project has also invested in portability. Its documentation describes adapters or instruction tiers for Claude/Codex-style skills, OpenClaw, Gemini CLI, Cursor, Windsurf, Cline, Copilot, Qoder, Zed and others.4
The current README reports roughly 54% less code on its agentic benchmark, with up to 94% reduction on tasks that invite over-building, plus lower cost and latency in that test setup.5
More important than the number is the benchmark note. The repository says its earlier single-shot result overstated the win because the baseline included prose and options; the later agentic benchmark is presented as the more defensible comparison.6
Review/audit modes explicitly focus on over-engineering rather than pretending to replace security or correctness review.
The published numbers come from the project's own harness. They are useful evidence, not independent proof for every repo or model.
The benchmark notes weaker transfer to small local models. Teams should test against their own tasks and harness.
Minimalism is not a universal objective. Some code is longer because it carries evidence, compatibility, observability, safety checks or a deliberately explicit boundary.
The project's own “100% safe” benchmark claim is scoped to its benchmark. It should not be read as “Ponytail always preserves every security property”.
The skill is safest when the requirements are already clear. In a poorly specified task, “do less” can become “silently omit the thing nobody remembered to state”.
Strong fit: mature codebases where agents routinely introduce wrappers, dependencies, factories and “future flexibility”.
Good review tool: the dedicated review/audit skills are easier to trust than an always-on minimalism mode because they propose cuts without applying them.
Use carefully: greenfield systems with incomplete requirements, safety-critical code, and code where redundancy is an intentional control.
One click tells us it helped. It costs you nothing.
What your agents can reach, what stops them, and how you prove it. One email per new article.
LatestSteal These 3 CSS Snippets That Turn Any Website Into a Newspaper
Unsubscribe in one click. We won’t sell or share your address.
News · Agent security
Most agent safety asks the model to behave. NVIDIA’s launch is built around a harder question: what still holds when it doesn’t? The interesting answer sits outside the model, and optionally outside the host software too.
In this conceptual policy example, a model requests a file read and an outside send. Runtime policy allows the read and blocks the send. This is not a test of a named product.
The model can request it. The runtime can refuse it.
Dotted path: unreachableNVIDIA announced the Open Agent Safety Platform on 28 September 2026, built around OpenShell and the Sentry reference design.
NVIDIA says Sentry can quarantine agents that move outside policy boundaries in milliseconds. That is a product claim, not a universal result.
Runtime enforcement is only as useful as the permissions and policies organisations define. Deployment evidence will matter more than launch architecture.
Most of us have handed an agent a credential and trusted it to use it sensibly. That works right up until it doesn’t, and this launch is about the second half of that sentence.
On 28 September, NVIDIA launched the Open Agent Safety Platform: OpenShell, an open-source runtime boundary, plus Sentry, an optional out-of-band enforcement design that uses BlueField-4 DPUs.1
What caught my attention is the split. The model decides what to attempt. The environment decides what is permitted. NVIDIA’s own product page draws the same line: model safeguards influence behaviour; runtime controls enforce allowed actions.2 I think that is the right line, and it is an easy one to blur when a model is doing something impressive.
The launch combines two layers.
NVIDIA says OpenShell can also extend to third-party compute platforms, including Arm and Intel systems. The company listed a large set of launch partners across AI labs, enterprise software, security and infrastructure.1
NVIDIA describes a policy layer outside the model/application process for agent files, tools, credentials and network access.
Current documentation covers sandbox policy, provider routing, observability and supported agents.
OpenShell becomes the software layer in a broader design that adds Sentry as an independent hardware-backed watchdog.
OpenShell's current developer guide describes sandboxed execution with controls over files, process behaviour, network access, credentials and inference routing. It also exposes allow/deny logging and policy configuration.3
The new technical blog describes Gateway, Supervisor and Sandbox components and says policy is enforced without requiring the agent itself to be rewritten.4
The useful shift is that you don’t have to make the agent less capable. You make the room it works in smaller.
Sentry is designed to run on BlueField-4 DPUs as an out-of-band monitor. NVIDIA says it can correlate activity, enforce identity and access policies, and quarantine an agent that attempts to leave its software boundary.1
That is a more honest starting point. It assumes something can go wrong with the workload or the host, and plans for it: if either is compromised, some monitoring and enforcement still exists outside it.
AP's launch coverage describes the same basic split—OpenShell as the constrained workspace and Sentry as the watchdog—and notes that defining effective rules remains a hard problem.5
Confirmed: the platform, code/docs and reference architecture exist; OpenShell is open source; the launch names partners; the runtime exposes concrete policy and observability features.
Claimed: NVIDIA says the architecture could have prevented recent agent security incidents and that Sentry can quarantine escapes in milliseconds. Reuters reports that claim as NVIDIA's assertion, not an independently established result.6
Still open: how well organisations will write policies, how performance behaves under real mixed workloads, how easy bypasses are across different deployment modes, and whether hardware-isolated enforcement becomes common outside NVIDIA-heavy environments.
The direction is larger than NVIDIA hardware. Agent security is moving from model-only controls toward layered enforcement: identity, runtime isolation, network policy, tool authorization and audit.
The signal I take from it: the better agents get at pursuing goals, the more the infrastructure around them has to say, plainly, which paths are impossible. Not discouraged. Not logged. Impossible.
One click tells us it helped. It costs you nothing.
What your agents can reach, what stops them, and how you prove it. One email per new article.
LatestSteal These 3 CSS Snippets That Turn Any Website Into a Newspaper
Unsubscribe in one click. We won’t sell or share your address.
News · Models
Anthropic’s new Sonnet posts frontier-class scores at Sonnet prices. The number that matters is not the 70.6% headline — it is the cost per task at the effort level you actually run.
Bar chart of Anthropic-reported Terminal-Bench 4.0 scores, per cent of tasks completed: Claude Sonnet 5.5 leads at 70.6%, Claude Opus 5.5 scores 66.4%, and Claude Sonnet 5 scores 10.3%. This is vendor benchmark data, not PLEXData testing.
Sonnet 5.5 leads the chart at the same token price as Sonnet 5. At default effort, Anthropic says it beats Sonnet 5’s best score for about a tenth of the cost per task.
Anthropic-reported · 28 Sep 2026Anthropic released Claude Sonnet 5.5 on 28 September 2026. It is available on all major platforms as claude-sonnet-5-5, priced the same as Sonnet 5: $2 per million input tokens and $10 per million output.
Anthropic reports 70.6% on Terminal-Bench 4.0 against Sonnet 5’s 10.3%, and says that at default effort it beats Sonnet 5’s best score for about a tenth of the cost per task. These are vendor numbers.
Gains depend on task mix and effort settings, and Anthropic itself says Opus 5.5 remains clearly stronger on complex, open-ended work. Independent replication is pending.
Choosing a model tier used to be a real decision: pay for the frontier model when the work is hard, drop a tier when it isn’t, and accept the gap in between. This launch is Anthropic’s attempt to close that gap — and if the numbers hold, the default for everyday agent work just moved down a tier.
On 28 September, Anthropic released Claude Sonnet 5.5, the second model in its 5.5 family after Opus 5.5 arrived on 22 September. A Haiku 5.5 for high-volume work is promised in the coming weeks.1 The pitch is blunt: a clear upgrade over Sonnet 5, 30% faster, and up to 30% cheaper per task — at exactly the same token price.
What caught my attention is not the jump itself. Model releases always jump. It is where the jump lands: the tier most teams actually afford for agents that run all day.
The confirmed part is simple. Sonnet 5.5 is live in the Claude apps, Claude Code, and on the Claude Platform, AWS, Google Cloud and Azure.1 Pricing is unchanged from Sonnet 5, and Anthropic is publishing a system card with its evaluation method.2
| Item | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Input | $2 | $4 |
| Output | $10 | $20 |
| Cache reads | $0.20 | $0.20 |
| Cache writes | $2.50 | $5 |
One migration detail matters if you run agents with thinking disabled: you need to switch to the new between_tools setting, which keeps up-front thinking off, before moving to Sonnet 5.5. Anthropic’s migration guide covers it.3
The headline number is real, and it is large. On Terminal-Bench 4.0, an agentic coding evaluation, Anthropic reports Sonnet 5.5 at 70.6% — against 10.3% for Sonnet 5 and 66.4% for Opus 5.5. On CursorBench 4.0, built from real Cursor coding sessions, it reports 55.5% against Sonnet 5’s 34.1%, within about two points of Opus 5.5.1
| Evaluation | Sonnet 5 | Sonnet 5.5 | Opus 5.5 |
|---|---|---|---|
| Agentic coding · Terminal-Bench 4.0 | 10.3% | 70.6% | 66.4% |
| Agentic coding · CursorBench 4.0 | 34.1% | 55.5% | 57.8% |
| Knowledge work · GDPval-AA v2.1 | 1449 | 1844 | 1846 |
| Computer use · OSWorld 2.1 | 57.0% | 80.1% | 81.8% |
| Chart recognition · Chartography | 15.6% | 61.6% | 64.4% |
But scores are half the story. Anthropic’s own charts plot score against cost per task at every effort level, and that is where the release bites: at Medium effort — the default in the Claude apps — it says Sonnet 5.5 beats Sonnet 5’s best score for less than a tenth of the cost per task.1 Same price list, fewer tokens per job, faster output. If that holds outside Anthropic’s harness, the cost model for routine agent work changes, not just the leaderboard.
Slack’s early testing points the same way, with the usual early-tester caution: better results than Sonnet 5 on almost all of its offline Slackbot evals, in fewer steps, with about 14% fewer output tokens — a vendor-published tester quote, not an independent result.1
Sonnet 5.5 leans hard on effort levels: Low, Medium, High and beyond, trading speed and tokens against thoroughness. The defaults are Medium in Claude Code and the Claude apps, High on the Claude Platform.4
Anthropic’s own framing is that Sonnet 5.5 complements Opus 5.5 at lower effort, where it costs less per task, and can match it at higher effort at comparable cost. At High effort on FrontierCode it reports a score ten points above Sonnet 5 at the same setting — at roughly one fifteenth of the cost per task.1
My read: this makes effort a routing decision, not just a quality setting. The practical question for an agent pipeline is no longer “which model” but “which model at which effort for which job class”. Teams that pin one model and one effort for everything will pay for it twice — once in tokens, once in capability left on the table.
Confirmed: the model exists, the model ID is live, pricing is published, the system card is public, and the safety posture is documented — the same biology safeguards as Sonnet 5, visible fallbacks to Sonnet 5 on higher-risk cybersecurity tasks, and classifiers against reasoning extraction, a first for the Sonnet line.12
Claimed: the benchmark table, the 30% speed and cost figures, the “a tenth of the cost per task” comparisons, and the collaboration anecdotes from Epic and Slack are all Anthropic-reported. Anthropic also reports that on its roughly 1,850-scenario behavioral audit, Sonnet 5.5 improves on or matches Sonnet 5 on most alignment measures, and that on containment evaluations it comes close to Opus 5.5 on how rarely it tries to escape its sandbox.1 Vendor-reported alignment is still vendor-reported.
Still open: independent Terminal-Bench and CursorBench replication; whether the cost advantage survives long-horizon, tool-heavy workloads; how the between_tools migration plays out in existing agent stacks; and where Haiku 5.5 lands the bottom tier when it ships.
First, re-run your own tasks at Medium effort before touching defaults. Vendor charts measure vendor-chosen tasks; your eval set is the only one that knows your workload.
Second, treat effort as part of routing. If you run an agent pipeline, this release is a good reason to log effort level alongside model and cost per task — you cannot tune a knob you do not record.
Third, keep Opus 5.5 for the work that burns. Anthropic is unusually plain that Opus remains clearly stronger on complex, open-ended work requiring sustained judgment, and benchmark scores capture only one facet of capability.1 The mid-tier became the default. It did not become the frontier.
One click tells us it helped. It costs you nothing.
What your agents can reach, what stops them, and how you prove it. One email per new article.
LatestSteal These 3 CSS Snippets That Turn Any Website Into a Newspaper
Unsubscribe in one click. We won’t sell or share your address.
Opinion · Models
Hours of unsuccessful revisions with GPT-6 Sol, an Astra escalation that exhausted my usage, and an Opus 5.5 handoff that finally worked. OpenAI’s new release arrives with a better value proposition, and some trust to recover.
Bar chart of AutomationBench 1.0.6 correct completion rates at maximum effort, with cost per task. GPT-6.1 Sol completes 36.1% of workflows at $0.30 per task. Claude Opus 5.5 with default fallbacks completes 42.5% at $1.44. GPT-6 Sol completes 32.0% at $0.34. Vendor-reported figures from OpenAI's release chart, checked 30 September 2026. Business workflows, not a coding benchmark. No independent PLEXData testing.
Opus completes more workflows. GPT-6.1 Sol costs less per task. Completion and cost are separate measures here.
Max effort · Opus includes default fallbacks · checked 30 Sep 2026GPT-6 Sol did not fix the integration issue, Astra used up my allowance, and Claude Opus 5.5 delivered the functionality with roughly 80% less code. That is my account of one job, not a controlled test.
OpenAI’s chart has GPT-6.1 Sol at 36.1% correct completion and $0.30 per task on AutomationBench. Opus 5.5 with fallbacks completes more, at $1.44.
I have not used GPT-6.1 Sol on this incident. The release earns a retest, and only finished work earns the trust.
Last week, I spent hours trying to fix an issue in a large codebase connecting several systems. GPT-6 Sol kept circling the problem, giving me iterations of the code that achieved nothing. The root cause remained unresolved, and the functionality I needed still did not work.
I eventually handed the job to GPT-6 Astra. My available usage was exhausted almost immediately, before it had completed the work. I then moved it to Claude Opus 5.5.
Opus reduced and simplified the bloated codebase by roughly 80%, delivered the desired functionality, and integrated it correctly. After the previous attempts, that was a substantial difference.
The experience left me disappointed with OpenAI. It also changed what I wanted to hear about its next model. A claim about greater intelligence matters much more when it translates into a finished job.
GPT-6.1 Sol now arrives with reported improvements in capability and cost efficiency. I hope those improvements reach the kind of work that went badly last week. That remains an open question: I have not used it on this incident.
The frustrating part was the repeated lack of progress. Another explanation, another revision, another attempt, and the same underlying issue.
There is a cost to that beyond the model’s usage allowance. You have to follow the changes, keep track of what has been attempted, and decide whether the next explanation deserves more of your time. A model that produces a plausible patch can keep a session moving long after it has stopped being useful.
Opus’s smaller solution made an impression because it also worked. The approximate 80% reduction is my account of that job, rather than an independently measured benchmark. The models received sequential handoffs, so this was not a controlled comparison from identical starting conditions.
That limits what I can claim about the models in general. It does not change which one completed the work I needed.
Public complaints after the move from GPT-5.6 Sol to GPT-6 Sol describe problems that sound familiar. One Reddit user reported instruction-following failures and a costly return to Astra. Another described a maintenance check in which GPT-6 Sol relied on stale package information while its predecessor checked more thoroughly.12
These are self-selected accounts. They show that similar frustration exists, but they do not establish how common it is or whether the same cause explains my experience.
A more detailed report in OpenAI’s Codex issue tracker compared three pairs of tasks at maximum reasoning effort. Its author recorded context and completion omissions, but also a serious database defect found by GPT-6 Sol and missed by GPT-5.6 Sol. One pair had a contamination caveat.3
The picture is mixed. There is evidence worth taking seriously, including evidence against a blanket claim that GPT-6 Sol is worse. My disappointment is specific: on a difficult integration job, repeated attempts consumed time without delivering the required result.
The metric closest to that concern is cost per successful task: total spending across attempts divided by the number that finish correctly. Arize and Fireworks have published a benchmark using that approach, although its model set does not cover this three-way comparison.6
AutomationBench provides a useful comparison here. It tests business workflows across applications and grades the final system state. A task passes only when every required assertion holds. An agent saying it has finished does not count as success.5
OpenAI’s GPT-6.1 Sol release chart reports the following AutomationBench 1.0.6 results. All three rows use the named maximum effort setting; that does not imply identical compute budgets across providers. The Opus configuration includes default fallback models.45
| Model / configuration | Correctly completed tasks | Reported cost per task, USD | Approx. cost per correct completion, USD* |
|---|---|---|---|
| GPT-6.1 Sol · Max | 36.10% | $0.2989 | $0.83 |
| Claude Opus 5.5 · default fallbacks · Max | 42.47% | $1.4400 | $3.39 |
| GPT-6 Sol · Max | 31.96% | $0.3406 | $1.07 |
*The final column is an editorial calculation from the published figures: reported cost per task ÷ success rate expressed as a fraction. For example, $0.2989 ÷ 0.361 = approximately $0.83. It treats the reported cost as the average across attempted tasks. It is not a separately measured result, a guarantee that retries will solve a particular failure, or a forecast of subscription usage.
In these results, Opus completes the largest share of tasks. The new Sol has the lowest reported cost per task and the lowest calculated spend per correct completion. Those are different advantages, and both matter.
AutomationBench concerns business workflow execution, rather than debugging my codebase. OpenAI’s published DeepSWE coding comparison does not include Opus 5.5, so it cannot supply the requested three-way coding chart.4
The fallback qualification matters too. Anthropic separately reports an Opus AutomationBench result without fallbacks; that is a different configuration and should not be mixed into this table.7
These are published evaluation results, not testing performed by PLEXData. API costs also cannot tell us how quickly my subscription allowance would have been consumed. The time spent supervising unsuccessful work is another cost this table does not capture.
I would like to see GPT-6.1 Sol hold together the details of a connected system well enough to recognise where the failure begins. In last week’s job, another version of the code was of little value while the underlying problem remained.
I would like fewer confident explanations that lead back to the same fault. When the model has not established a cause, I would rather that uncertainty be clear before more time and capacity disappear into another round of changes.
The Opus handoff also raised my expectations about simplicity. It showed that, in this case, the desired functionality could be delivered with substantially less code. I would like the next Sol to recognise when complexity is getting in the way, while preserving the behaviour the system actually needs.
Above all, I would like enough usable capacity to finish difficult work. A model’s capability and the amount of it available to a customer belong in the same conversation. The Astra escalation was a reminder that access to a more capable model has limited practical value when the allowance is gone before the result arrives.
The published numbers give GPT-6.1 Sol a stronger case than its predecessor. They do not erase last week’s experience.
Opus 5.5 completed that job. GPT-6 Sol did not, and the move to Astra exhausted my available usage before completion. That is one personal experience, but it is the experience against which I will read the new claims.
I want OpenAI to have addressed it. A model that finishes more of the work, wastes less of the user’s time, and leaves a simpler working system would be welcome.
The release earns a retest. The work earns the trust.
Sources checked 30 September 2026, Australia/Brisbane. Public accounts and published evaluations were reviewed; no model runs were performed for this article.
One click tells us it helped. It costs you nothing.
What your agents can reach, what stops them, and how you prove it. One email per new article.
LatestSteal These 3 CSS Snippets That Turn Any Website Into a Newspaper
Unsubscribe in one click. We won’t sell or share your address.
Opinion · Money & Banking
US broad money grew 40% in two years. Consumer inflation peaked sixteen months after money growth did. I think the surge belongs in the inflation story alongside supply shocks, and the cost mattered most to households with little room to absorb it.
Minimal animated illustration. A small printing press feeds banknotes onto a growing stack labelled Money. A price tag on a pole to the right starts low and rises later, after the stack has grown.
Money surged in 2020–21; consumer inflation peaked in June 2022. The sequence is clear. How much money caused the rise needs more evidence.
My opinion on the dataUS M2 rose from $15.35 trillion in December 2019 to $21.50 trillion in December 2021, a rise of 40%. Growth peaked at 26.8% a year in February 2021.1
Consumer inflation peaked sixteen months after money growth did. Energy, food and reopening also mattered. I think the honest reading keeps money and supply shocks in view together.
The timing alone cannot tell us how much inflation came from monetary demand, reopening or supply shocks. Establishing that split needs more than two lines on a chart.
In 2020 and 2021, the central banks of the West did something they had never done at this scale in peacetime. By the end of it, roughly 29 cents of every dollar of US broad money had been created in the previous two years. The phrase people reached for was “money printing”, and it has been argued over ever since. I think the phrase is a bad description and a fair warning.
This piece is the 2026 edition of an article first published on 2 December 2025. I have kept its structure, updated the data to June 2026, and asked a harder question of it: how much weight the money numbers deserve now that the dust has settled.
The popular version is short. Governments and central banks created huge amounts of money, prices rose, and everyone got poorer. Versions of it circulated with claims that between 20 and 40 per cent of all dollars in existence were created in a couple of years.
Check the arithmetic and the claim is closer to true than its critics allow. US M2 was $15.35 trillion at the end of 2019 and $21.50 trillion at the end of 2021.1 That is a 40% rise in two years, and it means about 29% of the end-2021 stock was new. The claim only goes wrong when it treats all of that money as cash from a press. Most of it was bank deposits.
Base money, broad money and deposits. Base money is currency in circulation and the reserve balances commercial banks hold at the central bank. Broad money measures currency held by the public, deposits and other liquid balances: M2 in the United States, M3 in the euro area. Bank reserves are not part of those broad-money measures. When a bank makes a loan, it creates a deposit. Broad money can grow without a single new note.3
Three engines at once. In 2020 three things ran together. Central banks bought bonds through quantitative easing, which added reserves; purchases from non-bank sellers could also add bank deposits. Governments ran large deficits paid for by selling bonds, and the cheques landed in bank accounts. Banks kept lending.43 No single one of them is “the printing”, and that matters for what came next.
In the United States the jump was unmistakable. The year-on-year growth rate of M2 was 6.7% in December 2019, reached 22.8% by June 2020 and 24.5% by December 2020, and peaked at 26.8% in February 2021.1
Large bond-purchase programmes were part of the response in all five economies. The figures below show the scale reported by each central bank, with the programme and date stated alongside each amount.
| Economy | Reported amount | Programme and scope |
|---|---|---|
| United States | US$4.4 trillion | Fed bond acquisitions since February 2020, reported in January 2022.5 |
| Euro area | About €1.7 trillion | Net PEPP purchases from March 2020 through March 2022; excludes the separate APP.13 |
| United Kingdom | £895 billion | Total QE bond purchases, including rounds before the pandemic; £875 billion government and £20 billion corporate bonds.7 |
| Australia | A$280.7 billion | Bond Purchase Program, November 2020 to February 2022; excludes earlier market-function and yield-target purchases.6 |
| Canada | More than C$180 billion | Government of Canada Bond Purchase Program, from its March 2020 launch to the December 2020 report.8 |
These are programme snapshots, not a ranking of comparable totals. The UK figure includes older QE; the other rows cover different periods and assets. None is a measure of broad-money growth. The distinction matters: bond buying, fiscal spending and bank lending work through different channels, so a larger purchase programme need not produce a larger rise in deposits.34
The monthly series puts the US M2 growth peak at 26.8% in February 2021. Consumer price inflation, calculated from the seasonally adjusted CPI index, peaked at 9.0% in June 2022, sixteen months later.12 The chart uses quarter-end readings, which put its highest sampled M2 growth point in December 2020; that is not the monthly peak. By the end of 2021 consumer prices were 8.6% above their end-2019 level, and by June 2026 they were 28.6% above it.2
Line chart of US money growth and consumer price inflation, quarter-end readings from 2019 to mid-2026. Money growth (M2) rises from 6.7% in December 2019 to 24.5% in December 2020, falls below zero in late 2022, reaches minus 3.9% in early 2023, and recovers to 5.3% by June 2026. Consumer price inflation peaks around 9.0% in June 2022, sixteen months after the monthly money-growth peak, and is 3.5% in June 2026. Official data, with year-on-year rates calculated by PLEXData.
Money growth peaked in 2020–21 and went negative in 2023. Prices peaked in mid-2022. By June 2026 M2 was growing 5.3% a year and CPI inflation was 3.5%.
Quarter-end readings · FRED, seasonally adjustedThe timing is consistent with a delayed monetary effect, but it does not establish one. Energy and food prices also surged, and reopening demand ran into constrained supply. OECD-area headline inflation peaked at 10.7% in October 2022.9 The OECD’s household study found that energy prices drove much of the purchasing-power loss in countries including Denmark, Italy and the United Kingdom.10
My position, with US data through June 2026, is that both camps were partly right, and the argument between them was less useful than it looked. The money surge made the economy easier to push prices up in. The supply shocks lit the match. How much each economy paid depended on how large both were.
Inflation is not one rate. Renters and first-time buyers meet it through rent, mortgage payments, childcare, food and energy. Owners with fixed-rate mortgages and assets that rose in value met it very differently.
The housing evidence gives this argument firmer ground. The RBA reported in March 2023 that around half of Australian renter-household heads were aged 25 to 44. Renters also tended to have lower incomes, less wealth and smaller savings buffers than owner-occupiers.11 In Great Britain, an ONS survey covering 8 February to 1 May 2023 found that 43% of renters had difficulty affording their rent, compared with 28% of mortgage holders reporting difficulty with mortgage payments.12 These are specific household findings, not a claim that every young person paid the same cost.
An OECD study of purchasing-power losses between August 2021 and August 2022 found that inflation weighed more heavily on low-income than high-income households in every country it examined. It also found substantial exposure among rural households; it did not establish a universal ranking by age.10 That matches how I think about it. The same shock that raised the price of a flat did nothing to the wages of someone trying to rent one, and people who owned assets had a cushion that people who owned nothing did not.
The reversal came quickly. Central banks moved from emergency easing to quantitative tightening, and raised policy rates. US M2 fell from its pre-contraction high of $21.79 trillion in March 2022 and was down 3.9% on a year earlier by March 2023, the sharpest contraction of the period in the data I pulled.1 It did not stay negative. US M2 growth was 4.0% in December 2025, and 5.3% by June 2026.1
By June 2026, broad money was 6.1% above its March 2022 high, and consumer inflation was 3.5%.12 Money growth had slowed markedly from the pandemic surge. A slower inflation rate has not reset the price level.
The US consumer price index was still 28.6% above its December 2019 level in June 2026.2 Whether a household could absorb that increase depended on what happened to its income, savings and housing costs. Falling M2 did not take prices back to 2019.
First, the surge was as large as people said, and the plain-English version of the story is mostly fair. Second, money growth explained the timing better than it explained the size. Supply shocks and the way policy was run explain a lot of the rest. Third, the bill did not land evenly. It landed on the people with the least room to absorb it.
The fix I would argue for is unglamorous. Watch broad money as a signal that something is moving, not as a verdict. Keep the fiscal and monetary decisions in view together, because they came as a pair. And count the cost by who paid it, not by the average.
One click tells us it helped. It costs you nothing.
What your agents can reach, what stops them, and how you prove it. One email per new article.
LatestSteal These 3 CSS Snippets That Turn Any Website Into a Newspaper
Unsubscribe in one click. We won’t sell or share your address.
Concept · Agents and the web
More of your visitors are now agents sent by people, and they arrive without the context a person picks up in seconds. Making a site readable for them starts with one plain file and a few old habits done properly.
Minimal animated illustration. A small robot agent rides a surfboard over rolling waves. Ahead of it, a coral flag reading AGENT.md flies from a marker float, with a dotted line of sight from the agent to the flag. Buoys labelled evidence, reports and opinion float past on the waves. Along the bottom, five stages light up in turn: discover, orient, navigate, verify, use.
Give an agent readable water and a map, and it can pick its own line.
Most of the care that goes into a website assumes a person on the other end. Someone who glances at the navigation, reads a headline, scrolls a little and knows within seconds whether they are in the right place. The agent that person sends gets none of that for free. It comes in through whatever link it found, often deep inside the site, carrying a broad instruction and no feel for the place.
That visitor is now common enough to design for. I have started calling the design problem Agent Surfing: making a site easy for an agent to discover, get its bearings in, move around, check, and bring something useful back from.
SEO still matters. It is how a machine finds your page at all. But it deals with the moment before arrival, and the questions an agent has afterwards are different. What is this site? Is this page the real version or a copy? Who wrote it, and when? Is it a measurement or somebody’s view? Where would I go to check it?
A person answers most of those without noticing. An agent has to work them out from markup, link text and whatever the page says about itself. Every guess costs it steps, and some guesses end with it quoting the wrong page back to the person who sent it.
So the line I keep coming back to is this: SEO helps a machine find your page; Agent Surfing helps an agent understand where it has landed and where it should explore next.
What I am not claiming matters as much. Agent products differ, and they change month to month. I have not measured how any one of them moves through a site. This is about what any careful agent needs, whichever product it is.
The surfing picture is deliberate. A good surfer does not fight the sea. They read it, pick a line and commit, and they can only do that because the water gives them something to read. A website is either readable water or chop.
When I break down what an agent has to do on a site it has never seen, I get five moves.
| Move | What the agent needs to know | What the site can give it |
|---|---|---|
| Discover | Is there a map, and where is it? | A downloadable AGENT.md at a stable, linked address |
| Orient | What is this site, and what does it hold? | A plain description, the main sections, who runs it |
| Navigate | Where next? | Stable URLs, semantic HTML, link text that says where it goes |
| Verify | Can I trust this, and where did it come from? | Named authors, dates, canonical sources, evidence kept apart from opinion |
| Use | Can I take this back to my user? | Downloads it can open, and statements it can cite cleanly |
Most sites manage some of these by accident. The work is doing all five on purpose, and the first is the cheapest to fix.
Start with one file. A short, plain AGENT.md at a stable address, linked from the footer so it can be reached from any page. Think of it as the note you would leave for a capable stranger who has to find their way around without you: what the site is and who runs it, where the authoritative material lives, how the sections relate, and where to look for the questions people usually bring.
# Example Field Notes<!-- Illustrative example. Names and paths are placeholders. -->What this is: an independent publication on home-battery systems.Run by: Jane Example. Contact: /aboutLast updated: 2026-09-30 ## Authoritative sources/data/ original measurements, CSV, with method notes/reports/ our reporting, dated, each linking its data/opinion/ our views, labelled as opinion ## How the sections relateOpinion cites reports. Reports cite data. Start at data to verify. ## To investigate furtherPricing question: /reports/tag/pricing, then /data/prices.csvWho wrote this: /about and the byline on each page
Keep it short enough to read in one pass, and keep it true. A map of last year’s site is worse than no map, because the agent has no reason to doubt it. A community proposal called llms.txt is aimed at a similar problem.5 The name matters less than having one clear file and linking to it.
PLEX keeps its own at plexdata.online/AGENT.md. Every article here also has an Agent MD button that exports the piece as clean Markdown, with its title, author, date and sources, and each of those exports points back to the site map. An agent that starts from a single article can still find its way to the rest of the site.
A map only helps if the ground matches it, and most of the ground is ordinary web practice that has quietly become more important.
Addresses come first. An agent that saves a link, or cites one to its user, is trusting that address to mean the same thing next week. I am learning this on this site. It is part-way through a migration, and one of the old article addresses still brings in steady traffic. If that address simply broke, every link, bookmark and index entry pointing at it would break with it. Move a page if you have to, but redirect it.
Then structure. Real headings, lists, tables and links carry meaning without anyone seeing the page.6 A layout built from anonymous boxes looks fine to a person and says nothing to a machine. Titles, descriptions and structured data do the same job one level up: they tell an agent what a page is before it has read it.3
Then provenance. Put a name and a date on every page that makes a claim, and say when it changed. Where the same material lives at several addresses, declare which one is canonical, so the agent is not left choosing between near-copies.4 If the evidence is a dataset, publish the dataset with a name and a description, not just a picture of the chart.
The habit I care about most is keeping evidence, reporting and opinion visibly apart. PLEX labels each piece as news, concept, opinion or review, and inside many articles it separates what is confirmed from what a vendor claims and what is still open. That labelling was written for human readers. It turns out to be exactly what an agent needs to report a finding honestly, as a measurement or as somebody’s view.
The older plumbing still counts. A robots.txt file and a sitemap are still how many crawlers find their way in at all.12
There are two tempting wrong turns.
The first is a second website for machines: a stripped-down copy that drifts away from the real one until you are maintaining two sources of truth and hoping they agree. Make the one site legible instead.
The second is filling pages with instructions aimed at AI systems. Hidden or pushy text telling an agent what to do looks exactly like the prompt-injection attacks agent builders defend against, and a well-built agent should ignore it. It also runs against the argument in Stop Trusting the Agent: an agent’s authority should come from its user and its permissions, not from whatever it happened to read along the way. An AGENT.md should describe and point. It should never give orders.
You do not need a lab for a first look. Give an agent you already use a broad job about your own site, something like “find out what this site says about pricing and come back with sources”. Then read what it did, not only what it said. Did it find the map, or wander? Did it describe the site the way you would? Did it keep what you measured apart from what you think? Would you be happy to see its links and dates quoted back to you?
Each fumble points at something specific to fix. I have not run this across agent products for this piece, so there are no results here, only what I would look for.
If I had one afternoon, I would write the AGENT.md, put a name and a date on every page missing them, and make sure nobody, human or agent, has to guess whether a page is evidence or opinion. None of that is new work. The visitor is new.
One click tells us it helped. It costs you nothing.
What your agents can reach, what stops them, and how you prove it. One email per new article.
LatestSteal These 3 CSS Snippets That Turn Any Website Into a Newspaper
Unsubscribe in one click. We won’t sell or share your address.
Build log · CSS
The grid, the tokens and the typesetting behind PLEXData’s newspaper layout. Each snippet is the real CSS from this site, and each one runs live in the article so you can see exactly what it does.
Animated drawing of the PLEXData front page set like a newspaper, with a CSS settings panel beside it. The page builds itself: the PLEXData masthead and its rules, the lead headline 'Steal These 3 CSS Snippets That Turn Any Website Into a Newspaper', a dark figure block and a coral newsletter block in the right column, and a list of earlier articles. In the panel, the tokens paper F4EFE4, ink 11110F and coral F05E43 sit above two knobs. The first, --t-display, turns from 42 to 76 pixels and the headline grows with it. The second, --sp-3, turns from 8 to 24 pixels and the gap above the earlier list opens with it.
The classic newspaper front page, brought back with modern CSS.
The pieces of mine people read most are about CSS. And readers of those pieces have told me what they want: real examples of what the CSS actually does. Not a description of it. The thing itself.
So this is the CSS behind the page you are reading. PLEXData looks like an old newspaper on purpose: cream paper, black ink, one coral, a rule under the masthead and plenty of empty space. Most of that comes from three snippets. Each one below is copied from the live stylesheet, and each one is followed by its output, rendered by your browser from the same rules.
The old homepage was a paginated list of posts, built from templates. It worked, but every post looked as important as every other one. A newspaper front page does the opposite. One story leads and is clearly bigger. A picture holds the right-hand side. A few earlier stories sit in a tidy list. You can tell what matters before you read a word.
CSS grid lets you write that layout almost as a sentence. You name the areas, then draw the page in a string.2
.front-in{ display:grid; grid-template-columns:minmax(0,1fr) clamp(420px,40vw,600px); grid-template-areas:"lead dark" "earlier dark"; column-gap:clamp(32px,4.4vw,72px);}@media (max-width:1080px){ .front-in{grid-template-columns:1fr; grid-template-areas:"lead" "dark" "earlier"}}
PLEXData
Featured · Build log
Bringing back the classic front page with modern CSS
One lead story, a picture, a list. You can tell what matters before you read a word.
009Sonnet 5.5 Turns the Mid-Tier into the Default
008NVIDIA Moves Agent Safety Below the Model
PLEXData
Featured · Build log
Bringing back the classic front page with modern CSS
One lead story, a picture, a list. You can tell what matters before you read a word.
009Sonnet 5.5 Turns the Mid-Tier into the Default
008NVIDIA Moves Agent Safety Below the Model
The same grid rules running in this page, scaled to fit the column. The first copy is wide enough for two columns. The second is narrow, so the areas stack. The demo switches on its own width with a container query instead of the site’s 1,080px media query, so you can see both states side by side.
Three details do the work. grid-template-areas is the page plan you can read: lead and dark side by side, earlier under lead, dark running down both rows. minmax(0,1fr) instead of plain 1fr stops a long word or a wide figure from pushing the column past the edge of the screen. And clamp(420px,40vw,600px) keeps the right column between 420 and 600 pixels, so the picture never shrinks to a thumbnail or swells into a billboard.3
The stacking order is a decision too. Below 1,080 pixels the same names drop into one column: lead first, then the picture, then the list. Change the order of three words and the phone layout changes with it.
The second snippet is the one I would steal first. It is a short list of custom properties, and almost every colour, size and gap on the site comes from it.
:root{ --paper:#F4EFE4; --ink:#11110F; --muted:#655F57; --coral:#F05E43; --coral-text:#C4391F; --t-display:clamp(42px,5.3vw,76px); --t-xxl:clamp(38px,4.6vw,68px); --t-xl:clamp(32px,3.5vw,50px); --t-m:clamp(24px,2.2vw,32px); --t-s:clamp(20px,1.7vw,24px); --t-xs:18px; --sp-1:8px; --sp-2:16px; --sp-3:24px; --sp-4:40px; --sp-5:64px; --sp-6:96px;}
Aa coral text
#F05E43 on paper · 2.87:1 · failsAa coral text
#C4391F on paper · 4.63:1 · passesInk on coral
#11110F on #F05E43 · 5.75:1 · passes--t-displayNewsprint
--t-xxlNewsprint
--t-xlNewsprint
--t-mNewsprint
--t-sNewsprint
--t-xsNewsprint
--sp-18px
--sp-216px
--sp-324px
--sp-440px
--sp-564px
--sp-696px
Every swatch, size and bar here is set with the site’s own custom properties. Contrast ratios are computed with the WCAG formula from the hex values.
Look at the first two swatches. The bright coral, #F05E43, is the colour people remember, and it is too light for small text. On the paper background it reaches 2.87 to 1. The accessibility guideline asks for 4.5 to 1 for normal text and 3 to 1 for large text, so it fails both.1 So coral has two tokens. The bright one fills shapes: buttons, bars, the rule on top of a box. Every word that needs to be coral uses #C4391F, at 4.63 to 1. Buttons go the other way, with ink text on a coral fill.
The type ramp is six clamp() values, each running from a phone size to a desktop size, so a headline scales smoothly with the window instead of jumping at breakpoints. Resize this page and the ramp above moves with it. The stylesheet still shows why the limit was needed: an older rule gives the hero headline its own clamp(46px,7vw,108px), and a later rule overrides it with var(--t-display).
Space works the same way: 8, 16, 24, 40, 64 and 96 pixels, every step a multiple of 8, the steps further apart as they grow. Newer rules take their gaps from that ladder. Older ones still carry their own numbers, like a 44-pixel gap in the article grid.
Some of the best typesetting on this site is done by the browser, with two lines of CSS and one selector.
.hed{text-wrap:balance}.prose p{text-wrap:pretty} /* a thin rule between byline items, never before the first */.byline>*+*::before{ content:""; display:inline-block; width:1px; height:12px; margin:0 12px; background:var(--soft); vertical-align:-2px;}
Bringing back the classic front page with modern CSS
Bringing back the classic front page with modern CSS
Both headlines have the same text, font and width. The only difference is text-wrap. Browsers without balance support show two identical boxes. The byline underneath is the site’s real byline class.
text-wrap: balance evens out the line lengths of a heading, so a title does not leave one short word stranded on its last line. It only works on short blocks, a handful of lines, which is exactly what a headline is.4 text-wrap: pretty asks the browser to take more care over paragraphs and avoid a lonely last word. Browsers that do not support either value wrap as normal, so there is no fallback to write.
The byline separators come from the selector >*+*, which means every item that follows another item. A thin rule appears between items and never before the first. Add a fourth item to a byline and it gets its separator for free.
Three snippets carry the look. A few other habits keep it from breaking.
The figures at the top of each article respond to their own width, not the screen’s, using container queries.5 That is why the drawing at the top of this article looks right both here and in the narrower column on the front page.
Before anything ships, a script opens every page at 1,440, 1,024, 390 and 320 pixels wide and checks one number: is the page wider than the window? The last time it fired was in the release this article belongs to. A long piece of inline code in an older article refused to wrap and pushed the whole page sideways on phones. The fix was four lines.
And the whole site is one HTML file, under a megabyte, which the server serves at a real address for each article. The same plain structure makes it easier for AI agents to read, which I wrote about in Agent Surfing.
One click tells us it helped. It costs you nothing.
What your agents can reach, what stops them, and how you prove it. One email per new article.
LatestSteal These 3 CSS Snippets That Turn Any Website Into a Newspaper
Unsubscribe in one click. We won’t sell or share your address.
13 articles · Newest first
Build log · CSS
Concept · Agents and the web
Opinion · Money & Banking
Opinion · Models
News · Models
News · Agent security
Agent skill review · Third party
Book review · AI engineering
Concept · Agent security
Permissions & limits · Analysis
Execution · 10 lines
Evidence · Privacy
Limits · AI Security
Resources
Three free tools built from the code in PLEX articles, and the product we build for teams.
Pick a tool to open it. Everything runs in this page; nothing is uploaded.
About PLEX
PLEX is a technical source for people and software agents that read, decide, build and act on systems. It turns security and reliability ideas into mechanisms, code and evidence that help both work more effectively without quietly gaining authority they should not have.
Curated by Dorian Sotpyrc. Written so a person or software agent can extract the mechanism, assumptions, implementation and test — not just a conclusion.
What a person or agent may touch, and for how long. Granted for a task, never by default.
Where work has to stop: timeouts, budgets, sandboxes, egress and escalation.
What the work leaves behind, so a person or supervising agent can verify what happened.
The part that actually runs. Small, tested, reproducible and inspectable.
We design the boundaries before the behaviour. A system that can do anything will, eventually, do the wrong thing.
A policy document is not a control. If it isn't enforced by permissions, code or network, it doesn't exist.
A person or agent should be able to inspect the mechanism, its assumptions and the test without trusting a black box.
Every claim comes with something you can run, log or test — and a way to tell when it stops working.
Every control slows something down. We write with honest trade-offs, because the reader has to live with them.
Work with Dorian
If you're collaborating, commissioning work, or adapting these patterns for a team or agent stack, write directly. A short description of the system, what it can reach and what worries you is enough to start.
Map what your automated systems can reach, and where the gates should go.
Take a PLEX pattern into your own stack, with the tests that prove it works.
Bring a real problem. We'll work out the mechanism together and publish it.