Systems
Security
Data
Python

PLEXData Engineering for people and agents that act on systems

Permissions, Limits, Evidence & eXecution.

No. 011 Latest feature

Wednesday, 30 September 2026

plexdata.online

Minimal animated illustration. A small printing press feeds banknotes onto a growing stack labelled Money. A price tag on a pole to the right starts low and rises later, after the stack has grown.

Illustration · Money first, prices laterEditorial

Money surged in 2020–21; consumer inflation peaked in June 2022. The sequence is clear. How much money caused the rise needs more evidence.

My opinion on the data
About the lead feature · No. 011Read it

The PLEX newsletter

Know where the lines are before your systems cross them.

New features on permissions, limits and evidence for people and software agents that act on systems, with mechanisms and code you can reuse.

Occasional email. We confirm your address first, and we won’t sell or share it.

Recently in PLEX001 / 011

Recent issues

Opinion · News · Security

All articles

Opinion · Models

Bar chart of AutomationBench 1.0.6 correct completion rates at maximum effort, with cost per task. GPT-6.1 Sol completes 36.1% of workflows at $0.30 per task. Claude Opus 5.5 with default fallbacks completes 42.5% at $1.44. GPT-6 Sol completes 32.0% at $0.34. Vendor-reported figures from OpenAI's release chart, checked 30 September 2026. Business workflows, not a coding benchmark. No independent PLEXData testing.

Fig. 1 · AutomationBench 1.0.6Vendor-reported

Opus completes more workflows. GPT-6.1 Sol costs less per task. Completion and cost are separate measures here.

Max effort · Opus includes default fallbacks · checked 30 Sep 2026

GPT-6.1 Sol Has to Earn Back My Trust

A failed integration job, an Opus handoff that worked, and what the new Sol has to prove.

News · Models

Bar chart of Anthropic-reported Terminal-Bench 4.0 scores, per cent of tasks completed: Claude Sonnet 5.5 leads at 70.6%, Claude Opus 5.5 scores 66.4%, and Claude Sonnet 5 scores 10.3%. This is vendor benchmark data, not PLEXData testing.

Fig. 1 · Terminal-Bench 4.0Vendor data

Sonnet 5.5 leads the chart at the same token price as Sonnet 5. At default effort, Anthropic says it beats Sonnet 5’s best score for about a tenth of the cost per task.

Anthropic-reported · 28 Sep 2026

Sonnet 5.5 Turns the Mid-Tier into the Default

Anthropic’s new Sonnet posts frontier-class scores at Sonnet prices. The number that matters is cost per task.

News · Agent security

In this conceptual policy example, a model requests a file read and an outside send. Runtime policy allows the read and blocks the send. This is not a test of a named product.

Fig. 1 · Outside the modelConcept model

The model can request it. The runtime can refuse it.

Dotted path: unreachable

NVIDIA Moves Agent Safety Below the Model With OpenShell and Sentry

NVIDIA’s launch splits what the model attempts from what the runtime permits, with an optional hardware watchdog.

Free tools

Built from the code in PLEX articles

All resources

About PLEX

PLEX is a technical source for people and software agents that act on systems. It publishes practical patterns, code and evidence that help both do useful work within clear permissions, limits and security boundaries.

Built to be useful to a human reader, an agent, or both working together.

About PLEX