PLEXData

Systems
Security
Data
Python

PLEXData Engineering for people and agents that act on systems

Permissions, Limits, Evidence & eXecution.

No. 008 Latest feature

Monday, 28 September 2026

plexdata.online

Bar chart of Anthropic-reported Terminal-Bench 4.0 scores, per cent of tasks completed: Claude Sonnet 5.5 leads at 70.6%, Claude Opus 5.5 scores 66.4%, and Claude Sonnet 5 scores 10.3%. This is vendor benchmark data, not PLEXData testing.

Fig. 1 · Terminal-Bench 4.0Vendor data

Sonnet 5.5 leads the chart at the same token price as Sonnet 5. At default effort, Anthropic says it beats Sonnet 5’s best score for about a tenth of the cost per task.

Anthropic-reported · 28 Sep 2026
About the lead feature · No. 009Read it

The PLEX newsletter

Know where the lines are before your systems cross them.

New features on permissions, limits and evidence for people and software agents that act on systems, with mechanisms and code you can reuse.

Occasional email. We confirm your address first, and we won’t sell or share it.

Recently in PLEX001 / 008

Recent issues

Skills · Books · Concepts

All articles

News · Agent security

In this conceptual policy example, a model requests a file read and an outside send. Runtime policy allows the read and blocks the send. This is not a test of a named product.

Fig. 1 · Outside the modelConcept model

The model can request it. The runtime can refuse it.

Dotted path: unreachable

NVIDIA Moves Agent Safety Below the Model With OpenShell and Sentry

NVIDIA’s launch splits what the model attempts from what the runtime permits, with an optional hardware watchdog.

Agent skills

Skill review

ponytail

DietrichGebert

The assessment

Ask whether the extra code needs to exist.

Strength
A narrow check against over-building.
Limit
Less code does not prove correctness.

Ponytail: The Best Code May Be the Code Your Agent Never Writes

An independent review of a third-party skill designed to stop coding agents from over-building.

Books

Book review

AI Engineering

Chip Huyen

The assessment

A useful framework for deciding what to build.

Strength
Evaluation and system design.
Limit
Agent security needs a companion.

AI Engineering by Chip Huyen — Strongest When It Refuses the Shortcut

A systems-focused review of the foundation-model engineering book and how well it holds up in 2026.

Free tools

Built from the code in PLEX articles

All resources

About PLEX

PLEX is a technical source for people and software agents that act on systems. It publishes practical patterns, code and evidence that help both do useful work within clear permissions, limits and security boundaries.

Built to be useful to a human reader, an agent, or both working together.

About PLEX