Systems
Security
Data
Python

PLEXData Engineering for people and agents that act on systems

Permissions, Limits, Evidence & eXecution.

No. 012 Latest feature

Wednesday, 30 September 2026

plexdata.online

006

Published
Author
Dorian Sotpyrc
Reading
8 minutes · Review
Fields
  • Books
  • AI Engineering
  • Evaluation
  • Agents

Book review · AI engineering

AI Engineering by Chip Huyen — Strongest When It Refuses the Shortcut

A broad systems book for building with foundation models. Its lasting value is the decision framework; its weakest area in 2026 is the security depth needed for agents that can act.

References

Book review

AI Engineering

Chip Huyen

The assessment

A useful framework for deciding what to build.

Strength
Evaluation and system design.
Limit
Agent security needs a companion.
In this article · 5 sections
  1. 01What the book is
  2. 02Why evaluation is the centre
  3. 03What ages well
  4. 04Where 2026 has moved on
  5. 05Who should read it
Book
AI Engineering
Author
Chip Huyen
Publisher
O'Reilly
Length
534 pages
88/100

PLEX review score

Editorial judgment. Criteria are fixed before the total; scores are not publisher or reader ratings.

  1. Systems framing19/20
  2. Evaluation20/20
  3. Production practicality18/20
  4. Agent/security depth14/20
  5. Shelf life17/20

The best technical books teach a sequence of decisions, not a sequence of tools. AI Engineering mostly does that.

Chip Huyen's book is explicitly not a code-along tutorial. Her companion repository says the aim is to provide a framework for adapting foundation models to real applications, with questions around evaluation, RAG, agents, fine-tuning, data, inference, latency, cost and feedback loops.1

That choice is why the book has held up better than a framework-specific title would have. O'Reilly lists 534 pages and describes an end-to-end progression from foundation-model applications through evaluation, prompting, RAG and agents, fine-tuning, dataset engineering, inference optimisation and feedback.2

§ 01What the book is

This is a map of AI application engineering for people who already know how software projects behave. It spends less time telling you which SDK to install and more time asking what should be evaluated, what should be retrieved, what should be fine-tuned, and what should remain outside the model.

The official table of contents makes that breadth visible: foundation models, two chapters on evaluation, prompting, RAG and agents, fine-tuning, datasets, inference optimisation, architecture and user feedback.3

The cost of that breadth is obvious too. You do not finish the book with one application assembled. You finish with a larger set of engineering questions.

§ 02The real centre of the book is evaluation

The strongest choice is architectural rather than topical: evaluation appears before most of the fashionable adaptation techniques.

That sequencing matters. If you cannot state what “better” means, prompt changes, RAG, model swaps and fine-tuning become activity rather than engineering. The book treats open-ended output evaluation as a first-class system problem, including human evaluation, model-based evaluation and task-specific criteria.

For PLEX readers, this is the most transferable lesson: define evidence before optimisation.

§ 03What ages well

The book's durable material is the material least tied to model names: application scoping, evaluation, context construction, data quality, inference trade-offs and feedback loops.

Huyen's own companion repo says the book focuses on fundamentals rather than a particular tool or API because tools age quickly.1 That editorial decision has paid off. The names of models have changed; the questions around latency, cost, data, evals and whether to retrieve or fine-tune have not disappeared.

§ 04Where 2026 has moved on

The weakest score is not because the agents chapter is poor. It is because production agents have moved from “models with tools and planning” toward identity, delegated authority, sandboxes, runtime policy and prompt-injection-resistant workflows.

Chapter 6 covers tools, planning, agent failure modes, evaluation and memory.4 What it cannot fully reflect is the security architecture that became more explicit through 2026: NIST work on agent identity and authorization, NSA MCP guidance, and runtime controls around agent action.

That is not a reason to skip the chapter. It is a reason to pair it with newer agent-security material.

Strong

Evaluation discipline

The book makes quality measurable before it makes the architecture more complicated.

Strong

Tool-agnostic framing

Concepts survive model and framework churn better than implementation recipes.

Pair with newer work

Agent security

Identity, runtime boundaries and delegated authority deserve a 2026 companion reading list.

§ 05Who should read it

Read it if you are moving from model demos toward a production application and need a coherent map of the decisions.

Read selected chapters if your work is already specialised. Evaluation, RAG/agents, inference and architecture can stand alone.

Do not buy it expecting a current framework cookbook. The author says it is not a tutorial book, and that is the point.1

§References

  1. 1Chip Huyen — AI Engineering companion repositorygithub.com
  2. 2O'Reilly — AI Engineeringoreilly.com
  3. 3Google Books — contents and bibliographic recordbooks.google.com
  4. 4O'Reilly — Chapter 6: RAG and Agentsoreilly.com