AI/ML engineers
For engineers building LLM, RAG, and agent systems who need to understand reliability beyond average benchmark performance.
A technical book written in public
AI systems can succeed repeatedly and still contain hidden failure modes. Their behavior depends on prompts, context, model configuration, tools, retrieved information, and other operating conditions.
develops a practical framework for measuring failure probability, discovering where risk concentrates, evaluating whether interventions work, and operating more reliable LLM and agent systems.
Starting with repeated trials and empirical failure estimation, the book builds toward conditional and latent failure models and applies them to tools, RAG systems, autonomous agents, guardrails, failure prediction, reliability budgets, monitoring, and production assurance.
This is a living technical book. Chapter 1 is available free now. Founding Reader access will open with Chapter 2, targeted for October 1, 2026. After that, new chapters are planned approximately every three weeks as the manuscript develops.
About the book
Conventional software can often be tested by asking whether a fixed input produces the expected result. LLMs and agents complicate that model: the same system can produce different outcomes across repeated executions, and its failure probability can change with prompts, context, tools, retrieval, model configuration, and deployment conditions.
treats reliability as an empirical engineering problem: estimate the probability of failure, understand how that probability changes under operating conditions, and determine whether interventions actually reduce risk.
The book develops that reasoning progressively:
and then applies it to tool use, retrieval-augmented generation, agents, failure propagation, guardrails, prediction, reliability budgets, monitoring, and production assurance.
Practical outcomes
Readers will learn how to:
Estimate AI failure probability using repeated trials instead of relying only on single executions or aggregate benchmark scores.
Quantify uncertainty and distinguish an observed failure rate from the underlying probability it is trying to estimate.
Validate human and LLM-based judges before relying on them as measurement instruments.
Measure how prompts, models, configurations, tools, retrieval, and operating conditions change failure risk.
Analyze reliability across RAG pipelines, tool-using systems, and long-running agent trajectories.
Test interventions, predict high-risk conditions, allocate reliability budgets, and monitor reliability after deployment.
Writing schedule
After Chapter 2, new chapters are planned approximately every three weeks.
This is a living technical book. Release dates may shift when experiments, validation, technical review, or editing require additional time.
Chapter 1
Free
Chapter 2
Target: October 1, 2026
Founding Reader access opens at
Chapter 3
Approximately three weeks after Chapter 2
Chapter 4
Approximately three weeks after Chapter 3
Chapter 5
Approximately three weeks after Chapter 4
The manuscript
Twenty-one chapters move from measuring failure to production reliability and assurance.
The reliability mental model, system boundaries, failure definitions, repeated execution, and the first reliability harness.
Bernoulli trials, binomial experiments, empirical failure rates, and estimating failure probability.
p_f = P(F = 1)Confidence intervals, zero-failure experiments, sample size, rare failures, and the difference between observed failure rate and underlying failure probability.
Rubrics, structured judgments, judge stochasticity, false positives, false negatives, and how judge error affects reliability estimates.
Blinded human review, confusion matrices, precision, recall, sensitivity, specificity, Cohen’s κ, disagreement analysis, bootstrap intervals, and deciding whether an automated judge is fit for use.
Conditional failure probability, prompt and context effects, model and decoding settings, system prompts, harness configuration, user state, and environment.
p_f(x) = P(F = 1 | X = x)Factors, controls, replication, randomization, blocking, factorial designs, ablations, frozen artifacts, and reproducibility.
Stress testing, interaction effects, stratification, high-risk regions, adversarial conditions, and response surfaces.
Single-run versus repeated-run evaluation, hidden failures, k-of-N outcomes, repeated-use risk, and cumulative failure probability.
P(at least one failure in N) = 1 − (1 − p)ᴺFoundation models, harnesses, prompts, orchestration, memory, guardrails, APIs, and the distinction between model-level and system-level failure.
Wrong-tool selection, incorrect arguments, tool execution failures, misinterpretation of results, false claims of successful action, and unsafe actions.
RAG pipelines, retrieval failures, context quality, stage-wise reliability, and propagation of upstream errors.
Sequential decisions, accumulating risk, per-step reliability, task completion probability, recovery, and dependence across steps.
Correlated failures, common-cause failures, dependencies, cascading actions, redundancy, recovery, and system reliability models.
Risk models for AI behavior, conditional failure prediction, embeddings, metadata, graph features, telemetry, and supervised models for failure prediction.
P(F = 1 | X)Risk-based sampling, ranking, top-k targeting, lift, failure-discovery recall, adaptive sampling, testing budgets, and stopping criteria.
Acceptable failure probabilities, severity classes, probability-consequence tradeoffs, component-level budgets, safety margins, escalation, and acceptance criteria.
Prompt changes, model changes, decoding, routing, retrieval, verification, guardrails, redundancy, retries, human review, and experimental verification of whether an intervention actually lowers failure probability.
p_fDrift, changing models, prompts, tools, user populations, retrieval corpora, distribution shift, rolling estimates, change detection, and statistical process control.
Detection, reproduction, diagnosis, traces, logs, root-cause analysis, failure taxonomies, corrective action, and mechanistic evidence where available.
Continuous testing, measurement, prediction, control, monitoring, judge validation, experiment provenance, audit trails, reliability reports, and retesting.
Pedagogical approach
The mathematics appears when the engineering problem requires it. Each chapter begins with a failure or reliability question, introduces the minimum statistical machinery needed to reason about it, implements the analysis, and ends with an engineering decision.
Audience
For engineers building LLM, RAG, and agent systems who need to understand reliability beyond average benchmark performance.
For practitioners designing evaluations, stress tests, human-validation studies, red-team exercises, and assurance programs.
For engineers bringing reliability, observability, fault analysis, and operational-risk thinking to probabilistic AI systems.
For people making deployment and governance decisions who need defensible evidence about failure probability, uncertainty, controls, and residual risk.
Editorial transparency
AI is used as an editorial and technical assistant, not as an autonomous author.
The book begins with the author’s own research, reasoning, experiments, and drafting. AI tools are used selectively to help brainstorm, organize and revise material, improve clarity, and assist with code and technical examples.
The author reviews and takes responsibility for the book’s technical claims, experimental results, interpretations, and final text. AI-generated material is treated as working material—not as an authoritative source.
Publishing model
The opening chapter is available now and will remain free.
Chapter 2 is targeted for October 1, 2026, and paid access will open at when it is published.
Founding Readers receive later chapters without paying future price increases as the manuscript grows.
The purchase includes all subsequently released chapters and the completed first edition.
Founding Reader access is not available yet. It will open with Chapter 2. This is a living technical book, so chapters may be revised, reorganized, expanded, or corrected as the work develops.
Purchasing and access
Yes. Chapter 1 is available free now and will remain free.
It opens when Chapter 2 is published, currently targeted for October 1, 2026. The date is a target and may shift if the chapter needs additional experimental, technical, or editorial work.
The current paid release, every subsequently released chapter, and the completed digital book. Early purchasers will not pay later price increases.
Founding Reader access will open at . The price may increase as the manuscript grows, but no later pricing tiers are being announced yet.
PDF will be included from the paid launch. EPUB support is planned for a later release.
No. This is a living technical book. Chapters, examples, organization, and wording may change as the manuscript develops.