Showing posts with label harness engineering. Show all posts
Showing posts with label harness engineering. Show all posts

Sunday, June 07, 2026

How 20-Year-Old Software Principles Govern Today's AI Agents

In our previous post, we explored the mathematical realities of agentic LLM architectures, detailing how deterministic harnesses and human-in-the-loop (HITL) evaluations prevent cascading failures. But if we step back from the probability formulas, a profound realization emerges: the architectural challenges we face with AI agents today are almost identical to the human management challenges we faced two decades ago.

In 2006, I co-authored a paper titled "Degree of Freedom - Experience of Applying Software Framework". Back then, we were not managing Large Language Models; we were managing global virtual system development teams. Yet, the core software engineering principle we applied then is exactly what we rely on now. We just use a different vocabulary.

Whether dealing with offshore human developers in 2006 or autonomous digital colleagues in 2026, unchecked freedom in a complex system inevitably leads to entropy.

The 2006 Problem: The Design-Implementation Gap

Twenty years ago, a major challenge in programming-in-the-large (PITL) was managing the control between upstream design teams and downstream implementation teams. Ideally, an implementation team would take a user-requirement-based design and execute it closely. In practice, this was incredibly difficult to enforce across different geographical locations and time zones.

We realized that the primary cause of this design-implementation gap was the degree of freedom granted to the implementation developers. Given system design, developers simply had too many ways to implement it.

When developers are given too much freedom, human nature and individual coding habits take over:

  • Developers would frequently bypass the structure of the original design by adding new classes to make their own coding easier.
  • Junior developers, failing to recognize necessary architectural layers, would mix persistence operations, presentation logic, and business operations all into a single server page.
  • Developers would bring in their own pre-existing habits, causing the implementation to drift further from the intended design.

Object-oriented programming was supposed to impose discipline, but we discovered that OO without strict discipline was just as chaotic as unstructured programming with goto statements.

The 2006 Solution: Limiting Freedom via Frameworks

To bridge this gap, we adopted and eventually open-sourced a software frameworks (Struts, Hibernate,  MVC pattern, etc). Our original goal was to standardize development and reuse major components. However, the most powerful benefit came as a surprise: the framework allowed us to efficiently manage the design-implementation gap.

The frameworks solved our problems by fundamentally limiting the developers' choices:

  • It constrained developers, forcing them to follow the framework's structure rather than improvising their own designs to make things work.
  • It homogenized the highly diversified coding styles and habits of developers, which was especially crucial in development centers with high turnover rates.
  • It made any deviations highly visible; if a developer tried to overload a single page with too much logic, the framework violation became immediately obvious during code reviews.

By limiting the degree of freedom, we standardized the process and ensured reliability.

Fast Forward to Today: Harness Engineering for Digital Colleagues

Today, we are building agentic pipelines where LLMs act as our "downstream implementation team." And once again, we are facing a massive design-implementation gap.

An LLM has a near-infinite degree of freedom. If you give an open-ended prompt to an LLM, it has billions of latent pathways it can take to generate an answer. Without constraints, it will "improvise"—which we now call hallucinating. It will mix logic, ignore the intended architecture, and output unstructured data, much like a junior developer in 2006 trying to cram all business and persistence logic into a single script.

In our previous post, we discussed how to fix this using Harness Engineering. A harness is simply the 2026 term for a framework.

When we build a deterministic harness around an LLM (such as enforcing strict JSON schemas, unit test validation, or routing the agent through a rigid LangGraph state machine), we are doing exactly what the frameworks did two decades ago. We are removing the agent's degree of freedom.

  • Instead of hoping the AI formats correctly, a structural harness forces it into predefined operational layers.
  • Instead of relying on the AI's "creativity," a deterministic evaluator acts as the ultimate code review, instantly catching any deviation from the required schema.
  • Instead of unconstrained generation, we limit its choices so that its output is homogenized, predictable, and manageable.

The Enduring Principle of Software Engineering

Technology evolves at a blinding pace, but the fundamental laws of systems engineering do not. The transition from human offshore teams to digital LLM colleagues has not changed the rules of the game.

Discipline is the bedrock of production software. Twenty years ago, we instilled that discipline by building frameworks to constrain human developers. Today, we instill it by building deterministic harnesses to constrain AI agents. The names have changed, but the secret to reliable architecture remains exactly the same: carefully control the degree of freedom.

Saturday, June 06, 2026

Beyond p^n: The Mathematics of Agentic Reliability and Harness Engineering

The transition from a linear Retrieval-Augmented Generation (RAG) pipeline to an agentic architecture is often framed as an upgrade in "intelligence." But under the hood, it is fundamentally an upgrade in probability.
If you have built a naive, multi-step LLM pipeline, you have likely encountered cascading failure. This post explores the exact mathematics of why linear pipelines collapse, how agentic retry loops fix them, and why your system is ultimately only as strong as your evaluation harness.

1. The Open-Loop Trap: The Math of "AND"

In a standard, fixed pipeline (e.g., Retrieve -> Rerank -> Synthesize -> Format), the architecture is an open loop. To get a successful final output, the LLM must succeed at Step 1, AND Step 2, AND Step 3.
If your underlying model has a zero-shot accuracy of p for any given task, the probability of successfully navigating an n-step pipeline is the product of those probabilities:
$$P(\text{success}) = p^n$$
Because p is a fraction, p^n decays exponentially. If your model gets it right 80% of the time (p = 0.8), a 4-step pipeline has a success rate of just 40.9%. The system is mathematically guaranteed to degrade as it scales.

2. The Agentic Loop: The Math of "OR"

Agentic architectures replace the "AND" with an "OR" by introducing a feedback loop. Instead of blindly passing data to the next node, an agent evaluates the result. If the result is flawed, it tries again, up to a limit of k retries.
Assuming the system can perfectly detect when an error occurs, the probability of failing a step is no longer (1-p). It is the probability of failing every single retry: (1-p)^k.
Subtracting this from 100% gives us the probability of succeeding at least once during those k attempts:
$$P(\text{step success}) = 1 - (1-p)^k$$
With p = 0.8 and k = 3, the success rate for a single step jumps from 80% to 99.2%. When you multiply these fortified nodes together, the pipeline stabilizes.

3. The Evaluator Bottleneck: Enter q

The previous equation makes a massive, often fatal assumption: that the agent perfectly recognizes its own failures. In reality, the evaluation step has its own probability of correctness, which we call q. When the agent generates an answer, the evaluator might falsely accept a hallucination (False Positive) or falsely reject a correct answer (False Negative).
When we account for an imperfect evaluator q, the math shifts from basic probability into a stochastic Markov chain. The master equation for a single agentic step becomes:
$$P(\text{step success}) = \left( \frac{p \cdot q}{p \cdot q + (1-p)(1-q)} \right) \left( 1 - (p(1-q) + (1-p)q)^k \right)$$

The False Positive Ceiling

Notice the fraction on the left side of the equation: $$\frac{p \cdot q}{p \cdot q + (1-p)(1-q)}$$.
This represents the absolute mathematical ceiling of your agent. Even if you give the agent infinite retries (k = infinity), it can never exceed this success rate. If both your generator and evaluator have an accuracy of 70% (p = 0.7, q = 0.7), your step's maximum possible success rate is capped at roughly 84%. False positives will eventually let bad data leak through the loop.

4. Harness Engineering: Breaking the Ceiling

To maximize the master equation, we must maximize q. This is the domain of Harness Engineering - the design of the environment that catches, tests, and evaluates the agent's output.
Harnesses generally fall into two categories:

Indeterministic Harness (LLM-as-a-Judge)

This relies on another prompt (or a separate model) to evaluate the output. "Does this summary look accurate?"
  • The Problem: q remains fractional. You are fighting probability with probability.
  • The Mitigation: If you must use an LLM judge, the task must be highly asymmetrical. Evaluating is easier than generating, so q is naturally higher than p. Furthermore, developers should route evaluations to heavier, highly capable models to push q as close to 0.99 as economically feasible.

Deterministic Harness (The Golden Rule)

The most robust agents do not use LLMs to check LLMs. They offload evaluation to strict, true/false code.
  • Format Checking: Did the model output pure, valid JSON, or did it break the code by adding conversational filler like "Here is your JSON:"?
  • Execution Checking: If the agent wrote a block of code, did it actually run successfully without crashing?
  • Rule Counting: If the prompt asked for exactly five product recommendations, does len(list) == 5?
By using code as the harness, q becomes 1.0.
When q = 1.0, the mathematical ceiling shatters. The left side of our master equation becomes 1, and the formula collapses back to the ideal 1 - (1-p)^k. A deterministic harness allows an agent to safely iterate until it hits a guaranteed success.

5. Human-in-the-Loop (HITL) as the Ultimate Harness

There are scenarios where a deterministic harness is impossible to build, and an indeterministic harness is too risky. If an agent is writing a sensitive client email or executing a live database mutation, a false positive is catastrophic.
This is where Human-in-the-Loop (HITL) architecture becomes a mathematical necessity.
HITL should not be viewed as a failure of automation; it is a dynamic routing strategy. When an agentic step involves high-stakes subjectivity, the retry loop routes the evaluation to a human operator.
  1. The human acts as an evaluator with a functional q ≈ 1.0.
  2. The human can either accept the output (Terminal Success), or reject it and provide targeted feedback.
  3. The LLM processes the human feedback on the next iteration, artificially boosting its generation probability (p) for the subsequent attempt.

Conclusion

Building a production-ready agentic system is not about choosing a smarter foundational model. It is about actively managing the mathematical probabilities of your pipeline. By keeping steps short (n), leveraging retry loops (k), strictly enforcing deterministic harnesses wherever possible (q = 1.0), and utilizing HITL for high-risk evaluations, you can engineer a system that naturally self-corrects and vastly outperforms traditional linear RAG.