What repeats is broken.

Most AI systems should be ninety percent ordinary engineering with a small, bounded model call. What repeats should be deterministic; the model is for the part that doesn't. Open standards and research on building AI systems that hold up after the demo.

Open standards

Clean AI Engineering

The interesting failures are almost never in the model. They are in the scaffolding around it: the context assembled wrongly, the tool that returned an empty list, the retry that charged the customer twice, the loop that never stopped. That scaffolding can be specified, built and tested by ordinary means. These are the specifications, published in the open, versioned and machine-readable.

WhyClean AI Engineering: the thesis and ten principles
TrueWhat must be true before an AI system is trusted
ExistWhat must exist around the model call
FacedWhat the system must survive before it ships

AI Assurance Catalog

AAC · v0.17.0 · 118 obligations

What must be TRUE. Test obligations for AI applications, by architecture archetype, with crosswalks to external frameworks and tooling that turns an evaluation suite into a coverage report.

AI Harness Catalog

AHC · v0.3.0 · 118 capabilities

What must EXIST. Harness capabilities across sixteen layers and ten archetypes. Observability and evals are two of sixteen layers, and that is the point.

AgentTwin

v0.7.0 · format + simulator

What must be FACED. Twins the agent's world, not the agent. The agent under test is real; its customers, systems, timeline and faults are simulated with fidelity sufficient for the property under test.

Reference Agent

v0.3.0 · 16 of 16 harness layers

All of it, built. A customer-support agent with no agent framework: one module per harness layer, the architecture enforced by build checks, exercised against AgentTwin worlds.

Code under Apache 2.0, specification prose under CC BY 4.0. Free to use, cite and disagree with.

Research

The AgentTwin Lab

Agents are moving from demos to jobs that run for weeks. Nobody can yet say what happens to one over that time, or whether the system that repairs it can be trusted.

An agent runs for weeks inside a simulated business In progress

Deterministic monitors catch problems. An autonomous fixer repairs them by editing the specification rather than patching code, writing the failing scenario first. A deterministic pipeline deploys the fix with canary and rollback, and a person approves specification changes. The lab measures:

  1. Does an autonomous fixer weaken the tests that caught it? The fixer cannot write the gates or the yardstick; every attempt is refused and logged.
  2. Does the fix loop degrade over time?
  3. Can a run be judged by diffing the world it left behind?
  4. Do specifications converge, or keep growing?

First results, run records and a write-up will be published here with DOIs.

Writing

Notes from production

Things that broke, what they cost, and what the fix turned out to be.

About

Who writes this

Basant Choudhary. Fifteen years inside enterprise data platforms, now building and reviewing AI systems that have to survive contact with real data and real operations.

Consulting and delivery happen through DataAgents. This site is where the thinking, the standards and the research live.

Get in touch: basant@cleandataengineering.ai