TechKiln

Only from things I ran

[ about ]

A kiln shows you what the clay was hiding

You cannot tell a sound pot from a flawed one by looking at it. You fire it. The heat finds the air pocket you could not see and the piece cracks along it. Production does the same thing to software, and it is the only test that counts.

What this site is

I build multi-agent systems — orchestration, evaluation harnesses, guardrails, the plumbing that decides whether an agent is allowed to do a thing. This site is where I write down what broke.

There is a great deal of material on getting an agent to work. There is very little on the harder half: keeping it honest once it is running unattended, and knowing when it has stopped working. That half does not demo well. It is mostly about failures that look exactly like success, which is why it tends to go unwritten.

How I write these

Only from things I ran. If I did not hit it myself, it does not go up — there is already enough writing that is a summary of other writing. Where I am unsure, I say so in the post rather than rounding it up to a conclusion. Where I was wrong first and right later, the wrong version stays in, because the wrong version is usually the useful part.

What I keep running into

A component fails in the direction that looks fine.

[ 01 ]

A gate that was never wired to anything still reports a verdict, and the verdict is pass.

[ 02 ]

A metric that scores well on a good result and also scores well on a bad one, so it can no longer distinguish between them.

[ 03 ]

A forecast that quietly stopped recording its uncertainty interval — leaving the point estimate, which is the one number that cannot be scored.

[ 04 ]

A log of what a tool recommended, read later as a record of what was actually done.

None of these announce themselves. Each one was found by accident, or by deliberately trying to make the check fail and noticing that it wouldn't.

Everything published is on the writing page. New pieces go out on theRSS feed.