The Method

Two Machines, One Referee

I pointed two AI models at my own company and did not let either one see the other’s work. One of them got it badly wrong. That is the part worth reading.

A few weeks ago I had my own websites reviewed. Not by me, and not by a friend being polite. I had one model walk the whole estate as a cold buyer, and a second one read the same estate the way a search engine reads it. They worked separately. Neither saw the other’s report until both were finished. Then a third one read both and had to decide.

That shape has a name here. We call it a triad, and it is how most of the real work gets done in this shop. Two specialists with deliberately different strengths, and one referee whose job is to reconcile them rather than add a third opinion.

Where it came from, which is not where you would guess

It did not come off a whiteboard. It came from me running two AI instances at once and relaying between them by hand, and noticing that the pair produced sharper work than either one alone — but only when somebody sat in the middle and reconciled them. I was that somebody. The theory came out of the practice, not the other way around.

If your two reviewers always agree, you did not build two reviewers. You built one, twice.

That line is the whole design. The value is not in the agreement. It is in the tension, and in having somebody whose actual job is to resolve it.

The rules all three work to

This is the part people skip, and it is the part that makes it work instead of making it noise.

  • Everybody puts a number on it. Nought to a hundred, on every claim, with a reason for that number and not a higher one.
  • Every reason has to say whether it was checked or guessed. No blending the two.
  • If a reviewer has thin information, it says so and gives a low number. It does not go quiet. Silence is not an answer.
  • When they disagree, the one who can show you the thing wins. Not the louder one. Not the more confident one. Not the one who spoke last.
  • When neither can show you anything, it comes to me as one decision in two lines. Not two essays.
  • Feeling sure is not a check. A page loading out of cache is not a check either.

What came back

The first review was uncomfortable in the way useful things are. One company, real products, real prices, a checkout that works — scattered across about twenty brand names and two websites that hand people back and forth. Three of my own front doors opened onto a staff login screen. A stranger could not tell in ten seconds what we sell or which door was theirs.

Fine. Fixable. We started fixing it.

Then I ran a second pair, and one of them broke

This is the actual reason for this post.

I wanted a genuine outsider in the second chair, so I pointed a different model at the same estate. It came back with a confidence of ninety-five and a disaster at the top of the list: my internal sales pipeline, it said, was sitting on the open web. Client names. Deal values. Bid numbers that nobody outside this house has any business seeing.

If that had been true, it was the worst thing anyone had ever found here.

It was not true. I had the referee check that one claim on its own, signed out, from an address a search engine is actually allowed to crawl. The page answers with a sign-in screen that says Locked · nothing here is public. The gate was doing exactly its job and always had been.

The model was not lying. It was looking through my own signed-in browser, because I had shared the tab with it.

I had handed it my session. So it saw what I see — the staff side — and reported it as what everybody sees. Every observation it made was real. It was standing on the wrong side of a door and did not know there was a door.

Why that was the most valuable hour of the week

Because it was caught before it cost anything. Nobody tore out a working security gate to fix a problem that did not exist. Nobody sent a panicked message to a client. The rule did its job: the reviewer that could show you the thing won, and in that case neither reviewer could — so the referee went and looked itself.

It threw out four more things from that same report while it was at it. The model had recommended taking down four live products with printed prices, on the grounds that they looked unmaintained. It had recommended folding our family store into our business store, which would undo a split we made on purpose. Good instincts, no evidence, wrong calls.

And it was right about something important, which is the other half of why you keep two. It ran ten cold searches before it was allowed to look at the sites at all. We showed up in two of them — once as a state business filing, once as a bare link with nothing attached. Its conclusion was that a buyer would decide this company does not exist. The first reviewer had reached the same verdict by a completely different route. Two methods, one answer. That one I believe.

What it is now

It is a room in our own operating system. Pick the subject, pick two angles that will actually disagree, pick the referee, and the room enforces the rules rather than trusting anyone to remember them. It will not let you load the same angle into both chairs. It will not accept a reviewer that returns no number. It works out the confidence band itself from whether they agreed and who held the evidence, so nobody can talk their way to a higher one.

The angles change with the job. A cold buyer and a search engine for a website. The picture and the sound for a film. The scoring rubric and the person who has to do the work for a bid. The referee changes too.

The honest part

I am not going to pretend a machine reviewed my company and handed me the truth. One of them handed me a five-alarm fire that was not burning. What I will say is that the method caught it, in about four minutes, because the method does not care how confident anybody sounded.

That is worth more to me than a review that agrees with me.