Two Machines, One Referee
A few weeks ago I had my own websites reviewed. Not by me, and not by a friend being polite. I had one model walk the whole estate as a cold buyer, and a second one read the same estate the way a search engine reads it. They worked separately. Neither saw the other’s report until both were finished. Then a third one read both and had to decide.
That shape has a name here. We call it a triad, and it is how most of the real work gets done in this shop. Two specialists with deliberately different strengths, and one referee whose job is to reconcile them rather than add a third opinion.
Where it came from, which is not where you would guess
It did not come off a whiteboard. It came from me running two AI instances at once and relaying between them by hand, and noticing that the pair produced sharper work than either one alone — but only when somebody sat in the middle and reconciled them. I was that somebody. The theory came out of the practice, not the other way around.
That line is the whole design. The value is not in the agreement. It is in the tension, and in having somebody whose actual job is to resolve it.
The rules all three work to
This is the part people skip, and it is the part that makes it work instead of making it noise.
- Everybody puts a number on it. Nought to a hundred, on every claim, with a reason for that number and not a higher one.
- Every reason has to say whether it was checked or guessed. No blending the two.
- If a reviewer has thin information, it says so and gives a low number. It does not go quiet. Silence is not an answer.
- When they disagree, the one who can show you the thing wins. Not the louder one. Not the more confident one. Not the one who spoke last.
- When neither can show you anything, it comes to me as one decision in two lines. Not two essays.
- Feeling sure is not a check. A page loading out of cache is not a check either.
What came back
The first review was uncomfortable in the way useful things are. One company, real products, real prices, a checkout that works — scattered across about twenty brand names and two websites that hand people back and forth. Three of my own front doors opened onto a staff login screen. A stranger could not tell in ten seconds what we sell or which door was theirs.
Fine. Fixable. We started fixing it.
Then I ran a second pair, and one of them broke
This is the actual reason for this post.
I wanted a genuine outsider in the second chair, so I pointed a different model at the same estate. It came back with a confidence of ninety-five and a disaster at the top of the list: my internal sales pipeline, it said, was sitting on the open web. Client names. Deal values. Bid numbers that nobody outside this house has any business seeing.
If that had been true, it was the worst thing anyone had ever found here.
It was not true. I had the referee check that one claim on its own, signed out, from an address a search engine is actually allowed to crawl. The page answers with a sign-in screen that says Locked · nothing here is public. The gate was doing exactly its job and always had been.
I had handed it my session. So it saw what I see — the staff side — and reported it as what everybody sees. Every observation it made was real. It was standing on the wrong side of a door and did not know there was a door.
Why that was the most valuable hour of the week
Because it was caught before it cost anything. Nobody tore out a working security gate to fix a problem that did not exist. Nobody sent a panicked message to a client. The rule did its job: the reviewer that could show you the thing won, and in that case neither reviewer could — so the referee went and looked itself.
It threw out four more things from that same report while it was at it. The model had recommended taking down four live products with printed prices, on the grounds that they looked unmaintained. It had recommended folding our family store into our business store, which would undo a split we made on purpose. Good instincts, no evidence, wrong calls.
And it was right about something important, which is the other half of why you keep two. It ran ten cold searches before it was allowed to look at the sites at all. We showed up in two of them — once as a state business filing, once as a bare link with nothing attached. Its conclusion was that a buyer would decide this company does not exist. The first reviewer had reached the same verdict by a completely different route. Two methods, one answer. That one I believe.
What it is now
It is a room in our own operating system. Pick the subject, pick two angles that will actually disagree, pick the referee, and the room enforces the rules rather than trusting anyone to remember them. It will not let you load the same angle into both chairs. It will not accept a reviewer that returns no number. It works out the confidence band itself from whether they agreed and who held the evidence, so nobody can talk their way to a higher one.
The angles change with the job. A cold buyer and a search engine for a website. The picture and the sound for a film. The scoring rubric and the person who has to do the work for a bid. The referee changes too.
The honest part
I am not going to pretend a machine reviewed my company and handed me the truth. One of them handed me a five-alarm fire that was not burning. What I will say is that the method caught it, in about four minutes, because the method does not care how confident anybody sounded.
That is worth more to me than a review that agrees with me.
I ran an adversarial review of my own web estate using a two-specialist, one-pacemaker topology. Both specialists ran sealed. The pacemaker reconciled them under a fixed output contract. One of the specialists returned a false blocker at confidence ninety-five, and the contract is the only reason it cost nothing.
The topology, and where it came from
The pattern is two analyst nodes homed on different angles of the same subject, deliberately configured so they can disagree, and one pacemaker node that reconciles, applies the decision rule and communicates. The pacemaker is a decision-maker, not a third analyst, and it is barred from substituting its own audit for the specialists’ work or from breaking a tie on its own conviction.
It came out of practice, not design. I was running two model instances concurrently and relaying between them manually, and the pair produced sharper output than either alone — conditional on a third party reconciling them. I was that third party. The topology is that loop, formalized.
Configuration is five levers, not weights: role, knowledge, tools, output contract, feedback loop. Nothing here is trained. Lever five is the only thing that makes it improve.
The output contract
This is the single most important technical detail, and the reason the pattern composes across layers.
- Specialist returns: stance, confidence nought to one hundred, and two to three reasons each tagged
[data]or[assumption]. Thin evidence returns low conviction, never silence. - Pacemaker returns: the concordance stated explicitly, concordance mapped to a confidence band, the decision threshold applied, and single-source findings that affect the plan spot-checked live and labeled verified.
- Concordance mapping: both agree, high. Split, lower, with the split stated. One specialist out of lens or absent, lower and flagged single-source.
- Split resolution: the specialist that can point at live evidence wins. Not the higher-confidence node, not the more recent, not the one that spoke last.
- Deadlock: neither side evidenced, it escalates as one founder decision, both positions in two lines.
- Verification is the named check. A cached two hundred is not verification.
Run one
Specialist A ran a cold-visitor conversion lens. Specialist B ran a crawler, entity and SEO lens with source access. They returned stances of No and Barely, at confidence seventy-four and eighty-two.
The concordant finding: a single legitimate entity with priced SKUs and a working checkout, fragmented across roughly twenty public brand names, two cross-redirecting origins, inconsistent Organization schema, and three public entry points resolving to a noindexed auth wall. The highest-leverage item was entity resolution — one Organization node with a stable @id, a full sameAs array, consistent NAP across every property, and a verified business profile. That is the E-E-A-T gap both human evaluators and ranking systems penalize first.
Run two, and the contamination
I swapped in a different model for the second chair to get genuine exogeneity, and pointed it at the same estate.
It returned confidence ninety-five and a P0 blocker: the internal pipeline board, publicly reachable, with client names and deal values quoted verbatim.
The pacemaker did not accept it. It re-tested that single claim unauthenticated, against the host where the same deployment is served under robots: Allow: / rather than the origin that carries Disallow: / — because a robots-blocked fetch proves nothing about public reachability. The path returned the sign-in screen. The edge middleware gate was behaving correctly and always had been.
Every observation it made was real. It was reading a privileged surface and reporting it as the public one, and nothing in its brief required it to assert its own authentication state before walking the estate. That is a lever-three fault — tool scope — which invalidated lever one, the role.
It also failed the contract on its own terms: its disclosure section describes those boards as internal while its top finding describes them as public. A node that contradicts itself between findings and disclosures has not met the contract, which alone is grounds to hold the section.
What the pacemaker kept, and what it discarded
Kept, and now two-source. The cold-search phase ran before the model touched any property, so it carries no session dependency. Ten queries; the entity surfaced twice — once as a state registry record, once as an orphan deep link with no entity attached. Its conclusion matched specialist A’s by a completely independent method. Also kept: store.html ships no server-rendered product entities. The catalog is built client-side from a cross-origin source, so the served HTML carries none of the trade or SKU terms. Re-verified unauthenticated.
Discarded. The pipeline exposure, falsified. The theme picker, which exists inside the authenticated hub and not on the public store. A recommendation to take down four live, priced product hosts as unmaintained. A recommendation to collapse the B2C catalog into the B2B origin, which would undo a deliberate audience split and violate a standing no-cross-store rule.
One new finding, surfaced incidentally during verification. A consumer product page ships with no price rendered, against a published pricing rule that every product carries a printed price. Neither specialist caught it. The pacemaker did, while falsifying something else.
The implementation
It is now a room in our operating system rather than a procedure somebody has to remember. A rig binds two lens configurations and a pacemaker configuration to a subject, each carrying its own copy of the five levers so editing a preset never rewrites a rig that already shipped.
Two invariants are enforced server-side rather than trusted:
- Loading the same angle into both chairs is rejected at write time. That is the one-analyst-twice failure, and it is the only way the topology silently stops working.
- A specialist payload with no confidence value is rejected. The contract permits low conviction; it does not permit silence.
The confidence band is computed from the concordance flag and the evidence holder, not typed by the operator, so nobody can argue a run into a higher band. Escalation fires when a split resolves with no evidence on either side, or when the weaker specialist falls below the rig’s threshold — the one value that is per-subject. Everything else is house law.
Lens and pacemaker shelves are data, not hardcoded. Swapping angles is a configuration change: cold visitor against crawler for a web estate, picture against sound for a cut, scoring rubric against delivery capacity for a solicitation.
The honest part
A model reviewed my company and handed me a P0 that did not exist, at a confidence higher than either node that was right. What the method bought was four minutes of falsification before anybody acted on it — because the resolution rule does not weight conviction, and the pacemaker is required to verify single-source findings that move the plan.
That is worth more than a review that agrees with me.
