Skip to content
How we prove it

It has to prove itself before it touches your business.

Most AI is sold on a demo. A demo is the vendor's best day. What matters is whether it can reproduce your work, on your records.

Every automation is given your own history and asked to produce the answer you already got.
If it does not match, it is not allowed to run. Not flagged for review, not run with a warning. Blocked.
How it earns the right to run

Match your history, or do nothing.

How an automation earns the right to run Your records become a test set with the answers you reached. The automation must reproduce those answers. If it does not match, it is blocked and produces nothing. If it matches, it runs as drafts, a named person approves each action, every action is written to an audit trail checked daily, and the test is re-run every night. NO MATCH MATCHES Your recordswork you already did A test set is builtwith your real answers The automation triesfrom scratch THE TESTDoes it match?your answer, not a demo BLOCKEDProduces nothingno warning, no override STEP 1Runs as draftsevidence attached STEP 2A person approvesper item, for money STEP 3On the recordchecked daily STEP 4Re-tested nightlya failure raises an alert

Pick a scenario, or any box.

Scroll the diagram sideways to see the whole path.

Whatever your business does

The test is always your own work.

Invoices and billsPayrollApprovalsEmailsCustomer calls and messagesLeases

The same rule on every workflow: it reproduces your own past before it runs, a person approves what it drafts, and the log keeps every decision. Eight industry templates to start from →

The six controls

What stands between an AI and your books

1

Tested against your history, not a demo

Your real invoices, approvals and emails, with the answers you actually reached. The automation has to reproduce them from scratch.

53 automations tested this way in our own family business11+ years of consecutive monthly payroll approvals
The detail

Before any automation goes live we build a test set from the work your business has already done: real invoices, real approvals, real emails, with the answers you actually arrived at. The automation is then asked to reproduce those answers from scratch.

In our own family's business this is 53 automations tested this way. The test set is built from that company's records, not a sample dataset, because an automation that works on somebody else's paperwork tells you nothing about yours.

The clearest example is payroll. Our family's manufacturing business has approved its monthly payroll by email every month since March 2015, more than eleven years without a gap, and every month since April 2018 has been pre-audited by an outside firm before payment. That is the record the payroll engine is tested against, and the test set grows with it every month.

2

Failing the test blocks it. There is no override

No pass, no output. A broken automation goes silent rather than confidently wrong.

The detail

Each automation refuses to produce output unless its test has just passed. It is not a warning in a log that somebody is supposed to read. The failing path simply produces nothing, so a broken automation is silent rather than confidently wrong.

Confidently wrong is the expensive failure. An automation that stops is an inconvenience; one that keeps going with a bad number is a payment you have to claw back.

3

A person approves every action that leaves the building

Nothing is paid, posted or sent because the machine decided it was right. Where money is involved, approval is per item.

The detail

Nothing is paid, posted, or sent because the machine decided it was right. Work arrives as a draft with the evidence attached, and a named person approves it. Authority is not delegated to the software, and the software cannot grant itself more.

Where money is involved the approval is per item: permission to post one document is never permission to post the next one. Closed accounting periods are refused outright, whatever anyone approves.

4

Everything that happened is on the record

Every action, approval and rejection is written down, and the trail is checked daily for tampering.

The detail

Every action, every approval, every rejection is written to an audit trail that is checked daily for tampering, with its own record kept off the database it describes. If the history were ever edited, the check fails.

The point is not that we expect tampering. It is that "trust us, the log is accurate" is not evidence, and an audit trail nobody verifies is decoration.

5

Re-checked every night, not just at launch

Every automation is re-run against its test set nightly. A failure raises an alert rather than waiting to be noticed.

The detail

Your business changes. A test that passed in March is a statement about March. So every automation is re-run against its test set nightly, the result is published to a dashboard, and a failure raises an alert rather than waiting to be noticed.

We learned this the hard way: a component of our own system failed silently for six weeks because nothing was watching the one signal that would have caught it on day one. Now something is.

6

We break our own safeguards on purpose

A test that cannot fail is worse than no test. We remove each safeguard deliberately and confirm the tests notice.

The detail

A test that cannot fail is worse than no test, because it manufactures confidence nobody checked. So we remove each safeguard deliberately and confirm the tests notice. If the tests still pass with a protection deleted, the tests were never testing it.

This has caught real gaps in our own work more than once, including two safety checks that appeared to be covered and were not, because a different check was quietly refusing first and hiding them.

Being straight with you

What we have not earned yet

Here is the other half, because you will find it out anyway.

  • , We have one deep deployment, not a hundred logos. Everything on this site runs on our own family's manufacturing business first. A real, complex company rather than a pilot, but one company.
  • , The tests prove reproduction, not judgment. Matching your past decisions shows the mechanics are right, not that it would make a good call on something new. That is why a person still approves every action.
  • , Automations that are only instructions have no test. Where a task is a written procedure rather than a calculation, there is nothing to reproduce. We say which is which.
Checkable from outside

Proof you can check without us

Everything above is verified on a client’s own records, so you take our word for it until you are one. This one you can check yourself, today. One public example: our lease-reading engine, run on commercial leases filed with the SEC.

148public leases: 52 office and laboratory, 96 retail
2,812attempts to read a term
960came back with a value, a page and the wording
1,663refused, with a specific reason
189came back empty. The number we work on
615automated tests run before anything ships

93 per cent is not an accuracy figure, and we will not present it as one.It says the engine rarely goes quiet on you. Whether a value is right is settled by opening the document at the cited page, which is why every value carries one.

The test set is public

Commercial leases are filed with the SEC as exhibits to public company filings. We pull 148 of them, 52 office and laboratory and 96 retail, and run the engine over every one. Anyone can assemble the same corpus from the same place and get the same documents. We are not asking you to trust a number from a slide.

Two corpora rather than one because the first was 37 office leases out of 52, and not one of them mentioned co-tenancy. An engine tuned on office space would have looked excellent and failed on a shopping centre.

We count what it refused, not just what it read

Across those 148 leases the engine makes 2,812 attempts to read a term. 960 come back with a value, a page, a clause and the wording it was taken from. 1,663 come back refused with a specific reason: the term is stated as a length from a commencement that has not happened, the escalation is CPI, the exhibit is a chain of amendments with no single operative figure.

189 come back empty with nothing useful said. That is the number we work on, and it is the number nobody publishes.

Why 93 per cent is not accuracy

93 per cent of those reads end in either a cited value or an explanation. That says the engine rarely goes quiet on you. It does not say the values are right. Whether a value is right is settled by opening the lease at the cited page, which is the entire reason every value carries one.

Every vendor in this category claims 95 to 99 per cent accuracy. None of those numbers has been independently measured, ours included. We would rather give you a number that means something narrower and is true.

The tests include the mistakes we made

615 automated tests run before anything ships. A large share of them exist because the engine got something wrong on a real lease and we pinned it so it cannot come back: a rent schedule whose first money column was monthly, so the annual rent came out twelve times too low; an expiry read from an amendment that had already been superseded; an exclusive-use clause that was actually a parking space.

Each of those produced a confident, well-formed, wrong answer with a citation to a real page. That is the failure worth engineering against, and a test suite that only proves the happy path would not have caught one of them.

Have leases? Bring one to the first call

Bring one of your own leases to the first call and we will run it in front of you. Check each citation against the document, then ask it something the lease does not clearly answer and watch it decline. If it declines on things your lease plainly states, you will see that too.

Ask us to prove it on your numbers.

Bring a month you have already closed. We will show you where we match what you did, and where we do not.

Start free → Talk to us