Skip to content

Methodology

What has to be true before we put a number on this site.

Most automation case studies lead with a large percentage and no way to check it. This page is the standard we hold our own write-ups to, published so that a reader can hold us to it as well.

BrisAI is a young studio and several of the systems on this site have not been formally evaluated yet. Those pages say so plainly, and carry no measured result rather than an approximate one.

We would rather publish three honest write-ups than five with a headline figure nobody can interrogate. When the evaluations described below have run, the numbers will appear here with the working attached.

Evidence labels

What each label commits us to.

Every case study carries exactly one of these. They are not synonyms, and none of them implies a client relationship unless it says so.

In-house evaluation
BrisAI built the system and tested it on a declared evaluation set or controlled process. No client outcome is implied.
Internal live system
BrisAI uses the system in its own operations and reports an observed period, volume and outcome.
Client case study
Results observed in a client setting, with the client named or clearly anonymised and permission recorded.

A demonstration run is not an evaluation

Where we have watched a system work end to end but not measured it against a held-out set, the page says “demonstration run” and shows what was observed with the caveat attached. It is evidence that a path works. It is not evidence of accuracy, coverage or speed, and we do not round it up into one.

Running an evaluation

Ten rules we follow before publishing a result.

  1. 01

    Decide what the number is for, first

    Before a test runs, we write down the business decision the result is meant to support. A metric with no decision behind it tends to be one that flatters the system.

  2. 02

    Freeze the run

    The workflow, prompts, model, settings and test data are fixed for a published run, and the versions are recorded alongside the result.

  3. 03

    Hold data back

    The examples used to build and tune a system are not the examples used to measure it. A development fixture set is not a benchmark, and we do not present one as if it were.

  4. 04

    Include the awkward cases

    Normal, ambiguous, malformed and adversarial inputs in realistic proportions. A test set made only of clean examples measures the wrong thing.

  5. 05

    Set pass and fail before looking

    What counts as correct, as a failure, and as acceptable-with-review is defined before results are seen.

  6. 06

    Compare against something fair

    The current manual process, a simple rules-based version, or whatever the system actually replaces — not against nothing.

  7. 07

    Count the review time

    Corrections, exceptions and failed runs are part of the outcome. A workflow that runs in 30 seconds saves nothing if a person spends two minutes fixing it.

  8. 08

    Repeat where output varies

    Where model variability could change the conclusion, the case is run more than once and the variation is reported.

  9. 09

    Save the evidence

    Machine-readable results and a written report, with a version and a date, kept so any published figure can be traced back to the run that produced it.

  10. 10

    Publish the limits

    Every result says what it does not cover. We do not generalise a finding beyond the sample it came from.

Claim standard

What travels with every published figure.

A percentage on its own is not a claim a reader can assess. Wherever we publish a number, these go with it.

  • The numerator and the denominator, not a rounded percentage on its own
  • The date of the run, or the observation period
  • Where the test data came from and what was included or excluded
  • What the result was compared against
  • The workflow, prompt and model versions in effect
  • Whether a person judged the output, or a script did
  • The limitations, next to the number rather than beneath the page

On money

Time and cost savings are modelled, not observed, until a real organisation has run the system and confirmed them. Where we show a figure of that kind, the assumptions behind it are visible and the wording is “could save at these assumptions” — never “saves”.