The importance of documenting decision-making processes when configuring AI
Ombudsman and ADR schemes have rules setting out their scope, but these rarely specify how caseworkers apply them in practice. Getting schemes to agree that in writing, and keeping the record current, is key to using AI well.
What the rules don’t say
At the start of a configuration project I ask the scheme for everything it has in writing: terms of reference, scheme rules, the boilerplate paragraphs caseworkers reuse in decisions. This material sets out what the scheme can decide but often says much less about how an experienced caseworker gets from a complaint to an outcome.
Many schemes will still take on a complaint brought too late, consider a case without a proper deadlock letter, or progress one where the requested remedy is not something they can offer. The caseworkers who have handled hundreds of these cases work to a more detailed test than the rules state: how late is too late, what counts as a deadlock letter, when an unavailable remedy is worth progressing anyway. None of that is written down. It lives in the heads of the people who have been there longest.
The same rule, applied two ways
Often once we start interrogating what the rules mean in practice, it becomes clear that they can be applied differently from one caseworker to the next. One caseworker treats a missing call recording as a point against the business. Another sees it as neutral unless the customer says the business promised something specific. Both readings may fit the scheme rules but can lead to very different outcomes.
Asked one at a time, two senior caseworkers tend to agree on most of the rules and part company when applying them to borderline cases. Schemes rarely know this about themselves before a configuration project, because day-to-day casework rarely puts the two approaches side by side.
The platform can only apply a test that has been documented. Schemes are often tempted to keep the vague wording, but that takes the inconsistency between caseworkers and feeds it straight into the platform. So, most of my configuration work goes on getting guidance drafted with enough specificity that the rules will be applied consistently across all cases.
Turning scheme rules into guidance the platform can apply
Drafting guidance involves having the most experienced decision-makers in the room, talking through their hardest cases and agreeing how to decide them.
Schemes write their rules for people who already know how to apply them. “Consider whether the business treated the customer fairly” works for a caseworker with ten years’ experience, but as an assessment question it gives no specificity for the platform to apply. We break it down into questions that point at evidence: “Did the business explain the cancellation charge before the customer signed, and where in the file is that explanation?” Each answer cites the document it relies on, so the reviewer can open it and disagree.
For the missing call recording, the scheme agrees the approach it wants the platform to take or, at least, the question it wants caseworkers to be answering. That reduces inconsistency, and where a case does sit in a grey area, brings it into the open.
Testing on historic cases
Before anyone uses the configured questions on a new complaint, we test them on cases with known outcomes. We ask for a mix of simple and complicated ones, so we can be sure the platform handles the simple cases easily and stress test it on the rest.
The awkward files are where ambiguous questions show up. If the structured analysis on a closed case leans away from the decision the scheme reached, we work with the scheme to find out why. Sometimes the guidance is missing a test. Sometimes the original decision was the exception to the rule, and the scheme has to say how that scenario should be handled in future.
Owning the guidance once it is in use
Someone in the organisation must own the guidance: a named person with the authority to sign off changes and settle arguments between teams. We can show a scheme where its guidance is unclear, but deciding what its rules mean is the scheme’s job.
Each change gets a version number, and each case records which version it went through, because the rules move. If there is a regulatory update or the scheme’s remit changes, the guidance needs updating in time for the day it takes effect.
In Ctrl AI the reviewer confirms or amends each assessment before anyone relies on it, and the platform records every review and override. In the first weeks the owner can learn a lot from that record: if reviewers keep correcting the answer to the same question, the fault is often in an unclear question or a gap in the guidance. The owner should agree the fix, put it into a new version and tell the reviewers who raised it what changed. Otherwise, each reviewer keeps a private correction in their head, and the scheme starts collecting unwritten tests again.
