Client

Self-initiated

Company

Docket

Project Type

AI review console · AML compliance

Year

2026

Docket

An AI-assisted source-of-funds review console for anti-money-laundering analysts at a GB-licensed remote casino. The model reads and cites. The human decides. The record keeps both, and that record is the product.

The docket. Every case names what was flagged, how many claims have an answer, and whose clock is running. You can batch the work, never the decision.


In six months, someone will read this file. You will not be there.

Docket is an AI-assisted source-of-funds review console for anti-money-laundering analysts at a GB-licensed remote casino. It records what was checked, not just what was decided.

This is a self-initiated project: no client, no users, no invented metrics. Every figure on this page was measured in the design file, and the parts that are not finished are listed at the end, because a case study is where overclaiming costs you.


The claim is the unit of work

One statement you can check against a source. Every claim carries the source, the verdict, and who asserted it: the model or the person.

No confidence percentage

The model reports the boundary of what it could read, not a score. Uncertainty sits on the claim, never on the case.

Five states for a wrong model

Override, a citation that will not resolve, conflicting sources, a changed rule, no model at all. The longest section of this page.

A record that answers why

Time-ordered, version-stamped, never re-scored. Dwell times survive as the evidence that verification actually happened.

The lane with criminal liability

Tipping off keys off state, not banned phrases. Two statutory clocks, counted differently, shown separately.

Two axes, no variants

Light and dark from one token collection, touch density from one variable. 22 screens drawn in both modes, no component gained a variant.


What Docket is

An analyst at a gambling operator picks up a flagged case: a player has deposited money whose origin has to be examined before it moves. The model reads the documents and drafts claims, each one citing the exact line it came from. The analyst answers the claims, decides the case, and the file that survives is a record of what was actually checked. Docket is the queue, the review view, and the sealed record, in light and dark, desktop and mobile, with the failure states designed first.


My Role

Product design, end to end: problem framing, research, competitive teardown, positioning, object model, voice system, brand, UI, states, accessibility audit and the developer handoff spec. Solo, over five days in August 2026. The domain is mine twice over: I designed money-sensitive flows for an iGaming operator and a payments marketplace before this.


DESIGNED

  • The docket, the review view, the sealed record

  • Five model-is-wrong states

  • The MLRO's escalations and reports screens

  • The request-evidence lane with the tipping-off gate

  • Queue edge states and responsive screens

SPECIFIED

  • A build-it-cold developer handoff

  • A 30-error matrix across 5 surfaces

  • A three-tier token contract, 2 modes

  • Authority and reporting invariants

  • A voice system with a testable verb law

VERIFIED

  • Three independent critic passes

  • Legal claims checked at legislation.gov.uk

  • 3,560 contrast checks, zero failures

  • 550 interactive targets, zero failures

  • A fourteen-item definition of done, run honestly


01. THE STAKE

The person who reads your work has no idea what you did

An analyst at a gambling operator picks up a case. A player has deposited €8,400 from an account that is not in their name. Before that money goes anywhere, the operator has to examine where it came from.

The analyst reads a bank statement, a payslip, a company registry entry. They form a view. They press Approve.

Months later a QA sampler, a supervisor or a regulator opens the same file. They can see who decided and when. They cannot see what was checked. So the analyst has two bad options: document everything, which is slow enough to fail the queue, or document lightly, which is indefensible when someone asks.

Both failures land on the same person.


FROM AN AML PRACTITIONER FORUM · THE THREAD HAD ZERO REPLIES

"We all know 'adequate mitigation' exists, but I can't be the only one who has noticed the huge variance in what people consider 'adequate.' Some analysts sign off on things others would never dream of approving."


02. THE GAP

Every audit trail records who decided. None of them records what was checked.

I walked three jobs end to end through Hummingbird, Unit21, Sumsub, WorkFusion and the status quo: flagged case to decision, request evidence and resume a week later, and reconstruct the file six months on.

They are good products, and several details are worth stealing outright. Hummingbird gets a flagged case to a decision in two screens, and pins case data beside the decision so the analyst never pays navigation cost to check a number they are about to write down. Unit21 hides the disposition button until the investigation checklist completes: gating by absence rather than by disabling, so the analyst cannot form the intention to skip. Sumsub blocks approval while any red flag is unresolved, forcing each contradiction to be dispositioned by a named human first.

What none of them has is an atom smaller than a container. Sumsub's unit is the document. Unit21's is the alert. Hummingbird's is the review task. WorkFusion's is the dossier. A per-assertion state, asserted, cited, verified, contradicted, unsupported, exists nowhere in the set.

That matters because a container cannot be verified. A document can only be attached. Only a statement can be checked.


THREE GAPS, RANKED BY CONFIDENCE

High, because the products' own data models preclude them: the claim as a first-class object with its own state; partial response as a designed state; and a record of which citations a human actually opened. Every vendor logs dispositions. None logs verification.


03. THE MOVE

Make the claim the unit of work

A claim is one statement you can check against a source. Not a document: a sentence.


"The €8,400 credit on 12 March originates from an account held by a third party, not the player."


Every claim carries three things that never separate: the source it was drawn from, the verdict a human gave it, and who asserted it, the model or the person. The record becomes a by-product of doing the work rather than a second job afterwards.

The model reads and cites. It does not recommend. There is no confidence percentage anywhere in the product and no pre-selected answer, because a pre-filled review is a pre-selected answer and an approval releases withheld money.


The review screen. Claims on the left, the document they were drawn from on the right, and the line the selected claim cites highlighted in the statement.


04. THE MECHANIC

Verification in under two seconds

Select a claim and the evidence pins to it: the document, the page, the line. Underneath the statement, one line that most products would never write.

That is the model reporting the boundary of what it could read, and then confirming it stayed inside it. Not a confidence score: a statement about where the machine stopped looking, which is the thing an analyst actually needs in order to know how much of the work is still theirs.

The four verdicts include Can't tell as a first-class answer. An analyst who cannot settle a claim from the source has told you something true, and a product that forces them to pick verified or contradicted has taught them to guess.


The claim, its citations, and the clause it exists to satisfy, in plain words, next to the decision it constrains.

The cited line, highlighted in the customer's own statement, with the OCR floor stated beneath it.


05. WHEN THE MODEL IS WRONG

The demo where the model is right is not the hard part

Every AI portfolio piece shows the happy path. It is invisible work, because nothing is at stake when the machine agrees with you.

Five screens exist for the case where it does not. They took the longest, and they are the reason the project exists.


1 · The model read it wrong

An override is an addition, not a replacement. The model's reading stays on the record beside the analyst's, because deleting it would destroy the only evidence that a human intervened, which is exactly what a supervisor or a regulator looks for.

Override and answer are one act. If you have replaced the claim with your own reading, you have already checked it.


The override, kept. The model's reading stays on the record beside the analyst's, with the reason and the time the source was open.


2 · The citation will not resolve

A citation that will not open is not an unsupported claim, and the console must not let one become the other. Unsupported means we looked and found nothing. Will-not-resolve means we could not look.

Collapsing them lets a plumbing failure be recorded as an analyst's finding: the most dangerous silent conversion in the product. The answer buttons are disabled. The duty to answer is not.


The file this claim was drawn from was replaced by the player. The answer buttons are disabled; the duty to answer is not.


3 · The sources conflict

The model cites both and adjudicates neither. Picking a winner here would be the model making the decision the product says it does not make, invisibly, inside a claim that reads as settled.

Resolving a conflict does not erase it. A customer's own documents disagreeing is a finding in itself, so the source not relied on stays on the record.


The payslip and the statement disagree. The model can cite both. It cannot tell you which is the truth, and it does not try.


4 · The rule changed

A claim carries the rule version it was drawn under, and never silently re-evaluates. Model output looks timeless; it is a reading of one rule set at one moment.

Which version governs is a policy decision, not the analyst's, so this is the one screen in the product where the console has an opinion about which button you should press, and says why.


The threshold moved after this claim was drawn. Nothing re-scores. Escalate is pre-selected, and the screen says why.


5 · The model is unavailable

The console degrades to manual; it does not stop. A vendor outage does not pause the obligation. Work done while the model is down is marked human-sourced, because a file whose claims were drawn by a person is a different artefact and that has to survive into the record.

Say what is missing, not that something failed. "Two documents have not been read" is actionable. "Service unavailable" is not.


No claims drawn since 14:12. The console degrades to manual, names the two unread documents, and marks new answers as read by you.


ONE RULE GENERATED ALL FIVE

The console never quietly resolves an ambiguity in its own favour. It does not turn a missing source into an unsupported claim, hot-swap a document under a reader, re-score claims under a new model version, or take the last write and drop the other. Every one of those is easier to build, and none of them can be defended afterwards.


06. THE RECORD

What it looks like six months later

Organised by what happened when, not by claim, because anyone reconstructing a decision reads time. Every model contribution is labelled with the version that made it and whether that version still runs.

What has changed since is shown as a delta and never applied: the rule set moved in August, and this decision is not re-scored under it. The dwell times survive, because they are the evidence that verification actually happened.

And one block that was not in the first version.


"Three claims did not close cleanly. They did not prevent approval, and this is why."


The first version of this record proved what was checked and never said why one contradicted claim, two unresolved ones and a document flagged for editing still added up to approval. That is the first question any regulatory review asks, and the file could not answer it. The gap came from an independent review, section 09 below, and closing it is the single biggest improvement in the project.

The product's own slogan was aiming one step short. What survives review is not what was checked. It is what was reasoned.


The sealed record. Recommended by the analyst, decided by the money laundering reporting officer, because the amount was above the analyst's authority limit. No money moves before that signature.


07. THE SUPERVISOR

Case-by-case handling hides the cause

The supervisor's question is not the analyst's. The analyst asks is this claim supported? The reporting officer asks do I trust this person's judgement enough to put my name beside it?

So the queue leads with the escalating analyst's own words rather than with case metadata. And above the queue sits the thing the screen exists for.


A PATTERN ACROSS THIS WEEK

"4 of 11 escalations this week name the same cause. The threshold moved from €8,000 to €5,000 on 14 August. Four analysts escalated deposits that fell between those two figures, because their own authority limits were written against the old number. That is one policy question, not four case decisions."


A supervisor tool that only shows the next item helps you process escalations. Showing the pattern helps you stop generating them, and case-by-case handling is structurally incapable of surfacing it.


The escalations queue leads with each analyst's own words. The pattern band above it turns four case decisions into one policy question.


08. THE PART THAT CARRIES CRIMINAL LIABILITY

Asking a customer for a document

When an analyst asks a player for evidence, they are writing to someone who may already be the subject of an internal report. Get the wording wrong and the offence is personal: tipping off under the Proceeds of Crime Act is committed by the analyst, not by the operator.

The first version of this screen blocked the phrase "anti-money-laundering review". It was specific, confidently written, and legally wrong. The offence requires a disclosure to have already been made. Telling a customer you are completing a review is lawful, and under the Money Laundering Regulations it is required.

The rebuilt version keys off state rather than strings. No report exists: the wording is fine, and the screen says why. A report exists: the block is correct, cites the right subsection, and routes the wording to the reporting officer.

A banned phrase is defeated by a synonym. A state is not.


No report on the case. The request is lawful, and the console explains the rule rather than blocking it.

A report exists. Now the offence bites, and the wording goes to the reporting officer first.


The same reasoning produced the lane that had been missing entirely: an internal disclosure to the reporting officer, their decision whether to report onward, and the two statutory clocks that follow. Seven working days of notice, then thirty-one calendar days of moratorium. They are counted differently, so they are shown separately and never as one total.


The internal disclosure. The claims travel with the report, and the screen states what freezes the moment you file.

The reporting officer's screen. Two statutory clocks, counted differently, shown separately, with what is frozen while they run.


09. THE REVIEW

What three independent reviewers found in my own work

Before publishing I ran three independent critic passes, with different lenses and no access to my decision log: a hiring manager, a UK money laundering reporting officer, and an engineer asked to build the thing cold from the spec.

Between them they found a misquoted regulation, a licence condition that does not exist, a legally wrong tipping-off block, an authority rule that my own hero screen contradicted, a case that was sealed in March and live in August, a person who was both a customer and an analyst, and a core data type used on every screen and defined nowhere.

I verified every legal claim against legislation.gov.uk and the Gambling Commission before changing anything, because a critic is an argument and not an authority. All three held.


WAS · THE CLAUSE THE PRODUCT QUOTES

"reg 33(1) — take adequate measures to establish the source of funds"

IS

reg 33(1) requires enhanced due diligence and enhanced ongoing monitoring. The source-of-funds wording belongs to the regulation covering politically exposed persons. Neither applies to this customer.

WAS · THE LICENSE CONDITION

LC 2.1.1

IS

No such condition exists. The anti-money-laundering condition is LC 12.1.1.

WAS · THE TIPPING-OFF BLOCK

Blocked routine wording.

IS

The offence needs a disclosure to have been made. Routine wording is lawful, and evidencing the request is required.



That is not a failure of the process. It is the process. What would have been a failure is running the review and publishing anyway.

The uncomfortable part: four of the engineer's seven blockers were introduced by me that same afternoon, while fixing other things, including one string that concatenated exactly the two things the rule I had just written forbids concatenating. A fix pass needs its own critic.


10. THE SYSTEM

Two axes, and no component variants

Three token tiers. Primitives are raw values and never theme. Semantic tokens carry meaning and hold no raw values at all: every one is an alias. Components bind only to semantic tokens, never to primitives, because a primitive reference cannot theme and cannot be reasoned about.

Light and dark are two modes of one collection, so in code they are one attribute on the root and one set of custom properties. Twenty-two screens are drawn in both. The first time dark mode was actually rendered it needed no manual fixes, which is the only real evidence that the architecture was doing its job rather than being asserted.

Touch density is a second collection holding a single variable. The control-height token aliases it, so switching the mode on a frame resized every button, nav item and field at once. No component gained a variant.


CSS PROPERTIES

:root { --control-height: 36px; } @media (pointer: coarse) { :root { --control-height: 44px; } }


Two axes, two independent switches, and neither one is a component's concern.


The master style frame. The palette has four sanctioned slots: the action you press, a measure's fill, a verdict badge, degraded system state. Nothing else carries colour.


The same review screen in Ink. One variable collection, two modes; the first dark render needed no manual fixes.


The docket on a phone. The table scrolls at full fidelity rather than losing columns.

The review view at touch density. One density variable resized every control; no component gained a variant.


The accessibility pass ran across all 44 frames with every exclusion named and counted: 31 components, 105 semantic tokens, 3,560 contrast checks and 550 interactive targets, zero failures.


44

FRAMES AUDITED

3,560

CONTRAST CHECKS

0

FAILURES


11. WHAT IS NOT DONE

Nine of fourteen

I ran the project's definition of done, fourteen items, pass or fail, each with its evidence, before writing this page rather than after. A case study is where overclaiming costs you, and the audit is what says which claims are supportable.

Nine pass. Three are open work: findability, empty and loading states for every list surface, and four remaining gaps in the handoff spec. Two cannot be closed at all, because nothing is built. The success metric is specified but not instrumented, and specified is not closed.

Those two are recorded as limitations with their closing conditions named. They are not counted as quiet passes.


The model reads. You decide. The record keeps both.

Self-initiated. No client, no users, no invented metrics. Every number on this page was measured in the file.


Other Projects

Let’s work together.

Complex product UX and design systems, shipped build‑ready and fast.

Follow Me

© 2026 Gmats.me