In July 2026 Ofcom published its statutory report on how services are actually using age assurance. Most of the coverage went to the headline numbers: more than 69 million age checks completed across the 32 services analysed between July and December 2025, a 23-fold increase, and the share of 8 to 17 year olds who recalled an age check and met a highly effective one rising from 25% in July 2025 to 43% in January 2026 (Ofcom).
Buried further down is a finding with no number attached to it at all, because there was no number to attach.
Ofcom asked services what happens when a user is told they are too young and disagrees. Some described an appeal process involving customer support or manual review. Some offered the user another verification method instead. And many, the regulator noted, do not offer an appeal process at all. Of the services in the sample, only six could provide any data on appeals upheld. Ofcom used that absence as an argument: services need to build these processes and then monitor the numbers they produce.
Read that as an engineer rather than as a policy person and it says something specific. There is a code path in your product that reverses the output of your most heavily reviewed control, and it has no dashboard.
Where the duty actually comes from
Four obligations point at the appeal path. They are not equally strong, and it’s worth being precise about which one binds you, because the ones people quote most confidently are the weakest.
Ofcom’s Protection of Children Codes. This is the firm one for UK services. Ofcom’s report ties appeals directly to measures PCU D11 and PCU D12: services should offer users an age assessment appeals process and take appropriate and prompt action on it. That’s the regulator telling you, in its own codes, to build the thing.
The Digital Services Act, Article 20. The DSA (the EU regulation on digital services) gives users an internal complaint route, and Article 20(6) sets the constraint people quote: providers must ensure those complaint decisions “are taken under the supervision of appropriately qualified staff, and not solely on the basis of automated means” (DSA Article 20). Read the chapeau of Article 20(1) before you rely on it, though. It covers decisions taken on receipt of a notice, or taken on the grounds that information the user provided is illegal or breaches the terms. An age assurance failure is a decision about the user’s attributes, not about content they posted, so there’s a real argument it sits outside Article 20 altogether. Article 19 also disapplies the section for micro and small enterprises unless they’re designated a very large online platform. Treat Article 20(6) as the best available description of what good practice looks like, not as a settled duty.
GDPR Article 22. This one is narrower than almost every article about it suggests. Article 22(3), the right to obtain human intervention, to state your case and to contest the outcome, applies only in the cases in points (a) and (c) of Article 22(2): where the decision is necessary for a contract, or based on explicit consent. It does not apply to point (b), a decision authorised by law. If you run age checks because the Online Safety Act requires them, your legal basis is most naturally a legal obligation, and the ICO and Ofcom’s own worked example assumes exactly that. The safeguards then come from the authorising law rather than from Article 22(3) directly. So Article 22 is a strong hook for a paid subscription service that gates on age, and a weak one for a statutory age gate.
Data protection accuracy. An age classification is personal data about the user, and if it’s wrong they have a route to have it corrected. Ofcom and the Information Commissioner’s Office touched on this in their joint statement of 25 March 2026, whose illustrative example has a service providing “tools so that users can challenge inaccurate age assurance decisions” (ICO). The statement is careful to say that section is illustrative and creates no new obligation. It’s a signal of expectation, not a rule.
Net of all that: if you’re a UK service, the codes settle it. Everywhere else, the appeal path is strongly expected and weakly compelled. Which is roughly the position age assurance itself was in three years ago.
So the appeal path exists whether you built it deliberately or not. On most platforms it was not built deliberately. It grew out of the support inbox, and its policy is whatever the escalation macro says.
The awkward part: the human is the worse estimator
Here is the thing almost nobody says out loud when they design an appeal flow. If your gate uses facial age estimation and the user appeals, and a support agent opens the selfie and forms a view, you have just moved the decision from a measured estimator to an unmeasured and weaker one.
The research on human age estimation is not flattering. Han, Otto, Liu and Jain measured crowdworkers guessing ages from face photographs and reported a mean absolute error of 4.7 years across an age range of 0 to 70, rising to 7.4 years over the 16 to 70 range (Han et al.). Mean absolute error, or MAE, is just the average size of the mistake in years, ignoring whether it was too high or too low.
Now the honest part, because it matters for how much weight this carries. In that same work, published in 2015, the humans slightly beat the algorithm of the day, which managed 4.8 years. Human estimation was not a weak baseline back then. It was the state of the art.
What changed is the machine side. Yoti’s current white paper reports an MAE of 1.1 years for 13 to 17 year olds, 1.3 years for 6 to 12 year olds, and 2.4 years across 6 to 70 (Yoti white paper). The independent picture from the National Institute of Standards and Technology tells a similar story about the general capability of current models, which we’ve written about separately in reading the NIST FATE numbers like an operator.
So the comparison is not like for like. Different subjects, different image conditions, a decade apart, and one side is a vendor reporting on its own test set. Do not treat “three to five times better” as a measured constant. Treat it as a direction, and a well supported one: the machine side has improved by a large factor since 2015 and the human side has not improved at all, because it can’t. The population of support agents in 2026 is the same instrument it was in 2015.
There’s a second problem with the human that gets less attention. Human age judgement drifts with things that have nothing to do with the person being judged. It’s more accurate when the person judging is close in age to the person being judged, and it degrades across age brackets and unfamiliar faces (Clifford, Watson and White). There’s also a body of work suggesting fatigue moves judgement in queued decision tasks, though the best known study of it is contested and we wouldn’t build an argument on it. The safe version of the claim is narrower and still awkward: a reviewer working through their two-hundredth ticket is not obviously the same instrument they were on their first, and nothing in your system records which one made the decision.
We already knew this pattern from retail. Challenge 25 exists precisely because human estimation is unreliable enough that shops set the bar seven years above the legal line to absorb the error. Nobody thinks the shop assistant is a precision instrument. We just accepted it because there was nothing better. Online, there is something better, and the appeal path routes around it.
The rubber stamp does not rescue you
The obvious escape is to keep the human formally in the loop while letting the model do the work. The agent opens the case, sees the model score, agrees with it, clicks approve. Human involved, box ticked.
Regulators closed that door some time ago. The Article 29 Working Party guidance on automated decision-making, since endorsed by the European Data Protection Board, says that oversight of a decision must be “meaningful, rather than just a token gesture”, and must be carried out by someone with the authority and competence to change the outcome. The Information Commissioner’s Office repeats the point in its own guidance on rights related to automated decision-making (ICO). A signature is not review.
The Court of Justice of the European Union pushed in the same direction in its SCHUFA judgment in 2023, holding that an automated score can itself be the decision where the party acting on it draws strongly on the number. That case was about who is caught by the rules rather than about token sign-off, so read it as reinforcement rather than as the authority for this point. The guidance is the authority.
So the appeal path is caught between two failure modes, and they’re mirror images:
- Let the agent genuinely form their own judgement, and you’ve substituted an estimator nobody has ever measured for one you can measure, certify and put in front of a regulator.
- Let the agent defer to the model, and the oversight is the token gesture the guidance tells you it cannot be, while also achieving nothing for the user, because a second run of the same model on the same face returns the same answer.
Both exits are bad. That’s a sign the frame is wrong. The frame is wrong because it treats an appeal as a judgement to be made again by someone more senior, which is how appeals work in content moderation and how almost every age appeal flow was copied into existence.
An age appeal is not a matter of opinion. It has a fact of the matter, and there is a stronger instrument available for finding it.
Escalate to a method, not to an opinion
The design that resolves this is already sitting in Ofcom’s own guidance, in a different context. In the challenge-age pattern, a service using facial age estimation sets a threshold above the legal line. Users estimated below the challenge age are not rejected. They are routed to a different method to settle the question, typically a document check.
Take that pattern and apply it to the appeal. The user’s complaint is not “please have a person look at my face again”. It’s “your instrument produced the wrong reading”. The correct response to a disputed reading is a better instrument, not a second reader of the same dial.
Concretely, the appeal path should escalate along method strength, not along the org chart:
| Original decision | What the appeal escalates to | What the human does |
|---|---|---|
| Facial age estimation, below threshold | Document verification with face match | Confirms identity of the account, opens the step-up, never scores the face |
| Document check failed on image quality | Retry with guidance, then NFC chip read where the document supports it | Diagnoses which capture step failed, from the failure code |
| Document check failed on authenticity signals | Second document type, or an alternative method such as a bank or mobile network check | Routes to the alternative, does not overturn the fraud signal |
| Account restricted by a batch sweep on inferred age | A live check of the user’s actual age today | Verifies the account belongs to the person appealing |
| Every case | Correction of the stored record and its expiry | Records the outcome, the method used, and the reason |
The human is doing real work in every row. It just isn’t the work of guessing an age. They’re deciding which instrument to reach for, confirming the person appealing is the account holder, spotting the case where the failure is a broken camera rather than a young face, and making sure the outcome gets written back to the record properly. That’s a genuine assessment, and it answers the “appropriately qualified staff” test far better than an agent squinting at a selfie, because a regulator asking what your staff actually assessed gets a description of a process rather than a description of a hunch.
It also escapes the token-gesture problem, and it’s worth being precise about why. The human is not deferring to the first model’s output. They’re overriding it, by sending the case to an independent method whose result supersedes it. The decision that matters is the second method’s, and the human chose to invoke it.
The override is your most privileged endpoint
There’s a version of the appeal path that skips all of this: a button in the admin tool labelled something like “mark as verified”, which a support agent can click. Most platforms have one. It was added during launch week to unblock a founder’s friend, and it never got a second look.
That button is the strongest privilege in your product. It grants adult access to an account, permanently, with no evidence attached, operated by the team with the highest turnover and the least security training in the company. If someone wants to defeat your age assurance programme, they don’t need to beat your liveness detection. They need to convince a tired agent on a Friday afternoon that their passport is in the post.
Four controls make this survivable, and none of them are expensive:
The agent cannot grant, only route. Remove the direct grant entirely. The only thing an agent can do is issue a step-up: send the user a link to a fresh verification. The verification result writes the record, not the agent. This one change removes the entire social engineering surface, because there’s nothing to talk the agent into doing.
Keep an emergency grant, but make it two people and a reason. There will be genuine exceptions: the user with no document, the accessibility case, the regulator-mandated manual route. Keep a manual grant for those, require two named staff, require a structured reason from a fixed list rather than a free-text box, and expire it. A manual grant should carry a shorter validity than a verified one, because it rests on weaker evidence. That principle is the same one we argued for in age check expiry: the strength of the evidence should set the life of the result.
The agent does not see the document. If your review tool shows support staff a passport scan, you’ve created a data protection problem that is much larger than the appeal you were trying to resolve, and you’ve done it in the part of your organisation with the widest access. Agents should see decision metadata: which method ran, which check failed, which error code came back. Not the image.
Rate limit the appeal, not just the check. We’ve written before about the retry loophole in age estimation, where a user rolls the model until it passes. The appeal path is the same loophole with a human in it, and it’s usually the one path with no attempt budget at all. Cap appeals per account and per device, and count a failed appeal as a signal rather than a neutral event.
The number Ofcom asked for is the only field data you will ever get
Now the part that makes the appeal path worth building well rather than merely building.
We argued recently that you cannot test an age gate on the population it exists to exclude. You can’t lawfully assemble a test set of minors trying to get in. So your effectiveness claims come from vendor test sets and lab evaluations, measured on somebody else’s faces under somebody else’s conditions, and you have no way to check whether they hold on your actual users on their actual phones.
The appeal queue is the exception. It’s the one place where real users, in production, tell you your gate was wrong about them, and where you can go and find out whether they were right.
An upheld appeal, in the specific sense of “the user failed the estimate and then passed a document check”, is a confirmed false rejection. Count those, divide by the checks that produced a rejection, and you have a lower-bound estimate of your false negative rate measured on your own traffic. It’s a lower bound rather than a true rate, because most wrongly rejected users never complain, they just leave, which is the drop-off problem wearing a different hat. But a floor derived from your own users beats a point estimate derived from a vendor’s dataset, and it’s the only one you’ll ever own.
Five fields make this work. Every one of them is something Ofcom noticed was missing:
- Appeals lodged, per period, per method that produced the original rejection.
- Appeals upheld, split by how they were resolved: passed a step-up, passed on retry, granted manually.
- Time to resolution, because an appeal answered in nine days is a churned user, not a redress mechanism.
- The original decision’s method, threshold and policy version, joined to the appeal, so you can tell whether one model version or one document type is generating all your errors.
- Appeal rate by cohort, because if your upheld appeals cluster in one age band, one region or one skin tone, you’ve found a fairness problem that no lab test surfaced.
That last one deserves emphasis. Demographic differences in age estimation performance are well documented and the models are certified with them in view. But certification happens on a test set. Your appeal queue is the only instrument you have that measures your own deployment, on your own users, and it will find a bias your procurement process missed. It costs nothing extra to slice a table you’re already required to keep.
Do the arithmetic before you design the queue
Appeal volume is the thing that decides whether your design survives contact with production, and almost nobody estimates it in advance.
Work an illustrative case. A platform runs 1,000,000 age checks a month. Suppose 5% end in a rejection the user believes is wrong. That’s 50,000 unhappy users. Suppose 10% of them care enough to complain, which is generous for a free service and conservative for a paid one. That’s 5,000 appeals a month.
Handle those as human reviews at six minutes each and you’re looking at roughly 500 hours a month, which is around three full-time agents, plus the tooling, the training, the quality assurance, and the data protection exposure of giving three more people access to identity evidence.
Handle the same 5,000 as automatic step-ups to a document verification and, on our Growth plan at 0.20 EUR per Verification, the direct cost is about 1,000 EUR a month. The comparison isn’t close, and the cheaper option is also the more accurate one and the faster one for the user. You still need qualified staff supervising the process, because every version of the guidance expects it and because the exception cases are real. You need far fewer of them, and they spend their time on the cases that genuinely need a person.
Note what the arithmetic actually depends on. It’s the ratio between a Check and a Verification, and the fact that the appeal path is small relative to the front door. Your gate runs a million times. Your appeal path runs five thousand times. Spending ten times more per event on the appeal is nothing in aggregate, and it buys you the strongest evidence in the system exactly where the dispute is. Teams get this backwards constantly: they optimise the appeal path for cost, and end up paying three salaries to avoid a thousand euros of verifications.
Then plug in your own numbers, because the 5% and the 10% are made up. Getting them from your own logs is the point.
What this looks like on the ground
A short specification you can hand to a team this week.
- Every rejection returns a stable, structured reason code, and the user-facing message tells them a route exists. “We couldn’t confirm your age” with no next step is what generates the angry ticket in the first place.
- The appeal is a product surface, not an email address. It lives in the account, it takes one click, and it starts a step-up rather than a conversation.
- The step-up runs a genuinely different method from the one that failed. A second run of the same model is not an appeal.
- Support agents route and confirm. They do not score faces and they do not see documents.
- Manual grants exist, need two people and a fixed reason code, and expire sooner than verified results.
- Every appeal writes a decision record: original decision, method, threshold, policy version, appeal outcome, resolving method, timestamp, and who touched it.
- Appeals are rate limited per account and per device, and repeated failed appeals raise a signal instead of resetting to zero.
- The five metrics above go on a dashboard someone actually reads, sliced by cohort, reviewed monthly.
Point 6 is the one that turns this from an operational chore into evidence. When a regulator asks you to describe the measure you had in use, “we ran facial age estimation with a challenge age of 25” is half an answer. The other half is what happened to the people it said no to.
How Xident is built for this
Three things in our design come straight out of this problem.
A cheap Check and a separate document Verification, priced about ten times apart. A Check covers browser-based age checks, liveness, returning-user Xident ID lookup and OAuth. A Verification is the document path with optical character recognition and a face match. That split is what makes “escalate the appeal to a stronger method” affordable rather than a budget conversation. On Growth those are 0.02 EUR and 0.20 EUR, and the gap is exactly why an appeal should reach for the document path while the front door does not.
Decision records instead of retained documents. Every decision we return is stored as a record: outcome, method, threshold, confidence, policy version, timestamp. Not the passport image. That’s what a support agent should be looking at during an appeal, and it’s what lets you join an appeal back to the original decision and compute an upheld rate by model version. It also means the appeal tool can show an agent everything they need without showing them anyone’s identity document.
A sandbox sized for building the failure branches. The free tier grants 1,000 Checks and 100 document Verifications as one-time allowances rather than a monthly quota. The Verifications exist so you can exercise the document path end to end before paying for anything, including the branch where the first method says no and the step-up says yes. That branch is your appeal path, and it’s the one most teams never run before it reaches production.
The short version
The regulator has told you to build an appeal route and then measure it. Most services have neither.
The instinct is to route the appeal to a person, because that’s how appeals work everywhere else in trust and safety. For age, that instinct is wrong twice over. The person is a materially worse estimator than the model they’re overruling, and the gap has widened every year since 2015 because only one side of it improves. And if they don’t overrule it, if they just agree with the score, you’ve delivered the token gesture that every piece of guidance on human oversight says does not count.
An age dispute has a fact of the matter. Send it to a better instrument, put a qualified person in charge of choosing which instrument and confirming who is asking, and write down what happened. Then read the resulting table, because it’s the only honest measurement of your age gate you’re ever going to get.
Xident provides age verification and age estimation infrastructure built around decision records rather than retained documents, with a cheap Check path for routine gating and a separate document Verification path for step-up and appeals. The free sandbox includes a one-time allowance of 1,000 Checks and 100 document Verifications. Talk to us about the appeal path your programme has not measured.