16 min read

Age Assurance Finally Has a Number, and a Court Didn't Set It

Meta's proposed settlement with state attorneys general puts hard false-positive ceilings on age assurance: 3% for 13-15 year olds and 10% for 16-17 year olds on commercially available methods, certified annually by an independent third-party tester and reviewed by an auditor with raw-data access for ten years. It is the first national accuracy floor for age checks in the US, and it was written into a private contract rather than a statute. Here is what the numbers actually demand, why the glide path quietly penalises buying over building, why they conflict with New York's per-year table, and why both schemes only constrain half the confusion matrix.

Editorial illustration on a deep slate-navy background: a horizontal age axis runs left to right with two error-tolerance bands stacked above it, a narrow emerald band over the younger segment and a wider amber band over the segment nearest the adult threshold. A single global cut line crosses both bands and clearly fails the narrow one. An audit seal glyph sits to the side, connected by a thin line to the bands. Abstract, no faces, no children, no brand marks, no readable text.

For three years the honest answer to “how accurate does an age check have to be” was that nobody had written it down. Ofcom asked for “highly effective.” The DSA asks for “proportionate.” Half the US states that mandate age verification specify a method and stop there. Vendors filled the gap with marketing numbers measured on their own test sets, and buyers had no basis to compare them.

That ended on 26 August 2026, and not in a courtroom or a legislature. It ended in a settlement agreement.

Meta’s proposed deal with a bipartisan coalition of state attorneys general, announced by California AG Rob Bonta alongside what his office puts at 51 offices, contains something no US statute except one has ever contained: a numeric ceiling on age-check error, with a compliance deadline, a testing requirement, and an auditor who gets to look at the raw data (California AG, ID Tech). The headline is the money, somewhere between $16.7 billion and $18 billion over ten years depending on which contingent portions you count. The part that will outlive the money is the spec sheet.

What the decree actually requires

The Age Verification Providers Association published the operative language, and it is worth reading precisely because the framing is easy to get backwards.

Commercially available age verification and estimation methods must achieve a false positive rate of no more than 3 percent for users aged 13 to 15 and 10 percent for users aged 16 to 17, certified by an independent third-party testing provider and reviewed annually by an independent auditor appointed jointly by Meta and the settling states (Biometric Update). Meta has one year from the effective date to get there.

A “false positive” here means a minor in that band who clears the check and is treated as an adult. If you come at this from a machine-learning background you will read “false positive” the other way round, as an adult wrongly flagged as a minor, so anchor on the plain-language version: these are ceilings on how many kids get through.

Proprietary methods, meaning Meta’s own in-house models, get a different schedule. They must reach 14 percent and 7 percent respectively within the first year, then 10 percent and 5 percent within the second. Under-13 detection and removal carries its own audited targets, and the auditor is not a light-touch reviewer. The State Committee, a bipartisan group of no more than six AG offices, jointly picks the auditor with Meta, and that auditor gets access to personnel, systems, raw and aggregated data, internal documents and internal communications, reports regularly, and can escalate directly to the attorneys general. The obligation runs for ten years.

Underneath the accuracy numbers sit the product terms that depend on them. Under-18 accounts get a default two-hour daily limit and an overnight block from midnight to 6am, both liftable only by a parent. Notifications go off between 10pm and 7am and during the school day, defined as 8am to 3pm from 15 August to 15 June. If TikTok and YouTube adopt comparable terms, the daily limit tightens to one hour and the block widens to 10pm to 7am, and roughly $5.3 billion of the payment turns on that happening. Texas settled separately the same day for more than a billion dollars with similar age-verification terms.

Read the structure rather than the dollar figures. Every one of those product obligations is downstream of a single classification decision. The decree is not really about screen time. It is about whether Meta can reliably place an account in one of three buckets, and the accuracy numbers exist because the attorneys general understood that a time limit for teenagers is worthless if you cannot tell who the teenagers are.

A national floor set by contract, not by statute

The interesting thing about this instrument is what kind of instrument it is.

Statutes go through legislatures, get challenged on First Amendment grounds, get enjoined, and apply within one state’s borders. A consent judgment binds one company, but this one binds it across the jurisdictions of roughly every state AG in the country, for a decade, with a monitoring apparatus attached. There is no severability fight and no preemption argument. Meta agreed to it.

That makes it the most operationally durable age-assurance standard in the United States, and it did not survive a single constitutional test because it never had to face one.

The second-order effect matters more than the first. Only one state has previously put numbers on age-assurance accuracy: New York, under the SAFE for Kids Act, which we broke down when the draft rules landed (Age Assurance Gets a Spec Sheet). Every other regulator has used adjectives. Regulators in Australia, the UK and the EU have spent two years being told by critics that age checks simply don’t work, with no agreed benchmark to argue against. They now have one, conceded by the largest social platform on earth, and they will ask for it. The AVPA’s read is that the settlement shifts the debate from “age checks don’t work” to “which vendor has been certified to do what,” and that seems right.

If you sell into this market, assume every serious RFP from 2027 onward asks for a false-positive rate by age band, from an independent tester, with a date on it. If you buy in this market, assume a plaintiff’s lawyer eventually asks you why your age gate performs worse than the one Meta agreed to.

Two published schedules that do not agree

Here is the first practical problem for anyone operating nationally. There are now two numeric standards in the US and they are not the same shape.

New York’s SAFE for Kids table is per-year and much more granular: no more than 0.1 percent of minors aged 0 to 7, 1 percent aged 8 to 13, 2 percent aged 14 to 15, 8 percent at age 16, and 15 percent at age 17. Meta’s is two coarse bands: 3 percent across 13 to 15, 10 percent across 16 to 17.

Line them up and they disagree in both directions. New York is stricter on the younger band, wanting 2 percent for 14 to 15 year olds where Meta’s decree allows 3 percent across 13 to 15. Meta is stricter at seventeen, capping the 16-17 band at 10 percent where New York allows 15 percent for seventeen-year-olds specifically. A model tuned to pass one can fail the other, in either direction, depending on where its errors cluster.

There is a subtler discrepancy in the denominator. New York’s rule explicitly excludes failures, refusals to provide data, and inconclusive outcomes from the calculation. The Meta language as published does not say. That sounds like a footnote and is not: if inconclusive results are excluded, a system can improve its measured rate simply by declining to decide more often, pushing users into a step-up flow and reporting a cleaner number on the ones it did classify. If they are included, an abstention counts against you and the incentive reverses. Two vendors can report the same headline figure while measuring different things. Ask which denominator, in writing, before you believe a number.

The workable posture is to design against the strictest applicable cell rather than an average, and to hold per-year error rates internally even where the external standard only asks for bands. You cannot decompose a band-level metric back into per-year rates after the fact, but you can always aggregate per-year rates up into whatever bands the next standard asks for. Instrument at the finer granularity now and the next table costs you a query rather than a re-test.

The glide path quietly penalises buying

The part of the schedule almost nobody has commented on is the asymmetry between commercial and proprietary methods, and it points the wrong way.

Third-party methods must hit 3 and 10 immediately, within one year. Meta’s own models get 7 and 14 in year one, and 5 and 10 in year two. In other words, the decree sets a materially looser bar for the in-house build than for anything purchased, and gives the in-house build an extra year to reach a standard the vendor had to meet on day one.

There is a defensible logic to it. A commercially available product can be tested off the shelf by an independent lab against a known corpus, so holding it to a firm number is easy. An in-house model trained on a live platform’s own data is harder to benchmark cleanly and was, presumably, further from the target when the deal was signed. The glide path is a concession to where Meta actually is.

The incentive it creates is still perverse, and it is worth naming. If you are a platform choosing between licensing a certified age-estimation vendor and building your own, this decree tells you the purchased option will be held to a stricter, earlier standard than the one you write yourself. That is backwards from how procurement risk normally works, and it is a reason a large platform might rationally keep the work in-house even where a vendor would do it better. We have written before about how badly the in-house build tends to pencil out once you count injection-attack defence, per-jurisdiction rules, and retention discipline (Build vs Buy). None of that changes. But an accuracy standard that is easier on the thing you built yourself is a standard that will produce more mediocre in-house models, and any regulator copying these numbers into statute should collapse the two schedules into one.

Both schemes constrain half the confusion matrix

The deeper limitation is shared by New York and by Meta’s decree, and it will cause real product damage before anyone corrects it.

Both instruments cap only the rate at which minors are passed as adults. Neither says anything about the rate at which adults are wrongly blocked as minors.

That is a one-sided constraint on a two-sided problem, and there is a trivial way to satisfy it. Raise the decision threshold. Push the challenge age up to 23, 25, 27. Every increment reduces the number of seventeen-year-olds who slip through and increases the number of nineteen and twenty-two year olds who get told to photograph a passport. The decree is satisfied. The audit passes. The cost lands entirely on adults, whose friction, abandonment and support tickets appear in no compliance report anywhere.

The scale of that cost is not hypothetical, because facial age estimation is genuinely imprecise near the adult threshold. NIST’s FATE evaluation has the best algorithms at a mean absolute error just under three years on some operational corpora, with teen subjects running closer to three to five years, and fewer than 35 percent of thirteen-year-olds estimated within one year of their true age by any algorithm tested (Biometric Update on NIST FATE). A three-year MAE at a seventeen-versus-eighteen boundary means the distributions overlap heavily. Getting the minor-pass rate down to 10 percent in the 16-17 band, purely by moving a threshold, drags a large population of young adults across with it. This is the failure mode we argued was the real one back in July (The Adult-Lockout Problem), and the new standards make it worse by pricing one error and not the other.

So treat the decree’s numbers as a floor on one axis and set your own ceiling on the other. If you are not tracking the rate at which adults are wrongly challenged, and the conversion cost of that challenge, you are optimising against the only metric anyone is measuring and paying for it somewhere you are not looking.

One threshold cannot satisfy a banded standard

The engineering consequence of band-conditional ceilings is that a single global cutoff will not get you there, and teams keep discovering this late.

Take the two bands seriously. To keep 13 to 15 year olds under 3 percent you need an aggressive threshold, because that population is far enough from adulthood that a well-calibrated model should separate it cleanly. To keep 16 to 17 year olds under 10 percent you need a threshold aggressive enough to catch a group that is, in the imaging sense, nearly indistinguishable from young adults. Those two requirements do not resolve to the same number, and the second one dominates: any threshold strict enough for the 16-17 band is far stricter than the 13-15 band needs, which means it also blocks an enormous number of adults you had no reason to block.

The way out is not one better model. It is layering, so that the expensive, high-assurance decision is reserved for the narrow region where the cheap signal is genuinely uncertain, and the cheap signal handles everything else. Three practical implications:

  • Measure per band, always. Aggregate error rates hide exactly the structure the standard cares about. Report a confusion matrix per age band, per method, and per demographic slice, and keep the per-year detail underneath it.
  • Make the gray zone explicit and route it, rather than guessing. Sessions the estimator cannot resolve confidently near the threshold should escalate to a stronger check, not get a coin-flip classification. That is the mechanism that lets you tighten the minor-pass rate without proportionally destroying adult conversion, and it is what an orchestrated waterfall is for (Orchestration and layered verification).
  • Keep the evidence for the auditor, not the biometric. An auditor with raw-data access wants the decision record: method, band, threshold, timestamp, outcome, and the test certificate behind the model. It does not want, and you should not be holding, the face images that produced them. Storing the evidence and storing the identity are different jobs, and conflating them is how a compliance win becomes the next breach (the retention problem).

The last point deserves emphasis because this decree makes it structural. Ten years of independent audit with access to raw data means whatever you retain, you retain in a form somebody external will read. Every extra field you keep is a field that has to survive a decade of review and whatever happens to your infrastructure in that time.

What this means if you are not Meta

You are not bound by this settlement. You will be measured against it anyway, through three mechanisms.

The contingent payment is the most direct. Roughly $5.3 billion turns on TikTok and YouTube adopting comparable terms, which makes the decree an explicit attempt to set an industry default rather than a company-specific remedy. Meta said as much publicly. If the two other large video platforms match, the terms tighten further for everyone.

The second mechanism is enforcement analogy. When a state AG opens an inquiry into a mid-sized platform in 2027, the question will not be whether you violated the Meta settlement. It will be why your age assurance performs worse than the standard the largest platform in the world accepted as reasonable. That is not a legal argument, it is a much more dangerous rhetorical one, and it works the same way in a plaintiff’s complaint.

The third is procurement. Third-party certification against a published corpus has just become the qualifying credential. A vendor that has been benchmarked on NIST FATE can point at a public result. A vendor that cannot will be asked why not, and “our internal testing shows” is not going to survive that conversation for much longer.

How Xident fits

Our architecture was built around the split that these numbers make unavoidable. A Check is the cheap, high-volume, privacy-preserving age-band signal: browser age check, liveness, returning-user lookup, OAuth. A Verification is the expensive document path with OCR and face match. The reason we price and meter them separately is that a banded accuracy standard forces you to spend assurance where the uncertainty actually is, and to stop spending it everywhere else.

Concretely, against a standard like this one:

  • The Check resolves the large majority of sessions that are not near the threshold, at roughly a tenth the cost of a document scan, without collecting an identity document to answer a yes-or-no question.
  • The gray zone near the adult boundary, which is where both the 16-17 ceiling and the adult-lockout cost live, escalates to a Verification rather than being forced into a guess. That is the only lever that improves both sides of the confusion matrix at once.
  • Every decision is logged as a decision: method, band, threshold, outcome, timestamp. The evidence is designed to be handed to an auditor without handing over the biometric that produced it, which is the same requirement Ofcom’s information requests exposed in the UK (evidence, not a pass/fail flag).
  • Per-band and per-year error accounting is the reporting granularity, not a bespoke export, because the next standard will slice the bands differently and re-testing to answer that is expensive.

We are not going to claim our numbers clear every cell of both tables in every deployment, because that depends on the method mix, the population, and the threshold a customer chooses. What we will say is that the ability to state those numbers per band, from an independent test, is now the entry ticket, and that any vendor answering an accuracy question with a single global figure has told you they have not measured the thing the standard measures.

The honest limits

The settlement is proposed, not final. It requires court approval through entry of a consent judgment, and terms can move before that happens. Reporting on the total varies between $16.7 billion and $18 billion because different outlets count the contingent portions differently, and the AG office count has been reported as 47, 51 and 52 across sources; treat the precise figures as provisional and the structure as reliable.

More substantively, nobody has yet demonstrated that these thresholds are reachable at Meta’s scale. TechCrunch’s critique is fair: the deal rests on technology that does not currently work especially well at the boundary it is being asked to police, and a 3 percent ceiling on the 13-15 band is a real engineering commitment rather than a formality. It is also possible the numbers get hit by throwing adults under the bus, which is the failure mode described above and the one to watch for in the first annual audit.

And a number is not a safety outcome. A platform can hit every cell in the table and still be a bad place for a fifteen-year-old to spend four hours a day. The accuracy floor makes the age signal defensible. What you do with the signal is a separate design question, and it is the one the rest of the settlement is actually about.

Share this article

Ready to implement age verification?

Get started in minutes with our simple SDK. Free trial includes 100 verifications.

Book a 20-minute demo