Summary
- The central confirmed finding comes from the GSA Office of Inspector General's March 2023 evaluation. It found that Login.gov had never included the physical or biometric comparison required for the IAL2 service it represented to customer agencies, that GSA did not promptly disclose known noncompliance, and that GSA billed more than $10 million to IAL2 customers through May 2022.
- The failure was not that Login.gov considered equity, privacy, or alternatives to automated facial matching. Those are legitimate public-service constraints. The failure was that a decision not to implement a required control did not trigger a simultaneous change to the assurance claim, interagency agreements, billing treatment, funding representations, and customer risk decisions.
- Repair evidence is substantial. GSA corrected public descriptions, expanded in-person proofing, piloted one-to-one facial matching, obtained an independent NIST SP 800-63-3 IAL2 assessment in October 2024, separated basic non-IAL2 verification from enhanced IAL2 verification, expanded authoritative-record checks, and documented current privacy practices. Oversight records also show the five 2023 OIG recommendations as closed.
- Repair is not the same as complete accountability. The public record does not quantify benefit applicants admitted or denied under an incorrect assurance assumption, identify every customer remediation decision, publish the billing review's customer-by-customer results, establish refunds or restitution, or disclose longitudinal pass, fail, fraud, and redress outcomes by proofing pathway and affected population.
- The durable test is whether every assurance assertion can be traced, at the time it is made, to a versioned standard, deployed controls, independent evidence, accurate customer terms, matched billing, timely exception disclosure, user-outcome measures, and a documented path for standards change. A later certificate proves a later service state; it cannot retroactively validate an earlier claim.
An assurance label is an operational promise
Digital identity discussions often collapse authentication, identity proofing, eligibility, and authorization into one idea: logging in. Login.gov does not make benefit or eligibility decisions, and its current public explanation says so. A person creates an account, authenticates with a password and another factor, and may then complete identity verification if the partner agency requires it. The partner agency decides what the person may access and whether the person qualifies for a service. That division of labor is important, but it does not make Login.gov's assurance representation incidental.
An Identity Assurance Level describes confidence that an applicant is associated with a real-world identity. It is different from an Authentication Assurance Level, which concerns confidence that the person presenting an authenticator controls it. A service can require strong multifactor authentication and still fail to establish the claimed real-world identity at the promised level.
The 2023 OIG report also found a separate AAL2 configuration issue that GSA later corrected, but the defining accountability issue here is IAL2: whether the identity evidence was not only collected and validated, but bound to the live applicant in the manner the governing standard required.
That label had at least five operational audiences. People encountered the proofing steps and surrendered sensitive information. Partner agencies used the resulting assertions to design access paths for benefits, records, financial transactions, and protected information. Procurement and program officials used interagency agreements to define what GSA would deliver. The Technology Modernization Fund Board evaluated a large funding proposal. Auditors and authorizing officials needed an evidence trail showing that the production service matched its stated control set.
For each audience, "IAL2" carried more information than "we perform several identity checks." It implied conformance to a published requirements set. If the service instead provided a useful but different bundle of document authentication, records checks, address confirmation, and fraud controls, the truthful product name needed to say so. Login.gov now does exactly that: its current service description distinguishes basic identity verification that is not IAL2-compliant from enhanced identity verification that is. That later separation illustrates the governance discipline missing during the reviewed period.
What the governing standard required at the time
Historical conduct should be evaluated against the standard then in force, not retroactively against a later revision. During the period covered by the OIG evaluation, NIST SP 800-63-3 and its companion volume SP 800-63A governed the relevant claim. The Rev.3 identity-proofing requirements described IAL2 as requiring evidence supporting the real-world existence of the claimed identity and verification that the applicant was appropriately associated with that identity. For the strongest identity evidence used at IAL2, verification needed strength classified as strong.
The accepted route was a physical comparison of the applicant with the photograph on the strongest evidence or a biometric comparison using the evidence.
NIST's Rev.3 implementation resource on identity verification made the control boundary especially clear. Identity validation asks whether evidence and attributes are authentic and accurate. Identity verification asks whether the valid evidence actually belongs to the applicant appearing for proofing. Comparing an identification card's data to authoritative or commercial records can support validation. It does not, by itself, bind the card to the live person presenting it. The missing physical or biometric comparison was therefore not a minor documentation defect. It was the step that linked evidence to applicant.
The standard did permit agencies to depart from a complete set of normative requirements in some circumstances, but that flexibility came with an evidence burden. The OIG quoted SP 800-63-3 as requiring comparability of an alternative, documentation of the departure, and details of compensating controls. The OIG found that GSA had no comparable alternative, no compensating controls, and no documented justification for the missing verification requirement. It rejected descriptions such as meeting the "spirit" or "flavor" of IAL2 as equivalent to meeting the standard.
The policy context also mattered. OMB Memorandum M-19-17, issued in May 2019, directed federal agencies to implement NIST SP 800-63-3 with the wider federal identity-management policy suite. A relying agency still had to perform its own digital identity risk assessment and choose an appropriate assurance level for each application. But that responsibility presupposed accurate service information. An agency cannot make an informed risk choice if the provider's label and deployed control differ.
This distinction prevents two opposite errors. It would be wrong to say that the old service performed no identity checks; official records describe document authentication, records checks, address confirmation, and fraud controls. It would also be wrong to say those checks made the missing verification control irrelevant. The confirmed issue was not whether Login.gov did something useful. It was whether it delivered the defined service it said it delivered.
The confirmed timeline shows repeated claim-control divergence
Login.gov launched in 2017 as a shared sign-in service. The OIG evaluation, opened in April 2022 after GSA's Office of General Counsel reported potential misconduct, examined program activity from May 2016 through December 2022. Its chronology shows that the mismatch was not discovered only after a sudden standards change.
In September 2018, Login.gov officials were discussing IAL2 services internally and with potential customers. Interagency agreement language began stating that the IAL2 service met NIST 800-63-3 for AAL2 and IAL2 and that TTS would provide those services on a reimbursable basis. The OIG found that 18 of 22 agreements executed from September 18, 2018, through July 7, 2021, said they included IAL2 services that met or were consistent with IAL2 requirements.
In July 2019, GSA's Chief Information Officer stated in Login.gov's FedRAMP agency authorization to operate that the system could support validation at IAL1 or IAL2. In November 2019, the CIO permitted customer deployment of IAL2 subject to conditions. Yet, according to the OIG, the production service available to customers never included a physical or biometric comparison. It used a third party to compare identification cards with information in commercial records. That could validate evidence and attributes, but it did not perform the required ownership verification.
Knowledge of the gap moved through the organization. The OIG reported discussions about inability to meet IAL2 at least as early as 2019. It said a TTS senior adviser informed the team in January 2020 that a biometric component was required. A consultant gave the same warning in August 2020. Internal discussions also recognized concerns about selfie technology, liveness detection, and differential rejection rates associated with physical traits such as skin color and tone. Those concerns were real governance inputs; they did not make the service conformant.
On June 24, 2021, the TTS Director and FAS Deputy Commissioner announced internally that TTS was suspending efforts to meet the biometric comparison requirement, citing equity concerns with liveness and selfie technology. The OIG found that customer agencies were not notified when that product decision was made. This date is the clearest accountability hinge. Once leadership decided not to deploy the control then understood to be necessary, the claim state should have changed everywhere: product catalog, authorization evidence, interagency agreement, billing code, sales discussion, funding proposal, and customer notice.
Instead, in September 2021 GSA submitted a Technology Modernization Fund proposal that said Login.gov provided identity services in accordance with M-19-17 and complied with NIST 800-63-3 for both AAL2 and IAL2. The proposal ultimately secured approximately $187 million for 2022 through 2025. The OIG did not treat the entire award as a measured loss; the funds had broader goals including adoption, equity, cybersecurity, and antifraud capacity. It did find that the compliance statement used to obtain the funding was inaccurate.
Customer notification came in February 2022. The OIG recorded customer reactions showing that agencies had understood they were receiving compliant IAL2 service and now had to determine what the change meant for production systems. Some asked whether the service had just been removed. One concluded from the clarification that its production system had not been IAL2-compliant since launch because facial comparison had never been in production. Those reactions are direct evidence of reliance and rework risk. They are not proof that a specific ineligible person received a benefit or that a named applicant was denied one.
GSA also reviewed its AAL2 representations, notified 51 customers in March 2022 and another in April, and changed settings to correct that separate issue. For IAL2, it did not yet have the capability to make the same kind of configuration fix. The January 2024 Login.gov privacy impact assessment later used precise language: the identity-verification service did not meet NIST IAL2 at that time, while listing the controls it did provide. That was a materially better disclosure pattern.
The billing timeline ran alongside the claim timeline. Starting in 2019, GSA charged agencies for services described as IAL2. The OIG found billings to 22 customer agencies and calculated $10,060,254 through May 2022: $459,877 in fiscal 2020, $4,288,990 in fiscal 2021, and $5,311,387 through May of fiscal 2022. Its footnote cautioned that the billed totals included identity-verification fees, platform fees, and authentication fees and that Login.gov did not differentiate all costs at a level that isolates a pure nonconforming-control amount.
The number is therefore a confirmed billing total, not a judicial damages award and not proof that every dollar purchased no value.
On March 7, 2023, the OIG issued five findings and five recommendations. It concluded that GSA misled customers, billed for IAL2 services not provided as represented, used inaccurate compliance language in the funding proposal, and allowed inadequate management controls and a hands-off culture. GSA management agreed with the findings and recommendations. The report was an inspector-general inspection and evaluation, not a criminal conviction or civil judgment, and its conclusions should be described in those terms.
Practical control was distributed, but responsibility was not diluted
The record supports a layered control map rather than a single-villain narrative.
Login.gov product and program leadership controlled the service state. The program controlled which proofing steps were deployed, which vendors and data sources were used, how evidence was bound to an applicant, what product names appeared in documentation, and what technical limitations were escalated. Product teams were also closest to evidence that the control set did not satisfy the asserted level.
Technology Transformation Services controlled commercialization and escalation. TTS leadership controlled the decision to suspend biometric implementation, partner communications, agreement templates, business operations, and the interface between program facts and external claims. Once the product decision and assurance claim diverged, TTS had practical power to stop the mismatch.
The Federal Acquisition Service controlled institutional oversight. FAS supervised TTS, signed or oversaw interagency relationships, and was responsible for management controls, financial planning, procurement, and program governance. The OIG specifically rejected the idea that program independence relieved FAS of responsibility. It found that inadequate FAS oversight enabled the hands-off culture and the misleading representations.
The CIO and authorizing process controlled formal risk acceptance. An authorization to operate is not an independent product certification, but it is a formal control checkpoint. When the 2019 authorization said the system could support IAL2, that statement became part of the institutional evidence chain. Authorizing officials needed a requirements-to-implementation map capable of detecting that validation was being substituted for verification.
Customer agencies controlled application-level assurance and authorization. Each relying agency had to assess the harm from identity errors, choose an assurance level, decide what attributes to request, and determine what a successful Login.gov assertion permitted. Those duties remain. They do not excuse inaccurate provider representations. Shared services exist partly so agencies can rely on a common control implementation rather than rebuild it; truthful service boundaries are therefore a prerequisite to customer responsibility.
The Technology Modernization Fund Board controlled funding, not service truth. The Board could evaluate and condition an award, but it depended on the accuracy of the proposal and supporting materials. Its approval did not transform a nonconforming service into a conforming one.
Component vendors controlled component performance. Commercial and government data sources could authenticate documents, validate attributes, perform facial matching, assess devices, or provide fraud signals. Login.gov, as the credential service provider presented to agencies, retained responsibility for assembling those components into a service that met the claimed level and for explaining the resulting data flows and limitations.
NIST, Kantara, OIG, and GAO controlled different evidence functions. NIST defined technical and process requirements; it did not operate Login.gov. Kantara later assessed a specified service against Rev.3. OIG reconstructed the historical claim and management failure. GAO assessed capabilities, agency experience, pilots, privacy controls, and follow-up actions. None of these bodies substituted for GSA's day-to-day duty to keep claims aligned with deployed controls.
This allocation matters because accountability can disappear when every entity points to another entity's role. The better rule is that each party answers for the decisions it could actually change. GSA controlled the product claim and evidence package. Customer agencies controlled use-case selection and downstream authorization. Assessors controlled the integrity and scope of their assessments. Users controlled almost none of those institutional choices.
Harm and cost begin before a confirmed fraudulent transaction
The public record does not establish a count of fraudulent approvals, wrongful denials, stolen benefits, or compromised accounts caused by the missing IAL2 verification step. It would be speculation to invent one. But absence of a quantified transaction loss does not make the failure harmless.
First, agencies lost decision quality. A relying party choosing a proofing service needs to know the likelihood and type of identity error the controls are designed to reduce. If a service represented that it linked evidence to a live applicant but did not perform the required comparison, the agency's residual-risk analysis started from a false premise. It could impose too little protection on a high-risk transaction or spend time designing compensating measures only after belated disclosure.
Second, agencies incurred investigation and migration costs. GAO later reported that the Small Business Administration paused adoption after the OIG finding and performed additional data calls and reviews on cost and security. Treasury told GAO that using a nonaligned service for applications requiring IAL2 would expose it to security risk. Such effects are operational costs even when they cannot be reduced to a public invoice.
Third, the billing created an auditable taxpayer interest. The confirmed $10,060,254 was not a clean measure of overpayment because the invoice bundles included services that were delivered. It was nevertheless money billed to customers under an IAL2 service relationship that the OIG found did not meet IAL2. A comprehensive billing review was therefore necessary to determine which charges matched delivered capability, which terms needed correction, and whether any financial adjustment was due.
Fourth, capability gaps can push agencies toward parallel systems. GAO's June 2025 comparison found that CFO Act agencies reported spending about $32.5 million on Login.gov and $209 million on commercial solutions from fiscal 2020 through 2023. GAO warned that proprietary pricing prevented a direct comparison. It also reported that much commercial spending reflected capabilities, including IAL2 and biometrics, that Login.gov did not offer during the period. These figures show the economic scale of fragmented identity proofing; they do not prove that Login.gov's misrepresentation caused the entire commercial spend.
Fifth, people bear access and privacy costs. A failed proofing attempt can delay access to records, applications, or benefits. A more invasive control can require an image of a face, identity document, Social Security number, device signals, and authoritative-record checks. A less effective control can expose public funds or personal accounts to impersonation. The accountability problem is therefore two-sided: prevent fraud without turning proofing errors, biometric bias, documentation barriers, or opaque data sharing into a denial mechanism.
GAO's October 2024 report captured both value and burden. Of 21 CFO Act agencies reporting Login.gov use, 16 cited improved operations, 11 improved user experience, and seven cost savings. Twelve cited the IAL2 alignment gap, nine cited technical problems, and eight cited cost uncertainty. One agency reported account-creation failure rates of 30 to 40 percent for its users; GAO did not attribute every failure to identity proofing or the historical IAL2 claim. That mixed evidence is more useful than a simple success-or-failure label.
Shared identity can reduce duplicated work and still impose serious control and access costs when its evidence is weak or its workflow fails.
Finally, credibility is itself infrastructure. Login.gov asks agencies to centralize a sensitive trust function and asks people to use one account across government. That model can reduce complexity, but it concentrates reliance on representations made by the shared provider. When assurance language proves unreliable, agencies may duplicate verification, delay adoption, or buy parallel services. The resulting cost is not merely reputational; it can undermine the economies of scale the service was created to deliver.
Equity concerns were valid constraints, not permission to mislabel the service
The OIG record should not be read as proof that GSA should have ignored disparate error rates or deployed any available facial technology immediately. NIST's continuing Face Recognition Technology Evaluation, updated through June 2026, reports demographic variation in false matches and false non-matches across algorithms. It also explains that image quality, lighting, exposure, and camera geometry can affect false-negative performance. Those facts support careful testing, accessible capture conditions, alternative paths, and ongoing outcome measurement.
The governance mistake was coupling a defensible product concern with an indefensible claim practice. If leadership concluded in June 2021 that the equity risk of the planned selfie and liveness approach outweighed its benefit, four truthful options remained. GSA could label the service as non-IAL2. It could develop an attended remote or in-person physical comparison. It could use a conforming external service for applications requiring IAL2 while retaining its own lower-assurance product. Or it could document a genuinely comparable alternative with compensating controls and obtain authoritative review before claiming conformance.
What it could not do was keep the IAL2 label unchanged while the required control was absent.
The later architecture demonstrates the value of multiple paths. GSA's April 2024 pilot announcement paired remote one-to-one facial matching with an upfront in-person option at more than 18,000 participating Postal Service locations. GAO reported that, in a 1,000-entity in-person pilot running from January through March 2024, 57 percent of those who successfully proofed chose the in-person path when offered at the start. GSA made the upfront option permanent in April. That is evidence that an alternative channel can be a core service path rather than an exception hidden after online failure.
Choice, however, must be measured rather than assumed. Proximity to a post office does not prove that a person can travel there during operating hours, has acceptable documents, can take time away from work or care, or can complete the agency process after proofing. A remote facial match does not prove equal operational performance across devices and populations. An in-person channel does not prove that staff apply comparison and exception rules consistently.
The correct control record includes pass and adjusted-fail rates by proofing type, abandonment points, suspected-fraud adjustment, wait time, redress time, accessibility outcomes, and demographic analysis performed with lawful privacy safeguards.
NIST's 2025 Revision 4 reinforces that approach. It permits IAL2 biometric, non-biometric, and digital-evidence pathways, requires comparable aggregate assurance, and adds stronger expectations for options, privacy, customer experience, and redress. It also says providers offering a non-biometric IAL2 pathway must communicate its use to relying parties. The lesson is not "biometrics are always required" or "equity overrides standards." It is that the provider must name the pathway, satisfy the applicable requirements, measure its performance, and tell customers exactly what assertion they are receiving.
Confirmed facts, supported inference, and unknowns
Confirmed facts in the public record are unusually strong. The GSA OIG found that the production service offered to customer agencies lacked the physical or biometric comparison required for its claimed Rev.3 IAL2 service.
It found that senior personnel learned of the gap at multiple points, that TTS suspended efforts to implement the biometric control in June 2021 without contemporaneous customer notice, that 18 of 22 reviewed interagency agreements contained misleading IAL2 language, that the September 2021 funding proposal included an inaccurate compliance statement, and that GSA billed more than $10 million to IAL2 customers through May 2022. GSA management agreed with the report's findings and recommendations.
It is also confirmed that GSA notified customer agencies in February 2022, corrected public descriptions, and later built new proofing paths. GSA announced a remote facial-matching pilot and expanded in-person proofing in April 2024. In October 2024, GSA announced that an independently assessed Rev.3 IAL2 service had entered general availability. The Kantara trust status list identifies Login.gov as a Rev.3 full service approved at IAL2 and AAL2. Current Login.gov materials distinguish basic non-IAL2 verification from enhanced IAL2 verification.
It is confirmed that oversight continued after the certificate. GAO's current recommendation page says it verified in March 2025 that the remote pilot was complete and that a third-party reviewer helped confirm the functionality worked as intended. It marks the pilot-completion and lessons-learned recommendations closed. It still lists the recommendation to address agency-reported technical challenges with mutually agreed time frames as open and partially addressed. Separately, GAO's 2025 backup-control recommendation now shows closed after GAO verified annual tests using June 2025 and June 2026 evidence.
Supported inference begins where the documents establish the premises but not a direct measured outcome. It is reasonable to infer that the inaccurate claim impaired customer risk decisions because agencies contracted for a defined assurance service and expressed surprise when told it was unavailable. It is reasonable to infer that some agencies incurred review, delay, integration, or replacement costs because GAO documented paused adoption and use of commercial products for capabilities Login.gov lacked.
It is reasonable to infer that a requirements-to-control gate tying product state to agreements and billing could have surfaced the mismatch earlier.
It is also reasonable to infer that current certification, service separation, in-person choice, passport validation, privacy documentation, and recurring oversight materially reduce the historical claim-control risk. These are observable changes in capability and evidence, not merely new language. But it would go beyond the evidence to infer that every historical customer has been made whole, every affected application has been risk-reassessed, or every current proofing pathway performs equitably in production.
Unknowns remain material. The public sources reviewed here do not identify every person who approved each historical representation, reproduce every interagency agreement, disclose all internal legal advice, or provide the complete billing-review workpapers. They do not establish whether GSA refunded or credited particular customers, what methodology it used to apportion bundled charges, or whether customer agencies changed authorization rules after notification.
The public record also does not quantify identity fraud, wrongful access, benefit loss, account compromise, false rejection, service abandonment, or delayed benefit receipt caused specifically by the missing comparison. It does not publish customer-by-customer reliance outcomes. It does not establish criminal intent, a False Claims Act judgment, or individual civil liability. Personnel departures reported in the OIG chronology are not, without additional evidence, proof of discipline or culpability for a specific act.
For the repaired service, the public record does not provide the full independent assessment report, all test cases, vendor-specific operating thresholds, production error rates, demographic outcome tables, redress resolution measures, or annual reassessment results. The March 2026 Login.gov privacy impact assessment names data categories, third parties, retention boundaries, and controls, but a PIA is not a public performance audit. A trust mark is strong conformance evidence for its assessed scope and date; it is not a perpetual warranty against configuration drift or a substitute for user-outcome evidence.
The oversight and legal record is broader than one report
The 2023 OIG evaluation produced five recommendations: establish adequate TTS management controls; document policies, decisions, procedures, essential transactions, and records; comprehensively review IAL2 billings; establish internal compliance reviews; and adopt a policy that clearly tells each customer whether Login.gov meets applicable NIST standards and the services specified in the agreement. Oversight.gov currently marks all five closed: management controls, documentation and records, billing review, internal compliance review, and customer notice.
Closure is important, but its evidentiary meaning is bounded. It shows that the responsible oversight process accepted agency action sufficient to close each recommendation. The public status pages do not publish the billing review itself, customer-level adjustments, test samples for internal reviews, or all new management-control artifacts. A durable accountability file should preserve both facts: the recommendations are not open, and some underlying outcome evidence remains nonpublic.
The legal and policy setting includes more than NIST conformance. The current PIA identifies statutory authorities including federal cybersecurity requirements, the E-Government Act, federal information-policy law, and GSA's authority to provide services to executive agencies. It maps the Login.gov system to a Privacy Act system-of-records notice. It says current identity verification may collect name, date of birth, address, Social Security number, state ID or passport data and images, a self-photograph when required, and device or behavioral signals for fraud mitigation.
It also identifies AAMVA, the Department of State, LexisNexis, and Socure as current identity-verification services or sources.
Those disclosures make data governance part of assurance accountability. A stronger proofing pathway may reduce impersonation risk while expanding collection and third-party processing. The current PIA says Login.gov does not store state ID images after authenticity processing and that certain third-party and anti-fraud data are retained for limited periods, generally not exceeding one year unless an investigation requires longer. These are official descriptions of current practice. Durable proof requires audit evidence that the described deletion, access, purpose limitation, vendor restriction, and incident controls operate as stated.
GAO adds a continuing control record. Its 2024 review identified agency benefits and technical challenges and recommended dated action on reported issues, completion of the remote pilot, and documented lessons learned. As of July 15, 2026, two are closed and the technical-challenge recommendation remains partially addressed. Its 2025 review found Login.gov largely implemented selected data-protection practices but had not demonstrated annual backup-integrity testing. The current GAO page records closure after evidence of tests in June 2025 and June 2026.
That is a useful model of repair evidence: policy, test output, recurrence, and independent verification.
Repair evidence is real, but it must be read by date and scope
The repair chronology begins before certification. GSA corrected websites and customer communications after the 2022 disclosure. Its January 2024 PIA explicitly said the service was not IAL2 at that time. GSA completed an upfront in-person proofing pilot in March 2024 and made that path generally available in April. Its April 2024 announcement described a separate remote facial-matching pilot due to begin in May and said GSA would seek independent assessment.
On October 9, 2024, GSA announced general availability of an independently certified IAL2 option. The remote workflow added one-to-one matching of a live selfie with the photograph on a user-provided ID. GSA said it did not use one-to-many identification and did not use the images for unrelated purposes. The certification also covered an in-person option. This matters because it tied a new claim to a deployed control and an outside assessment before broad availability.
Dates reconcile an apparent conflict in the record. GAO's October 16, 2024 report said Login.gov was not yet IAL2-compliant as of July 2024 and described the remote pilot as incomplete based on evidence gathered before the October certification. GSA's announcement concerns the later service state. The sources are sequential, not contradictory. This is exactly why assurance statements need effective dates and service versions.
Current product separation is another repair. The partner site describes three services: authentication, basic identity verification that does not meet IAL2, and enhanced identity verification that is IAL2-compliant. The assurance-level guidance for agencies tells partners to perform a digital identity risk assessment and identifies higher-risk use cases that may warrant IAL2. It also says GSA does not choose the level for the partner. This makes the customer decision boundary more visible than it was in the historical agreement language.
Current user-facing documentation also shows multiple routes. The identity-verification help page says users may provide a driver's license, state ID, or passport book, a Social Security number, and a U.S. phone number or mailing address; some users are asked for a selfie, and an in-person post-office path may be available. GSA's August 2025 passport-based verification announcement describes checking passport attributes through a privacy-preserving Department of State interface. These are capability and access improvements, although each introduces its own availability, privacy, and error questions.
Scale makes control durability more consequential. GSA's December 2025 roadmap reported more than 100 million accounts, more than 500 million annual sign-ins, over 700 live sites and services, and 54 agency and state partners. These are program-reported metrics, not independently audited counts in this article. At that scale, even a low failure or false-acceptance rate can affect many people, while a common corrective control can benefit many agencies.
The December 2025 roadmap also says Login.gov completed the independent Rev.3 assessment and is working toward the newly released NIST SP 800-63-4. That is an appropriately bounded statement: working toward a newer standard is not the same as already certified to it. The roadmap is forward-looking and labels estimates as subject to change. Its plans for mobile driver's licenses, inherited proofing, alternative paths, additional vendors, and fraud-signal sharing are not proof those capabilities are complete until release and operating evidence exist.
The repair record therefore has multiple strengths. The strongest are corrected claim language, deployed service separation, a completed pilot, independent Rev.3 assessment, public privacy documentation, closed OIG recommendations, and GAO-verified recurring backup tests. Weaker but useful evidence includes roadmaps, agency press releases, and planned capabilities. Missing evidence includes detailed customer billing outcomes, the full certification assessment, longitudinal production measures, demographic results for the actual service, redress efficacy, and a completed Rev.4 conformance record.
The counterfactual was disclosure plus service separation
A credible counterfactual does not assume that GSA could have produced a perfect, equitable biometric service overnight. It asks what controls were available when leaders knew the existing service did not match the claim.
At the latest, the June 24, 2021 decision to suspend the planned biometric requirement should have triggered a claim-control incident. The service owner could have frozen new IAL2 representations; changed the catalog to "identity verification, not IAL2"; notified existing agencies; amended interagency agreements; reviewed authorization language; stopped or segregated IAL2 billing; corrected the pending funding proposal; and given each relying agency a transition package.
That package could have offered continued use for lower-risk transactions, external IAL2 service for higher-risk transactions, an attended or in-person alternative under development, and a documented timetable for reassessment.
An independent conformance gate before the first IAL2 customer deployment would have been stronger still. The assessor would have traced each Rev.3 requirement to a production control, implementation evidence, test result, exception, and compensating-control analysis. The missing applicant-to-evidence comparison would have appeared as an unresolved mandatory control rather than being obscured by a long list of other identity checks.
The cost of this counterfactual would not have been zero. GSA might have lost or delayed customers, borne amendment and review costs, accelerated an in-person channel, or funded commercial capacity. Agencies might have had to change integrations. But those are the costs of accurately allocating a known risk. The actual path transferred some of that cost to agencies later, after they had made plans and entered agreements under the stronger label.
The current service taxonomy validates the central design idea. GSA now offers a basic non-IAL2 product and a distinct enhanced IAL2 product. It now presents in-person and remote options. That does not prove the exact current architecture was feasible in 2021, but it shows that usefulness did not require collapsing every service into one assurance claim. A truthful lower-assurance service can coexist with a separately evidenced higher-assurance one.
The broader comparator is not a particular private vendor. Commercial alternatives have different prices, privacy models, populations, data sources, and transparency. The useful comparison is procedural: a provider that gates claims on independent evidence and immediately discloses a material control departure gives customers a chance to choose. A provider that preserves the claim while changing the control makes the choice for them.
A durable control-evidence accountability test
Login.gov should be judged neither by the historical failure alone nor by the latest certificate alone. A durable accountability test asks whether the assurance system keeps claims, controls, customer decisions, and user outcomes aligned over time.
1. Versioned claim test. Every assurance claim should identify the standard revision, service variant, proofing pathway, effective date, and material exclusions. "IAL2" without "Rev.3 enhanced remote" or an equivalent service identifier is too coarse once multiple standards and pathways coexist. NIST published SP 800-63-4 in August 2025, superseding Rev.3. A Rev.3 trust mark remains evidence of what was assessed; it should not be presented as silent proof of Rev.4 conformance.
2. Requirement-to-control test. A maintained matrix should map every normative requirement to deployed code, process, vendor component, responsible owner, test, evidence location, exception, and last validation date. Validation and verification must occupy separate rows so a records check cannot be mistaken for applicant binding. Changes to any mapped component should trigger impact analysis.
3. Pre-claim gate test. No public page, agreement template, authorization statement, funding proposal, sales artifact, or protocol value should assert an assurance level until an accountable official and independent assessor have approved the evidence package for the production configuration. The gate should fail closed when a mandatory control is suspended.
4. Contract and billing parity test. Product identifiers in the technical request, interagency agreement, service catalog, invoice, and customer report should resolve to the same assurance definition. Billing systems should separate authentication, basic verification, enhanced proofing, platform, and support charges sufficiently to support review and correction. A control downgrade should automatically stop higher-assurance billing until customers accept amended terms.
5. Disclosure and escalation test. A material control departure should have a notification clock, named recipients, severity rules, and a preserved decision record. Notification should state what changed, when, which assertions may be affected, what interim controls exist, what customers must reassess, and how billing will be handled. Equity or privacy concerns should be documented as reasons for a design decision, not used to blur the resulting assurance state.
6. Independent and continuous evidence test. Certification should be scoped, dated, renewable, and supplemented by configuration monitoring. It should cover policy and implementation, not only a demonstration flow. Annual reassessment, significant-change review, control sampling, and public status of the trust mark should make drift visible. GAO's closure of the backup recommendation after observing evidence from two annual test cycles illustrates the difference between a written policy and operating proof.
7. Relying-party decision test. Agencies should receive enough evidence to conduct their own risk assessment: assurance level, pathway, attributes, known limits, fraud model, privacy allocation, redress route, performance measures, and change history. The provider should obtain affirmative customer acknowledgment when a service boundary changes. Customer responsibility is meaningful only when the customer receives accurate facts.
8. User outcome and equity test. NIST SP 800-63A-4 calls for options, comparable assurance, notices, redress, and customer-experience controls. The wider Rev.4 suite identifies measures including overall and pathway-specific pass rates, fail rates, and fraud-adjusted fail rates. Login.gov should publish or provide independently reviewable aggregates for completion, abandonment, false acceptance, false rejection, manual review, wait time, accessibility, and redress, with careful demographic analysis and privacy protection. Vendor laboratory results alone do not establish end-to-end public-service outcomes.
9. Data-governance test. Every collected attribute, document image, biometric sample, device signal, and behavioral feature should have a purpose, lawful authority, sharing map, retention period, deletion proof, access log, and vendor restriction. Changes in fraud models or data sources should trigger privacy and bias review. Users and agencies should be able to understand which party has which data and what recourse exists when it is wrong.
10. Retrospective correction test. When an assurance claim is found inaccurate, repair should look backward as well as forward. The provider should identify affected agreements, assertions, applications, invoices, funding statements, and users; notify the right parties; support risk reassessment; correct charges where warranted; preserve records; and publish an appropriately aggregated closeout. A new conforming product is forward repair. It does not answer every question about the old product.
11. Standards-transition test. A shared identity provider needs a public transition plan whenever NIST issues a new revision: applicability date, gap assessment, temporary status of old certifications, customer migration rules, protocol changes, training, reassessment, and final conformance evidence. Login.gov's roadmap says Rev.4 work is underway. The durable gate remains open until current claims and independent evidence are explicitly mapped to the superseding revision.
12. Governance survival test. Controls must survive personnel changes, urgency, funding cycles, and product culture. Decision records, internal compliance review, audit access, records management, and clear FAS-TTS-program ownership should prevent a hands-off structure from recurring. A system is not durable if truth depends on one employee remembering to challenge a label.
Applied as of July 15, 2026, the result is mixed but not ambiguous. The historical claim-control test failed. The later Rev.3 capability and independent-assessment test passed for the scoped enhanced service. Product labeling and channel choice improved. The OIG recommendations are closed, and several GAO follow-ups have operating evidence. But the current GAO technical-challenge recommendation remains partially addressed, Rev.4 transition is still described as work in progress, and public evidence for retrospective billing outcomes and end-to-end user performance remains incomplete.
Accountability follows the ability to keep the assertion true
The most durable lesson from Login.gov is not that public agencies should avoid ambitious shared services or that facial matching is the only route to trustworthy identity. It is that an assurance assertion is a controlled output. It must change when the underlying control changes, and it must be supported by evidence before agencies, funders, or users are asked to rely on it.
GSA's later repair record deserves weight. The service now distinguishes non-IAL2 and IAL2 offerings, supports remote and in-person pathways, has an independent Rev.3 trust mark, publishes a more detailed privacy record, and is moving toward Rev.4. Those steps show that security, access, privacy, and honest product boundaries can be pursued together.
The remaining accountability question is whether that discipline is permanent and retrospective. Agencies and the public need evidence not only that today's enhanced flow contains the missing comparison, but that claims are versioned, customer terms and invoices match production, exceptions trigger immediate disclosure, old errors are reconciled, and real users can complete or challenge the process without hidden exclusion. That is the control-evidence test. Login.gov passes it only when the assertion, the deployed service, the contract, the invoice, the independent assessment, and the lived result all describe the same system.

