Summary
A distribution is not an individual result. A model can hold national or centre-level patterns close to prior years while remaining uncertain for an individual candidate whose counterfactual exam performance cannot be observed.
The policy objective came before the code. Ministers cancelled exams and directed a standardised approach; Ofqual designed the regulatory and statistical method; awarding organisations implemented it. Accountability has to follow that chain without collapsing it into one actor or one algorithm.
Equalities analysis answered a bounded question. Ofqual's later work did not find systematic bias against candidates with protected characteristics or disadvantaged backgrounds. That does not prove every result was reliable, nor does it settle every concern about centre history, cohort size or an atypical student.
Remedy had to be feasible at results-day speed. Appeals, autumn exams and later publication were parts of the control environment, but they could not by themselves make a contested grade immediately usable for university, college, training or employment decisions.
The reversal was a policy and legitimacy event. On 17 August, students were moved to the higher of their centre assessment grade or calculated grade. It is inaccurate to describe every changed grade as correction of a coding defect.
Future model governance must join technical and human evidence. Purpose, assumptions, data lineage, individual exception testing, implementation assurance, communication and a usable remedy need one accountable operating design.
The emergency removed the normal evidence, not the need for a defensible outcome
On 18 March 2020, the government announced that summer examinations would not take place. The public-health decision solved one immediate risk but removed the ordinary measurement event at the centre of GCSE, AS and A-level awarding. Students had prepared for examinations that would now produce no scripts, marks or externally moderated performance evidence. Yet universities, colleges, training providers and employers still needed results, and students needed to progress without waiting indefinitely.
The Direction issued to the Chief Regulator of Ofqual is the correct starting point for authority. It records ministerial directions under section 129(6) of the Apprenticeships, Skills, Children and Learning Act 2009. The direction required an approach standardised across centres and sought a grade distribution with a similar profile to previous years. Those were policy constraints, not details invented by a programmer.
There was no knowable single answer for what every student would have achieved. An exam result is itself an observation under particular questions, marking and performance conditions; a cancelled exam leaves even that observation unavailable. A centre assessment grade was a professional judgement about a counterfactual. A calculated grade was a model-supported estimate using centre evidence and historical patterns. Neither could be checked against the missing summer examination.
This uncertainty matters because institutional language can make an estimate sound like a recovered fact. The system was not reconstructing grades that already existed in a database. It was allocating consequential outcomes under an emergency policy using evidence with different scopes and limitations. Good governance would therefore state, for each evidence type, what it could support, what it could not support and how disagreement would be handled before results were issued.
Continuity was a legitimate objective. A delayed or absent grade could also harm an individual. The accountability failure is not that decision-makers acted under uncertainty; they had no alternative to uncertainty once exams were cancelled. The test is whether the operating system made that uncertainty visible and controlled it at the same level where consequences fell.
The policy problem contained objectives that could not all be maximised
The chosen arrangement tried to do several things simultaneously. It aimed to enable progression, preserve the value of qualifications, maintain standards over time, treat centres consistently, limit excessive aggregate inflation and remain fair to students. Each objective was reasonable. Together they created trade-offs that no statistical technique could erase.
Ofqual's announcement of exceptional arrangements followed consultation and confirmed that calculated grades would be awarded. It also pointed to an autumn exam series. This established an emergency architecture: centres would provide evidence, awarding organisations would calculate grades under Ofqual's rules, and a later examination route would exist for some students.
The central conflict was between system comparability and individual estimation. If centres' submitted grades were accepted unchanged, differences in how generously schools judged students could produce inconsistent outcomes and a large rise in grades. If historical centre performance constrained results, the system could limit those differences but would make a current student's outcome partly dependent on previous cohorts. Rank order preserved a centre's relative judgement, yet it did not establish the grade boundary that any one student would have crossed in an exam.
That conflict should be expressed as a decision record, not hidden inside a model specification. Commissioners needed to say which harms they were prepared to tolerate, how much deviation from historic patterns was acceptable and which individual cases required a different route. Without that statement, aggregate stability could be mistaken for the definition of fairness rather than one component of it.
There were alternatives, but none was costless. Unmoderated centre grades risked inconsistent generosity. Delayed examinations risked disruption and unequal readiness. Standardised grades created individual reliability and acceptability problems. A common assessment assembled at short notice would have raised safety, access and validity questions. Accountability does not demand a perfect option. It demands a transparent comparison that shows why the selected option remained suitable as evidence accumulated.
Centre grades and rank orders were human evidence with their own lineage
Schools and colleges were asked to provide a centre assessment grade for each student and rank students within each grade for a subject. The CAG represented the grade a student was judged most likely to have achieved had teaching continued and exams taken place. The rank order told the system which student should sit nearer a boundary when the model allocated a different grade distribution.
The formal requirements for calculating summer 2020 results translated policy into obligations for awarding organisations. This is where data governance becomes operational. A calculation is only as accountable as the definitions, submission rules, validation, version control and exception handling around its inputs.
Centre evidence was not raw fact. Teachers could draw on classwork, mock examinations, non-exam assessment and their knowledge of a student's performance. They also had to imagine a counterfactual continuation of teaching. Different subjects and centres possessed different amounts and qualities of evidence. Some students were improving quickly; some had recently moved; some were private candidates; some had unusual patterns of prior attainment. The submission process compressed these contexts into a grade and rank.
That compression is not inherently wrong. Structured professional judgement is often the best available evidence. But the data record should preserve the origin of the judgement: which evidence categories were available, who reviewed them, when the decision was made and whether the centre identified an exceptional case. A head-of-centre declaration can confirm process and good faith; it cannot turn a prediction into an observed examination result.
The model also inherited dependencies from centre practice. If CAGs were optimistic on average, accepting them without control could change overall standards. If a rank order was wrong, standardisation could allocate the available grades to the wrong students. If a centre had weak historic comparators, the model needed a fallback. Data governance therefore covered both the statistical engine and the distributed human system that produced its inputs.
Historical performance governed groups more strongly than individuals
Ofqual's Direct Centre Performance approach predicted a distribution for a subject within a centre. It used the centre's historical performance and adjusted for prior attainment of the current cohort. Teachers' rank orders then helped assign grades within that distribution. The method was built to control the results profile of groups, not to predict each student's independent examination mark.
The 319-page interim report on summer 2020 awarding explains the model options, testing and selected approach. It states that historical performance in the subject and changes in cohort prior attainment informed the predicted distribution. It also explains that CAGs received more weight for smaller entries because statistical prediction was less reliable there.
This design created a clear level-of-analysis boundary. Centre history could support an expectation about a distribution when enough comparable data existed. It could not establish which grade a particular student would have achieved on a particular set of papers. Rank order added individual placement, but it was still a centre judgement rather than an external measurement of distance from a grade boundary.
The distinction can be represented as three separate questions. First, what national outcome profile would preserve qualification standards? Second, what grade distribution was plausible for this subject in this centre? Third, what outcome was defensible for this student? Evidence that performs well for the first or second question cannot automatically answer the third.
An accountable system would attach reliability labels to each transition. It would show when a distribution relied heavily on history, when current prior attainment altered that prediction, when CAGs dominated because a cohort was small and when a student's position was close to an allocated boundary. Those labels would not produce certainty, but they would make the system's evidence proportionate to the consequence.
Historical data also contain the effects of earlier educational conditions. Using them does not prove that every historical pattern is unjust, and it does not by itself establish discrimination. It does mean that commissioners should examine whether reproducing a pattern is consistent with the current policy purpose and communicate that dependency plainly.
Small cohorts showed that the model already recognised limits
Statistical models usually become less stable as the relevant sample becomes smaller. Ofqual's approach treated small subject cohorts differently. For very small entries, centre assessment evidence carried decisive weight; for somewhat larger small cohorts, the method tapered between centre judgements and the statistical distribution. This was a substantive governance choice, not a minor implementation detail.
The published Ofqual board minutes for 2020 show repeated emergency meetings through model development, delivery and the August crisis. The record supports institutional chronology and board involvement. It should not be used to infer the unrecorded motive or personal culpability of any board member, minister or official.
Small-cohort treatment demonstrates that decision-makers understood that evidence strength varied. The same principle should extend beyond cohort size. A student whose prior-attainment pattern was atypical, a newly established centre, a subject with unstable historical entries or a centre undergoing rapid change could also weaken the link between group history and an individual outcome.
An exception framework should define these conditions before release. It could trigger targeted review where the model changed a grade by more than a set amount, where historical comparability was poor, where a rank position sat near a model-created boundary or where the centre documented an unusual evidence profile. Reviewers would not simply substitute preference; they would assess whether the model's assumptions held in that case.
The operational challenge was scale. Millions of grades had to be processed across awarding organisations within a fixed timetable. That made universal manual review unrealistic. It did not make risk-based review impossible in principle. A system can prioritise the cases with the weakest evidence or greatest consequence, document what was not reviewed and reserve capacity for rapid remedy.
Model testing showed performance, but the metric did not settle the decision
Ofqual tested alternative models on earlier data, including an exercise using 2019 outcomes as a known comparison. The selected model performed accurately under stated measures and was assessed for protected-characteristic effects. That work was serious and relevant. It is also easy to overread.
The written statement from Ofqual's chair to the Education Committee later explained that the approach had failed to win public confidence even though substantial technical and equalities work had been done. It also located the standardised-teacher-assessment policy in the Secretary of State's direction while accepting collective institutional failure to design an acceptable mechanism. That is a stronger boundary than a story of one autonomous regulator or one minister acting alone.
A measure such as accuracy within plus or minus one grade can be suitable for comparing models. It does not mean that an outcome one grade away is harmless to an individual. A single grade can affect whether an offer condition is met, whether a student enters a course or whether another decision-maker treats the qualification as evidence of readiness.
Back-testing also differs from the live case. Earlier exam years contained actual results and historic centre relationships. Summer 2020 used teacher evidence collected under exceptional conditions and applied a model to a cohort whose exams would never reveal the counterfactual truth. Predictive performance on a previous distribution informs model selection; it cannot validate each live grade after the fact.
Governance should therefore maintain a metric catalogue. Each measure needs a definition, population, tolerance and decision use. Aggregate grade distribution, centre-level error, proportion within one grade, subgroup gap and individual exception rate answer different questions. A dashboard that puts them together without labels can create false assurance.
The approval record should also state the residual risk. If the model is expected to be less reliable for unusual students, leaders must decide whether appeal, manual review or a different grade rule is capable of controlling that risk. Technical validation is complete only when the remaining error can be governed in the real service.
Equalities findings did not certify every individual outcome
The most sensitive evidentiary boundary concerns protected characteristics. Before results, public debate questioned whether standardisation could disadvantage particular groups. Ofqual conducted aggregate analyses, and later work examined student-level data. Those analyses matter. They should be reported accurately, neither dismissed nor expanded beyond their design.
The official student-level equalities analyses concluded that there was no evidence that calculated grades or final grades were systematically biased against candidates with protected characteristics or from disadvantaged backgrounds. The report compared relationships between characteristics and outcomes across approaches and years. It did not claim that every individual grade was correct.
This distinction prevents two opposite errors. It would be wrong to state, against the official analysis, that the model was proven to have produced systematic protected-group bias. It would also be wrong to say that an absence of detected systematic bias proved individual reliability or eliminated all fairness concerns. A subgroup analysis can miss harms that do not align cleanly with the categories tested, and a result can be unreliable for a person without generating a population-level disparity.
The later evaluation of centre assessment grades examined teacher judgements and factors associated with them. That evidence shows why replacing standardisation with CAGs was not a move from a biased machine to a neutral ground truth. Centre judgements also contained uncertainty, variation and possible differences in generosity.
Equalities governance should operate across the whole chain: evidence available to teachers, centre judgement, rank order, historic data, model allocation, awarding-organisation implementation, appeal access and progression consequences. A protected-characteristic check at the model-output stage is necessary but not sufficient to assess the service.
The right conclusion is bounded. The official analyses did not find systematic protected-characteristic or disadvantage bias in the calculated or final grades. Public concern about individual treatment, historic-centre dependence and remedy can still be legitimate. Maintaining both statements is not equivocation; it is faithful evidence governance.
Outliers made individual reliability an explicit known risk
An outlier was not a problematic person. It was a student for whom the model's group assumptions might fit poorly—for example, because prior attainment made the student unusual within a centre. Ofqual examined such cases during and after model development. The existence of that work confirms that individual reliability was a recognised issue, not merely a criticism invented after results day.
Ofqual's research on standardisation and outliers explains that an atypical profile could make a calculated grade unreliable. It also records the later apology and the decision to use the higher of the CAG or standardised grade. The report should not be converted into a claim that all atypical students were wrongly graded.
Outlier analysis needs a defined operating response. Detecting a case is only the first step. The system must determine whether the case is excluded from the model, reviewed with additional evidence, flagged for the centre, or left to appeal. It must also define who can see the flag and whether it is available before a progression decision is made.
There is a risk of circularity. A model based on centre history can identify a student as unusual precisely because the student differs from that history. Treating difference as evidence against the student's potential would make improvement difficult to recognise. Conversely, automatically accepting every exceptional claim would undermine comparability. The control is evidence-based review, not a presumption in either direction.
Individual reliability also has a communication dimension. A student seeing a grade lower than a centre judgement may reasonably ask what evidence changed it. An answer that only describes national standard maintenance does not address the personal decision. The service needs an explanation that connects the student's rank, the allocated centre distribution and any relevant exception rules without exposing other students' personal information.
Grade gaps showed why two plausible estimates could still conflict
The policy reversal created a valuable comparison between CAGs and calculated grades. Later analysis found that, for most A-level entries in its dataset, the two were the same, while a substantial minority differed. Most discrepancies involved a CAG higher than the calculated grade. Those patterns describe disagreement between estimates, not verified model errors.
The later analysis of grading gaps is unusually clear: it is impossible to know, for an individual candidate, whether the CAG or calculated grade better reflected what the student would have achieved had exams taken place. That statement should govern every later interpretation.
The same report found no strong basis for turning grade discrepancies into a universal protected-group causation story. It focused on the incidence of differences and candidate characteristics, using linked data with limitations. Its value is to show where outcomes diverged and how those divergences were distributed, not to recreate missing exam performance.
A data-governance record should retain all three values where applicable: the submitted CAG, the calculated grade and the final awarded grade. They must not be overwritten into one field. The system should also preserve why the final rule selected one value, the date of the change and any later review. Without that lineage, researchers and students cannot distinguish original model output from final qualification outcome.
This is an important lesson for any public model. When a rule changes, the new outcome may be a remedy, a policy compromise or a risk-control choice. It is not necessarily a technical correction. Audit trails must preserve the status of each value rather than allowing the final database state to rewrite the history of the decision.
Exam boards implemented the calculation, so assurance had to cross institutions
Ofqual developed requirements and demonstrator code, while four exam boards were responsible for the production systems that calculated results. Each awarding organisation had its own technology and operational environment. A common method therefore had to be translated into separate implementations while preserving consistent outcomes.
Ofqual later published code used to support grade calculation. The publication contained an important limitation: Ofqual's code was not the final code used by exam boards. It demonstrated how the regulatory requirements could be implemented; awarding organisations produced the final operational code for their systems.
This prevents the convenient but inaccurate phrase “the algorithm” from carrying the whole explanation. There was a policy direction, a regulatory method, common requirements, data supplied by centres, implementation by awarding organisations, quality assurance and release operations. A defect could arise at any interface, but an outcome differing from a CAG is not itself evidence that a coding defect occurred.
Cross-institution assurance requires more than testing each component separately. It needs shared test cases, reconciled output distributions, version identifiers, input validation, exception logs and signed release decisions. If an awarding organisation interprets a rule differently, Ofqual needs a route to detect and resolve the difference before grades reach students.
The annual report and accounts for 2020–21 describes the External Advisory Group, model testing, data-protection work and exam-board quality assurance. It also recognises that the approach was not widely accepted and harmed public trust. These facts can coexist: substantial governance activity occurred, and the end-to-end system still failed its legitimacy test.
Results-day release converted statistical uncertainty into immediate consequences
A model may be evaluated over weeks; a student experiences the result at a particular moment. On 13 August, AS and A-level results entered university admissions, course offers, school decisions and family plans. The system's time horizon narrowed from policy design to urgent individual action.
Ofqual's summer 2020 results analysis later updated the interim data and reported the final outcomes after the policy change. It is evidence about distributions and grade relationships. It should not be used to claim that a national uplift repaired every progression consequence or that every initial result caused a loss.
Release readiness should include consequence testing. Decision-makers need to ask what happens when a student misses an offer, how rapidly a centre can identify a data error, whether universities can hold places, what explanation is available and how the service handles surges in calls. A technically correct batch can still be operationally unsafe if the remedy arrives after the decision it is meant to protect.
Communication also affects legitimacy. Before results day, public explanations emphasised maintaining standards and broad accuracy. Students who received unexpected outcomes wanted individual reasons. If the system cannot answer that question, it should say so directly and explain the alternative evidence and remedy. Assurance language framed only at aggregate level can sound evasive even when statistically correct.
The relevant control is a release gate joining analytical, operational and public criteria. It should require sign-off on output distributions, high-impact exceptions, appeals capacity, stakeholder coordination, communications and the ability to pause. A model with irreversible immediate effects needs stronger readiness evidence than a model used for internal planning.
Appeals were part of the design, but timing and grounds mattered
Ofqual planned appeals for errors and for cases where the standardisation model used inappropriate data for a centre. Students could also sit examinations in the autumn. These were meaningful safeguards, but neither automatically provided an individual, timely reconsideration of the merits of a grade.
The Education Committee's report, Getting the grades they've earned, raised concerns about transparency, accessibility and the burden of proving bias or discrimination. Those were parliamentary conclusions and recommendations made before results, not judicial findings. Their relevance is that the remedy risk was visible in advance.
An appeal must identify who may bring it, on what grounds, using which evidence, within what time and with what effect on a progression decision. A centre-led appeal can protect consistency and privacy, but it may leave a student dependent on the centre whose judgement is also part of the evidence. An autumn exam can offer a fresh measurement, but months later may not preserve an August opportunity.
The distinction between data error and model disagreement is essential. A transposed rank, wrong historical record or implementation error can be corrected. A claim that the model's assumptions do not fit an individual needs a different review. A belief that the CAG is a better estimate is another category again. Combining them under one “appeal” label obscures the decision standard.
The later Government and Ofqual responses to the Committee record how the reversal overtook parts of the planned calculated-grade appeals system. They also preserve separate institutional responses. The record should not be read as proof that every original concern was accepted or that subsequent changes resolved every consequence.
Remedy is part of model validity because a known error profile may be tolerable only if review works at the necessary speed. If the service relies on appeal to catch predictable individual anomalies, appeals capacity, evidence access and decision timing belong in the pre-release assurance case.
The 17 August reversal changed the decision rule, not the historical facts
After A-level results did not command public confidence, Ofqual announced on 17 August that students would receive their CAG or calculated grade, whichever was higher. GCSE results were then issued on that basis. The change reduced uncertainty for students and altered aggregate outcomes.
The statement from Ofqual's chair is the primary public record of the decision. It states the policy problem, acknowledges the absence of an easy solution and sets out the higher-of-two rule. It does not identify a universal coding fault, nor does it assign personal guilt.
The reversal must be classified correctly. It was a policy and regulatory response to legitimacy, operational and individual-outcome concerns. It did not establish that every calculated grade was wrong. Some calculated grades were higher and remained final; some matched the CAG. Nor did it establish that every CAG was the true counterfactual exam result.
Versioned decision records matter here. Before 17 August, the calculated-grade rule controlled. After the announcement, the higher-of-two rule controlled. Data extracts, public statistics, appeals and later research must state which grade status they use. Otherwise the phrase “2020 grade” can refer to a CAG, calculated output or final award without warning.
Political direction and regulator independence also require precision. Ministers set the emergency policy objective and issued statutory directions. Ofqual exercised regulatory and technical responsibilities. Awarding organisations implemented results. The later reversal involved government and Ofqual. Accountability is distributed but not diluted: each institution should explain its own decision, evidence, escalation and sign-off.
Progression effects made the service more than a grading calculation
Qualifications operate inside a network. Universities make conditional offers, schools allocate sixth-form places, employers assess applicants and students decide whether to enter clearing, defer or choose another route. A grade can be statistically ordinary at population level and still be decisive at the point of use.
The House of Commons Library briefing on A-level results and university admissions documents the rapid effect on admissions and the capacity pressures created when final grades changed. It is a parliamentary research synthesis, not a determination of liability in an individual case.
This downstream network should have been represented in the governance model. Each consumer needed contingency rules for delayed, appealed or revised grades. Data exchanges required version control so that a university did not act on an obsolete result. Students needed to know which institution could change which record and whether a place would be held.
Impact measurement should also avoid overclaiming. A changed grade may have affected a decision, but the frozen official evidence does not establish the outcome of every application or attribute every disappointment to model code. Conversely, counting how many students ultimately entered higher education would not show that every initial decision was harmless.
The control objective is recoverability. Before release, the system should identify high-consequence dependencies, establish update channels and test how quickly a corrected or changed rule reaches them. Public models become infrastructure when other institutions rely on their outputs. Infrastructure governance includes the ability to roll forward, not merely calculate.
Transparency after the event was necessary but could not substitute for pre-release challenge
Ofqual published a detailed interim report on results day, later released research, board minutes and demonstrator code, and supported further evaluation. This created an unusually substantial record. Transparency improved retrospective scrutiny and institutional learning.
The Office for Statistics Regulation's review, Ensuring statistical models command public confidence, provides the strongest cross-system assessment. It found that the regulator and awarding teams acted with integrity, while concluding that model governance must be open and trustworthy, rigorous throughout, and directed to public value.
OSR also identified limited human review of individual outputs before results day and an overreliance on aggregate analysis. It observed that equalities analyses did not show widening gaps, while public concern about socioeconomic effects remained. These findings do not certify or invalidate each result. They show a mismatch between assurance performed and the level at which trust was demanded.
Publication timing affects the value of challenge. Releasing hundreds of pages alongside results permits later examination but gives external experts little opportunity to influence deployment. Some pre-release secrecy was intended to prevent gaming or premature estimates. That legitimate concern could have been managed through confidential independent review, staged publication of assumptions or an expert red-team process.
Transparency is therefore more than open-sourcing code. It includes purpose, governance, alternatives, data definitions, known limitations, model cards, equality and individual-reliability evidence, implementation differences, decision minutes, appeal grounds and post-release incidents. Code without these materials can invite false certainty about where the decisive choices occurred.
A stronger accountability design begins with separate assurance claims
Future emergency awarding should begin with an explicit claim map. The first claim concerns continuity: results can be produced in time to support progression. The second concerns standard maintenance: distributions remain interpretable against earlier years. The third concerns centre consistency. The fourth concerns protected-characteristic effects. The fifth concerns individual reliability. The sixth concerns remedy. Each needs its own evidence and owner.
The model should then have a purpose boundary. If its strongest evidence supports group distributions, that should be stated. It should not be represented as an individual prediction engine without individual validation. Where a final grade joins model output and professional rank, the documentation should describe the contribution and uncertainty of each.
Data lineage should preserve CAG, rank, centre-subject cohort, historical years, prior-attainment inputs, model version, awarding-organisation implementation, calculated grade, exception flags, appeal status and final grade. Access controls must protect students and peers, but privacy should not become an excuse for an unexplainable decision.
Risk-based human review should be designed before deployment. Triggers could include large deviations from CAG, poor historical comparability, outlier indicators, data-quality warnings and especially consequential boundaries. Review should test assumptions rather than re-run the same formula. Its scale and limitations should be reported.
Commissioners should conduct acceptability testing with students, centres and downstream users. This is not a popularity vote on arithmetic. It tests whether affected people understand the purpose, view the evidence relationship as legitimate and can use the remedy. A technically sound method that cannot be accepted may still fail the policy need.
The release gate should be independent enough to stop or modify deployment. It should review the full chain, including ministerial constraints, regulator method, exam-board implementation, service capacity and downstream coordination. No single assurance score should average away a failed individual-remedy gate.
Institutional learning requires precise responsibility rather than a search for one culprit
The 2020 system was shaped by an unprecedented emergency, limited data, compressed time and competing objectives. Those conditions explain difficulty; they do not remove accountability. Equally, a serious outcome does not justify personal attribution unsupported by official findings.
Responsibility should be recorded by decision. Ministers owned cancellation and policy direction. Ofqual owned regulatory design, model requirements, oversight and public assurance within its remit. Awarding organisations owned their implementations and operational calculations. Centres owned submitted judgements and rank orders under the specified process. Downstream institutions owned how they handled changing grades within their authority.
This allocation allows learning without pretending that every actor controlled the whole system. It also prevents the word “algorithm” from shielding policy choices. A model does not decide its objective, choose which harms count, set an appeal timetable or determine when to reverse policy. People and institutions do.
Evidence boundaries must remain durable. The equalities analyses did not show systematic protected-group or disadvantage bias; they did not prove individual accuracy. The reversal changed the awarding rule; it did not label every initial grade a software defect. Board records show governance activity; they do not establish personal motive. Parliamentary criticism and OSR lessons are authoritative within their scopes, not findings of individual wrongdoing.
The final test is operating capability. Before another emergency, the system should be able to show who can authorise an alternative, what evidence supports each claim, how individual exceptions are found, how every implementation is reconciled, how a student can obtain timely review and how downstream users receive an updated outcome. That is what converts a lesson into resilience.
Conclusion
Ofqual's 2020 grading model was not simply a failed piece of code. It was one component of a public decision system created after examinations—and the ground truth they normally supplied—disappeared. The system had to reconcile progression, standards, centre consistency, equality, delivery speed and individual fairness with evidence that could not establish every counterfactual result.
The model and associated analyses could support claims about distributions, centres and protected-characteristic patterns. They could not remove uncertainty about an individual. Public confidence broke where the system's strongest aggregate assurances met a student's immediate need for a personally defensible outcome and usable remedy.
The accountable response is not to reject statistical models or to declare centre judgement infallible. It is to govern the whole chain: ministerial purpose, regulator design, human inputs, historical data, model assumptions, awarding-organisation implementation, individual exception review, release consequences, appeals, reversal and learning. Each claim must stay at the level its evidence can support.
That boundary is the enduring lesson. A public model earns legitimacy not because its national totals look plausible, nor because its code runs as specified, but because decision-makers can explain what it knows, recognise what it cannot know and provide an effective path when an individual bears the cost of that uncertainty.

