Summary
- England’s NHS Test and Trace programme was created at emergency speed and assembled laboratories, logistics, digital systems, national contact centres and local public-health teams into a service whose remit kept changing as the pandemic changed.
- The accountability problem was not merely that large sums were authorised. It was whether decision-makers could connect procurement choices, contract outputs and operational activity to the public-health outcomes that justified emergency expenditure.
- Official records require several distinctions: budget is not expenditure; a direct award is not inherently unlawful; a late publication breach does not prove a contract was improperly awarded; a distributed test is not necessarily a used test; and an operational metric is not causal proof of infections prevented.
- National Audit Office and parliamentary findings identified weak business-case documentation, unused capacity, incomplete end-to-end performance visibility, data and local-integration problems, reliance on external advisers, and later inventory and financial-control difficulties.
- The programme also delivered extraordinary scale. Accountability therefore cannot be reduced to a verdict that speed was either wholly justified or wholly wasteful. The more useful question is which controls could have preserved emergency speed while producing timely, decision-grade evidence.
- The durable model is a traceable chain from emergency need to award rationale, deliverable, service use, public behaviour, outcome hypothesis, uncertainty, asset disposition and exit. Without that chain, spending totals and activity dashboards can coexist with weak public assurance.
1. The accountability test was an evidence-chain problem
NHS Test and Trace was announced in May 2020, when England needed testing and tracing capacity far beyond the system that existed before the pandemic. It had to mobilise laboratories, sample transport, sites, call handlers, data platforms, digital products, consultants and local public-health relationships while scientific knowledge and policy requirements were moving. That context matters. A conventional procurement timetable could not simply be imposed on an emergency whose costs were being measured in transmission, illness and disruption.
But urgency did not remove the need for accountability. It changed the form accountability needed to take. The central question became whether the Department of Health and Social Care could preserve a usable chain from a public-health need to a procurement decision, from that decision to a delivered capability, and from the capability to a measurable contribution to control of infection.
The National Audit Office’s interim report described a service built from scratch at great speed but found that not all objectives had been achieved and that the initial delivery model had not been supported by a documented business case until September 2020. It also examined testing and tracing capacity, use and performance rather than treating procurement volume as an outcome in itself. That is the right analytical starting point: the interim report is evidence of a control challenge under emergency conditions, not a claim that scale-up had no value.
An evidence chain has at least eight links: the emergency condition; the objective; the chosen delivery model; the award and price rationale; the contractual output; actual use; the behavioural and epidemiological mechanism; and the outcome, with uncertainty recorded at every step. If one link is absent, an impressive count at another link can mislead. Contract value does not show service use. Tests processed do not show how quickly a person changed behaviour. Contacts reached do not show whether they isolated. A fall in infections does not, by itself, isolate the contribution of one programme.
2. Budget, expenditure, commitment and value are different numbers
The programme’s public debate was frequently organised around a headline allocation of tens of billions of pounds. That headline was politically understandable but analytically blunt. An authorised budget sets a spending ceiling or planning envelope. Expenditure records resources consumed during a period. A contractual commitment may span years or may not be fully drawn. Inventory purchases create assets before those assets are consumed, written down or disposed of. An underspend can reflect over-budgeting, changed demand, delayed delivery, cancellation or cost avoidance; it is not automatically a saving and not automatically a failure.
The National Audit Office’s progress update reported that NHS Test and Trace spent £13.5 billion in 2020-21 against a budget of £22.2 billion, an £8.7 billion underspend. Of the expenditure it described, £10.4 billion related to testing, £1.8 billion to support for local authorities managing outbreaks and £0.9 billion to tracing. The same report found improved performance in parts of the service but said the government still had more to do to demonstrate an effective and integrated system. It was a progress review rather than a complete value-for-money conclusion. Those boundaries in the NAO update are essential.
A responsible evaluation should therefore maintain a reconciliation table. It should show the approved budget, revised forecast, cash expenditure, accruals, contract ceilings, orders placed, inventory acquired, inventory consumed, write-downs, cancellation liabilities and transfers to successor bodies. Each number should have a period, scope and accounting basis. Without that table, critics can describe the full budget as if it were spent, while defenders can describe an underspend as if it proves efficiency.
The accountability issue is not semantic. If public managers cannot reconcile these categories, they cannot know whether capacity was deliberately reserved, accidentally idle, no longer required or paid for without usable output. Nor can Parliament compare alternatives. Precise financial language is part of operational control.
3. Emergency direct awards were a lawful route, not a blank cheque
Public procurement law can permit contracting authorities to act without a conventional competition when extreme urgency caused by unforeseeable events makes ordinary timelines impracticable. The existence of that route is not evidence of wrongdoing. During the first pandemic wave, speed and scarce global supply created circumstances in which direct awards could be defensible. The governance question is whether each use of the exception was proportionate, documented, time-limited and revisited when conditions changed.
The Public Accounts Committee’s first report considered the early operating model, use of private suppliers, tracing capacity and performance. It reported that a large share of early contract value had been awarded directly under emergency measures and questioned whether the programme had demonstrated an effective balance between national and local capability. The committee also highlighted underused call-handler capacity and weak evidence of overall effectiveness. Those are parliamentary findings and judgments contained in the committee’s report; they should not be recast as a judicial determination that emergency awards were illegal.
The control design for a direct award should be more explicit, not less. A short emergency award record can state the precise urgency, why the requirement could not be delayed, the suppliers considered, available price benchmarks, conflicts checked, the deliverable, acceptance evidence, termination rights, publication deadline and the date on which competition will again be feasible. The record need not be a long conventional business case. It must be contemporaneous enough that a later reviewer can distinguish a reasoned emergency choice from retrospective justification.
4. Publication duty and award propriety must not be collapsed
Transparency is a separate control from the legal basis for choosing a supplier. In litigation brought by the Good Law Project, the High Court considered the Secretary of State’s failure to publish contract award notices within the required period. The judgment recorded that there was no dispute that the legal obligation to publish within 30 days had been breached in many cases. That finding matters because delayed publication prevents timely scrutiny of suppliers, values, scopes and amendments.
The same judgment also sets an important limit. The litigation was not a general adjudication on whether each individual direct award was substantively improper, nor did it establish that every contract made under emergency procedures was corrupt or poor value. The High Court judgment supports a precise conclusion: compliance with transparency deadlines was materially deficient. It does not support a blanket conclusion about every underlying award.
The distinction reveals why procurement assurance needs multiple ledgers. One ledger records legal authority and award rationale. A second records publication and redaction decisions. A third records performance, payments and changes. A fourth records conflicts, gifts and supplier relationships. Passing one control does not compensate for failing another. A legally defensible direct award can still be published late; a promptly published notice can still conceal weak performance management.
For an emergency programme, publication timeliness is itself an operational metric. Departments should monitor the percentage of notices published within the legal deadline, the age of overdue notices, amendment publication, and reconciliation between finance systems and the public contract register. A dashboard of tests and contacts that omits procurement disclosure creates a partial picture of accountability. Transparency cannot wait until the emergency is over, because the opportunity to challenge extensions, duplication and changing requirements exists while the programme is running.
5. Contract notices show purchased scope, not delivered public value
Individual contract notices are useful because they make the programme’s component parts concrete. A notice for Boston Consulting Group described digital, technology and programme support with a stated value of £4,996,056 under a call-off arrangement. The published notice is evidence that a defined service was procured through a stated route at a stated value.
It is not evidence, on its own, that every deliverable was accepted, that the price was competitive, that the work was duplicated, or that a particular epidemiological outcome followed. Those questions require the contract schedule, change control, invoices, deliverable acceptance, dependency map and outcome evaluation. Treating a notice as proof of either success or scandal asks the document to do a job it was not designed to do.
This distinction should shape the public contract register. A notice is the first layer. The operational layer should identify the accountable senior owner, service component, start and end dates, deliverables, milestone status, payments, amendments, dependencies, data handled, intellectual-property position, knowledge-transfer requirement and exit plan. The evaluation layer should connect the component to a testable service hypothesis. For digital programme support, that might be a reduction in deployment lead time, an improvement in data completeness or removal of a known bottleneck. The indicator must be defined before the result is known.
Consultancy can be rational during a surge. External teams can add scarce expertise and delivery capacity faster than permanent recruitment. The accountability risk arises when an authority buys named inputs without a measurable transfer of capability, permits rolling extensions without a market test, or cannot identify which internal team will own the work after departure. The relevant question is not “consultants or no consultants.” It is whether temporary external capacity was governed as a bridge to a sustainable operating model.
6. Digital products required product controls as well as procurement controls
A separate notice for Zühlke Engineering described work associated with the NHS COVID-19 app and recorded a contract value of £10,255,842 under a call-off arrangement. The contract notice establishes a procurement event and a scope category. It does not show how much of the app’s adoption, reliability or public-health effect can be attributed to that supplier.
Digital products in an emergency need an assurance model that bridges contract management and product management. Contract milestones can be met while the product still fails users. Code can be delivered while integration, accessibility, security or operational support remains incomplete. Conversely, a product may be valuable even when direct epidemiological attribution is difficult.
Managers therefore need layered measures: release reliability, defect severity, accessibility conformance, latency, uptime, security findings, user adoption, notification completion, integration quality and the behavioural pathway through which the product is expected to help.
The product backlog should also expose the ownership boundary. Which requirements were set by policy? Which were constrained by public-health guidance, privacy design or operating-system interfaces? Which decisions belonged to the supplier, the department or another public body? When responsibility is distributed, a generic claim that “the app failed” or “the app succeeded” obscures the decisions that can actually be improved.
Exit evidence is equally important. Source-code access, documentation, architecture decisions, automated tests, deployment pipelines, incident history and support knowledge must be transferable. A department that receives working software without the ability to operate or change it has purchased dependency. Procurement accountability for a digital service is therefore not complete at acceptance. It ends only when ownership, continuity and the future cost of change are visible.
7. Research and advisory outputs need a decision trail
McKinsey’s published notice described “Voice of the Customer” support and a stated value of £463,200 under a call-off contract. The notice demonstrates that advisory research formed part of the programme’s commercial ecosystem. It does not reveal whether a recommendation changed policy, improved service design or merely restated information available elsewhere.
Advisory work should be governed through a decision-use register. For every commissioned output, the authority can record the question asked, decision deadline, method, evidence base, deliverable, reviewing official, recommendation, decision taken, reason for acceptance or rejection, and follow-up result. This is especially important for customer research. A polished presentation can be delivered exactly to contract while sampling, interpretation or timing makes it operationally irrelevant.
Public reporting need not expose personal data or commercially sensitive material. It can still describe the decision problem, methodology, top-level conclusion and implementation status. The purpose is not to force every adviser’s paper into the public domain. It is to make visible whether commissioned knowledge entered a real decision process. Procurement of advice becomes accountable when the authority can show not only what document arrived but what decision it was meant to improve.
8. Integration assurance is a service property, not a supplier label
A Capgemini contract notice described quality assurance and integration support with a stated value of £4.5 million under a framework call-off. The notice helps identify a purchased integration function. It cannot establish that the entire end-to-end service was integrated, because integration depends on interfaces among laboratories, booking systems, couriers, results channels, contact tracing, local authorities and users.
An end-to-end assurance map should name each handoff, data owner, service-level expectation, failure mode and recovery path. For example, a laboratory turnaround target may be met from sample receipt to result, while the person’s full journey from symptom onset to isolation remains too slow. A tracing centre may reach a high share of contacts whose details it receives, while delays or incomplete records upstream reduce public-health value. Component performance can look strong even when the pathway is weak.
Supplier accountability follows the map. A contract should specify interface obligations, evidence needed for acceptance and the process for resolving shared failures. Without that structure, each supplier can satisfy a local specification while the public experiences a broken service. Integration governance is the discipline that turns a portfolio of purchased components into an accountable public system.
9. Capacity was an option whose value depended on use and flexibility
Emergency services often have to buy capacity before demand is known. Idle capacity is not automatically waste: a fire service is not judged only by the percentage of every day spent fighting fires. Reserved laboratory slots and call handlers can have option value when demand is volatile and delay is costly. Yet option value must be explicit. Otherwise, managers cannot distinguish prudent resilience from a poorly designed fixed commitment.
The Public Accounts Committee’s first Test and Trace inquiry examined the national tracing model, including the recruitment of thousands of health professionals and call handlers and very low reported utilisation during part of June 2020. Its inquiry publications page brings together the report, government response and oral and written evidence. The committee’s concern was not merely that utilisation fell below 100 per cent; it was that the programme could not convincingly show how purchased capacity related to demand, performance and the developing role of local teams.
A capacity contract should define three states. Baseline capacity serves expected demand. Surge capacity can be activated within a specified time. Strategic reserve is intentionally idle but available against severe scenarios. Price, training, mobilisation and release terms should differ across those states. If all capacity is bought as a fixed block, the authority bears the full forecast risk. If everything is bought on demand, the supplier may be unable to respond during a national surge.
Utilisation also needs the right denominator. Logged-in time, handled cases, successful contacts and productive public-health actions are different measures. A low percentage can arise because cases were lower than forecast, data arrived late, shifts were mismatched, workflows failed or reserve capacity was deliberately maintained. Each explanation implies a different remedy. A dashboard should link demand forecast, rostered capacity, available capacity, productive work, exception time and unit cost. This makes resilience auditable without pretending that the optimal target is constant maximum use.
10. Operational statistics were essential, but they were not an impact evaluation
NHS Test and Trace published extensive weekly statistics on tests, cases transferred to tracing, people reached, contacts identified and turnaround times. Those releases supported public scrutiny and operational management. Their value should not be understated. In an emergency, consistent time series can reveal bottlenecks, regional variation and changes in service performance.
The official methodology, however, explains that the statistics were produced from operational systems and that definitions and processes changed. It also notes that routine contact tracing ended on 24 February 2022. The methodology is therefore part of the evidence, not a technical appendix to ignore. It sets the population, exclusions, revisions and limitations within which the numbers can be interpreted.
An operational measure answers “what moved through the service?” An outcome measure asks “what changed for people or public health?” An impact evaluation asks “what changed because of this service compared with a credible alternative?” The three questions overlap but are not interchangeable. A high contact-reach rate can coexist with slow case transfer. Fast notification can coexist with low adherence to isolation. A rise in testing can reflect a change in prevalence, policy or eligibility rather than service effectiveness.
Good dashboards preserve these layers. The first layer reports volume, timeliness, quality and equity. The second follows the behavioural sequence: receipt of result, disclosure of contacts, successful notification and support for isolation. The third estimates outcomes such as transmission interrupted, while showing assumptions and uncertainty. Governance fails when a ministerial target from the first layer is presented as proof from the third.
Method changes should be versioned, with comparable series identified and breaks clearly marked. Denominators, missing data and revisions should be visible. That discipline makes it harder to select a favourable metric after the fact and easier for decision-makers to understand whether an apparent improvement reflects the service, the case mix or the measurement system.
11. A weekly release can demonstrate performance without proving effectiveness
The weekly release covering 31 December 2020 to 6 January 2021 reported detailed test and tracing activity during a period of intense winter transmission. The release illustrates both the strength and the limit of contemporaneous reporting: managers and the public could inspect concrete turnaround and reach measures, but the release was not designed to estimate the counterfactual number of infections prevented.
The distinction matters because the programme operated inside a changing system. Legal restrictions, public caution, vaccination, variants, school policy, workplace practices, prevalence and support for isolation all changed the likelihood that a test or tracing call would alter transmission. A person reached quickly might already have isolated; a person reached slowly might nevertheless avoid contacts. Positive cases not transferred, contacts not named and people unable to comply all sit outside a simple reach percentage.
An outcome architecture should begin with a causal diagram. Testing can shorten the time from infection risk to knowledge. Notification can change behaviour. Support can make isolation feasible. Those behavioural changes can reduce exposure, which can reduce onward transmission. Each arrow needs an indicator and a known uncertainty. Randomised or quasi-experimental methods may be feasible for particular interventions, but not every system-wide question has a clean experimental design. Where causal identification is weak, the programme should say so and use sensitivity ranges rather than a single heroic estimate.
Operational releases should also serve local management. National averages can conceal laboratory routes, regions or groups with materially different experiences. Publishing distributions and high-percentile delays is often more informative than a mean. A service that is fast for most users but extremely slow for a critical minority can still fail its public-health purpose. Measurement becomes accountable when it helps managers find and change those failure modes, not merely defend an aggregate.
12. National scale and local knowledge had to be designed as one system
Central mobilisation offered purchasing power, standardised platforms and the ability to create nationwide capacity. Local directors of public health and their teams had contextual knowledge, established relationships and the ability to pursue complex cases. Treating these as rival delivery philosophies misses the governance problem. The programme needed a designed handoff between national scale and local action.
The Public Accounts Committee’s later Test and Trace inquiry returned to performance, contracts, local integration and the programme’s plans. Its publication record includes the committee report, government response, correspondence and evidence. The committee questioned whether the programme had learned enough, reduced reliance on consultants and demonstrated the effectiveness expected from its scale. These are parliamentary conclusions that should be considered alongside, not substituted for, the underlying operational evidence.
Local integration requires more than passing a spreadsheet down a chain. The operating model should define when a case moves from national to local handling, the minimum data package, lawful data access, response times, ownership of failed contacts, feedback on disposition and the route for outbreak escalation. Both sides need a shared case identifier and an auditable status. Otherwise, people can be counted twice, disappear between systems or receive inconsistent contact.
Funding design also affects accountability. Short, uncertain grants make it difficult for local teams to retain trained staff or build durable analytical capacity. Conversely, devolving funds without common service definitions can make national comparison impossible. A balanced model establishes minimum national data and service standards while allowing local teams to adapt channels and interventions. It evaluates local success against population need rather than imposing one operational template.
The same lesson applies to suppliers. A national contractor should not be rewarded only for handing off a record; the quality and timeliness of the handoff matter. Shared service-level objectives can prevent each organisation from optimising its own queue at the expense of the full pathway.
13. Test distribution, registration, use and inventory were separate facts
Mass availability of lateral-flow devices was intended to identify infections among people who might otherwise not know they were infectious. Procurement and distribution at scale could therefore have public-health option value even when not every test was used immediately. At the same time, very large flows of physical stock created forecast, storage, shelf-life and accounting risks.
The NAO reported that by 26 May 2021, 691 million lateral-flow devices had been distributed in England while results for 96 million—about 14 per cent—had been registered. That gap was an accountability signal, but it was not a direct count of 595 million unused tests. Some tests could have been used without registration, held in inventory, passed through organisational reporting routes or lost to follow-up. The correct conclusion is that the evidence system did not provide a complete bridge from distribution to registered use and outcome.
Control should begin with a stock-flow model: contracted, ordered, received, quality-released, stored, dispatched, received by channel, issued to user, result registered, returned, expired, damaged, written down and disposed. Reconciliations should be performed by batch and location where feasible. Forecast scenarios should record demand assumptions, policy triggers and shelf life. When policy changes, managers should rapidly identify cancellable orders, redeployment options and residual liabilities.
The programme set out its later-year approach in the NHS Test and Trace programme for 2021 to 2022, including continued testing and tracing activity as the pandemic evolved. Its published cost descriptions should be read by period and scope, especially during the Omicron response, rather than added casually to earlier budget headlines.
Inventory accountability is not a demand for perfect foresight. A pandemic stock decision can look excessive after a favourable scenario and inadequate after an adverse one. The standard is whether assumptions, expiry risk, option value and decision gates were documented; whether stock was visible; and whether managers adapted commitments as uncertainty resolved.
14. Financial accounts exposed a transfer-of-evidence problem
The creation of the UK Health Security Agency changed organisational responsibility while the programme was still operating. Assets, liabilities, contracts, staff, data and control evidence had to move from the Department of Health and Social Care and predecessor arrangements into a new body. Such a transition is not administrative housekeeping. It is a high-risk event in the accountability chain.
DHSC’s 2021-22 annual report and accounts described the transfer of Test and Trace balances, including inventory and accruals, and the audit difficulties associated with evidence supporting those amounts. The departmental accounts provide accounting evidence within a defined reporting framework. They should not be translated into a claim that all transferred stock was missing, unusable or fraudulently acquired.
The control failure to avoid is “responsibility transferred, evidence left behind.” Every balance should have a transfer pack containing ownership, valuation basis, location, batch detail, age, condition, restrictions, source-system reconciliation, contract link and named accepting officer. Accruals need the underlying goods or services, acceptance status, invoice expectation and estimation method. Data assets need lawful basis, quality profile, retention rule and operational owner.
A joint transition board should track exceptions until both organisations agree resolution. Unresolved items should remain visible rather than being absorbed into a general opening balance. Internal audit can sample the chain before legal transfer. Contract novations and delegations should be reconciled to payment authority so that a supplier is neither paid twice nor left without a responsible counterparty.
The wider lesson is that machinery-of-government changes do not reset accountability. A successor agency inherits not only operational capability but also the need to explain what was bought, held, consumed and owed. If evidence architecture is weak before transition, reorganisation can make reconstruction more difficult precisely when scrutiny increases.
15. Audit qualification was a boundary on assurance, not a verdict on every item
UKHSA’s first annual report and accounts recorded major audit limitations concerning Test and Trace inventory transferred into the agency and inventory consumption. The Comptroller and Auditor General could not obtain sufficient appropriate evidence over specified balances, including £794 million of transferred inventory and £3.305 billion of consumption in the relevant accounts. UKHSA’s 2021-22 annual report also placed those figures in the context of establishing a new agency and inheriting programme responsibilities.
An inability to obtain sufficient audit evidence is serious. It means the auditor cannot provide the intended level of assurance over the affected balances. But it does not, by itself, prove that the entire amount was wasted, stolen or physically absent. Audit language is calibrated. Accountability analysis should preserve that calibration because exaggeration makes it harder to identify the specific control that failed.
The likely governance focus is inventory traceability and evidence retention across systems and organisations: what was received, where it was held, how consumption was recorded, what estimates were used, and whether references supported the ledger. Remediation should therefore be transaction- and batch-level. Reconciliations, warehouse confirmations, consumption-method validation, exception sampling and clear write-off authority can improve assurance. A general promise to strengthen controls is not enough.
Audit findings also need operational owners and dates. The finance team cannot reconstruct physical inventory alone. Commercial teams hold contract and order records; logistics providers hold warehouse and movement data; programme teams know distribution channels; technology teams understand interfaces; and local bodies may hold evidence of receipt or use. A remediation plan should map each evidence gap to its source and to a senior responsible owner.
The distinction between an evidence gap and a proven loss is more than defensive language. It supports proportionate action. Suspected fraud requires investigation. Obsolete stock requires disposal governance. Unsupported accounting estimates require audit remediation. Conflating them can send resources toward the wrong control.
16. Losses and special payments needed programme-level visibility
UKHSA’s advisory board finance report in September 2022 discussed financial-control challenges, the anticipated audit position and losses and special payments associated with Test and Trace activity. The finance report is valuable because it shows governance bodies considering these issues during the agency’s early operation, rather than only after final accounts appeared.
Losses and special payments are not one homogeneous category. They can include obsolete or damaged stock, cancelled commitments, fruitless payments, compensation and other transactions that require specific approval and disclosure. Each category has a different root cause and preventive control. Aggregating them into a large headline may attract attention, but it does not tell managers what to change.
The programme should maintain a loss-event register linked to contracts, inventory batches and policy decisions. Each event should record amount, accounting category, cause, approval authority, recoverability, supplier action, lessons and control change. Trends matter: repeated small write-offs from the same process can reveal a systemic problem before a major loss occurs. Near misses should also be captured, because a control that narrowly prevented payment or expiry may fail next time.
Governance boards should see both gross and net exposure. Expected supplier credits, insurance recoveries, redeployment and residual asset value should be identified without being recognised prematurely. Decisions made under uncertainty should be evaluated against the information available at the time, while failure to update a decision after new information emerged should be treated separately.
17. Later parliamentary scrutiny focused on institutional control
In 2023, the Public Accounts Committee examined UKHSA’s governance and financial management. It highlighted the agency’s weak initial controls and the inability to properly verify large Test and Trace inventory-consumption figures transferred into its accounts. The committee’s findings show that the accountability question outlived the operational peak of the programme.
Again, attribution matters. These are the committee’s conclusions based on the audit and evidence it considered. They are not a criminal finding, and they do not mean every UKHSA function was uncontrolled. Their significance is institutional: emergency delivery and organisational transition had produced evidence and governance weaknesses serious enough to constrain public assurance.
The institutional remedy should combine formal governance with operational evidence. A board and committees are necessary but cannot compensate for ledgers that do not reconcile. Conversely, a detailed ledger without clear decision rights can accumulate unresolved exceptions. The agency needs named ownership for financial reporting, inventory, commercial assurance, data quality and programme outcomes, plus a route for cross-functional risks that no single team can resolve.
Institutional legitimacy depends on the ability to state uncertainty plainly. A public body gains credibility when it identifies what is known, what cannot be verified, what is being reconstructed and when a decision will be made. Overclaiming certainty after an audit limitation weakens trust; treating the limitation as proof of total failure weakens analysis.
18. Winding down required an accountable exit, not just a lower run rate
Routine contact tracing ended in England in February 2022 as policy shifted and the programme moved away from the universal service built during the emergency. Winding down a large temporary system creates its own procurement risks: termination charges, stranded inventory, data retention, supplier knowledge loss, staff transition, unresolved invoices and unclear ownership of capabilities retained for future outbreaks.
DHSC’s 2020-21 annual report and accounts documented the programme in its earlier expansion phase and placed Test and Trace spending within departmental financial reporting. The accounts provide a baseline against which later transfers and closure decisions can be reconciled. A complete lifecycle record should bridge that expansion evidence to subsequent UKHSA accounts rather than treating each reporting year as an isolated story.
Every temporary contract should have an exit schedule before extension is approved. The schedule should state notice periods, data return or deletion, asset ownership, licences, staff dependencies, final acceptance, open claims, knowledge transfer and the retained surge requirement. Contracts that support essential capability may need to move to a competitively procured standby model rather than simply expire. Others should close promptly when demand changes.
Inventory exit needs decision gates based on shelf life and plausible scenarios. Stock can be retained, redeployed, donated where lawful and safe, sold, recycled or destroyed. Each route requires evidence of authority, condition and value. A late decision can convert useful option value into avoidable disposal cost.
The programme should also preserve a reactivation package: architecture, supplier market map, tested specifications, data standards, legal templates, lessons, training materials and scenario-based capacity plans. The goal is not to freeze the emergency model. It is to ensure that a future response can mobilise faster without repeating the same evidence gaps. Accountability reaches maturity when closure produces reusable capability as well as a final set of accounts.
19. The effect on continuity extended beyond the programme boundary
Testing and tracing were public-health services, but their continuity effects reached employers, schools, care settings, supply chains and small businesses. A timely and credible result could influence whether people attended work, isolated, sought support or warned contacts. Delayed or confusing journeys could impose uncertainty on households and organisations. Those mechanisms are plausible and important, but programme-wide evidence does not justify assigning every period of economic disruption to Test and Trace performance.
For small and medium-sized enterprises, the relevant accountability questions include accessibility, turnaround distribution, clarity of guidance, availability of support and predictability of changes. A national average can conceal the operational reality of a small employer whose staff cannot work remotely. Service continuity should therefore be evaluated by user segment and vulnerability, not only aggregate throughput.
The correct control response is not to load every social outcome into each supplier contract. It is to define which service properties the contract can influence and how those properties contribute to wider outcomes. A laboratory can be held to sample quality and turnaround; a logistics provider to collection integrity and delivery time; a call centre to timely, accurate and accessible contact; the public authority to the integrated pathway and policy design.
This allocation of accountability prevents two opposite errors. Suppliers should not be blamed for outcomes outside their control, and the authority should not hide behind fragmented contracts when no one owns the end-to-end effect.
20. A stronger emergency-procurement operating model
The evidence suggests a practical control model that can preserve speed. First, create a contemporaneous emergency decision record for every exceptional award. It states need, urgency, route, supplier selection, price evidence, conflicts, deliverables, publication deadline and review date. Second, maintain a public-facing contract register reconciled to finance, with amendments, payment status and accountable owners.
Third, classify purchased capacity as baseline, surge or reserve and price it accordingly. Demand forecasts should be versioned, with scenarios and release gates. Fourth, map the full service journey and instrument every material handoff. Contract acceptance should test interfaces and complete user journeys, not only local components.
Fifth, separate output, service outcome and public-health impact. Dashboards should disclose definitions, denominators, missingness, revisions and causal limits. Evaluation plans should be agreed early enough to collect comparison data. Where a credible counterfactual is unavailable, managers should report ranges and contribution evidence rather than claim precise causality.
Sixth, establish a physical-and-financial asset ledger. Orders, receipts, locations, issues, results, expiry, write-down and disposal should reconcile by batch where proportionate. Seventh, require a transition pack for every transferred balance, contract, dataset and capability. An accepting officer must own each unresolved exception.
Eighth, govern external expertise as temporary capability. Every advisory or digital engagement needs decision use, internal ownership, documentation, knowledge transfer and an exit test. Ninth, use shared objectives for national and local teams so that handoffs do not become accountability gaps. Finally, track loss events and near misses to verified control changes.
These controls do not require a return to peacetime bureaucracy during a crisis. Most can be expressed as short structured records and automated reconciliations. Their value is cumulative: they allow leaders to make fast decisions while preserving the evidence needed to change course, challenge suppliers, transfer responsibility and explain results.
21. The durable lesson
NHS Test and Trace demonstrated that a state can assemble enormous operational capacity in a short time. It also demonstrated that scale, expenditure and activity are not self-interpreting. Public accountability depends on the evidence connecting them.
The fairest conclusion is neither that emergency procurement was intrinsically improper nor that emergency conditions excused later control weaknesses. Direct awards could be lawful and necessary. Consultants could supply scarce capacity. Reserved resources could have insurance value. Mass testing could create options under uncertainty. Each proposition still required documentation, performance evidence, review gates and an accountable exit.
Official findings also need disciplined boundaries. A late publication breach is not proof that every award was unlawful. A committee criticism is not a court judgment. Unregistered tests are not all proven unused. An audit limitation is not proof that every pound was lost. A budget is not expenditure. Preserving those distinctions does not minimise accountability; it makes accountability credible.
For future emergencies, the decisive capability will be institutional memory encoded in systems rather than remembered by individuals. A government should be able to retrieve the reason for an award, the contract version, delivered output, payment, service use, outcome hypothesis, uncertainty, asset location, transfer record and exit decision. That chain should survive supplier turnover, departmental reorganisation and the end of the crisis.
The programme became a public-procurement accountability test because the public needed two things simultaneously: rapid protection and evidence that extraordinary powers and resources were being governed. Emergency speed and rigorous measurement are not opposing values. When evidence is designed into the operating model, it is what allows speed to remain legitimate.

