Summary
- On 18 March 2018, a developmental Uber automated driving system controlling a modified Volvo XC90 struck and killed Elaine Herzberg in Tempe, Arizona. The system first detected her 5.6 seconds before impact but repeatedly changed its classification, discarded useful tracking history after reclassification, and did not determine that she was in the vehicle's path until 1.2 seconds before impact.
- The sole vehicle operator was streaming a television programme on a personal phone and looked away from the road for about five of the final six seconds. The National Transportation Safety Board determined that her failure to monitor was the probable cause. It also found that Uber ATG's inadequate risk assessment, ineffective operator oversight and failure to manage automation complacency contributed, all consequences of an inadequate safety culture.
- The human was not merely an independent cause outside the system. Uber ATG relied on the operator as the primary countermeasure when an emergency exceeded automated braking specifications, while Volvo forward-collision warning and automatic emergency braking were disabled in autonomous mode. About five months before the crash, ATG had consolidated two operators' responsibilities into one role and rarely reviewed inward-facing video that could have exposed distraction.
- Arizona had welcomed autonomous testing without a special permit. The NTSB found that state oversight was limited public evidence and recommended a safety-focused application and expert review before testing. Later Arizona law created notification, certification, enforcement and law-enforcement interaction provisions, but the visible framework still depends substantially on operator or company attestations rather than publication of a case-specific independent safety approval.
- Repair evidence is material but bounded. ATG restored two operators, added real-time attention monitoring, changed path prediction, removed the one-second action-suppression logic, re-enabled compatible Volvo collision warnings and emergency braking, reduced public-road testing speed and built a safety management system. The NTSB closed its recommendation to the former ATG in 2024, after the unit had been sold to Aurora. Closure does not reproduce every validation result or prove how the original organisation would perform today.
- Important governance gaps remained visible as of 15 July 2026. NHTSA's crash-reporting order supplies mandatory incident data and enforcement leverage, but the NTSB recommendations to require and evaluate pre-test safety self-assessments remain on NHTSA's open list. The durable test is therefore whether a developer can prove the fallback chain before exposing the public: independent emergency control, realistic human-performance assumptions, active attention safeguards, controlled change approval, state and federal review, and auditable evidence from scenarios that stress all layers together.
The driver-versus-system framing fails the control test
The NTSB investigation page states the official causal conclusion plainly. The probable cause was the vehicle operator's failure to monitor the driving environment and the automated driving system because she was visually distracted by her personal phone. The same conclusion identifies contributing factors within Uber ATG: inadequate safety risk assessment, ineffective oversight of operators and no adequate mechanism for addressing automation complacency. The pedestrian's unsafe crossing and impairment, and the Arizona Department of Transportation's limited public evidence oversight, were further contributing factors.
These findings do not divide neatly into one guilty human and one neutral machine. A developmental automated vehicle is a socio-technical test operation. Its safety case includes sensors, perception software, motion planning, braking logic, base-vehicle systems, route and speed restrictions, the human-machine interface, the operator, supervision, change control, incident review and permission to use a public road. If the developer assigns emergency recovery to a person, that person's likely performance under prolonged passive monitoring becomes a design input.
If the developer removes a second person or disables an independent brake, the remaining operator carries more of the risk whether or not a policy says that phones are prohibited.
That is why the case is an accountability test rather than a contest over which single element caused the last second of failure. Criminal law, an NTSB safety investigation, corporate safety management and vehicle regulation answer different questions. The NTSB's final report HAR-19/03 is a fact-finding safety report, not a judgment allocating civil damages or criminal guilt. A criminal case can establish the operator's legal responsibility without validating ATG's system architecture. A repaired software version can improve entity handling without proving that supervision and public-road approval were adequate.
Each conclusion must remain within the authority and evidence of its source.
The distribution of consequence makes that discipline essential. Herzberg lost her life. Her family and community bore harm they did not control. The operator faced prosecution and years of public scrutiny. Safety drivers were asked to manage a low-frequency, high-consequence monitoring task. Engineers, managers, Volvo, Uber, Arizona agencies, federal regulators and later autonomous-vehicle developers controlled different preventive layers. Road users had no practical way to inspect those layers before sharing the test environment.
The core accountability question is therefore operational: who could change what before the crash? Uber ATG could change classification and prediction logic, emergency response, testing speed, staffing, training, video review and release criteria. Volvo and ATG together controlled the integration that deactivated factory collision-avoidance systems. The operator controlled attention, steering and braking during the trip. Arizona controlled the terms under which testing occurred on its roads. NHTSA controlled federal defect authority and guidance but did not preapprove ATG's safety case.
The pedestrian controlled where and when she crossed, but not the approaching system's design or test permission. Consequences overlapped; control did not.
Who had practical control
Uber ATG had the broadest direct control over the test operation. Its software detected entities, assigned classifications, predicted movement, planned a trajectory and issued steering and braking commands. Its managers selected the operational design domain, including mapped routes, lighting and weather conditions and a maximum testing speed. Its operations organisation hired and trained operators, wrote phone and driving policies, chose whether one or two people would staff a vehicle, stored inward-facing video and decided how often supervisors would inspect it.
Its safety and leadership structure determined whether hazards would be identified before public-road release.
The NTSB vehicle automation report identifies the specific Krypton platform and software version fitted to the crash vehicle. It also records that the system was intended to operate autonomously only on designated mapped routes while a human oversaw both the system and the environment. That division matters: the ADS performed the driving task until it did not have an adequate response, while the person was expected to detect the failure and take over rapidly.
The operator controlled the immediate manual fallback. ATG instructions required monitoring, hovering hands over the steering wheel and a foot over the brake, tagging unusual events and intervening in emergencies. The operator also controlled the personal phone and violated the no-phone rule. Phone and application evidence showed that video was streamed throughout the trip. Her final downward gaze removed the last available human response during the critical interval. Recognising this agency is not the same as treating the operator as an unlimited guarantee against foreseeable system limitations.
ATG's managers controlled the conditions under which that fallback was expected to work. The operations factual report shows that a written no-phone policy existed and that violation could lead to termination. It also shows control weaknesses behind the rule: the policy was not standalone, ATG had no record that this operator individually acknowledged it, and supervision did not turn stored inward-camera evidence into regular checks. A policy establishes a standard. It is not an operating safeguard unless training, observation, escalation and consequences make adherence visible.
Volvo controlled the original vehicle and its production safety systems, while ATG controlled their integration with the developmental stack under an arrangement involving both companies. The factory vehicle had forward-collision warning and automatic emergency braking. Those functions were active in manual mode but automatically disengaged in autonomous mode. The NTSB record says simultaneous operation was considered incompatible because ATG and Volvo radars could interfere and because the brake module was not designed to prioritise two braking commands. This was not a mysterious component failure on the night.
It was a known integration decision that removed one independent intervention layer while the ADS was active.
Arizona controlled admission to public roads. The state's 2015 executive order required a licensed operator able to take control, proof of financial responsibility and compliance with traffic laws, and created an oversight committee. Yet ADOT's December 2016 statement welcoming Uber said there were no special permits or licences for autonomous testing. The state therefore set broad conditions without a pre-test technical assessment of ATG's hazard analysis, fallback architecture or operator-monitoring plan.
NHTSA controlled federal motor-vehicle defect authority and national guidance. At the time, its ADS 2.0 policy was expressly voluntary. The official ADS 2.0 page explains that a developer did not have to submit a self-assessment, wait for approval or meet a new compliance mechanism before testing. Federal authority could respond to an unreasonable safety defect, but guidance did not amount to a prospective approval of this developmental test system.
Finally, neither Herzberg nor other road users controlled any of those safety decisions. The public road functioned as part of the validation environment, yet a person encountering the vehicle could not know its emergency braking limit, whether factory AEB was active, how often the operator was monitored or whether an independent reviewer had challenged the removal of redundancy. That asymmetry creates the strongest case for auditable public assurance before testing, not merely investigation after harm.
Timeline: policy, operating changes and the final seconds
Arizona's framework preceded Uber's arrival. Executive Order 2015-09 encouraged testing and established basic operator and insurance conditions. In December 2016, after a dispute stopped Uber's testing in California, Arizona publicly emphasised that it did not require a special autonomous-testing permit. The state sought investment and technical development while relying substantially on general vehicle rules and developer responsibility.
Uber ATG initially used two people in each test vehicle. One monitored the road and prepared to take control; the other documented system behaviour and the environment. About five months before the fatal crash, ATG introduced a tablet interface and consolidated those responsibilities into a single operator. The NTSB found that the interface reduced the complexity of logging events, but one person still had to watch the road, supervise the ADS, be ready for an emergency and sometimes look away to use the tablet.
Removing the second operator also removed a person who could detect danger, remind the driver of duties or provide a second attention channel.
The change should be read alongside the system architecture. In autonomous mode the factory forward-collision warning and emergency brake were inactive. ATG's ADS limited extreme braking and used an action-suppression interval to guard against false-positive hazards. If the system judged that a collision could no longer be avoided within its braking and jerk limits, its design did not apply maximum braking merely to reduce impact severity. It was expected to warn the operator and begin a gradual slowdown. The operator was thus not one defence among several equally capable layers.
In important emergency cases, the operator was the primary final defence.
On 1 March 2018, Arizona issued Executive Order 2018-04, expanding the framework for autonomous operation with and without a person present and describing public safety as a priority. The order required compliance with applicable laws, a minimal-risk condition after an ADS failure for driverless operation, insurance and a law-enforcement interaction plan. It did not require an independent technical panel to approve a developer's testing safety case before road use.
On Sunday 18 March, the modified 2017 Volvo XC90 began the relevant test trip in Tempe. The road was dry and illuminated by street lighting. The vehicle was on its second loop of an established route and had been in autonomous mode for roughly 19 minutes when it approached northbound Mill Avenue at about 45 mph. Herzberg began crossing from the centre median outside a crosswalk while walking a bicycle.
The ADS first detected her by radar 5.6 seconds before impact at approximately 44 mph. It initially classified the return as a vehicle and did not place it on the SUV's path. Lidar then classified it as an unknown entity, then as a vehicle. Between 3.8 and 2.7 seconds before impact, classification alternated between vehicle and other. At each classification change, the system did not carry forward the entity's tracking history, weakening the path prediction. At 2.6 seconds it classified the entity as a bicycle, but initially treated it as static and outside the SUV's path.
At 1.5 seconds before impact, classification returned to unknown. The entity was partly in the SUV's lane, and the ADS planned a path around it. At 1.2 seconds, the lidar classification changed again to bicycle and the predicted path finally occupied the SUV's lane. The earlier steering plan was no longer feasible. The situation became hazardous and the one-second action-suppression interval began. At 0.2 seconds before impact, suppression ended and the ADS generated a slowdown and auditory alert. The NTSB record could not determine whether the slight communication delay allowed that slowdown to start before the operator took control.
The operator had looked down for about five of the final six seconds. The human performance factual report documents her training, duties, gaze and phone evidence. Across the trip she looked toward the lower centre console for 34 percent of the time. On the first loop at the same road section, one downward gaze lasted 26.5 seconds. In the final approach, she raised her gaze about one second before collision and steered left 0.02 seconds before impact. She braked only after impact.
The timing does not prove that a person should have reacted at the first radar return. The NTSB analysed geometry and ordinary scanning behaviour, and said even an attentive operator might not have noticed Herzberg at 5.6 seconds while the SUV was leaving a curve. Once the vehicle reached the straight section 3.9 seconds before impact, however, she would have entered the normal field of view of an attentive driver. Braking tests supported a window of roughly two to four seconds to detect and initiate a response. The NTSB concluded that an attentive operator likely could have avoided the collision or mitigated its severity.
The crash occurred at 9:58 p.m. Herzberg died at the scene; the operator was not injured. Mechanical inspection did not identify a major defect in the vehicle's base braking, lighting, suspension, electrical systems, wheels or tyres. The ADS reported no sensor or system fault during the trip. These negative findings are important because they narrow the explanation away from a random mechanical failure and toward expected system behaviour, human performance and the way the test operation was governed.
On 19 March, ATG stopped public-road testing in all operational centres. On 26 March, Arizona's governor sent Uber a suspension letter preserved in the NTSB docket, directing ADOT to suspend Uber's ability to test and operate autonomous vehicles on Arizona public roads. The letter called the incident a failure to comply with the state's safety expectation. It was a decisive enforcement response, but it occurred after the permissive admission decision and after the fatality.
ATG did not resume public-road ADS testing until 20 December 2018, and then only on a one-mile Pittsburgh loop with a 25 mph limit. The company had altered technical, operational and organisational controls. In November 2019, the NTSB adopted its final report and recommendations. The criminal case continued separately: a 2020 grand-jury indictment accused the operator of negligent homicide, an accusation rather than a conviction. In July 2023, a filed plea agreement resolved the case through a guilty plea to endangerment, and the county attorney reported three years of supervised probation.
That disposition establishes personal legal accountability under the agreement; it does not decide whether ATG's design and governance controls were adequate.
The technical causal chain was a sequence, not one missed classification
It is tempting to summarise the crash by saying the vehicle did not see a pedestrian. That is inaccurate. Its sensors detected an entity several seconds before impact. The failure emerged in how detections were represented, classified, carried through time, converted into possible motion and then connected to braking authority.
Detection was the first layer and it operated. Radar produced the initial return; lidar and camera systems contributed to the environmental representation. The NTSB found no relevant sensor malfunction. The problem was not simple sensor blindness, and darkness alone did not explain the event. The forward fleet camera made the scene appear darker than investigators observed at the location, while sight-distance work found no obstruction that would have prevented an attentive operator from seeing the crossing pedestrian during the relevant approach.
Classification was unstable. The same road user moved among vehicle, other and bicycle labels. Classification matters because the system associated different goals with different classes. An unknown entity could be treated as static. A bicycle detected in a travel lane could be assigned a goal consistent with the lane, even when the observed person was crossing it. Pedestrians outside a crosswalk were not assigned an explicit crossing goal in this software version. The system therefore embedded assumptions about where a pedestrian should be and how a classified entity should move.
Tracking continuity was weakened by the architecture. When classification changed, the system no longer used the entity's prior tracking history for the new trajectory. A road user's physical continuity did not guarantee continuity in the model. This meant multiple seconds of observations did not accumulate into one robust estimate of a person steadily traversing the road. The system repeatedly restarted part of its reasoning as the label changed.
Prediction translated that unstable model into an unsafe path assessment. For most of the approach, the entity was treated as static or moving in the adjacent left lane. Only at 1.2 seconds did the ADS place the bicycle-classified entity fully in the SUV's lane. By then, steering around it was no longer possible under the generated plan.
Emergency response added another delay. ATG had designed a one-second suppression interval to verify a potential hazard, calculate an alternative or allow the operator to intervene. The concern was legitimate in isolation: a developmental system that reacts violently to false positives can create its own collisions. But the safety question is not whether false braking mattered. It is whether the combined architecture preserved a safe response when classification stabilised late. In this case, one second consumed most of the remaining time.
Braking limits then prevented a strong mitigation response. The ADS treated situations requiring braking beyond its defined deceleration or jerk specifications as emergencies for human intervention. If maximum braking within those limits could not avoid collision, the automated system was not designed to apply maximum braking solely to reduce severity. It would warn and initiate a gradual slowdown. No meaningful automated emergency brake compensated for late prediction.
The base vehicle could not supply an independent rescue because its forward-collision warning and AEB were disengaged in autonomous mode. The integration rationale involved radar interference and unresolved priority between two braking command sources. Those are real engineering constraints. They also created a known common operational gap: when ATG's stack was active, the production safety layer that might otherwise warn or brake was not available.
Finally, the human fallback did not operate. The operator was looking at a phone, ATG was not watching in real time, a second operator was absent and the ADS did not provide a useful warning until the collision was effectively unavoidable. The causal chain was therefore detection without stable interpretation, interpretation without timely trajectory recognition, recognition followed by suppression, limited public evidence automated mitigation and an unavailable human fallback. No single link explains the whole result.
The safety driver was part of the engineered system
The phrase "safety driver" can conceal an important allocation decision. The person was expected to remain vigilant while the automation handled normal driving for long periods, then detect a rare failure and intervene within seconds. That is not ordinary manual driving. It is supervisory control, and its feasibility must be assessed with human-performance evidence.
The NTSB concluded that prolonged visual distraction was a typical effect of automation complacency and that humans perform poorly as monitors of highly reliable automation. The problem does not excuse deliberate phone use. It changes what a responsible designer must anticipate. A prohibition may deter some misconduct, but a risk assessment should assume that passive vigilance degrades, rare failures are difficult to detect and a single person may violate a rule. High-consequence design does not treat perfect compliance as an independent redundant layer.
ATG possessed useful monitoring evidence before the crash. Inward-facing cameras recorded the operator. Supervisors could have reviewed footage for phone use, hands-away behaviour and attention trends. The NTSB found that review was rare. That is a distinction between data collection and control. A camera that produces evidence only after a fatal crash is an investigation tool. A camera connected to alerts, sampling, coaching, suspension and trend analysis is a preventive control.
Removing the second operator intensified the risk. The documented vehicle-operator duties placed road monitoring, intervention and event notation in the operating role. The tablet reduced the notation burden but did not eliminate glance demand. A second person could have handled logging, cross-checked attention and acted as a social reminder. Two people are not automatically safe, and both can be distracted, but eliminating one without a validated replacement removed diversity from the fallback.
The right accountability question is not whether ATG was allowed to make operations more efficient. It is what evidence supported the change. A defensible change record would identify the hazard controlled by the second operator, quantify the new tablet's residual workload, analyse vigilance duration, test takeover performance after long uneventful periods, set maximum shift and route exposures, define attention thresholds and require approval from an independent safety function. The public NTSB record found no adequate safety risk process performing that role before the crash.
This is also why the operator's misconduct and system-design responsibility are compatible. The operator had a duty to watch the road and did not. ATG had a duty to assess whether the fallback was robust against foreseeable human failure and did not do so adequately. The former explains why the last human intervention failed. The latter explains why one person's attention had become such a critical single point of failure.
Safety culture turned separate weaknesses into one operating condition
The final report did not use "safety culture" as a vague criticism. It tied the term to missing structure and ineffective execution. At the time of the crash, ATG had no corporate safety plan assigning risk responsibilities across departments, no safety division and no dedicated safety manager. The head of operations, without a background in safety management, carried the safety-manager responsibility in addition to operational duties.
That structure mattered at several decision points. Who could independently challenge deactivation of production AEB? Who owned the hazard created by resetting tracking history after classification changes? Who approved the one-second action suppression and the choice not to mitigate an unavoidable crash with maximum available braking? Who signed off on removing the second operator? Who reviewed attention data and stopped an operator or a route when policy violations appeared? Without clear, competent and independent ownership, each issue could remain inside its local engineering or operations rationale.
Policy execution showed the same weakness. ATG had a cell-phone rule, a disciplinary schedule and a drug-testing policy. Yet it lacked a record of individual acknowledgement for the phone rule and implemented drug testing sporadically. The crash operator had not received pre-employment or random testing and was not tested under the company's own post-crash policy. The NTSB found no evidence that operator impairment contributed to the crash, so drug testing is not a causal explanation. It is evidence that written safety commitments did not reliably become operating practice.
Crash and incident data also required organisational learning. The automation report lists 37 prior autonomous-mode crashes and incidents from September 2016 to March 2018, most involving another vehicle striking the test SUV. That history does not establish a prior warning identical to Herzberg's crossing. It does show that ATG had a growing field record requiring structured classification, exposure-normalised analysis and hazard review. Raw event counts cannot prove safety without mileage, scenario and severity context; nor should a lack of an identical fatal event be treated as proof that the architecture was safe.
The external culture review commissioned after the crash recommended an SMS, senior safety leaders, internal communication, training reform, unannounced operator checks and stronger operational risk controls. Because ATG commissioned it, it is not equivalent to a regulator's enforcement finding. The NTSB independently evaluated the crash record and reached aligned conclusions about the pre-crash safety framework. The public NTSB docket makes the underlying factual reports, company submissions and attachments inspectable, allowing readers to separate company representations from investigator findings.
Public-road governance left the developer to grade critical safeguards
The Tempe operation sat between traditional divisions of authority. NHTSA regulates motor vehicles and equipment at the federal level. States regulate drivers, registration and road operation. A developmental ADS complicates that split because the machine performs the driving task while a company employee supervises it, and because the modified vehicle may comply with base vehicle standards even when experimental software creates new behaviour.
ADS 2.0 identified twelve safety elements, including the operational design domain, entity and event detection and response, fallback, validation, human-machine interface, data recording and post-crash behaviour. The guidance itself encouraged a public voluntary safety self-assessment but said it was not exhaustive, mandatory or subject to federal approval. That made the document a transparency mechanism, not a gate.
The NTSB found that voluntary submissions had limited safety value where developers could omit them and NHTSA did not evaluate adequacy. It issued recommendations H-19-47 through H-19-52. H-19-47 asks NHTSA to require developers testing on public roads to submit safety self-assessments. H-19-48 asks the agency to evaluate them for appropriate safeguards, including operator-engagement monitoring. H-19-49 and H-19-50 ask Arizona to require a risk-focused application and expert review before granting a test permit. H-19-52 asks ATG to complete a safety management system.
Arizona's later statutory framework is more explicit than the 2016 no-special-permit position. Arizona Revised Statutes section 28-9702 allows operation with a licensed human able to resume the driving task. For driverless operation it requires a law-enforcement interaction plan and a written statement acknowledging federal compliance when required, a minimal-risk condition after relevant ADS failure, traffic-law capability, title, registration and insurance. ADOT can issue a cease-and-desist letter if required documents are not submitted.
Other sections create registration-suspension procedures when evidence indicates a vehicle may be unsafe.
Those are real legal controls, and the enacted 2021 Chapter 117 moved key requirements from executive policy into statute. The current ADOT pages distinguish testing with a safety driver from testing or operation without one. Yet the visible provisions do not themselves publish an independent expert determination that a specific developer's entity-response, fallback, operator-monitoring and change-control evidence is adequate before testing. Comparing the public framework with H-19-49 and H-19-50 supports an inference that attestation and enforcement remain more visible than case-specific expert preapproval.
It does not prove what non-public engagement ADOT may conduct for a particular operator.
In 2023, Governor Katie Hobbs rescinded Executive Order 2018-04 because subsequent legislation addressed the subject. The rescission order confirms the legal transition; it does not erase the historical role of the 2018 order at the time of the crash. This temporal distinction prevents two errors: applying today's statute backward to Uber's 2018 test, or treating Arizona as if no reform followed the NTSB finding.
Federal evidence controls also changed. NHTSA's Standing General Order on crash reporting, first issued in 2021 and later amended, requires identified ADS and Level 2 entities to report qualifying crashes. NHTSA can use those reports for follow-up, defect investigation and enforcement. The order improves timeliness and comparability relative to ad hoc notification. It operates after a crash or reportable incident, however; it is not a substitute for prospective evaluation of a public-road test plan.
As of the article's access date, NHTSA's open NTSB recommendations page still listed H-19-47 and H-19-48. The agency's March 2020 response asked that they be treated as open acceptable responses while relying on voluntary assessments and related initiatives. The defensible conclusion is not that federal oversight stood still; mandatory crash reporting is substantial. It is that mandatory pre-test submission and substantive federal evaluation, the specific controls the NTSB requested, were not shown as completed by the agency's own open-recommendation list.
Legal accountability resolved one case, not the architecture
The operator's criminal case is important because it rejects a narrative in which automation removed all human duty. The grand jury charged negligent homicide in 2020. The later plea replaced the trial risk with an agreed endangerment conviction and supervised probation. The Maricopa County Attorney's 2023 announcement stresses the responsibility of a person behind the wheel despite available technology.
That resolution has limits. A guilty plea does not establish every factual allegation in the earlier indictment. It does not adjudicate the technical design, ATG management's reasonableness, Arizona's regulatory duty or a hypothetical allocation of civil fault. It also does not convert an NTSB contributing factor into a criminal element. Using the case as proof that only the operator mattered would exceed the court record.
The converse error would be to treat the NTSB report as a corporate liability verdict. The Board determines probable cause and issues safety recommendations. It does not award damages and its process is not an adversarial trial. Its systemic findings are authoritative for transportation-safety analysis but do not themselves prove a criminal offence or civil claim.
There is a broader accountability problem when legal visibility concentrates on the last human actor. A driver has a name, a phone record and a precise moment of inattention. System decisions are distributed across requirements, software versions, integration meetings, staffing budgets, safety review and state policy. They can be causally important without mapping cleanly to one individual offence. A mature accountability regime must preserve the driver's duty while making institutional decisions equally traceable.
That means recording named approval for safety-critical changes, not attributing decisions only to a corporate unit. It means retaining the hazard analysis that justified disabled AEB, the validation that supported action suppression, the performance evidence for single-operator staffing, the review rate for attention video and the state basis for allowing public-road exposure. Without those records, criminal law can resolve personal conduct while system learning remains incomplete.
Impact extended beyond the collision
The direct impact was irreversible: Herzberg was killed. Public analysis should not turn her into a test scenario or reduce the event to a technical milestone. She was a person using a public road who had not consented to become part of a developmental system's validation. The fact that she crossed outside a crosswalk and that toxicology supported possible impairment belongs in the causal record, as the NTSB concluded. It does not transfer control over Uber's safeguards to her.
Her family faced loss and a long public process in which video, toxicology, software and driver conduct were repeatedly examined. The primary official sources used here do not provide a reliable public measure of the family's total economic or non-economic harm, and no such figure should be invented. Fatality is not made more precise by attaching an unsupported monetary estimate.
The operator faced criminal prosecution, a guilty plea and probation. Other safety drivers inherited a changed profession: more training and monitoring, but also clearer evidence that supervising automation can produce personal legal exposure. Employers that call someone a fallback must disclose the actual limitations, train for rare emergencies, measure attention and give the person enough authority and redundancy to perform the role.
Uber ATG stopped public testing, lost permission to operate in Arizona, incurred investigation and redesign work and reduced the scale and speed of its resumed programme. Uber's registration statement filed before its public offering disclosed the fatal crash and the risk that autonomous-vehicle failures could create liability, scrutiny and reputational harm. The SEC-filed S-1 is a company disclosure to investors, not independent proof of safety, but it confirms that the consequences became material corporate risk.
Volvo's role required clarity because the vehicle carried its base architecture while ATG controlled the developmental system. In its NTSB party submission, Volvo emphasised the human-factors challenge of keeping a supervising driver ready as fallback. A party submission advocates the party's position and should not be treated as a Board finding. It is useful evidence that the base-vehicle manufacturer recognised fallback readiness as a design issue, not merely an operator character issue.
Arizona experienced a legitimacy cost. The state had advertised its low-regulation environment as an advantage and then suspended Uber after the fatality. The issue is not that innovation and safety are mutually exclusive. It is that public assurance weakens when admission is easy, technical review is opaque and decisive enforcement comes only after harm. Later statutes can improve authority, but trust also requires disclosure of what was reviewed and why a test was permitted.
The autonomous-vehicle sector faced increased scrutiny over public-road testing, safety drivers and voluntary self-certification. That effect cannot be quantified from this case alone. Technology investment, public opinion and deployment schedules respond to many events. The supported inference is narrower: the crash supplied concrete evidence that a developmental ADS could detect a road user yet fail to produce timely avoidance, and that a human fallback could fail under an inadequately managed monitoring task.
Repair evidence: what changed and what the evidence proves
ATG's first control was a stop. It ended public-road testing the day after the crash and did not resume until December. A suspension prevents immediate recurrence while investigation and redesign occur. It is not, by itself, evidence that the later system is safe.
Technical changes directly addressed the recorded sequence. ATG changed sensor fusion and path prediction so prior positions remained relevant when classification changed. The updated system could assign a midblock crossing goal to a pedestrian. It removed action suppression, permitted braking to mitigate a crash even when full avoidance was impossible and increased permitted jerk. It changed ATG radar frequency and worked with Volvo on braking-command priority so forward-collision warning and pedestrian AEB could remain active during autonomous testing.
These changes have strong causal fit. Each maps to a failed layer: continuity, crossing prediction, delayed response, severity mitigation and independent braking. ATG replayed crash sensor data through a September 2018 software version and reported that the new version would have classified the pedestrian earlier, predicted the crossing conflict and initiated controlled braking more than four seconds before the original impact time. This is useful regression evidence, but its source and method matter. It was an ATG simulation reported to investigators, not a public independent reproduction with full software, scenario and acceptance details.
Operational changes also matched the findings. ATG restored two trained mission specialists, separated road monitoring in the driver seat from other duties, expanded distraction and emergency-manoeuvre training, and added real-time inward-camera attention monitoring. A several-second gaze away could trigger an in-vehicle alert and a supervisor report. NTSB concluded that those measures began to address the oversight deficiencies.
The word "began" is appropriately cautious: installation and procedure show control design, while long-term alert performance, false-negative rates, supervisor response and behaviour under pressure show control effectiveness.
Organisational repair included a safety management system with policy, risk management, assurance and promotion; dedicated safety leadership; revised testing governance; and more formal release review. Uber ATG's party submission to the NTSB described its internal and external reviews and claimed broad improvements. As a company submission, it proves what ATG represented and what materials it supplied, not independent effectiveness.
External evidence became stronger later. The NTSB's current automated-vehicle investigative outcomes page says H-19-52, the recommendation to complete the SMS, was closed on 26 April 2024. Recommendation closure is meaningful: the Board judged the recipient's responsive action sufficient for its tracking purpose. The page also notes that Uber ATG had been acquired by another developer. Closure is not a certification that the original Uber unit still operates, that every pre-crash decision was repaired or that all test evidence is public.
Ownership changed the entity being assessed. Uber announced the sale of its ATG business to Aurora in December 2020 and completed it in January 2021. Uber's 2020 Form 10-K records the transfer and continuing commercial relationship. A repaired control embedded in ATG may have transferred with people and assets, changed under Aurora or been retired. Public evidence should not casually describe 2026 Uber as operating the same Tempe development programme.
Uber now acts partly as a platform integrating autonomous partners. Its April 2026 autonomous mobility and delivery safety guidelines describe partner safety-plan assessment across planning, demonstration and operation. They call for performance reporting and reference standards, but also say assessment does not prescribe one metric and threshold for every partner. This is evidence of a current governance framework, not proof that a specific partner passed a disclosed test or that the Tempe controls were independently reproduced.
Regulatory repair is similarly mixed. Mandatory crash reporting gives NHTSA data it lacked in 2018, and Arizona now has statutes governing autonomous operation and enforcement. Yet pre-test federal safety self-assessment remains voluntary and the relevant NTSB recommendations remain open. GAO's automated-vehicle oversight review, updated through March 2026, continued to report that DOT lacked an overall plan meeting comprehensive-planning principles. The federal system has more evidence after incidents, but the public record does not show one uniform approval gate for developmental ADS public-road safety cases.
Confirmed facts, supported inference and public unknowns
The confirmed facts are extensive. The ADS detected Herzberg 5.6 seconds before impact and did not identify a collision path until 1.2 seconds before impact. Classification changed repeatedly and tracking history was not carried through those changes. Action suppression then lasted one second. Factory forward-collision warning and AEB were disabled in autonomous mode. The operator was streaming video and looked down during the critical approach. ATG had removed a second operator and rarely reviewed inward-camera footage.
The NTSB identified operator distraction as probable cause and ATG safety culture, risk assessment, oversight and complacency controls as contributing factors. Arizona suspended testing, ATG stopped and later resumed under changed conditions, and the operator ultimately pleaded guilty to endangerment.
It is also confirmed that ATG changed the relevant technical and operating controls. The NTSB documented re-enabled compatible base-vehicle safety systems, altered path prediction, removed suppression, added mitigation braking, restored a second operator, added real-time attention monitoring and developed an SMS. The NTSB later closed H-19-52. NHTSA now mandates qualifying ADS crash reports. Arizona enacted autonomous-vehicle provisions. Uber sold ATG to Aurora.
Several conclusions are supported inferences rather than directly observed facts. Retaining track history and treating midblock crossing as a possible goal should reduce the specific perception and prediction failure seen in Tempe because those changes target its mechanism. Independent AEB and removal of suppression should improve mitigation opportunity. Two operators and real-time attention alerts should reduce the chance that one long gaze defeats the whole fallback. The NTSB and ATG simulation support those propositions, but the public cannot calculate exact risk reduction.
It is a supported inference that removal of the second operator required a stronger change-safety case than the public record shows. The NTSB found increased workload and reduced redundancy, and it found inadequate risk assessment. That does not prove the business motive for the staffing change or identify every manager who approved it. The public record does not establish whether cost, fleet scaling, confidence in the tablet or another factor dominated the decision.
Important technical details remain unknown publicly. Full source code, model parameters, training data, false-positive distributions, simulation corpus, release criteria and complete scenario test results are not in the docket. The public cannot reproduce the updated software replay or inspect every case in which the revised system braked unnecessarily. It cannot verify the attention monitor's sensitivity across lighting, eyewear, head pose and operator diversity, or determine how often supervisors acted on its reports.
The complete internal decision record is also unavailable. Public materials do not name every approver for disabled AEB, action suppression, braking limits, one-operator staffing or road release. They do not expose every dissent, risk acceptance, schedule pressure or escalation. Absence of that evidence is not proof of improper motive; it limits attribution.
Repair durability is partly unknowable because the organisation changed ownership. The NTSB closure establishes an accepted response to H-19-52, while the 2021 sale means later operating evidence belongs to a different corporate structure. Uber's current platform framework cannot be used as a simple continuation of ATG's post-crash controls. Aurora's later performance cannot be attributed automatically to Uber, and Uber's current partner governance does not validate the 2018 programme.
Aggregate impact is not public. The official records used here do not quantify all family loss, employee consequence, programme cost, municipal impact or sector-wide delay. They do not support a single dollar total. Nor do they establish how public acceptance would have evolved without the crash.
The public evidence cannot prove a no-crash counterfactual with certainty. An attentive operator likely could have avoided or mitigated impact, according to NTSB testing. Updated ATG software reportedly would have braked much earlier in replay. Those are strong bounded conclusions. They do not establish exactly where Herzberg or the vehicle would have stopped in every plausible variation, or that no different failure would have occurred.
Counterfactuals: where an earlier control could have broken the chain
The most immediate counterfactual is an attentive operator. The NTSB's sight-distance and braking analysis supports a likely avoidance or mitigation window after the SUV entered the straight road segment. This is the strongest individual counterfactual because it uses measured geometry and vehicle braking. It still does not justify making attention the only safety case.
A second counterfactual is active operator monitoring. If the inward-facing camera had been connected to a reliable live alert before the crash, the operator's repeated and sometimes very long downward gazes might have triggered intervention earlier in the trip. A supervisor could have warned or removed her, or the system could have required a safe stop. This is supported by the pattern in recorded footage and the form of ATG's later control. The exact alert threshold and response outcome remain unknown.
A third is retention of the second operator. A passenger-seat operator might have managed notation, noticed the driver's gaze or identified the crossing hazard. The NTSB treated the second person as lost redundancy and ATG restored the role. It cannot be assumed that any second person would have remained attentive or intervened in time. The counterfactual is that two independent attention channels reduce single-point risk, not that two people guarantee safety.
A fourth begins in software. If subject identity and prior motion had persisted across classification changes, the observed steady crossing could have influenced trajectory prediction before 1.2 seconds. ATG's changed logic and replay support that inference. A robust architecture would also test uncertain or changing classifications conservatively: an unknown moving entity near a reachable path should retain collision potential even if its semantic label is unstable.
A fifth concerns emergency control. Removing the one-second suppression interval would have preserved scarce response time once the hazard was identified. Allowing maximum feasible braking for mitigation could have reduced impact even when full avoidance was no longer possible. Re-enabled Volvo AEB might have supplied an independent warning or brake. The public record does not contain a complete reconstruction of how the production AEB would have classified this exact crossing with ATG sensors operating, so claims of certain prevention would be too strong.
A sixth is lower test speed. ATG restarted at 25 mph rather than the 45 mph operating limit in Tempe. Lower speed reduces stopping distance and impact energy and gives a human more time. It can also change route representativeness and delay testing at higher-speed conditions. A responsible expansion would require staged evidence: prove detection, prediction, mitigation and takeover at lower speed, then approve higher speed only when scenario-based limits are met.
A seventh counterfactual moves outside the company. A safety-focused Arizona application could have required ATG to disclose the disabled base AEB, emergency braking limits, one-second suppression, single-operator workload and attention-monitoring plan. An expert panel might have required changes or confined testing. That outcome is plausible because those are exactly the safeguards the NTSB later highlighted. It is not certain that a reviewer would have discovered every interaction or denied the route.
An eighth is federal evaluation. Mandatory VSSA submission and substantive NHTSA review could have created a national minimum and supported state decisions. The limits are real: a high-level self-assessment may not reveal source-level failure, and regulators need skilled staff and test criteria to challenge it. A document gate without evidence and expertise can become compliance theatre. The useful counterfactual is a review tied to hazards, validation results, change records and enforceable conditions.
The final counterfactual concerns stop authority. Any engineer, operator or supervisor who sees an unmitigated hazard should be able to pause testing without schedule penalty. Public evidence does not identify a specific pre-crash employee who raised and lost such a challenge. The lesson is structural: stop-work authority, protected reporting and independent safety sign-off make weak signals more likely to interrupt deployment before the road supplies the proof.
A durable accountability test for developmental automated driving
The first test is a named control map. For each dynamic-driving function and failure mode, the developer should identify the system component, human role and manager with authority. Detection, classification, tracking, prediction, braking, fallback, remote support, operator monitoring and state liaison need owners. Shared responsibility without named decision rights becomes unowned risk.
The second test is hazard continuity across software boundaries. When classification changes, uncertainty and motion history should not disappear. Validation should include unusual road users, people outside expected crossing locations, partial occlusion, low light and entities whose semantic class remains uncertain. The safety requirement is not only correct labelling; it is conservative collision response under uncertainty.
The third test is independent mitigation. A developmental stack should not leave a single human as the only effective emergency layer when a collision becomes unavoidable. Base-vehicle AEB, an independent monitor, a safety kernel or another diverse path should retain authority to warn, slow or stop. If integration disables a production safeguard, a documented equivalent or stronger replacement should be required before road testing.
The fourth test is realistic human performance. The safety case should model vigilance decrement, automation complacency, glance behaviour, workload, takeover time and rule violations. Training and policy remain necessary, but neither counts as independent redundancy. Shift duration, route monotony, HMI glance demand and alert design should be tested under long uneventful runs followed by rare hazards.
The fifth test is active attention assurance. Inward cameras and telemetry should produce defined alerts, safe-stop logic, supervisor response, coaching and removal criteria. The developer should measure missed detection and nuisance alerts, not merely announce installation. Privacy retention and access controls should be explicit because the system records workers continuously.
The sixth test is safety-critical change control. Removing a second operator, altering speed, disabling AEB or changing suppression logic should trigger a fresh hazard analysis and independent approval. The record should state the previous safeguard, replacement, evidence, residual risk, approver, expiry and rollback condition. Operational efficiency is not evidence of equivalent safety.
The seventh test is scenario-based validation before public exposure. Closed-course and simulation work should cover foreseeable edge conditions and combined failures, with clear pass thresholds and version traceability. Replay of the crash is useful regression testing, but an organisation must also test adjacent scenarios that could defeat the fix. Passing one historical case can overfit a safety response just as easily as a model.
The eighth test is exposure-aware field evidence. Developers should publish miles by operational design domain, disengagement and intervention definitions, crash and near-miss categories, attention alerts, safety-driver hours and material software changes. Counts without denominator or scenario mix mislead. Confidential technical detail can be protected while still disclosing enough for independent trend analysis.
The ninth test is prospective public authority. A state application should identify the test area, speeds, vehicle configuration, fallback, operator plan, remote support, incident response and evidence supporting release. A qualified multidisciplinary group should review it, impose conditions and revisit approval after safety-critical changes. Federal review should provide a consistent baseline rather than leaving every state to infer adequacy from voluntary marketing-level documents.
The tenth test is incident evidence and escalation. Crash reporting should preserve common-clock sensor, control, video, operator, software-version and decision data. Near misses and attention violations should feed the same learning process. NHTSA's standing order improves crash visibility, but organisations should not wait for a reportable collision to investigate leading indicators.
The eleventh test is independent assurance. Internal simulation and policy documents show intent. External assessors, regulators and reproducible tests provide stronger evidence. Assurance reports should explain scope, exclusions, failed tests and unresolved hazards. A recommendation's closure is meaningful but should not be stretched beyond the specific action assessed.
The twelfth test is continuity through corporate change. When a programme is sold, reorganised or integrated into a platform, safety cases, hazard logs, validation assets, incident records and open obligations must transfer with clear ownership. The public should be able to distinguish the original developer's repairs from a successor's controls and a platform's partner requirements. Otherwise accountability dissolves when the org chart changes.
Tempe remains a defining case because every major layer left evidence. Sensors detected the road user. Software repeatedly changed what that detection meant. Emergency logic consumed time and withheld stronger mitigation. Base-vehicle collision safeguards were unavailable. The sole operator was distracted. Monitoring data existed but was rarely used. A second operator had been removed. Safety management was immature. State admission lacked a safety-focused technical review.
The operator's responsibility is confirmed, not erased, by that system account. The institutional lesson is that a person cannot be labelled a safety layer and then treated as independent of design. Whoever makes the human the fallback controls the workload, information, authority, monitoring, redundancy and time available to that human. Whoever permits public testing controls the evidence threshold before involuntary road users bear the residual risk.
Repair therefore requires more than a better model or stricter phone rule. It requires an auditable chain in which uncertain perception remains conservative, emergency braking can mitigate, human limits are engineered into the safety case, changes receive independent review, authorities examine evidence before exposure and every claimed fix is tested across the whole fallback system. The standard is not that no developmental vehicle can ever fail. It is that no known weakness is silently transferred to one person, one disabled safeguard or an uninformed public.

