Summary

  • The original TFTP wording allowed both sides to retransmit their current packet after receiving an old duplicate; one delayed ACK could therefore make all later DATA and ACK packets appear in pairs and deepen the congestion that caused the delay.
  • RFC 1123 broke the loop by forbidding the DATA sender from retransmitting current DATA merely because a duplicate ACK arrived; RFC 1350, authored by Karen Sollins, carried the corrected protocol forward while explicitly crediting Noel Chiappa for the 1992 fix.

Imagine a transfer that is correct byte for byte. Its block numbers arrive in order. Its final file matches the source. Nothing has been corrupted. Yet from one moment onward, every data packet crosses the network twice and every acknowledgement returns twice. The extra traffic can extend the delay that triggered it, produce still more late packets and eventually make the otherwise correct transfer time out.

That is why the Sorcerer’s Apprentice syndrome is more revealing than an ordinary packet-loss bug. It did not need a malicious peer, a damaged payload or an impossible state. Each endpoint followed a superficially reasonable recovery rule. Together, those local acts of caution created a self-sustaining global error.

Karen Sollins is the named author on both RFC 783, published in 1981, and RFC 1350, published in 1992 as the later TFTP Revision 2 specification. The protocol itself was not a solo invention. RFC 1350 says Noel Chiappa originally designed TFTP and redesigned it with Bob Baldwin and Dave Clark, with comments from Steve Szymanski. It names a wider group behind subsequent modifications and says specifically that Chiappa performed the May 1992 revision that fixed the Sorcerer’s Apprentice bug and other document problems.

That attribution matters. Sollins’s place in the story is not that of a lone programmer discovering one defect. It is the work of protocol authorship across time: preserving a deliberately small system, explaining why it is small, absorbing implementation experience and publishing a rule precise enough to stop compatible machines from amplifying one another’s caution.

One late acknowledgement was enough to start the pattern

TFTP uses UDP and supplies its own reliability. In the base exchange, a sender transmits one DATA block and waits. The receiver acknowledges that block before the sender advances. RFC 1123 described the effective window as one 512-octet segment. This stop-and-wait rhythm is slow on a path with a large bandwidth-delay product, but it is easy to understand and small enough for a bootstrap program.

The old failure begins with timing, not loss. Endpoint A sends DATA block X. Endpoint B receives it and sends ACK X, but that acknowledgement is delayed in the network. A’s timer expires, so A sends DATA X again. B receives the duplicate and sends ACK X again.

The original ACK now reaches A first. A correctly advances and sends DATA X+1. Then the second ACK X arrives. Under the faulty rule, A treats the old duplicate as a reason to resend its current packet, DATA X+1. B acknowledges both copies. The first ACK X+1 advances A to DATA X+2; the second causes DATA X+2 to be sent again. Unless a packet is lost in just the right place to interrupt the pattern, the rest of the transfer proceeds in pairs.

The file can still be right because duplicate block numbers can be recognised. Operationally, however, the transfer is wrong. RFC 1123 warned that excessive retransmission could make it time out. It also pointed to the hostile feedback mechanism: delay is often caused by congestion, and the duplicate traffic creates more congestion. A rule meant to recover from uncertainty converts uncertainty into load.

The name comes from the scene in which an apprentice animates brooms to carry water and cannot stop the multiplication. The metaphor is apt because no single packet is catastrophic. The damage lies in repeated obedience to a rule after the state that justified it has passed.

The repair made recovery asymmetric

RFC 1123 made the correction a requirement: the side originating DATA must never retransmit its current DATA packet merely because it received a duplicate ACK. A sender still needs a retransmission timer. If the expected acknowledgement never arrives, time remains a valid reason to try again. What disappears is the causal link between an acknowledgement for old state and fresh transmission of current state.

The receiver may answer a duplicate DATA block with another ACK. That response can be useful if the earlier ACK was genuinely lost. But the repeated ACK must be harmless at the sender. The repair therefore does not ban duplication everywhere. It assigns authority. A timer governing outstanding DATA may initiate recovery; an old ACK may not create new DATA work.

This is a small state-machine change with a broad engineering lesson. Idempotence at one endpoint is not sufficient if its duplicate response triggers a non-idempotent action at the other. A message may be safe to repeat locally and unsafe to treat as a new command remotely. Distributed protocols need both halves of that statement.

RFC 1123 placed adaptive timeouts and at least exponential backoff beside the duplicate fix. Those mechanisms solve related but different problems. A better timer reduces premature retransmission and backs away under trouble. It does not, by itself, break a loop after a delayed packet has produced a duplicate. The state transition has to be corrected as well.

Trivial meant constrained, not disposable

RFC 1350 explains TFTP’s name without apology: it is a very simple file-transfer protocol. It can read and write files, but cannot list directories and, in the base design, provides no user authentication. Simplicity was not a claim that the protocol’s work was unimportant. It was a resource decision.

RFC 906 proposed TFTP for bootstrap loading because both a booting machine and its server could implement it easily. A small client could live in ROM or EPROM, obtain the first program image and hand control to software capable of richer configuration. The protocol’s low performance mattered less when its job was to make the next layer possible.

This environment magnifies the value of exact specification. Code burned into firmware may be harder to inspect or replace than an ordinary application. A vague recovery rule can survive across vendors and device generations. Interoperability can preserve a defect as efficiently as it preserves a feature.

Later TFTP work added option negotiation, block-size controls and timeout or transfer-size options. RFC 2347, for example, defined a separate option-acknowledgement packet and required unsupported options to be omitted rather than silently changing the base exchange. Extensions improved usefulness, but they did not repeal the underlying obligation to distinguish the expected block from an old duplicate.

Nor did standard status turn TFTP into a general secure file service. RFC 1350 records the absence of user authentication. RFC 1123 recommends configurable pathname access control and says broadcast TFTP requests should be ignored because weak access controls can create a significant security hole. A protocol can be appropriately minimal inside a controlled boot path and dangerously exposed elsewhere.

Sollins’s broader work makes the maintenance role visible

MIT CSAIL’s biography of Sollins places TFTP inside a longer research career concerned with network-based systems. It records a mathematics degree from Swarthmore, graduate degrees in computer science from MIT and work on distributed name management, authentication, global naming, long-lived information systems and extreme scaling. It also notes that she served as a senior programme director for networking research at the US National Science Foundation in 1999 and 2000.

Those subjects share a question that the duplicate-packet loop makes concrete: how does a distributed system preserve meaning when messages, implementations and institutions outlive the moment in which they were designed? The answer is not merely to add features. Sometimes durability depends on removing one ambiguous permission.

The historical record also resists a clean hero narrative. Chiappa is credited with the original design and the May 1992 bug fix. Baldwin, Clark, Szymanski and many reviewers shaped the protocol. Robert Braden edited RFC 1123, where the failure sequence and mandatory correction were laid out with unusual clarity. Sollins authored the durable TFTP specifications that carried this collective design into an Internet Standard.

That division of credit is itself useful. Infrastructure is often maintained through documents whose author, original designer, bug finder, implementer and operator are different people. What matters is whether the document preserves enough provenance to show why a rule exists. Without that explanation, a future simplification may reintroduce the very behaviour compatibility was supposed to prevent.

A correct result is not enough evidence

The Sorcerer’s Apprentice transfer could produce the right file. That makes a final checksum necessary but insufficient evidence of a sound implementation. Validation must also examine how much traffic was generated, which state transitions occurred and whether delay or reordering turned one recovery event into a persistent pattern.

The practical test is simple to state. Delay an ACK until after the sender retransmits the corresponding DATA. Deliver both acknowledgements. The next DATA block should be sent once, not once for each copy of the old ACK. Repeat the experiment with duplicated DATA, lost ACKs, option negotiation and boundary block numbers. Observe the timeout algorithm independently.

For operators, the present relevance is conditional rather than universal. A device may use TFTP only during manufacturing, emergency recovery or network boot. Another may have removed it. A third may expose a server with permissions broader than its boot function requires. The protocol name in a product sheet does not reveal which code path runs or whether the corrected state machine is present.

Sollins’s TFTP record therefore leaves a disciplined question for modern infrastructure: when an old message returns, does the system recognise it as evidence about old work, or mistake it for authority to create new work? The difference is one comparison in a small protocol. Under delay, it is the difference between recovery and multiplication.

Sources