Summary
- RFC 1025 made compatibility a set of observable exchanges: opening, carrying data, closing, repeating without a reset, and reaching other implementations through a gateway.
- Its points organized useful evidence, but did not turn a passing score into universal correctness, a performance ranking, or proof of deployment.
Correctness was still an argument
In September 1987, Jon Postel published a short document with an unusually candid account of how TCP and IP had been tested while both the software and specifications were still developing. When few implementations existed, RFC 1025 says, the practical way to judge whether one was “correct” was to run it against another and argue from what happened. The test could change the implementation. The discussion could also change the specification.
That sentence puts the bake-off in a different place from a modern certification lab. The object under examination was not a sealed product facing a finished rulebook. It was a working protocol, shared by a small set of implementers, whose behavior had to become legible across machines before its written description could settle. In that setting, a failed exchange was evidence, but it was not automatically a verdict. The participants still had to decide whether the implementation had misunderstood the rule, whether the rule was incomplete, or whether the test had asked the wrong question.
RFC 1025 is itself a retrospective, not a minute-by-minute record of every event. It says an early test list appeared in IEN 69 in October 1978; four TCP implementations were demonstrated in Reston on 4 December 1978; six implementations met at USC’s Information Sciences Institute on 27–28 January 1979; and a distributed bake-off took place over the network in April 1980. The 1987 memo reproduces, with slight editing, the procedures, tests and scoring used for that 1980 event. The dates and the changing shape of the exercise matter: this was a practice assembled through repeated encounters, not a single official exam handed down after the protocols were finished. RFC 1025, IEN 69, IEN 77
A conversation had several parts
The first division was almost disarmingly small. A TCP earned a point for opening a connection to itself, another for sending and receiving data, and another for closing gracefully rather than crashing. Repeating the exchange without reinitializing the TCP was worth two more points. A complete conversation through a testing gateway was worth five.
The sequence separated events that a single “connected” label would flatten. An open showed that two endpoints could establish state. Data transfer showed that state could carry something useful. Graceful close tested how the relationship ended. Repeating without a fresh initialization asked whether the implementation remained usable after the first cycle. The gateway added an intermediary and a second network boundary. A successful first handshake could not answer all of those questions.
The middleweight division made the change from self-test to interoperability explicit. Opening, exchanging data and closing with another TCP were each worth two points; repeating without reinitializing earned four; and completing the exchange through the testing gateway earned ten. RFC 1025 notes that these opportunities applied for each distinct TCP contacted: the example gives as many as twenty points per other implementation. The aim was N-squared connectivity—testing the mesh of possible relationships, not merely demonstrating that one favored pair could communicate.
That goal changes the unit of evidence. A result belonged to a pair of implementations, under a particular route and test condition. If A talked to B, the result did not establish that B talked correctly to C, or that A could keep working through a gateway that changed packet delivery. The matrix exposed seams that a single showcase could hide.
The gateway was allowed to make the path unpleasant
The memo’s memorable device was the “flakeway,” a deliberately unreliable gateway with adjustable percentages for dropping datagrams, corrupting them and passing them on, or reordering them before delivery. The name is playful; the purpose is serious. It made the path itself part of the experiment. A connection that survived only on a clean route was a narrower result than one that continued to exchange data when packets disappeared, changed or arrived in a different order.
The checksum rule made that experiment harder to game. RFC 1025 says checksums had to be enforced: no points were awarded if the checksum test was disabled. That constraint connected an apparent success to the mechanism meant to detect corruption. The test was not satisfied by turning off the guard that could reveal the fault.
The scorecard then widened beyond the ordinary conversation. Its heavyweight TCP division assigned points for simultaneous connections to multiple peers, urgent data, sequence-number wraparound and a “Kamikaze” segment combining many header features at once. It even distinguished legal blows—segments meeting the specification—from dirty blows that violated it: the table awarded 30 points for knocking an opponent out with legal segments and 20 for doing so with dirty ones. The phrasing sounds like a contest because it was designed to make implementation limits visible in adversarial exchanges.
It does not mean that an invalid segment became a normal operating condition or that one crash proved a protocol universally unsafe.
There was also a separate host-and-gateway IP division. Its points covered fragmentation and reassembly, source routes, return routes, routing advice, source-quench messages, service markings and options. It rewarded finding gateway behavior that failed to reduce time to live, forwarded a datagram whose time to live was zero, or mishandled a checksum. This separation reflected the different jobs at stake. TCP managed a host-to-host conversation; IP moved datagrams across interconnected networks and gateways. RFC 793, RFC 791
The proposed test list kept drilling into lifecycle and edge conditions: open one connection repeatedly; open several at once and check that data remains separated; crash a local TCP and try the same connection again; connect to a socket that refuses service; send to a zero-window receiver; push data quickly through a “fire hose”; test urgent delivery; and exercise sequence numbers around their wrap boundary. A final case combined a nasty segment with a half-open connection just as the sequence number approached wraparound. The list was not elegant because networks were not elegant.
It tested combinations where apparently independent rules could collide.
A point total had a boundary
The scorecard also included points for the longest conversation, the most simultaneous connections and even excuses. Those touches made the bake-off memorable, but they complicate any attempt to treat its total as a single ranking. Some points measured basic lifecycle behavior; others counted features, gateway handling or the ability to keep many conversations active. They were not interchangeable measurements of one property.
RFC 1025 says so in its final section. The tests above checked basic operation and tricky cases, but did not consider performance or whether more recent ideas had been implemented. It names the John Nagle procedures, Van Jacobson’s slow-start and round-trip-time work, and the SQuID procedures as examples that needed separate attention. It then lists possible performance exercises: transferring a one-megabyte file over Ethernet or ARPANET with FTP or NETBLT, and measuring an echo character’s round trip. The memo cautions that performance results depend heavily on the test environment.
This is a useful boundary. A score for opening, using and closing connections through selected peers did not say how fast a system would perform on a different path, under another load, or with mechanisms absent from the test. Nor did a successful pairwise exchange establish that every host used the same software, that every deployment passed, or that all corner cases had been exhausted. The evidence had scope, and the document named some of it.
The test was part of the protocol’s working life
The Internet’s early implementation culture did not wait for a clean separation between specification and operation. The bake-off gave participants a repeatable surface on which disagreements could be observed: this endpoint opened, that one did not; this pair survived a reset; this gateway mishandled a header; this damaged packet was caught—or was not. Those observations could drive a code repair, a clarification, or a change in the written rule.
That feedback loop is more historically revealing than the novelty of the point system. The score did not make the network correct by arithmetic. It made specific behavior discussable between people responsible for implementations. A shared rule became meaningful when independent systems could exercise it, and a test became useful when its limits were visible enough that engineers could argue over what a result did and did not establish.
The later RFC preserved a snapshot of that practice. It treated compatibility as a mesh of conversations, recovery as a sequence of distinct lifecycle events, and failure handling as something to test through the path as well as at the endpoints. Its own caveats kept the matrix from pretending to be a complete measure of performance or every new TCP idea. Read that way, RFC 1025 is not a victory lap and not a modern conformance certificate. It is a record of how a young Internet tried to turn “it works” into a question that other implementations could challenge.
Sources
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
