Summary
- RFC 3607 presented a 2003 Informational thought experiment about using a large population of Internet hosts for cryptanalytic work. Its numerical outcomes depend on stated assumptions about population, participation, work rate, coordination and the cryptographic target.
- A candidate that passes a public verification test proves only that the candidate satisfies that test. It does not prove which hosts worked, whether work was authorized, how much unique search space was covered, which failures or duplicates occurred, or who should be credited or held accountable.
The striking object in RFC 3607 is not a program. It is a claim ladder. A challenge is named. A very large population of connected computers is imagined. A fraction of that population is assigned a rate of work. The work is aggregated, and a result is expected after a calculated interval. At the end, anyone can test a candidate against a public condition. That last test is crisp. Almost everything before it is conditional.
Published in September 2003 as an Informational RFC, “Chinese Lottery Cryptanalysis Revisited: The Internet as a Codebreaking Tool” revisited an older idea about exploiting widely distributed computing capacity. It used the rapid historical spread of Code Red as evidence that Internet hosts could be reached at enormous scale. The document then explored what such a population might mean for cryptanalytic economics. It did not create an Internet standard, authorize deployment, or establish a measurement system for actual participating machines.
That status matters. An Informational model can expose a risk without proving that every input has been observed. The model asks what follows if its assumptions hold. It does not transform assumptions into receipts merely because the resulting number is memorable.
The first boundary is between a historical population measurement and a future compute fleet. CAIDA documented more than 359,000 Code Red-infected hosts in under fourteen hours. That is evidence about a particular worm, observation method and period. It does not establish that the same hosts were simultaneously available, computationally homogeneous, persistently reachable, under one controller or capable of completing a different workload at the estimated rate.
Population is not capacity. A host can appear in a scan and disappear before useful work. Two observations can refer to one machine behind changing addresses. One observation can conceal many devices behind translation. Some machines may repeat the same assignment, fail silently, return corrupt output or be removed. Counting visible endpoints and multiplying by a benchmark is an estimate, not an execution ledger.
The second boundary is consent. A distributed-computing result can come from volunteers, contracted infrastructure, compromised systems or some mixture. The mathematical candidate does not contain that history. It does not say whether machine owners opted in, whether an operator exceeded authority, whether a third party suffered damage, or whether the computation violated a policy or law. Verification of output cannot retroactively authorize input.
This distinction is easy to lose when the challenge itself is public. A public target permits anyone to know what a successful answer would look like. It does not grant anyone control of other people’s processors. Challenge authorization and compute authorization are different records with different issuers.
The third boundary is unique coverage. Suppose a coordinator claims that a large fraction of a finite space was examined. A trustworthy record would need assignment identifiers, non-overlapping ranges or equivalent work units, issue and completion times, worker attestations, result validation, duplicate detection, timeout handling and reassignment history. Without those records, aggregate work rate can hide repeated regions, abandoned regions and fabricated completions.
A valid candidate does not repair that gap. Search can end early when a candidate is found, so the result says nothing about exhaustive coverage. A candidate can also arrive through an unreported shortcut, prior knowledge or an independent party. The verifier can establish that it works without learning how it was obtained. Result validity and process completeness are orthogonal.
The fourth boundary is attribution. An encrypted or anonymous return channel may protect the person who submits a candidate. That may be part of the threat model, but it prevents the result channel from serving as a provenance channel. A service can know that the answer is valid while remaining unable to prove who commissioned the work, who operated the fleet, which institutions supplied resources or which victims bore the cost.
Credit and responsibility therefore cannot be inferred from possession of the result. The first publisher may not be the discoverer. The coordinator may not own the machines. The machine owner may not know that work occurred. The organisation that verifies the candidate may have had no role in generating it. Each relationship needs evidence outside the candidate itself.
The fifth boundary is time. RFC 3607 used the cryptographic and network conditions discussed in 2003. Its comparisons belong to that period and to the algorithms, key sizes, hardware assumptions and attack models it names. NIST later withdrew DES, while current AES specifications describe different key sizes and security assumptions. Later RFCs revised guidance on MD5 and SHA-1. A historical arithmetic exercise cannot be pasted onto a modern algorithm by preserving only the host count.
Cryptanalytic cost changes with specialised hardware, algorithmic advances, energy, bandwidth, memory, coordination overhead and defensive migration. The meaning of a result also changes when the target is a password-derived secret, a random key, a protocol transcript or a flawed implementation. “Internet scale” is not one reusable multiplier.
RFC 3766 is useful here because it treats key length through attack cost and system context, not as a slogan. RFC 4086 makes a related point about randomness: entropy must be understood at the source and construction, not inferred from a label. RFC 7696 turns the operational consequence into algorithm agility. Systems need a governed way to migrate when assumptions weaken. None of these documents licenses an unscoped claim that a single historical result proves the insecurity of every present deployment.
The evidence architecture should make each transition explicit. First, preserve the exact challenge statement, algorithm identifiers, target material, success predicate and publication time. Second, record the authority that issued the challenge and the authority that approved the resources. Third, maintain a work ledger with unique assignments, worker identity or bounded attestation, issue time, completion time, validation status and retry lineage. Fourth, keep the candidate, verification software, verifier identity and verification transcript. Fifth, record the decision that followed and the scope of systems to which it applied.
Absence should remain visible. If worker identity is unknown, say unknown. If consent cannot be demonstrated, do not replace the gap with an aggregate estimate. If unique coverage cannot be reconstructed, report the claimed arithmetic and the missing ledger separately. If the result came through an anonymity-preserving channel, do not manufacture an attribution from timing or possession.
This is not an argument against models. RFC 3607 made a strategic risk legible: a huge pool of networked general-purpose computers could change the economics of work that institutions once imagined required dedicated machines. The lesson survives precisely when the model remains conditional. Inflating it into a historical fact about an actual fleet makes the warning easier to quote and harder to audit.
Nor is it an argument that output verification is weak. Public verification is powerful. It lets independent parties reject false candidates without trusting the submitter. But that strength has a defined domain. It establishes a relationship between a candidate and a predicate. It does not establish the social, operational and legal history of computation.
The durable operating rule is therefore simple: verify the result, but do not let the result impersonate the missing ledger. The candidate can be valid while the compute history remains unknown. Both statements can be true at the same time, and responsible leadership must preserve both.
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
