Summary

  • Archie turned remote FTP directory inventories into searchable names and locations; it did not move the listed files into the index.
  • Its May 1991 manual described nightly collection from a subset of sites but an approximately monthly refresh for each site, so a result reported an observation whose age mattered.

The date was part of the answer

A search result looks like a statement about the present. In Archie’s case, it was closer to a dated note: this name appeared at this path on this archive the last time the system collected that site’s directory. The distinction was easy to miss if a user saw a plausible filename and treated the match as proof that the file could still be fetched.

The service built its index from anonymous FTP. A manual dated 20 May 1991 says Archie’s database subsystem knew about roughly 600 FTP archive sites. Each night it connected anonymously to a subset and fetched a recursive directory listing—or, where one existed, a file containing that listing. Each site was updated about once a month. Archie stored compressed listings at McGill’s quiche.cs.mcgill.ca, where the Internet community could retrieve them through anonymous FTP. The crawler did not need to copy every archive’s software, documents and data to McGill. It collected the names and paths that made those remote holdings easier to locate. (1991 Archie manual)

The user-facing commands exposed the difference between a name index and a file server. prog searched names and returned information such as the archive host, path, file size and modification date. site could print a complete listing for a known archive. The list command showed when a site’s inventory had last been updated. After finding a match, the user still had to connect to the archive and transfer the file. In 1992, RFC 1325 described the same handoff: Archie returned an archive name, IP address and location; the user’s next step was at the remote site. In 1994, RFC 1689 called Archie a “secondary source” and explicitly described using FTP to fetch a match from the archive itself. (RFC 1325; RFC 1689)

That split gave the index a clock. A directory could change between collections. A file could move, be renamed or disappear while Archie continued to report the last listing it had gathered. The May 1991 manual did not hide the age entirely: it exposed last-update information. But it also did not claim that a search result was a live check. A result was evidence that a name and location had been observed; only the archive’s current state and a successful transfer could establish that the file was still there and usable.

The boundary was also visible in the system’s coverage. The same manual listed “Only UNIX sites” as a limitation and said users could not constrain a search to particular sites. Those are not small interface omissions. They shaped which archives could appear and what a user could ask the index to distinguish. A central directory made a large set easier to search, but its view was not the Internet’s complete file system and it did not make unlike archives uniform.

In May 1991, this was an expensive service to keep local. The manual estimated a database of about 70 MB and said updates and searches put noticeable load on the Sun 4/280 hosting it. Archie was still described as developmental; its software was not yet being released to outside sites, and distribution to other servers was a long-term plan. The collection schedule, storage footprint and query load therefore belonged to the same design problem. Crawling more often might have reduced the age of observations, but each refresh also required network work at the archives and processing at the index host.

The service grew quickly enough that its counts must be read as snapshots, not as a single timeless total. Emtage and Deutsch’s Winter 1992 USENIX paper includes a table dated 30 October 1991: 1,025 sites were known, 886 were indexed, and the database held 1,502,976 file references alongside 686,104 unique filenames. The difference between references and unique names matters: a filename was not necessarily one globally unique object. The May 1992 RFC later described roughly 1.5 million filenames at around 900 archives and nine Archie servers worldwide. It said a user could choose a closer server to relieve some load on McGill. That documents broader access and a load-reduction goal; it does not establish that every server held an identical or equally fresh listing set. (Emtage and Deutsch, 1992 USENIX paper; RFC 1325)

By the 1993 update recorded in RFC 1689, the commercial Archie server system had about 27 installations, not all publicly available. The path from one McGill machine to more servers widened access, but it did not collapse discovery and delivery into one authority. Archive operators still controlled the bytes; Archie assembled a useful, dated account of where names had been seen. (RFC 1689)

Archie’s contribution was not that a central answer made a remote file current. It made the Internet’s public file collections legible enough to search without pretending to own them. That design shifted the hard question from “Where might this file be?” to “How old is this observation, and can the archive still serve it?” Search was a route to evidence, not the evidence’s source.

Sources