Summary

  • Parnas compared two KWIC systems that could use the same algorithms and even produce identical runnable representations. Their difference was where design knowledge lived: process-step interfaces exposed shared formats, while information-hiding modules confined volatile decisions.
  • The change table was concrete. Moving lines out of memory or changing character packing touched every module in the first decomposition but only Line Storage in the second; changing the representation of circular shifts touched three modules in the first and one in the second.
  • Information hiding is not secrecy, access control, object orientation or automatic speed. The paper itself warns about procedure-call overhead and calls one of its own supposedly hidden interfaces a design error for revealing an unnecessary ordering.

Begin after the acceptance test

Imagine a small test harness. Feed both implementations the same titles. Check every circular shift and the final alphabetical order. Both pass. If correctness at this moment were the only criterion, the architectures would be tied.

Parnas deliberately denied the reader that easy result. In On the Criteria To Be Used in Decomposing Systems into Modules, published in Communications of the ACM in December 1972, he says both schemes work. They may share algorithms and data representations. After assembly, their runnable forms could even be identical. The distinction appears in the representations used to change, document and understand the system: how work is assigned, which decisions cross interfaces, and who must be consulted when one of those decisions changes.

That is why the useful experiment starts only after output equivalence has been established. Replace the line representation, stop holding all lines in memory, calculate shifts on demand, or spread alphabetization across requests. Count the work assignments that must understand the new choice. The resulting blast radius is an architectural property that a snapshot acceptance test cannot see.

KWIC was a real indexing problem before it was a design parable

KWIC means Key Word in Context. Hans Peter Luhn’s 1960 paper described a machine-generated index in which a keyword is displayed with its surrounding title words, allowing a current index to be produced mechanically. Parnas did not invent that information-retrieval form. He chose a compact, already intelligible system on which competing decompositions could be compared without burying the argument under a large application.

His simplified KWIC accepts an ordered set of lines. Each line contains words, and each word contains characters. Repeatedly move the first word to the end and a line yields its circular shifts. The system outputs all shifts of all lines in alphabetical order.

The example was intentionally small—Parnas estimated that a good programmer could build it in a week or two. The point was not that KWIC demanded industrial-scale project management. It was that a small system made hidden assumptions visible enough to inspect.

The first cut followed time

The conventional decomposition mirrored the processing sequence. Input read lines and stored them in core. Circular Shift constructed an index of shifts. Alphabetizing reordered that index. Output formatted the result. Master Control sequenced the other four.

This arrangement looked modular. Each task had a name; the program was divided into manageable pieces. Yet the interfaces were not merely statements of what service one task needed from another. They included core layouts, indexes, pointer conventions and table formats. Input packed four characters into a word. Circular Shift and Alphabetizing knew how earlier arrays were represented. Output used arrays produced by both the alphabetizer and the input stage.

The processing steps were separate, but knowledge of the storage choices was shared. A flowchart divided control while leaving several changeable decisions in the seams.

The second cut followed decisions

The alternative decomposition did not map each module to one phase. Line Storage owned the representation of lines and exposed operations for addressing characters, words and lines. Input used those operations. Circular Shifter presented the appearance of a collection of shifted lines without requiring clients to know whether the shifts were stored, indexed or calculated when requested. Alphabetizer exposed an ordering operation; Output consumed the resulting abstractions; Master Control coordinated the work.

Here “module” meant a responsibility assignment, not a subroutine. That distinction is central. One module could contain several callable routines. Conversely, an efficient assembled program could contain code contributed by more than one module. A source file, class, package, process and team can sometimes implement a module boundary, but none is automatically the boundary Parnas meant.

The criterion was knowledge. Each module was characterised by a difficult or likely-to-change design decision that it knew and others did not need to know. Its interface should reveal as little about that decision as clients require.

Five changes exposed the difference

Parnas’s comparison is more useful than the slogan because it identifies the affected sets.

First, an input-format change remained inside Input in both decompositions. Information hiding did not win every row of the table; the conventional design had already placed that choice well.

Second, suppose the system could no longer keep every line in memory. In the first decomposition, every module used the core line format, so every module changed. In the second, only Line Storage knew whether lines were resident, paged or otherwise represented.

Third, replace the decision to pack four characters into one machine word. Again, the first design spread that format across every program, while the second confined it to Line Storage.

Fourth, change circular shifts from an index into the original lines to explicitly stored lines—or calculate every requested character on demand. Circular Shift, Alphabetizer and Output knew the representation in the first design. In the second, Circular Shifter alone owned it.

Fifth, stop alphabetizing the whole list once. Search for the next item as needed, or distribute sorting work across the period when the index is produced. In the first design, Output expected a finished index. In the second, clients could not detect when the alphabetization occurred, so the alternative stayed inside Alphabetizer.

These are not promises that every future requirement will change one file. They are evidence that named assumptions can be assigned deliberately. The first decomposition optimised the picture of execution. The second optimised the distribution of knowledge about volatility.

An interface can hide the mechanism and still reveal too much

Parnas did not present the second design as flawless. His Circular Shifter interface specified that shifts of earlier lines came first and that each line’s original version preceded its successive rotations. Clients did not need that order. It prevented an implementation that generated shifts directly in alphabetical order and made the Alphabetizer nearly empty.

He classified that choice as a design error. The method of storing or calculating shifts was hidden, but an unnecessary ordering guarantee had escaped. This is a stricter test than “access the data through functions.” A wrapper around a representation can reproduce all of its volatility in method names, return shapes, enumeration order or timing guarantees.

The question for an interface review is therefore not only whether internal fields are private. It is whether each public fact is required by a client. Information hiding reduces the number of people and components entitled to rely on a decision; it does not praise opacity for its own sake.

What the criterion does not mean

The word “hiding” invites a security reading that the paper does not support. It is not encryption, authorisation, sandboxing or runtime isolation. A secret can be securely encrypted and still be represented by a schema that every module must understand. A volatile decision can be fully public and still be confined behind one programming interface.

Nor is information hiding simply object orientation. An abstract data type or a class can be a useful mechanism for controlling access to a representation. It can also expose implementation order, persistence details or a changeable taxonomy. Parnas’s criterion applies in languages without classes and across work assignments larger than a single type.

“High cohesion” and “low coupling” are helpful later vocabulary, but they are not substitutes for naming the hidden decision. Compile dependencies are different again. A build system may recompile a client after a library changes even when the client’s source and reasoning remain valid. Conversely, a service may deploy independently while its consumers still depend conceptually on an undocumented ordering or error convention.

Performance is not automatic. Parnas warned that the second design could be slower if every fine-grained operation became an elaborate procedure call. He sketched specialised transfers and assembly-time insertion as ways to retain the design boundary without paying conventional call overhead. The trade-off must be engineered; the principle does not make it disappear.

Parnas later found another leak in KWIC

Later reflection sharpened the lesson. Parnas observed that the original example still let modules assume that a string was a sequence of characters. That shared assumption could prevent representing frequently used strings by compact integers. A celebrated information-hiding example had not hidden every representation decision that might matter.

This admission protects the idea from ritual. A module diagram is not evidence that the right information is hidden. Designers have to identify future decisions under uncertainty, and they can be wrong. The remedy is not to declare an API permanent. It is to keep revisiting what clients genuinely need to know.

The idea grew through a community

The 1972 paper built on Parnas’s 1971 work about information distribution, which described connections between modules in terms of assumptions they make about one another. His 1976 work on program families widened the time horizon: if related programs will exist, their common properties should be designed before each version is treated separately.

In 1985, Parnas, Paul C. Clements and David M. Weiss jointly published The Modular Structure of Complex Systems. They distinguished module structure, uses structure and process structure, and described a module guide that helps maintainers find the parts they must understand without reading irrelevant details. That is a continuation, not a footnote: it prevents a work-assignment structure from being confused with call order or runtime processes.

Later researchers used KWIC to compare shared-data, pipe-and-filter, event-based and other architectural styles. Abstract data types, encapsulation mechanisms, software architecture and product-line methods developed the conversation further. The 1972 result remains powerful without claiming ownership of everything that followed. Its durable contribution is the experiment: hold observable output constant, perturb a decision, and see how far the need to know travels.

Sources