Summary

  • RFC 1595 divided SONET/SDH performance evidence by Section, Line, Path and Virtual Tributary, by near and far end, and by current or completed 15-minute interval. A bare integer discarded the coordinates that gave it meaning.
  • Unavailable time began with the first of ten contiguous severely errored seconds. If the run crossed the 900-second boundary, a later poll could legitimately change SES and UAS values in the previous interval.
  • RFC 2558 and RFC 3592 exposed two implementation choices: publish performance counts promptly and correct them retroactively, or hold one-second observations in a ten-stage delay line. Immediacy and finality could not both be free.

The closed window that was still open to evidence

Imagine a 15-minute performance interval ending while a SONET Path has accumulated only part of a severe-error run. The next interval starts. A manager polls the row for interval 1—the period that has just closed—and receives a set of counts. A few seconds later, the tenth contiguous severely errored second arrives.

The physical seconds have not moved. Their classification has.

Under the rule captured by RFC 1595, unavailable time begins at the onset of ten contiguous severely errored seconds, not at the tenth second. All ten belong to unavailable time. If the run began before the quarter-hour boundary, some of the newly classified unavailable seconds belong to the completed interval. The RFC therefore warned that successive GETs of the previous Path, Line or VT interval could return different SES and UAS values when the first read occurred in the opening seconds of the new window.

This was not a vague disclaimer about imperfect hardware. It followed from the object model itself. The agent could expose a row as “most recently completed” before it possessed the future observations needed to settle the status of its final seconds.

Ten seconds reached backward

The unavailable-time rule worked in both directions. At the Line, Path and VT layers, ten contiguous severely errored seconds made the interface unavailable from the first of those seconds. Once unavailable, ten contiguous seconds without an SES restored availability from the first clean second; those ten clean seconds were excluded from unavailable time.

While the layer was available, its applicable error counters advanced. While it was unavailable, only the unavailable-seconds counter advanced at that layer. The ten-second sequence was therefore not merely an alarm debounce. It decided which ledger received the preceding seconds.

That distinction matters. A management system might see an errored second immediately, but it could not know immediately whether that second would remain an ES or SES observation or later fall inside unavailable time. “Observed now” and “finally classified now” were different propositions.

At a 900-second rollover, the difference crossed a storage boundary. Interval 1 was already addressable as history. Yet the state machine continued; it did not reset when the counters rolled over. The previous row could still inherit the result of a sequence completed in the new interval.

A count had a layer before it had a meaning

RFC 1595 was careful about where an error existed. SONET/SDH equipment did not terminate every layer at every point. A regenerator might terminate only the Section and remain transparent above it. Add-drop multiplexers and digital cross-connects terminated Lines. Terminal multiplexers could also terminate Paths and Virtual Tributaries or Virtual Containers.

The MIB reflected that architecture through interface entries and their stack. Section coding violations came from the B1 byte. Line violations came from B2. Path violations came from B3. A floating VT used BIP information in V5. Section status could expose loss of signal or frame; Line and Path surfaces had their own alarm and remote-defect conditions.

The common names—coding violations, errored seconds, severely errored seconds—did not erase these origins. Two values could be numerically equal and operationally incomparable because they belonged to different layers, rates or termination points. A collapsed “circuit errors” total might be convenient, but it could not preserve the localization built into the standard.

The same caution applies to ifOperStatus. RFC 1595 linked a combined Medium/Section/Line entry's down state to Section or Line defects and linked a Path entry's down state to Path status. “Down” was a useful projection for the interfaces table. It was not, by itself, the physical address of the fault.

The far end was a different witness

The MIB also separated near-end and far-end statistics. Far-end Path performance was derived from the far-end block error indication in the G1 Path Overhead byte. It used familiar metric names, but the measurement arrived through a different channel and described the remote side's report.

RFC 2558 later made an absence rule explicit. Far-end one-second statistics were to be marked absent during a second when an incoming defect existed at the same layer or a lower one. The lower-layer fault could destroy the path by which the remote evidence became interpretable.

Zero and absent are not the same value. A zero can say that the relevant counter observed no qualifying event under its rules. Absent says that the observation was not available for that second. Replacing absence with zero makes the impairment look like proof of remote cleanliness precisely when visibility is weakest.

Revise the answer or delay it

RFC 2558, which replaced RFC 1595 in 1999, described the timing problem more explicitly. An agent that updated performance statistics in real time had to be prepared to correct earlier ES, SES, SEFS, CV and UAS values once the ten-second availability classification became known. When the sequence crossed the 15-minute boundary, those corrections could reach the previous interval.

The alternative was a ten-element delay line. Each second's near- and far-end violation counts and alarm flags entered the line; performance counters were updated only after the state associated with that observation was known. This kept the published allocation correct without later revision, but it made the counters lag physical time by ten seconds.

Neither representation was inherently dishonest. One offered lower latency with a provisional edge. The other offered stable allocation with explicit delay. The dangerous design was the collector that documented neither, sampled once, stripped the poll time and later compared the stored integer as though it were an immutable event total.

RFC 2493 generalized the surrounding bookkeeping for 15-minute histories. It distinguished current, historical and optional aggregate tables, included elapsed time and valid/invalid interval concepts, and allowed fewer complete rows after a restart or where proxy data was unavailable. “Previous 24 hours” was therefore a capacity envelope, not a promise that 96 complete and valid intervals always existed. RFC 1595 itself required at least four completed intervals and gave 32 as a default for its tables.

The trap could name an earlier time

The later SONET MIB text carried the same temporal distinction into notifications. A link-down trap was to be sent only after the agent knew for certain that unavailable state had been entered. Its effective time, however, referred to the first unavailable second—ten seconds earlier. Link-up followed the same logic.

Emission time and event-effective time were both true, and they answered different questions. One showed when the agent possessed enough evidence to notify. The other showed where the state machine placed the transition. If a collector retained only arrival time, it shifted the impairment forward. If it retained only the effective time, it could falsely imply that the agent knew and reported the transition at once.

An audit trail needed both, together with the policy that created the difference.

Even “severe” needed provenance

Standards evolution added another reason not to detach a counter from its semantics. RFC 2558 recorded that several standards used different thresholds for declaring a severely errored second. It added sonetSESthresholdSet so an agent could identify which interpretation was in force. An agent did not have to support every set, and changing the selected set invalidated previously collected SES statistics.

The number did not become less exact. Its category had changed. Historical comparison required the threshold choice and the time at which it applied. Without those fields, a trend line could show an artificial improvement or deterioration caused by a semantic change rather than a change in the optical signal.

RFC 3592 replaced RFC 2558 in 2003 while retaining the central structure: layer-specific observations, near/far-end performance, 15-minute history and the correction-or-delay problem. The lineage matters because it shows that the awkwardness was not an editorial accident quietly forgotten. The later standard explained and preserved the trade-off.

What the record can prove

A defensible SONET performance snapshot identifies the managed instance and stack position; Section, Line, Path or VT layer; near or far end; metric definition; threshold set; current or historical interval number; interval start and rollover; elapsed time; valid and invalid interval counts; agent epoch; poll time; and whether values are provisional, correctable or delayed.

That record can establish what the agent exposed under a named standard and configuration. It cannot by itself prove the failed component, customer impact, repair action, vendor conformance or end-to-end restoration. A later revision can establish that the agent reclassified the window; it does not rewrite what happened on the fibre. A ten-second lag can establish a chosen reporting method; it does not prove the device is slow or faulty.

Lu Heng's Running-Code Primacy makes the operative test concrete: the management symbol must stay tied to what the implemented agent observed, classified and exposed. Minimum Initial Specification explains why the common layer should standardize the thin receipt while leaving polling, retention and response choices visible as local decisions. Why BTW.Media Exists supplies the editorial obligation: report the revision, latency and missing evidence instead of turning the first available number into a campaign.

The quarter-hour was useful because it ended. It became trustworthy only when the system also said whether its ending values were final.

Sources and evidence limits

The technical account rests on the RFC Editor records and texts for RFC 1595 and its full specification, RFC 2558 and its full specification, RFC 3592 and its full specification, and RFC 2493 with its 15-minute performance-history conventions. The three Lu Heng essays supply interpretive discipline, not SONET facts.

The document lineage is precise. RFC 1595 was published on the Standards Track in March 1994; RFC 2558 replaced it in March 1999, and RFC 3592 replaced RFC 2558 in September 2003. The security treatment also changed: RFC 1595 did not discuss security, while RFC 2558 warned that GET and SET access could expose sensitive configuration and control information.

These sources establish document identity, standards lineage and object semantics. They do not establish a named implementation, deployment prevalence, circuit condition, outage, customer impact, notification delivery, collector behaviour, security posture or successful remediation. The opening is a logical reconstruction of the specified timing rule, not a reported incident.