Summary

  • RFC 9737 lets a pNFS client use an all-zero anonymous stateid during the metadata server's recovery grace period to report data-server I/O errors observed during the outage.
  • The zero value admits the report; it does not restore the expired layout, authorize ordinary I/O, prove the report true, or prove that an unreported mirror is correct.
  • A clean reclaim, an explicit error, silence, a changed mirror set and an older server lead to different rebuild decisions and must remain different evidence states.

The metadata server restarted. Its old layout stateids were no longer valid. Yet the clients had continued writing directly to data servers while the metadata plane was absent. One client knew that a mirror had rejected a write. The server needed that fact before choosing which copy would repair the others.

The normal credential for saying it had disappeared with the restart.

RFC 9737 resolves that paradox with an identifier that looks like no authority at all: an all-zero anonymous stateid. During the recovery grace period, an NFSv4.2 metadata server must accept that value in LAYOUTRETURN so a client can submit the error it observed during the outage. The same value after grace must be rejected with NFS4ERR_NO_GRACE. A nonzero layout stateid during grace receives NFS4ERR_GRACE.

The zeros do not mean “trust this client with the old layout.” They mean “hear this narrowly scoped recovery evidence now.”

The data plane outlived its coordinator

Parallel NFS separates metadata authority from data access. The metadata server supplies a flexible file layout; the client can then read or write the data servers directly. With client-side mirroring, RFC 8435 makes the client responsible for updating all mirrored copies and reporting I/O errors back to the metadata server.

That report can be precise. ff_ioerr4 identifies a byte offset and length, a stateid, a device, an operation and an error status. RFC 8435 nevertheless calls these indications hints to the metadata server. The client is an important witness, not an omniscient integrity oracle.

A restart breaks the ordinary evidence route. NFS recovery gives clients a grace period to reclaim open state using CLAIM_PREVIOUS, then close the reclamation phase with RECLAIM_COMPLETE. But the old layout stateids are invalid, and the client cannot simply obtain a new layout during grace to report what happened under the old one. Before RFC 9737, the MDS could know that writes occurred without knowing whether all mirrors received them consistently. It therefore had to assume inconsistency and rebuild.

The extension preserves the report without pretending to preserve the old authority. This is a subtle but powerful separation. Evidence may survive the credential that originally framed it. The recovery protocol can admit that evidence through a new, narrower rule instead of restoring a stale credential wholesale.

One zero value, two hard temporal boundaries

The anonymous stateid works only while the recovery grace period is active. That clock is not administrative decoration. It is part of the meaning of the message.

Inside grace, zero is admissible for this LAYOUTRETURN. Any other stateid is rejected because the server cannot treat pre-restart layout state as current. Outside grace, zero is rejected because the exceptional recovery channel has closed. The server also must not increment the sequence ID of the returned layout state when responding to this anonymous form. The report enters the recovery decision without masquerading as an ordinary transition of live layout state.

Its operational value lies in narrowness. The common rule is strict, tiny and locally testable: one value, one operation, one phase, one purpose. Implementations remain free to choose how reconstruction proceeds, but they may not stretch the anonymous value into general authority.

An audit should therefore record more than stateid=0. It needs the restart epoch, the active grace interval, the file and mirror-set identity, the client's recovered write intent, the device and byte range, the operation and status, the exact acceptance response and the server's later decision. Without the clock and context, zero says almost nothing.

Silence is not the same as a clean report

RFC 9737 creates a decision table with three evidence states that dashboards often collapse.

If the client reclaims the file and reports no errors, the metadata server must not resilver it. If the client reports an error, the file must be resilvered. If the client neither reclaims the file nor reports an error before grace ends, the server must resilver because the client may have restarted and lost its state.

The second and third branches both end in reconstruction, but for different reasons. An explicit error is positive evidence about a problem the client observed. Silence is missing evidence. Treating silence as “no error” would turn a crashed or disconnected client into a clean bill of health.

The first branch also needs discipline. “No error reported after reclaim” is sufficient for the protocol's no-resilver decision. It is not a byte-level proof that every byte on every mirror is correct. It says the recovering client did not encounter and report an error through the specified path. Other corruption mechanisms, stale reads, application-level invalidity and observations held by another client remain outside that receipt.

The recovery decision must honor exactly what the system observed, not upgrade a negative report into universal truth.

A correct report can become stale before it is heard

The evidence is also bound to a mirror topology. If the returned layout does not match the file's current mirror instances, the MDS must ignore the LAYOUTRETURN and resilver. The client may be perfectly accurate about the former set. Accuracy about a superseded mirror set does not authorize a present decision.

This is the part most likely to disappear in a generic recovery event. A log line can say “client reported device error” while omitting that the device set had changed. The operator then sees a report and a rebuild and assumes the first caused the second. In fact, the server may have rejected the report as stale and chosen the conservative path independently.

The evidence chain needs both identities: which mirror set the client observed, and which mirror set the MDS considered current. Only their match gives the report decision value.

Reconstruction has its own authority and order

An accepted report is not a completed repair. RFC 9737 describes a broad sequence: fence the file, record that resilvering is required, release the write intent, and start reconstruction only when no outstanding write intents remain. The MDS must not rebuild while a client still holds one.

What happens to access during copying is an implementation decision. The server may block I/O, force it through the MDS, or insert a proxy that updates the new copy while reconstruction runs. If it continues access, the client must see the same layout set as before the restart, and a proxy used for the copy must remain until grace finishes.

Those choices belong to local running code, but the outcome needs independent verification. The selected source mirror must be readable. The copied byte ranges must converge. Later reads must return the intended data. The application must accept it. A LAYOUTRETURN response proves none of those stages.

Compatibility fails safely, but not cheaply

An older metadata server does not have to understand the new anonymous-stateid behavior. It returns NFS4ERR_BAD_STATEID; the client should fall back to the old behavior and stop trying to report the outage error through this path. That avoids pretending the evidence was accepted.

The cost is conservative rebuilding. Compatibility preserves safety by spending I/O, time and capacity rather than inventing confidence. Operators need to see that distinction. A successful fallback does not mean the extension worked; it means the system reverted to a more expensive recovery policy.

The useful operational record therefore has separate fields for client support, server support, zero-stateid attempt, response, fallback, recovered file, error array, mirror-set comparison, rebuild decision, fencing, outstanding write intents, copy progress and post-repair verification.

The all-zero stateid is not a loophole in NFS authority. It is the opposite: a carefully bounded way to preserve evidence without preserving obsolete power. The identifier may be empty, but the recovery decision it informs is not. Leadership should insist that storage systems retain the difference between a report being heard, a report being applicable, a repair being ordered and the data being proven usable.

Sources