要約

  • RFC 3139は、Gold serviceのようなnetwork intentが、topology、status、capabilityを使って各vendor deviceのlocal configurationへ変換される必要を示した。
  • network全体では複数translatorが連携できるが、一つのdeviceには同時に一人だけが操作できた。policy承認、candidate生成、write、feedback、実trafficは別のreceiptである。

Goldはrouter commandではなかった

operatorは短い言葉で期待を示せる。このcustomer groupにはGold serviceを提供する。しかしdeviceは商業上の色を理解しない。interface、classifier、queue、scheduler、bandwidth、failure pathを具体化しなければならない。

同じ目的でもvendorごとにparameterが違い、一台に多数のcommandが必要になる。backup pathはmain pathと同じprofileでは足りないかもしれない。topology変更はpolicyの対象deviceを変える。

RFC 3139が示したのは、この抽象の価値と限界だった。high-level policyはnetwork behaviorを表すが、device stateでもdelivery evidenceでもない。

2001年6月のInformational RFCはprotocolを選ばなかった。COPS/PIB、SNMP/MIBや個別技術の解決策がfragmentする懸念から、1999年の議論後に共通requirementを整理した。文書のMUSTはfuture systemへの要求であり、当時のdeploymentが満たした証明ではない。

intentとdeviceの間に三つの表現があった

RFCはhigh-level configuration data、network-wide configuration、device-local configurationを分けた。最初はmodelと期待、次は複数deviceへ展開できる共通表現、最後は一台固有の最細粒度である。

configuration-data translatorが層を渡る。humanでもcentral softwareでもintermediate entityでもdevice内でもよい。固定されたのはfunctionであり、物理配置ではない。

変換にはpolicyだけでなくtopology、status、performance、monitoring、capabilityが必要だった。古いpath mapから生成したcommandはsyntaxが正しくても配置を誤る。model上のqueueをfirmwareが持たない場合もある。

error-free conversionに必要な情報がなければ、solutionはそれを検出し、対処しなければならない。停止、延期、範囲縮小、人手確認のどれかはlocal policyである。不明を成功として扱うことだけは許されない。

pipelineは複数でも、device writerは同時に一人

複数translatorはtandemに動けた。business intentをnetwork-wide modelへ、そこからtechnology profileへ、最後にvendor commandへ変換できる。

しかしRFC 3139は、特定deviceのある瞬間には一つのtranslatorだけがoperateできるとした。networkには多段処理があっても、deviceに競合authorを置けない。

これはlock protocolそのものではない。leader election、lease、fencing、crash recoveryは定義しない。守るべきinvariantを示した。異なるrevisionとsnapshotから二つの合理的candidateが生まれ、interleaveすると誰も望まない状態になる。

一方がqueue limitを変更し、他方が古いGold profileをrestoreすれば、新classifier、旧scheduler、失われたexpiryが残り得る。各commandの成功はcoherent final stateを証明しない。

concurrent shared writeによるmisconfigurationを排除するrequirementも同じ境界である。writer identity、policy revision、input snapshot、対象device、排他intervalを残す必要がある。last-write-winsは順番を決めても理由を回復しない。

localに正しい二台の間でnetworkは壊れる

複数deviceが一緒に変わるとき、complete/partial configurationを同時または必要な同期順序でadd、modify、delete、dump、restoreできなければならない。

partialが常に悪いわけではない。inactive objectは先に置ける。しかしrouteがfilterより先なら漏れ、新markが未対応coreへ入ればserviceが消え、migrationの半分だけならloopやblackholeになる。

RFCはdata-specific error検出とrecovery、不適切なpartial configurationの防止を求めた。ただしdistributed atomic commitやuniversal rollbackを保証しない。

orchestrator batch、各device accept、stored state、applied state、traffic outcomeを分ける必要がある。あるactorのreceiptを別actorの証明にしてはならない。

fast failoverは事前のconfigurationだった

failure時に大量downloadせず切り替えるため、複数device-local configurationを事前provisionするrequirementがあった。速いfuture actionは過去の準備で成り立つ。

preloadはactivateではない。activateもcurrent suitabilityではない。topology、customer、capabilityが変われば、昨日のbackupは今日のriskになる。

redundant platformとnetwork elementも必要だった。そこでwriter ownershipが問題になる。new controllerはどのsnapshotを継ぎ、old leaderはどうfenceされるか。二つのlive instanceだけでは安全な冗長性にならない。

feedbackはdevice loopだけを閉じた

deviceはconfiguration confirmation、status、monitoring、eventを返す必要があった。feedbackがなければ一方向publishである。feedbackがあればreconcileできる。

ただしconfirmationのverbが要る。parse、validate、candidate storage、commit、effective、restart persistence、forwarding applyのどれか。successの終点は同じでない。

local configuration/statusをnetwork-wide contextで解釈する必要もあった。queueが正しくinstalledでもpath変更で無意味になる。local差異は異種hardware上の正しい共通intentかもしれない。

全deviceがconfirmしてもend-to-end serviceではない。trafficが別pathを通る、classifierが対象packetを外すことがある。Gold installedとGold experiencedを分ける。

configurationには有効期間があった

effective timeとexpiration timeがrequirementだった。expireするitemとnever-expire itemがある。future valueはvalidでもinactive、expired valueはstoredでもauthorityを失う。

clock skewでcontrollerとdeviceの現在が違う。old snapshot restoreで期限切れruleが復活する。receiptにはauthored time、effective、expiry semantics、clock basis、active intervalが必要である。

event-driven provisioningでもevent identity、policy revision、acting writer、recovery lifetimeが要る。これがなければfast feedbackがoscillationになる。

traceabilityはcorrectnessだった

access control、authentication、integrity、replay protection、必要なprivacy、host/user granularityのtraceが求められた。well-formed configurationでもunauthorized writerならinvalid transitionである。

過去に合法だったsigned policyのreplayも危険である。authenticationはsender、integrityはbytes、freshnessとauthorizationは「今、このactionを許されるか」を答える。

drift調査にはfinal diffだけでなく、policy、network-wide projection、snapshot、device owner translator、candidate、response、error、次writer前のfeedbackが必要になる。

evolutionは意味の欠落を隠してはならない

data model、message、typeの進化をfleet replacementなしに扱い、MIB/SMIの経験を活用することも要求された。

new fieldをold translatorが知らず、deviceがpartial capabilityだけを持ち、controllerがrender不能部分をomitすることがある。syntax-valid outputでもintentは欠ける。

capability discovery、unknownの明示、visible degradationが必要である。RFC 3139はold deviceがfuture policyを全て実現すると約束しなかった。

RFC 3535のworkshop、その後のNETCONF、NMDAはdatastore、lock、validate、commit、intended/applied/operationalを具体化した。問題の継続を示すが、2001年にRFC 3139がそれらを実装済みだった証拠ではない。

network policyは何が真になるべきかを言う。disciplineあるtranslation、exclusive write、coordinated apply、feedback、observationだけが、何が真になったかを示す。