Reconciler

The reconciler expresses deployment and destruction as Yieldstar async generators. Notation owns desired-state decisions and provider lifecycle; the caller's Yieldstar runtime owns durable execution, waiting, and shared state.

Deploy flow

deploy takes the deployment hold, walks dependency levels in order, decides an action for every resource, executes provider calls as durable steps, persists the result in a resource store, and deletes registered orphans.

ConditionDecision
Not in statecreate
In state, params changedupdate
In state, params unchanged, no driftnoop
In state, but deleted from the providerdrift-recreate
In state, provider state differs from stored statedrift-update
In state, not in graphdelete-orphan

Dry-run deploy performs decisions and emits lifecycle events without provider mutations or state mutations. When drift detection is enabled, it can still call provider read operations to decide whether a nominal noop has drifted.

Destroy flow

destroy is a first-class durable operation. It takes the same deployment hold as deploy, deletes desired resources in reverse dependency order, deletes registered persisted orphans, and conditionally removes each resource store only after the provider delete succeeds or reports that the resource is already absent.

Waiting and replay

Provider calls are stable durable steps, but provider acknowledgement and the Yieldstar heap checkpoint are not atomic. If the process crashes between them, replay repeats the call, so provider create, update, and delete operations must be idempotent. Event subscribers must likewise tolerate duplicate delivery when a crash occurs before the event checkpoint.

A resource operation throws ResourceOperationPendingError when it has not finished. The error gives the reconciler a delay and optional callback context. The runtime stores the context, waits without keeping the process busy, and calls the same operation again. See Operation errors for the complete API.

Each attempt, delay, event, state read, state write, and hold change has a stable step key. A resumed execution must use the same execution ID. A new deploy or destroy must use a new execution ID.

State and the deployment hold

Each resource is stored under notation/resource-state with a deployment-scoped ID. Conditional updates and deletes compare the snapshot's UUIDv7 instanceId and version, so a stale execution cannot modify a deleted and recreated store.

Deploy and destroy take an exclusive hold on the deployment through one notation/deployment-hold store per deployment. store.take suspends a competing execution as a durable waiter and wakes it after the holder releases. A waiter that finds the hold already taken when it inspects it emits reconciler.hold.waiting naming the holding execution ID, so a wait behind a crashed execution is visible instead of silent; a holder that appears only between that inspection and the take suspends the waiter without the event.

A failed or suspended execution keeps its hold, which is what makes resuming it safe. The hold of an execution that will never be resumed is cleared with clearDeploymentHold from @notation/reconciler/durable — the only supported way out of that state.

Events

The durable workflows emit these events:

EventWhen
reconciler.deploy.decisionAfter deciding what action to take for a resource
reconciler.drift.detectedWhen drift is found between stored and actual state
reconciler.operation.lifecycleWhen an operation starts, finishes, skips, or fails
reconciler.hold.waitingWhen another execution holds the deployment
reconciler.orphan-deletion.skippedWhen no registered class can delete an orphan

Lifecycle events cover create, read, update, and delete with start, success, error, skip, or dry-run status.