Lifecycle Runtime — Workflow Registry

> Canonical source: workflow_runtime.py loads an item's immutable > workflow_id / workflow_version_id pin. Registry definitions are served by > yoke workflows definition get.

This document describes the runtime contract. Definitions own ordered stages, labels, terminal stages, gates, policies, entry surfaces, and registered skill bindings. Live transition, frontier, scheduler, QA, approval, and delivery paths all interpret the item's pin.

<!-- BEGIN GENERATED: field-note-directive --> When you hit a recipe gap or notice a minor bug best held as a supporting record, file a field-note immediately — before retrying, before moving on. yoke ouroboros field-note append --kind <failed|new|unclear|observation> --evidence '...' Run yoke ouroboros field-note append --help for the worked failure modes and decision tree. <!-- END GENERATED: field-note-directive -->

Item stage authority

Do not copy a progression into operator logic or documentation. Use yoke workflows definition get for current definitions. A live item loads the exact pinned version, so a newer one cannot alter work in flight, and a forward move it does not declare is refused before the target stage materializes. Backward rework uses definition.stages order and needs no declared edge: yoke_core.domain.workflow_declared_transitions.undeclared_forward_transition checks only forward moves. The holder uses yoke lifecycle transition to return to the requested earlier stage; ordinary claim and target-stage checks still apply.

Exceptional Item States

These are reachable from multiple points and are not part of the normal forward progression:

  • cancelled
  • stopped
  • failed

> Item-level blocked is not a lifecycle status. It is an orthogonal > flag-and-reason pair on the item that preserves the lifecycle position > (cross-reference: see your items packet stanza for the > blocked/blocked_reason columns). Set it via > yoke items block PREFIX-N --reason "<reason>"; clear via > yoke items unblock PREFIX-N. > The board renders blocked items in their own section and the frontier > routes them to WAIT. The doctor health checks HC-blocked-status-drift > and HC-blocked-flag-consistency surface leftover flags and > legacy status='blocked' rows. Epic-task blocked stays > as a status. Full architectural-why (yoke source repo): > docs/archive/decisions/blocked-flag-retirement.md.

Canonical Epic Task Progression

Epic tasks mirror the implementation-family vocabulary:

planning
-> plan-drafted
-> refining-plan
-> planned
-> implementing
-> reviewing-implementation
-> reviewed-implementation
-> polishing-implementation
-> implemented
-> release
-> done

Task exceptional states:

  • blocked
  • stopped
  • failed

Epic tasks do not use item-only statuses such as cancelled.

Ownership Boundaries

Definition-bound segments

At a live item stage, the owner is the registered skill binding whose half-open interval contains that stage: from_stage_id <= current_stage < through_stage_id. Stage names do not select the skill by themselves, and a workflow name is never a substitute for reading the pinned version.

The current built-in definitions reuse stage ids such as idea, refining-idea, planning, and planned, but each definition decides which of those stages exists, their order, and the skill that owns the segment.

idea -> refine handoff: two-layer guard against title-only dispatch

/yoke idea writes the row in two phases — items add lands the PREFIX-N row with empty spec, and body-and-sync.md writes the structured spec fields a few seconds later. The window between the two phases is unprotected unless both layers below hold:

  • Layer 1 — claim-on-create (live-race fix). infer-and-create.md

step 5b acquires a draft work claim with reason draft-in-progress immediately after items add returns the PREFIX-N id, and body-and-sync.md step 10b releases it with reason idea-complete once the spec/body, AC normalization, and every enabled File Budget or path-claim artifact has landed. The release path canonicalizes idea-complete → handed_off for schema storage and preserves the original intent on the WorkReleased event. While the draft claim is held, the scheduler filters the row out via the standard live-claim conflict gate. Held duration is recorded on the IdeaClaimHeld event for doctor and Ouroboros observability.

  • **Layer 2 — body-completeness skip on the frontier (structural

defense).** yoke_core.domain.frontier_compute calls yoke_core.domain.idea_body_completeness.is_idea_body_incomplete for every status='idea' row and pushes the title-only ones into blocked with reason idea-incomplete. This catches every tail case Layer 1 cannot reach: a /yoke idea session that crashes between the two phases (the title-only body remains, while the active claim stays held until explicit release or authorized termination); a manual python3 -m yoke_core.cli.db_router items add from ad hoc tooling that bypasses the claim convention; any future /yoke idea variant that forgets to acquire the claim. The doctor health check HC-incomplete-idea-bodies reports items in this state so the operator can rescue or freeze them.

Implementation and review

  • implementing means work is actively being built.
  • reviewing-implementation means coding/self-verification is complete and the branch is in the deliberate review/fix loop.
  • reviewed-implementation means meaningful implementation review passed and the work is queued for finishing polish.

When a definition declares this loop, its active skill binding drives it. For example, current definitions bind either implement, conduct, or a direct skill across implementation work; the stage name alone does not choose one.

Claim continuity across transient SessionEnd. A Claude Desktop SessionEnd event (laptop sleep, app reload, idle timeout) never destroys mid-flight claims: the hook runs the non-destructive end_session_if_empty, which ends a session only when it holds nothing — no active claim, session-owned document lock, keep-alive hold, pending launch delivery, or in-flight wake delivery. A session holding any of those is reported as skipped and stays live. Destructive ends are explicit operator calls (session-end --release-claims); releases record agent_presence_evidence on the terminal events. On reactivation, conditional auto-reacquire restores prior session_ended claims within session_reactivation_reacquire_window_s when no conflicting holder exists. The stale-session sweep can end idle claimless sessions; it preserves every active work-claim holder regardless of heartbeat, park, or process state. See docs/harness-substrate.md for the full contract.

Polish handoff

For definitions that declare and bind these stages:

  • polishing-implementation means the registered polish segment owns the finishing pass.
  • implemented means the branch is implementation-complete and ready for the next definition-bound handoff.

Merge and deployment

Current definitions whose policies.delivery is release_stage bind usher across their delivery tail. Their run-backed path uses implemented -> release -> done; their no-flow path closes from implemented -> done. Definitions with continuous_slice_actions or after_merge_action have a different tail and do not inherit those stages.

Read the pinned definition before invoking usher; it owns a boundary only when an active skill binding says so.

What The Statuses Mean

Status Meaning
idea Filed but not yet shaped into an execution-ready item
refining-idea The item is being clarified and tightened
refined-idea Idea-level shaping is complete
planning Planning or decomposition has started in a workflow that declares this stage
plan-drafted An initial plan or generated-task decomposition exists
refining-plan Plan is being revised after critique/simulation
planned The plan is accepted and ready for the next bound skill
implementing Engineering work is actively in progress
reviewing-implementation Review/fix/verify loop is in progress
reviewed-implementation Implementation review passed; ready for polish
polishing-implementation Finishing pass is in progress
implemented Implementation complete; ready for usher/merge/deploy handoff
release Deployment run is actively executing
done Delivery complete. When an item finished is the item_status_transitions row that put it here, not merged_at, which dates the code landing that close-out can trail by a day
cancelled Item was intentionally abandoned
~~blocked~~ Not a lifecycle status for items. Items use an orthogonal blocked flag that preserves lifecycle status (cross-reference: see your items packet stanza). Epic-task blocked is a status.
stopped Work halted unexpectedly or intentionally paused. Terminal for resource release, so claims and lanes are freed, but a pause is not an ending: it never counts as finished work and stays out of Done counts and "finished" card text
failed Work concluded in failure and needs intervention

QA And Lifecycle

QA evidence is recorded in qa_requirements, qa_runs, and qa_artifacts, not in lifecycle status names. A transition is QA-gated only when the target stage in the item's pinned definition references the qa_verification gate. Project and item attachments materialize the requirements for that transition; the stage name alone does not imply a fixed QA recipe.

Post-Merge Behavior

No-flow / internal delivery

For a current release_stage definition whose delivery does not require a run-backed deployment:

implemented -> done

The code is already live once merged, so release is skipped.

Run-backed deployment flows

For a current release_stage definition enrolled in a deployment run:

implemented -> release -> done

Operationally:

  • item stays implemented until the deployment run actually begins execution
  • item moves to release while the run is executing
  • item moves to done when the run succeeds and blocking post-deploy/manual-acceptance requirements are satisfied

Terminal items are immutable

Once an item reaches a terminal stage its records are frozen. This is deliberate, not an oversight, and it is enforced structurally rather than by convention: the ordinary scalar-write path requires the item's work claim, and a work claim cannot be acquired against a terminal item (INVALID_CLAIM: PREFIX-N is terminal at workflow stage 'done'). --force bypasses the frozen-item and gate guards, not the claim check.

The same stance governs the adjacent record types — an unsettled QA record blocks the terminal transition rather than being corrected afterward, so the repair happens while the claim is still held. Prefer that shape whenever a value must be right before an item freezes: gate the transition, do not reopen the record.

A closed, verdict-less QA run stops blocking when its successor has completed passing evidence or a recorded discharge; post-deploy proof needs accepted completion-member copies. An unsettled successor is named in the refusal. Live runs and active plan executions still block; the old run stays unchanged.

Ad hoc write SQL against the authoritative DB is not an escape hatch here; it is banned by the governed-mutation contract.

The one exception: an unrecorded merge timestamp

A branch that lands outside the merge boundary — a hand merge outside Yoke, for example — leaves items.merged_at unset, and the item can then reach a terminal stage with no record of when it merged; terminal records being immutable, nothing could repair that afterward. One narrow human-only surface exists for exactly that gap:

yoke items merge-provenance operator-correct PREFIX-N --merged-at YYYY-MM-DDTHH:MM:SSZ --reason TEXT

It fills an unset value on an already-terminal item and nothing else, refusing a hook context (human-only), a non-terminal item, an item whose merged_at is already set, and a timestamp that fails to parse or lies in the future. --pr-number N repairs the same provenance's other half — which pull request carried the landing — for the landings close-out could not read for itself; that replacement is verified merged against GitHub first, and the predecessor's stamps drop with it. Each accepted correction emits its WARN event (OperatorMergedAtCorrection, OperatorLandingPullRequestCorrection) carrying the operator reason before the write lands, so the ledger records the action even if that write then fails. yoke items merge-provenance operator-correct --help has both workflows.

A live item never needs this: yoke merge item PREFIX-N is the merge boundary and stamps merged_at from the landing merge commit's own time rather than from when close-out ran, so a re-entered close-out writes the same answer and an item that lands a second time reports that landing instead of its predecessor's. Both boundaries do this: the standalone engine reads the merge it just made, and the queue close-out waits until its merge-group receipt names the merge before stamping, so a queue project is not the one left dating an item to a landing it has replaced. Where the control plane observed the landing itself, that observation wins over the commit: a queue merge commit is created when the train forms and merged minutes later, so merge_queue_landed_at is the truer moment and the supersede prefers it. On a merge-queue project close-out need not be the process that sees the merge — the pull request number is recorded when that pull request opens, and the control-plane observer stamps merged_at and the merge-queue landing columns from GitHub's own merge time. A worker whose wait died therefore leaves a recorded landing rather than an untouched-looking item: the fleet report reports it as landed without close-out, and re-entering yoke merge item PREFIX-N closes out from those recorded facts only when the current candidate is already on the base. Recording a landing never advances the stage; close-out stays a claim-holding step. Which pull request the item points at is settled there too: the marker names the pull request the item last armed, and a lane whose commits reach the base under a sibling pull request leaves its own open forever, so close-out repoints the marker at the pull request that carried the merge the base holds the candidate it is closing out under — GitHub's record of that merge, else the subject GitHub wrote on it. It asks that of the candidate in hand, and the evidence record — the commit, the merge and the file set — describes that same landing: a lane that landed twice has both candidates on the base, each under a merge of its own, so close-out keeps the recorded merge receipt only when that receipt landed under the same merge as the candidate. A lane fast-forwarded onto the base, or a rebased copy whose shas the base never saw, names no merge of its own, and there the receipt is still the only identity there is. An ordinary landing finds the marker already right and writes nothing, and a landing with no merge commit to read — a fast-forward or squash — or an unreadable checkout or provider keeps its marker and names the operator repair above.

Entering a release_stage release wait requires a recorded landing or attested no-change evidence (GATE_MERGE_UNRECORDED), and epics refuse done without merged_at (GATE_EPIC_MERGE). A merge-free definition records no landing, so the correction surface above remains its repair.

Registered Skill Boundaries

Commands do not own global status ranges and do not apply by item type. Each immutable workflow version binds registered skill ids to contiguous stage segments. For a live item:

  1. Run yoke workflows item get PREFIX-N to read its workflow id, logical

version, and current stage.

  1. Run yoke workflows version get WORKFLOW VERSION to read that exact

definition.

  1. In ordered stages, find the one skill_bindings row whose interval

satisfies from_stage_id <= current_stage < through_stage_id.

  1. Invoke /yoke <skill_id> and let the target-stage gate references and

structural gates govern each move; a crossed segment comes back as skill_handoff, naming the next skill's command and claim.

The registered skills have these behavioral contracts; their source and target stages always come from the binding:

Skill id Segment behavior
refine Critique and improve the artifact selected by the pinned policies
shepherd Run quality-gated planning for a compatible generated-task policy
implement Drive a single implementation lane and its review loop
conduct Drive generated task lanes and their integration/review loop
polish Perform the definition-bound finishing pass
usher Merge and deliver a release_stage workflow
dash, blitz Execute their definition-bound direct-work segments

Worktree shape also comes from policies.worktrees and policies.generated_children. A single-lane skill keeps implementation and review in one claimed worktree; a task-graph skill provisions the registered worker/integration lanes; a none policy provisions no lane at all, so the item runs in place under the session's existing write authority and can live in a project that is a bare folder with no git repository.

policies.delivery of merge_free pairs with that: done is the recorded floor attestation — the agent account plus the observed changes — and no merge SHA is required or expected.

A binding's through_stage_id is a handoff boundary. The next skill starts as a fresh command entrypoint and acquires its own claim; the prior skill does not carry claim ownership across the boundary.

Claim release at handoff — visible failure

The implementation-skill finalize step that hands the claim across a binding boundary is best-effort: when it cannot release (cross-session mismatch, claim already terminal, item never claimed, or the underlying domain validator raised), the transition remains committed. The failure is visible as a Warning: claim release failed for PREFIX-N (intent=X, exit=Y) line and an ItemClaimReleaseFailed event carrying the item, caller, holder, failure reason, target stage, and release intent. Operators investigating a retained claim should query the events ledger first: yoke events query --item PREFIX-N --event-name ItemClaimReleaseFailed.

Routing And Explicit Staffing

The pinned workflow binding selects the skill for an item's live stage. Steering owns staffing; the shared scheduler computes the runnable frontier. The canonical sources are:

  • charge-frontier.md — frontier computation, status-to-adapter mapping, ranking
  • yoke_core.domain.scheduler_routing — the next_step function that turns a status into a command
  • yoke_core.domain.session_launch_mandate — resolves assigned item routes from pinned workflow bindings

Agents reading the lifecycle should treat those files plus the item's pinned definition as authoritative for "which command runs next?" The tables here describe skill behavior and shared stage meaning; they do not define an item's stage graph.

See Also

Lifecycle Runtime — Workflow Registry

Lifecycle Runtime — Workflow Registry · Yoke