Yoke documentation
Lifecycle Runtime — Workflow Registry
> Canonical source: workflow_runtime.py loads an item's immutable > workflow_id / workflow_version_id pin. Registry definitions are served by > yoke workflows definition get.
This document describes the runtime contract. Definitions own ordered stages, labels, terminal stages, gates, policies, entry surfaces, and registered skill bindings. Live transition, frontier, scheduler, QA, approval, and delivery paths all interpret the item's pin.
<!-- BEGIN GENERATED: field-note-directive --> When you hit a recipe gap or notice a minor bug best held as a supporting record, file a field-note immediately — before retrying, before moving on. yoke ouroboros field-note append --kind <failed|new|unclear|observation> --evidence '...' Run yoke ouroboros field-note append --help for the worked failure modes and decision tree. <!-- END GENERATED: field-note-directive -->
Item stage authority
Do not copy a progression into operator logic or documentation. Use yoke workflows definition get for current definitions. A live item loads the exact pinned version, so a newer one cannot alter work in flight, and a forward move it does not declare is refused before the target stage materializes. Backward rework uses definition.stages order and needs no declared edge: yoke_core.domain.workflow_declared_transitions.undeclared_forward_transition checks only forward moves. The holder uses yoke lifecycle transition to return to the requested earlier stage; ordinary claim and target-stage checks still apply.
Exceptional Item States
These are reachable from multiple points and are not part of the normal forward progression:
cancelledstoppedfailed
> Item-level blocked is not a lifecycle status. It is an orthogonal > flag-and-reason pair on the item that preserves the lifecycle position > (cross-reference: see your items packet stanza for the > blocked/blocked_reason columns). Set it via > yoke items block PREFIX-N --reason "<reason>"; clear via > yoke items unblock PREFIX-N. > The board renders blocked items in their own section and the frontier > routes them to WAIT. The doctor health checks HC-blocked-status-drift > and HC-blocked-flag-consistency surface leftover flags and > legacy status='blocked' rows. Epic-task blocked stays > as a status. Full architectural-why (yoke source repo): > docs/archive/decisions/blocked-flag-retirement.md.
Canonical Epic Task Progression
Epic tasks mirror the implementation-family vocabulary:
planning
-> plan-drafted
-> refining-plan
-> planned
-> implementing
-> reviewing-implementation
-> reviewed-implementation
-> polishing-implementation
-> implemented
-> release
-> done
Task exceptional states:
blockedstoppedfailed
Epic tasks do not use item-only statuses such as cancelled.
Ownership Boundaries
Definition-bound segments
At a live item stage, the owner is the registered skill binding whose half-open interval contains that stage: from_stage_id <= current_stage < through_stage_id. Stage names do not select the skill by themselves, and a workflow name is never a substitute for reading the pinned version.
The current built-in definitions reuse stage ids such as idea, refining-idea, planning, and planned, but each definition decides which of those stages exists, their order, and the skill that owns the segment.
idea -> refine handoff: two-layer guard against title-only dispatch
/yoke idea writes the row in two phases — items add lands the PREFIX-N row with empty spec, and body-and-sync.md writes the structured spec fields a few seconds later. The window between the two phases is unprotected unless both layers below hold:
- Layer 1 — claim-on-create (live-race fix).
infer-and-create.md
step 5b acquires a draft work claim with reason draft-in-progress immediately after items add returns the PREFIX-N id, and body-and-sync.md step 10b releases it with reason idea-complete once the spec/body, AC normalization, and every enabled File Budget or path-claim artifact has landed. The release path canonicalizes idea-complete → handed_off for schema storage and preserves the original intent on the WorkReleased event. While the draft claim is held, the scheduler filters the row out via the standard live-claim conflict gate. Held duration is recorded on the IdeaClaimHeld event for doctor and Ouroboros observability.
- **Layer 2 — body-completeness skip on the frontier (structural
defense).** yoke_core.domain.frontier_compute calls yoke_core.domain.idea_body_completeness.is_idea_body_incomplete for every status='idea' row and pushes the title-only ones into blocked with reason idea-incomplete. This catches every tail case Layer 1 cannot reach: a /yoke idea session that crashes between the two phases (the title-only body remains, while the active claim stays held until explicit release or authorized termination); a manual python3 -m yoke_core.cli.db_router items add from ad hoc tooling that bypasses the claim convention; any future /yoke idea variant that forgets to acquire the claim. The doctor health check HC-incomplete-idea-bodies reports items in this state so the operator can rescue or freeze them.
Implementation and review
implementingmeans work is actively being built.reviewing-implementationmeans coding/self-verification is complete and the branch is in the deliberate review/fix loop.reviewed-implementationmeans meaningful implementation review passed and the work is queued for finishing polish.
When a definition declares this loop, its active skill binding drives it. For example, current definitions bind either implement, conduct, or a direct skill across implementation work; the stage name alone does not choose one.
Claim continuity across transient SessionEnd. A Claude Desktop SessionEnd event (laptop sleep, app reload, idle timeout) never destroys mid-flight claims: the hook runs the non-destructive end_session_if_empty, which ends a session only when it holds nothing — no active claim, session-owned document lock, keep-alive hold, pending launch delivery, or in-flight wake delivery. A session holding any of those is reported as skipped and stays live. Destructive ends are explicit operator calls (session-end --release-claims); releases record agent_presence_evidence on the terminal events. On reactivation, conditional auto-reacquire restores prior session_ended claims within session_reactivation_reacquire_window_s when no conflicting holder exists. The stale-session sweep can end idle claimless sessions; it preserves every active work-claim holder regardless of heartbeat, park, or process state. See docs/harness-substrate.md for the full contract.
Polish handoff
For definitions that declare and bind these stages:
polishing-implementationmeans the registeredpolishsegment owns the finishing pass.implementedmeans the branch is implementation-complete and ready for the next definition-bound handoff.
Merge and deployment
Current definitions whose policies.delivery is release_stage bind usher across their delivery tail. Their run-backed path uses implemented -> release -> done; their no-flow path closes from implemented -> done. Definitions with continuous_slice_actions or after_merge_action have a different tail and do not inherit those stages.
Read the pinned definition before invoking usher; it owns a boundary only when an active skill binding says so.
What The Statuses Mean
| Status | Meaning |
|---|---|
idea |
Filed but not yet shaped into an execution-ready item |
refining-idea |
The item is being clarified and tightened |
refined-idea |
Idea-level shaping is complete |
planning |
Planning or decomposition has started in a workflow that declares this stage |
plan-drafted |
An initial plan or generated-task decomposition exists |
refining-plan |
Plan is being revised after critique/simulation |
planned |
The plan is accepted and ready for the next bound skill |
implementing |
Engineering work is actively in progress |
reviewing-implementation |
Review/fix/verify loop is in progress |
reviewed-implementation |
Implementation review passed; ready for polish |
polishing-implementation |
Finishing pass is in progress |
implemented |
Implementation complete; ready for usher/merge/deploy handoff |
release |
Deployment run is actively executing |
done |
Delivery complete. When an item finished is the item_status_transitions row that put it here, not merged_at, which dates the code landing that close-out can trail by a day |
cancelled |
Item was intentionally abandoned |
~~blocked~~ |
Not a lifecycle status for items. Items use an orthogonal blocked flag that preserves lifecycle status (cross-reference: see your items packet stanza). Epic-task blocked is a status. |
stopped |
Work halted unexpectedly or intentionally paused. Terminal for resource release, so claims and lanes are freed, but a pause is not an ending: it never counts as finished work and stays out of Done counts and "finished" card text |
failed |
Work concluded in failure and needs intervention |
QA And Lifecycle
QA evidence is recorded in qa_requirements, qa_runs, and qa_artifacts, not in lifecycle status names. A transition is QA-gated only when the target stage in the item's pinned definition references the qa_verification gate. Project and item attachments materialize the requirements for that transition; the stage name alone does not imply a fixed QA recipe.
Post-Merge Behavior
No-flow / internal delivery
For a current release_stage definition whose delivery does not require a run-backed deployment:
implemented -> done
The code is already live once merged, so release is skipped.
Run-backed deployment flows
For a current release_stage definition enrolled in a deployment run:
implemented -> release -> done
Operationally:
- item stays
implementeduntil the deployment run actually begins execution - item moves to
releasewhile the run is executing - item moves to
donewhen the run succeeds and blocking post-deploy/manual-acceptance requirements are satisfied
Terminal items are immutable
Once an item reaches a terminal stage its records are frozen. This is deliberate, not an oversight, and it is enforced structurally rather than by convention: the ordinary scalar-write path requires the item's work claim, and a work claim cannot be acquired against a terminal item (INVALID_CLAIM: PREFIX-N is terminal at workflow stage 'done'). --force bypasses the frozen-item and gate guards, not the claim check.
The same stance governs the adjacent record types — an unsettled QA record blocks the terminal transition rather than being corrected afterward, so the repair happens while the claim is still held. Prefer that shape whenever a value must be right before an item freezes: gate the transition, do not reopen the record.
A closed, verdict-less QA run stops blocking when its successor has completed passing evidence or a recorded discharge; post-deploy proof needs accepted completion-member copies. An unsettled successor is named in the refusal. Live runs and active plan executions still block; the old run stays unchanged.
Ad hoc write SQL against the authoritative DB is not an escape hatch here; it is banned by the governed-mutation contract.
The one exception: an unrecorded merge timestamp
A branch that lands outside the merge boundary — a hand merge outside Yoke, for example — leaves items.merged_at unset, and the item can then reach a terminal stage with no record of when it merged; terminal records being immutable, nothing could repair that afterward. One narrow human-only surface exists for exactly that gap:
yoke items merge-provenance operator-correct PREFIX-N --merged-at YYYY-MM-DDTHH:MM:SSZ --reason TEXT
It fills an unset value on an already-terminal item and nothing else, refusing a hook context (human-only), a non-terminal item, an item whose merged_at is already set, and a timestamp that fails to parse or lies in the future. --pr-number N repairs the same provenance's other half — which pull request carried the landing — for the landings close-out could not read for itself; that replacement is verified merged against GitHub first, and the predecessor's stamps drop with it. Each accepted correction emits its WARN event (OperatorMergedAtCorrection, OperatorLandingPullRequestCorrection) carrying the operator reason before the write lands, so the ledger records the action even if that write then fails. yoke items merge-provenance operator-correct --help has both workflows.
A live item never needs this: yoke merge item PREFIX-N is the merge boundary and stamps merged_at from the landing merge commit's own time rather than from when close-out ran, so a re-entered close-out writes the same answer and an item that lands a second time reports that landing instead of its predecessor's. Both boundaries do this: the standalone engine reads the merge it just made, and the queue close-out waits until its merge-group receipt names the merge before stamping, so a queue project is not the one left dating an item to a landing it has replaced. Where the control plane observed the landing itself, that observation wins over the commit: a queue merge commit is created when the train forms and merged minutes later, so merge_queue_landed_at is the truer moment and the supersede prefers it. On a merge-queue project close-out need not be the process that sees the merge — the pull request number is recorded when that pull request opens, and the control-plane observer stamps merged_at and the merge-queue landing columns from GitHub's own merge time. A worker whose wait died therefore leaves a recorded landing rather than an untouched-looking item: the fleet report reports it as landed without close-out, and re-entering yoke merge item PREFIX-N closes out from those recorded facts only when the current candidate is already on the base. Recording a landing never advances the stage; close-out stays a claim-holding step. Which pull request the item points at is settled there too: the marker names the pull request the item last armed, and a lane whose commits reach the base under a sibling pull request leaves its own open forever, so close-out repoints the marker at the pull request that carried the merge the base holds the candidate it is closing out under — GitHub's record of that merge, else the subject GitHub wrote on it. It asks that of the candidate in hand, and the evidence record — the commit, the merge and the file set — describes that same landing: a lane that landed twice has both candidates on the base, each under a merge of its own, so close-out keeps the recorded merge receipt only when that receipt landed under the same merge as the candidate. A lane fast-forwarded onto the base, or a rebased copy whose shas the base never saw, names no merge of its own, and there the receipt is still the only identity there is. An ordinary landing finds the marker already right and writes nothing, and a landing with no merge commit to read — a fast-forward or squash — or an unreadable checkout or provider keeps its marker and names the operator repair above.
Entering a release_stage release wait requires a recorded landing or attested no-change evidence (GATE_MERGE_UNRECORDED), and epics refuse done without merged_at (GATE_EPIC_MERGE). A merge-free definition records no landing, so the correction surface above remains its repair.
Registered Skill Boundaries
Commands do not own global status ranges and do not apply by item type. Each immutable workflow version binds registered skill ids to contiguous stage segments. For a live item:
- Run
yoke workflows item get PREFIX-Nto read its workflow id, logical
version, and current stage.
- Run
yoke workflows version get WORKFLOW VERSIONto read that exact
definition.
- In ordered
stages, find the oneskill_bindingsrow whose interval
satisfies from_stage_id <= current_stage < through_stage_id.
- Invoke
/yoke <skill_id>and let the target-stage gate references and
structural gates govern each move; a crossed segment comes back as skill_handoff, naming the next skill's command and claim.
The registered skills have these behavioral contracts; their source and target stages always come from the binding:
| Skill id | Segment behavior |
|---|---|
refine |
Critique and improve the artifact selected by the pinned policies |
shepherd |
Run quality-gated planning for a compatible generated-task policy |
implement |
Drive a single implementation lane and its review loop |
conduct |
Drive generated task lanes and their integration/review loop |
polish |
Perform the definition-bound finishing pass |
usher |
Merge and deliver a release_stage workflow |
dash, blitz |
Execute their definition-bound direct-work segments |
Worktree shape also comes from policies.worktrees and policies.generated_children. A single-lane skill keeps implementation and review in one claimed worktree; a task-graph skill provisions the registered worker/integration lanes; a none policy provisions no lane at all, so the item runs in place under the session's existing write authority and can live in a project that is a bare folder with no git repository.
policies.delivery of merge_free pairs with that: done is the recorded floor attestation — the agent account plus the observed changes — and no merge SHA is required or expected.
A binding's through_stage_id is a handoff boundary. The next skill starts as a fresh command entrypoint and acquires its own claim; the prior skill does not carry claim ownership across the boundary.
Claim release at handoff — visible failure
The implementation-skill finalize step that hands the claim across a binding boundary is best-effort: when it cannot release (cross-session mismatch, claim already terminal, item never claimed, or the underlying domain validator raised), the transition remains committed. The failure is visible as a Warning: claim release failed for PREFIX-N (intent=X, exit=Y) line and an ItemClaimReleaseFailed event carrying the item, caller, holder, failure reason, target stage, and release intent. Operators investigating a retained claim should query the events ledger first: yoke events query --item PREFIX-N --event-name ItemClaimReleaseFailed.
Routing And Explicit Staffing
The pinned workflow binding selects the skill for an item's live stage. Steering owns staffing; the shared scheduler computes the runnable frontier. The canonical sources are:
- charge-frontier.md — frontier computation, status-to-adapter mapping, ranking
yoke_core.domain.scheduler_routing— thenext_stepfunction that turns a status into a commandyoke_core.domain.session_launch_mandate— resolves assigned item routes from pinned workflow bindings
Agents reading the lifecycle should treat those files plus the item's pinned definition as authoritative for "which command runs next?" The tables here describe skill behavior and shared stage meaning; they do not define an item's stage graph.
See Also
Lifecycle Runtime — Workflow Registry