Project Lifecycle¶
How a greenlit objective becomes one trackable initiative: the project knows the plan it is executing, the plan's items know the tasks implementing them, and the project's status advances from that work.
This page owns the project side of the graph. Plan Review owns the plan's authoring and review phases; Engine owns the task lifecycle.
The graph¶
Three entities, linked by scalar foreign keys pointing upward only.
flowchart LR
Project -->|plan_id| Plan
Plan -->|project| Project
Plan -->|parent_task_id| Task
Task -->|plan_id| Plan
Task -->|plan_item_id| PlanItem
Task -->|project| Project
Plan --> PlanItem
| edge | field | notes |
|---|---|---|
| Project to Plan | Project.plan_id |
the plan being executed |
| Plan to Project | Plan.project |
set at plan creation, immutable |
| Plan to Task | Plan.parent_task_id |
the objective task the plan decomposes; a real FK, ON DELETE RESTRICT |
| Task to Plan | Task.plan_id |
stamped at dispatch |
| Task to PlanItem | Task.plan_item_id |
stamped at dispatch |
| Task to Project | Task.project |
set at intake, immutable |
Plan.parent_task_id is the one downward-pointing edge, and the only one the
database enforces, because it is the one whose violation strands a row an
operator cannot reach: an orphaned plan cannot be approved (its parent 404s),
superseded, or deleted. Deleting a task a plan references is refused
(409) rather than allowed to orphan it; the exit is DELETE /plans/{id}. See
Plan review.
No entity stores a collection of its children, or of its participants.
Reverse lookups are indexed queries: TaskFilterSpec(plan=...) for a plan's
tasks, TaskFilterSpec(project=...) for a project's tasks,
PlanFilterSpec(project=...) for a project's plan history, and
initiative_contributors(...) (engine/initiative/contributors.py) for the
agents that worked an initiative, which unions the assignees of its tasks that
left the queue with the project's recorded lead.
This is a deliberate correction, made twice for the same reason.
Project.task_ids existed as a stored tuple of child ids and was never
populated by anything, so the dashboard showed a task count of zero next to a
full task list. Project.team was the same shape one field over: every
production path that mints a project created it empty, so "who is on this
initiative" read as nobody, and the retrospective silently discarded every
learning it held for a non-lead agent.
A collection embedded in a row has to be written by every actor that creates a
child, in the same transaction, forever; a scalar upward key is written once by
the actor that already owns the write. Task.assigned_to is written by the
actor that made the assignment, on the row it already owns, which is why
contributors derive from it, together with the lead the project already
records. Because that write lands at ASSIGNED, before anything runs, the
derivation drops the statuses that prove no execution happened, or a queue
nobody has started would read as a roomful of contributors. Both dead fields
were removed rather than filled in.
Project.plan_id always names the one plan the project is working
through. Every earlier revision stays reachable through
PlanFilterSpec(project=...), which returns superseded plans too.
Re-planning a dispatched initiative¶
A plan under review is edited in place: same entity, bumped version, back to
pending review. Once a plan is dispatched that is no longer possible, because
its items are already building and rewriting them would leave running tasks
implementing items that no longer exist. Revising a dispatched plan is
therefore a re-plan (POST /plans/{id}/replan), which:
- retires the current revision to
SUPERSEDED; - cancels the in-flight tasks it dispatched, through their audited lifecycle transitions, so no live work points at a withdrawn revision;
- opens a successor plan entity carrying the objective and framing forward,
entering
PENDING_REVIEWbecause its items hold no approval; - repoints
Project.plan_idat the successor.
The ordering protects one invariant: a project never has two live plans. If a
later step fails, the recoverable state is an initiative whose plan is
superseded and whose successor is missing, which an operator resolves by
planning again. The alternative ordering would leave two live plans with
Project.plan_id naming one arbitrarily and the rollup deriving status from a
revision the operator had already abandoned.
The successor is not dispatched by the re-plan. It goes through review like any other plan, and approval activates the project and repoints it through the same path first dispatch uses, so there is one dispatch path rather than two.
Status lifecycles¶
Both Plan and Project have a real transition table on the shared
core/state_machine.py, exactly as Task does. Illegal jumps are impossible
rather than merely unwritten.
stateDiagram-v2
[*] --> PLANNING
PLANNING --> ACTIVE: plan approved and dispatched
ACTIVE --> INTEGRATING: every plan item done
INTEGRATING --> EVALUATING: assembly job passed its gate
EVALUATING --> COMPLETED: every success criterion met
INTEGRATING --> ACTIVE: an item regressed
EVALUATING --> ACTIVE: an item regressed
ACTIVE --> ON_HOLD: operator pauses
ON_HOLD --> ACTIVE: operator resumes
ON_HOLD --> CANCELLED: operator cancels a paused project
PLANNING --> FAILED: its plan failed terminally
ACTIVE --> FAILED: its plan failed terminally
INTEGRATING --> FAILED: its plan failed terminally
EVALUATING --> FAILED: its plan failed terminally
FAILED --> PLANNING: a fresh plan is written
FAILED --> ACTIVE: a fresh plan is dispatched
PLANNING --> CANCELLED
ACTIVE --> CANCELLED
FAILED --> CANCELLED
COMPLETED --> [*]
CANCELLED --> [*]
The project mirrors its plan's tail stage by stage, so the cockpit distinguishes
an initiative still building from one whose pieces are being assembled, and both
from one awaiting a verdict. INTEGRATING and EVALUATING also carry the
ON_HOLD and CANCELLED edges, omitted above for readability; every other edge
in core/project_transitions.py is shown.
ACTIVE -> COMPLETED does not exist, and neither does the plan's
EXECUTING -> COMPLETED. Delivery has exactly one predecessor, which is what
makes the tail structural rather than a convention: see
Initiative Tail.
ON_HOLD has no direct hop to COMPLETED: an operator who paused an initiative
must resume it before it can finish, so work never completes out from under a
deliberate hold. Resuming returns to ACTIVE, from which the tail is
re-derived, rather than dropping the operator back into a half-finished stage
whose gate has already run.
Every state has an exit a writer can always take¶
A transition table full of legal edges is not the same as a reachable exit. The
task machine had one edge out of FAILED, BLOCKED, INTERRUPTED and
SUSPENDED, and it was ASSIGNED, which Task's validator refuses without an
assignee. A task that failed before assignment therefore could not leave any
state: the raw ValidationError escaped as a repeating 422, and its plan and
its project became undeletable behind it.
So each StateMachine declares unconditional_targets: the statuses a writer
can always reach because they need nothing the entity might lack (a reason the
writer authors, never an assignee or a non-empty item DAG). Every stuck task
status carries a direct CANCELLED (abandonment needs no assignee) and every
tail plan status carries FAILED. check_lifecycle_exit_reachable.py walks
those hops alone, breadth-first, so plain terminal reachability cannot
satisfy it.
A park is a destination, never a corridor¶
The same declaration carries a second claim about the entity's world that the
edges cannot express. HopRules.no_transit_states names the statuses a
multi-hop walk (StateMachine.path_to, which resolves a Kanban column move
into a legal route) may finish on or start from, but never route THROUGH: for
tasks that is BLOCKED, AWAITING_INPUT, AUTH_REQUIRED, SUSPENDED and
INTERRUPTED.
Each of them means "something must change before this moves again", and each
takes its meaning from a reason column the walker never sets: a route that
transits BLOCKED records a park that never happened and leaves a
blocked_reason of None behind it, which reads as a park nobody can explain.
Successor ordering already prefers the ordinary route, but order only breaks
ties: once the route through a park is strictly shorter, breadth-first search
takes it and the preference is silently lost. The walk therefore discovers a
no-transit state (so a longer route back to it is not explored) and never
expands it.
Every status change is recorded¶
The status says where an initiative is. The append-only
lifecycle_transitions ledger says how it got there and who moved it:
PlanService.sync_status and advance_project_status each append a row per
persisted hop, including the intermediate hops of a multi-hop walk that a log
line loses. GET /plans/{id}/transitions reads it back, which is what makes
"only the evaluate stage writes COMPLETED" provable from persisted state
rather than from a container's stdout.
Every write path is covered; the write itself is not guaranteed. The append runs after the status write has committed, so it cannot be rolled back into it, and a ledger outage that outlasts the retries leaves a status with no row behind it. That gap is bounded and loud rather than silent: the append is idempotent on the row's own id (a retry after a lost response is the success it was, not a duplicate-key failure), a row that still does not land is logged at ERROR with every field it would have carried so it can be reconstructed, and a run of them says so explicitly instead of reporting each as an isolated blip. Read the ledger as complete up to a reported outage, not as complete by construction.
That reporting covers the persistence failures the append is allowed to
absorb, which is every one except a critical exception. MemoryError,
RecursionError and an exception group carrying either are re-raised the
moment they surface, ahead of the ERROR line and ahead of the outage streak,
because a process out of memory is not a ledger outage to report and carrying
on to format a log record is how the report itself fails. A gap left that way
is announced by the crash rather than by the ledger.
The ledger repository is a required collaborator, not an optional one, on
both LifecycleLedger and PlanService: a service built without it would
persist the status and emit the transition log line while the durable row
silently never happened, which is a ledger that looks complete and is not.
build_plan_service is the single construction site that binds it.
The append runs after the status write commits, so a ledger outage is logged at ERROR with every field the row would have carried and never reported to the caller as a failed transition: the move already happened.
A failed project mirrors a failed plan, and nothing else¶
ProjectStatus.FAILED exists, and it is reached from exactly one place: the
rollup mirroring PlanStatus.FAILED. Nothing else derives it.
That narrowness is the whole design, because the argument against a derived
failure is still correct for every other source. A completion-oracle REJECT
routes a task back to IN_PROGRESS for rework, not to failure. A task that does
reach FAILED stays reassignable (FAILED -> ASSIGNED in the task state
machine). Derived from task outcomes, a project failure would flap the moment
the work was retried, and it would assert a judgement the system is not entitled
to make. So failed and blocked TASKS still surface as derived counts on the
progress view and move no project status at all.
A terminally-failed PLAN is different in kind: it does not flap, it carries the
reason somebody can read, and the project it belongs to has nothing left in
flight. Leaving that project at PLANNING was not neutrality, it was the
dashboard asserting that planning was still happening.
FAILED is therefore not terminal. A fresh plan walks the project back out
(FAILED -> PLANNING or -> ACTIVE), and an operator can still cancel it. It
is a park, not an ending, and ending an initiative remains a human act that
CANCELLED expresses.
The general rule survives, with its scope stated: a status is what the organisation decides or what its plan has already recorded; "some work went wrong" is signal.
Completion¶
An item is done when:
Every item being done is the start of the tail, not the end of the plan. A
set of individually-verified pieces has not been shown to work together, so the
plan moves to INTEGRATING and delivery becomes the evaluate stage's verdict.
An itemless plan never self-advances: it has delivered nothing, so "every item
is done" being vacuously true must not read as progress.
A project is COMPLETED when the plan it is executing is, and nothing in the
rollup can write COMPLETED onto a plan. See
Initiative Tail.
Decision items are included deliberately. They never dispatch as tasks
(decomposition_from_plan strips them before dispatch), but an unresolved
decision is real work the operator still owes, so an initiative cannot complete
around one.
How this composes with the verify gate¶
The rollup reads persisted Task.status, never execution outcomes.
That single choice is what keeps an initiative from completing on unverified
work. Under the wired agent runtime a task reaches COMPLETED through
ReviewGateService._apply_decision, which runs the full gate chain (build/test
oracle, completion-oracle peer review, output policy, red team, vision).
Requiring COMPLETED therefore inherits every one of those gates without the
rollup making a single oracle call.
This used to be a property of which writers were wired rather than a structural
invariant, because two paths reached COMPLETED without the oracle chain. Both
are now fenced:
LifecycleAdvancingExecutionService(workers/execution_service/_lifecycle.py), the lifecycle-only baseline the app self-constructs when no agent runtime is installed, refuses to advance a plan-linked task out ofIN_REVIEW. Stopping there is honest for a boot with no runtime to verify anything; jumping toCOMPLETEDwas a lie. A directly filed task keeps the baseline's full happy path.- The coordination parent rollup (
engine/coordination/parent_rollup.py) reads each subtask's persisted status rather than theDispatchResultoutcomes it used to derive from. Those outcomes report execution success before verification, so a task parked inIN_REVIEWawaiting the oracle counted as completed. Immediately aftercoordinate()most subtasks are thereforeIN_REVIEWand the parent staysIN_PROGRESS; the initiative rollup re-derives the parent on every later task event, so it lands its terminal status once the gate has ruled on each child.
The objective task itself is held open for exactly as long as its plan is: every item passing its own gate does not deliver the objective, the tail does.
Rollup¶
ProjectRollupService (engine/initiative/rollup.py) registers as a
TaskEngine observer, so it observes every task status write regardless of which
path produced it: the review gate's decision, the execution loop's failure
handling, an operator cancellation.
It recomputes; it does not accumulate. The event is only a trigger. On each one the service re-queries every task for the plan and derives plan and project status from scratch, then writes under optimistic concurrency.
Two properties follow, and both are the reason for the design:
- Idempotent, therefore self-healing.
TaskEngineobservers are explicitly lossy: a bounded queue, drained at shutdown, so events can be dropped or redelivered. A full recompute means the next event repairs any drift and a duplicate event changes nothing. An incremental counter would be corrupted permanently by a single dropped event, which is why there is no reconciler worker: correctness does not depend on delivery. - Verification-derived, not execution-derived, as above.
Writes are version-guarded (expected_version) with a bounded retry, and a
per-plan in-process lock serialises same-process recomputes. A losing write
re-reads and recomputes rather than clobbering the winner.
The rollup also fires the loop's detached tails, each failure-tolerant and each unable to block or fail it:
- Integrate and evaluate, while the plan reads as
INTEGRATINGorEVALUATING. Both stages are idempotent, so firing on every recompute rather than on an edge is safe and needs no "already started" flag to keep in step with reality. An unwired stage parks the plan visibly instead of completing it. See Initiative Tail. - Auto-replan, while the plan reads as stalled: outstanding items exist and none of them can advance without a new decision.
- The SHIP retrospective, on the edge a project first reaches
COMPLETED(and only that transition, never a recompute over an already-terminal project), so finished work feeds a retrospective back into org and agent memory. See the "Retrospective Capture on SHIP" section of memory-learning.md for the capture pipeline.
Where linkage is written¶
At approval, in _dispatch_approved_plan
(api/controllers/_plan_review_resume.py), and before any task is filed:
Project.plan_idis repointed and the project goesPLANNING -> ACTIVE.- The plan goes
APPROVED -> SKELETON. decomposition_from_planstampsplan_id+plan_item_idonto every child task it builds.
The ordering is load-bearing: a rollup event fired mid-preparation would
otherwise observe a project still PLANNING with its plan already staged.
The coordinator is not called here. Approval's job ends at making the graph
durable; it then asks the rollup to recompute, and the rollup opens the contract
stage. Only when that contract job passes its review gate does the plan reach
EXECUTING and its units get dispatched, so nothing builds against a contract
that does not exist yet.
Task.plan_item_id makes the previously implicit correlation explicit. The child
task id is still minted deterministically from the plan item id
(subtask_uuid), but reading the graph no longer requires knowing that trick.
Operator surface¶
GET /projects/{id}/progress returns the initiative view: plan status, every
plan item with its task status, derived counts (total / done / failed /
blocked), and the critical path.
The critical path is the longest dependency chain through the plan's item DAG
(engine/initiative/critical_path.py): the chain that sets the delivery date, so
shortening any other chain does not bring the plan in sooner. It is computed
server-side, which keeps the dashboard a pure API consumer and makes the same
view reachable by any API client rather than existing only in the browser.
The project detail page renders it as the initiative cockpit, subscribing to the
plans channel alongside projects and tasks so a plan status change
refreshes it live. A project with no plan yet returns the same shape with an
empty item list rather than a 404, so the view is stable across an initiative's
whole life.
Persistence¶
projects.plan_id, tasks.plan_id, and tasks.plan_item_id are nullable TEXT
columns; tasks.plan_id is indexed because the rollup and the progress endpoint
both query by it. plans.parent_task_id is the one enforced reference
(REFERENCES tasks (id) ON DELETE RESTRICT, indexed (parent_task_id, id)).
projects records created_at / updated_at, which is what lets conversational
intake bound project reuse by age rather than by an in-process cache. The plans
status CHECK carries the full enum including executing and completed. SQLite
and Postgres are in parity, with one yoyo revision per backend.
Deleting a project resolves its children first and only then removes the row:
every non-terminal plan is retired (SUPERSEDED, or FAILED with "project deleted"
when it has no items, because the items CHECK forbids superseding an itemless
plan) and every non-terminal task is cancelled, each through its own audited
transition. The cascade and the delete are separate audited operations rather
than one transaction, because the task transitions emit domain events that
cannot be rolled back; consistency comes from idempotent forward recovery
instead, so re-issuing a failed delete re-runs the cascade as a no-op over the
already-resolved children.
Forward recovery is what makes a partial cascade safe, not a licence to
delete past one. A plan the initiative rollup is writing concurrently gets a
bounded re-read budget, and exhausting it aborts the delete with a 409 rather
than counting the plan retired: plans.project carries no foreign key, so a
project removed over a plan still live leaves an orphan nothing can reach.
Contention is transient, so repeating the delete is the resolution.
Retiring a child is not removing it. A terminal status stops a plan advancing;
it does not stop the row existing, and every listing endpoint still returns it
naming a project id that no longer resolves. Worse, the cascade retires a plan
with items to SUPERSEDED, which is the one status DELETE /plans/{id}
refuses, so a retire-only cascade makes whether an operator can ever clean up
depend on whether the plan happened to have items. The cascade therefore
removes the retired plan and cancelled task rows too, and each removal writes
a tombstone (below).
Before each child is removed, every pending approval that decides about it is expired. An approval is a question about something that exists; once the row is gone the queue still offers approve and reject, and answering drives the resume path at an id that resolves to nothing. Retirement runs first and the delete is conditional on it, so a decision landing mid-flight refuses the delete with a 409 rather than racing the resume path; anything the refused attempt already expired is put back. See Plan Review.
The workspace goes too. The database forgets a deleted project's workspace
immediately, because the project_workspaces row cascades on its foreign key.
Disk does not, and a live run finished holding 24 trees under the workspace
root, two of them belonging to projects deleted through the dashboard during
that same run. They are not merely wasted space: planning recall spans every
project the organisation has run, so a tree that outlives its project stays
available as evidence for a plan that should never have seen it, which is
exactly how one plan came to assert another project's seven files as its own
foundation. The removal happens BEFORE the row is deleted, so a tree that
cannot be removed takes the delete down with it and leaves a project the
operator can retry; the reverse order would report a deletion over files that
are still there. Only the managed base_root/projects/<project_id> directory
is removed, and a failure is a typed refusal rather than a raw filesystem
error, so one project in a bulk selection cannot end the whole request after
its earlier rows are already gone. Emits PROJECT_WORKSPACE_DISCARDED.
Data retention. A task cannot be pinned by the records that describe it:
spend, metrics, approvals and decision records all name the task they are
about, and a foreign key on each makes every one of them a reason the task
can never be removed. Those pins are gone. To keep the identifiers resolvable
rather than dangling, every operator deletion writes a row to
deleted_entities naming the entity, what it was called, who removed it and
when. The identifier is derived from the kind and the entity id, so a
re-issued teardown writes the same row rather than a second copy. GET
/tasks/{id} reads it: an id that is no longer a row answers 404 with what it
was and who removed it, which is what the surviving cost, metric, and decision
rows resolve through. Tombstones are subject to retention like any other
append-only record (purge_before).