Skip to content

Latest commit

 

History

History
1265 lines (1107 loc) · 73.9 KB

File metadata and controls

1265 lines (1107 loc) · 73.9 KB

Issue and PR Scheduler

  • Status: active, volatile architecture and operations reference
  • Owner: ClawSweeper maintainers
  • Source of truth: .github/workflows/review-plan.yml, .github/workflows/sweep.yml, .github/workflows/target-fanout.yml, .github/workflows/audit.yml, planner/runtime source, config/automation-limits.json, and focused scheduler tests
  • Last verified: openclaw/clawsweeper@647503ec44b8e777dd172adf974a945367da0d19
  • Update when: cadence, fanout, admission, retry, publication, apply, or state-writing behavior changes

Read when changing .github/workflows/review-plan.yml, .github/workflows/sweep.yml, .github/workflows/target-fanout.yml, .github/workflows/audit.yml, src/clawsweeper.ts planner selection, review or audit cadence, dashboard capacity fields, or GitHub Actions concurrency for issue/PR review and apply.

The global worker budget comes from config/automation-limits.json; see Automation Limits for the derived lane limits and GitHub variable overrides.

Repair and automerge jobs also carry the canonical job_intent frontmatter described in ClawSweeper Orchestration. Workflow inputs can still override live-worker caps, but when they do not, repair:dispatch derives the priority lane from job_intent instead of relying on workflow-specific defaults.

Exact-review completion sends review_failure_reason for every diagnostics manifest with failure.stage: agent_input_scan and retryable: false, using failure.reason_code verbatim. The accepted scanner reasons are scanner_unavailable, scanner_failed, findings, deadline, staging_limit, incomplete_source, source_drift, unsafe_path, and unsupported_content; source_incompatible remains a separate terminal source-preparation reason. The shared allowlist lives in src/exact-review-failure-reason.ts. Exit codes 78 and 79 still identify incomplete_source and findings without a manifest. A matching non-retryable review_failure detail is accepted; retryable or mismatched detail is rejected. Diagnostic detail alone does not suppress retry: without review_failure_reason, failed completion still retries. During Worker/workflow deploy skew, an HTTP 400 invalid_review_failure_reason response triggers one immediate compatibility retry without the terminal reason or its dependent status receipt, preserving review_failure diagnostic detail and emitting a workflow warning. An older Worker can then complete the lease under its existing retry policy; compatible Workers receive the terminal reason on the first request. Other errors retain the existing failure and retry handling. Terminal completion does not turn the failed workflow green. An automatic PR review with a recorded head SHA and source_incompatible retains a parked queue item: the exact source pin cannot improve on that head, so scheduled intake must not repeat its preparation. Scheduled PR intake binds the live head, base, draft state, and body identity before dispatch takes its lease snapshot. This state has no timed retry. Existing parked-item reconciliation recovers changed head, base, or body identity and removes closed targets; an explicit maintainer re-review can retry unchanged source. The failure remains visible in lifecycle and Bay status. Terminal input-scanner refusals instead retain the existing queue item as scanner_refused, for both issues and PRs. Automatic events, source changes, close/reopen, and failed-shard recovery cannot release this hold. It has no timed retry, expiry, closed-target cleanup, or operator source-drift recovery. A fresh verified re-review command or a newly dispatched explicit item request replaces the failed request's authority. Verified command timestamps must be strictly later than the refusal and failed command timestamps. GitHub timestamps have second precision, so all same-second requests stay held regardless of origin or comment ID; a new command or edit in a later second can retry. Replaying an old command or workflow run does not. The manual API also accepts named request IDs; use a new unique ID for each explicit retry. The failed ID remains fenced on the held decision even after delivery receipts expire. An intentional EXACT_REVIEW_RETRY_POLICY_EPOCH change permits the next automatic admission without inheriting failed command context; ordinary deployments, model changes, and prompt changes do not.

Untargeted workflow-dispatch sweeps, including automatic continuations, feed the same queue as scheduled intake. Their branch, prompt, and timeout options remain supported, but a broad sweep is not an explicit retry of every hold. The hold applies to newly observed terminal refusals; previously deleted rows are not reconstructed. Deploy Worker enforcement before relying on updated producer routing. Already-running direct shards from older workflow revisions remain a rollout limitation. Bay uses the terminal failure lifecycle rather than displaying the retained hold as active review work; queue reason counts remain visible and the public surface remains observer-only.

Native blob-fetch transport failures, including the hydration deadline killing a Git fetch, remain source_preparation / review_blobs_unavailable with retryable: true. They enter the existing bounded retry schedule without a terminal scanner reason. Every retry must still prepare complete input and pass the canonical input scan; scanner refusals and staging limits remain terminal until the explicit release described above.

Scheduled and manual explicit queue admissions use the same exact-event review step. Aggregate shard recovery uses its per-item terminal ledger instead of producing a queue-level failure_reason from the shard's process exit. When a completed review has handed off to pending or active publication, a late failed-shard recovery remains queued until publication settles or parks. Its request is preserved because it may refer to a newer source. Explicit re-reviews, source changes, and publication source-drift or artifact-retention recovery retain their normal admission paths. Bay continues to show the publication as the item's current work until that handoff releases the deferred recovery.

ClawSweeper has three issue/PR scheduler paths:

  • exact event review for one target issue or pull request
  • hot intake for new or recently active queue edges
  • normal backfill for due backlog review

The lanes share report storage and apply rules, but they intentionally do not share throughput. Event review and hot intake keep new maintainer-visible work fast. Manual normal backfill has a configured ceiling of 89 concurrent Codex review shards within the 104 background slots available after reservations. Scheduled hot intake and normal backfill share a 32-slot queue cap. Manual normal review has an active floor of 38 requested shards: due items win first, and if fewer than 38 items are due, the planner can fill from older eligible reviews. The smaller worker allowance always wins. Scheduled planning does not use this floor.

Scheduled reviews can reuse exact unchanged inputs through structural or content caches. Changed PR content goes to Codex, including source comments and formatting. See Review Cache for admission, freshness, and runtime packaging rules.

Control-plane workflow retries

Shell calls to the control plane in review-plan.yml, sweep.yml, exact-review-reconcile-run.yml, and exact-review-dead-letter-reconcile.yml use scripts/control-plane-curl.sh. Each request retries connection failures and HTTP 5xx up to four attempts. A valid Retry-After delay (seconds or HTTP-date) is capped at 60 seconds; otherwise the waits are 2, 4, and 8 seconds. Each attempt emits a notice. The helper preserves the caller's final curl exit code, HTTP status output, and body handling. HTTP 4xx responses are not retried; callers retain their explicit lease-conflict, supersession, and deployment-skew handling. This includes enqueue, claims, review/status heartbeats, completion, lifecycle receipts, terminal-finalization operations, reconciliation, and the DLQ health probe. Existing typed batch clients retain their separate lease-aware retry policy. Fence and reservation failures are attributed to the existing queue_completion_failure infrastructure category, including a review that never starts because its status fence is unavailable.

Signed shell requests use control_plane_signed_post <url> <body> [curl options...] from the same helper. Shell steps compute the x-clawsweeper-exact-review-signature header only there: sha256= plus the hex HMAC-SHA256 of the exact body bytes, keyed by CLAWSWEEPER_WEBHOOK_SECRET. It sends the body with --data-binary and content-type: application/json through control_plane_curl, so signed calls get the same retries.

Before a job's source checkout, the workflow downloads this single helper from raw.githubusercontent.com, pinned to GITHUB_REPOSITORY and GITHUB_SHA, with three curl retries into RUNNER_TEMP. The bootstrap fails if the download fails, is empty, or does not define control_plane_curl. These steps source the temporary copy; after the full checkout, steps source the repository copy. Claimed-lease completion and terminal retry steps select the repository copy only when the source checkout succeeded. They use the validated temporary copy when checkout failed or was skipped, including direct-lifecycle recovery. The bootstrap never changes workspace Git configuration: an early sparse checkout can otherwise leave later checkouts sparse and omit local actions.

After setup, event-review-apply builds every lease heartbeat body with node dist/repair/exact-review-queue-request.js heartbeat --phase <review|status|finalizing>. The command reads the lease tuple from the EXACT_REVIEW_* and GITHUB_RUN_* environment, validates it, and prints the JSON body for control_plane_curl. --review-acknowledgement-comment-id is accepted only with --phase status, and --generation-start only with --phase review, as the queue requires. An invalid tuple stops the step before any request.

After checkout, the direct lifecycle step in event-review-apply and the receipt steps in event-review-publish build their lifecycle bodies with exact-review-queue-request.js lifecycle <router-receipt|canonical-receipt|terminal-disposition>. The command reads TARGET_REPO, ITEM_NUMBER, FENCE_KEY and REVISION, and accepts only the outcomes and terminal kinds that the queue accepts. A receipt id is --receipt-id-prefix plus the run id and run attempt.

The direct-lifecycle replay in event-review-publish runs before checkout. Its bootstrap step also downloads src/repair/exact-review-queue-request.ts, pinned to GITHUB_SHA like the curl helper, into RUNNER_TEMP/exact-review-queue-request.mts. The runner Node strips the types, so the source file runs without a build; it imports only Node built-ins. The replay builds each body before its side effect, so a body error stops the step before the router dispatch or any queue write.

The legacy-event-queue-intake job has no checkout. Its bootstrap step downloads the same command source. enqueue route prints the queue path for the repository_dispatch client payload in CLIENT_PAYLOAD: branch-authority when the payload names no branch, source-authority for an edited pull request with its complete source tuple, and enqueue otherwise. enqueue body prints the request body. Both reject an invalid target repository or branch before any request.

event-review-apply claims and completes its lease with the same command. Its bootstrap step downloads the command source. claim body reads the dispatch tuple (QUEUE_LEASE_ID, ITEM_KEY, QUEUE_LEASE_REVISION); an older dispatch without a tuple claims by lease id only. complete body reads the claim outputs and the results of the review steps, and a protocol 1 claim completes by lease id only. On HTTP 409, claim conflict and complete conflict read the response in RESPONSE. They print the error, and the step stops without an error, only when another run or a newer revision owns the lease: lease_not_active, lease_already_claimed, lease_decision_unavailable or stale_run_attempt for the claim, and lease_superseded for the completion. Every other 409 fails the step. Like the curl helper, the completion runs src/repair/exact-review-queue-request.ts from the checkout only when the checkout succeeded, and the downloaded copy otherwise.

event-review-publish claims and completes its publication lease with the same command. claim body --require-tuple fails when the dispatch does not name its tuple, because a publication dispatch always names it, and claim conflict classifies its 409. complete publication reads the publisher claim outputs and the publication result. It accepts only the completion kinds and reason codes that the queue accepts, and it adds the lifecycle disposition that the result implies. The publication completion has no safe 409, so every non-2xx response fails the step. The completion uses the same checkout-or-download rule. Its checkout is main, like the rest of the publisher.

event-review-terminal-finalization builds all its bodies with the same command. Its claim uses claim body --require-tuple and claim conflict before checkout. terminal-finalization <attempt|skip|retry> reads the claimed lease tuple from the EXACT_REVIEW_* and GITHUB_RUN_* environment. retry carries only the tuple. The requeue step that sends it uses the same checkout-or-download rule as the other completions. lifecycle command-ack-failed and lifecycle command-ack-observed read the lifecycle target like the other lifecycle records. --status-marker and --status-comment-id address the command status comment; an empty value means no such address. The attempt, skip and observed bodies need at least one address. Every step that runs before checkout now builds its queue request body with the command.

The two enqueue steps that run after checkout use the built command. enqueue publication ("Queue durable exact review publication" in event-review-apply) reads the claim outputs, ARTIFACT_NAME, GITHUB_SHA and the LIVE_* result flags. A protocol 1 claim sends a null lease revision and claim generation. enqueue source-drift ("Queue fresh review after source drift" in event-review-publish) reads the claimed decision and the producer run in PRODUCER_RUN_ID and PRODUCER_RUN_ATTEMPT. It keeps the failed_review_shard_recovery and command_proof_result source actions and sends every other action as source_drift_requeue. No step in sweep.yml builds a queue request body inline now.

The terminal-run observer (scripts/review-run-observer.mjs) uses plain Node after checkout and retries its telemetry POST up to three times. Each attempt retains the 20-second deadline. Connection resets and other recognized transient transport failures, timeouts, HTTP 408/429, and HTTP 5xx can retry; other HTTP 4xx responses and configuration or certificate failures remain terminal. Fallback waits are one and two seconds. A valid Retry-After value (seconds or HTTP-date) replaces that wait, capped at ten seconds, so publication has at most 60 seconds of request time and 20 seconds of backoff. GitHub job discovery is unchanged and remains subject to the enclosing workflow timeout.

The observer serializes and signs once, then reuses those exact bytes. The existing telemetry owner records the first (run_id, run_attempt) tuple and ignores duplicates, including retries after a committed response is lost. Responses are cancelled before retry. An acknowledged write whose response cleanup fails is reported separately and is never replayed. Retry diagnostics include the attempt, status or transport category, and delay; upstream response bodies are not logged. Exhaustion remains a failed workflow. This improves completeness of the existing review-observability data used by OpenClaw Bay; its schema, observer-only UI, and mutation boundaries do not change.

Workflow

Explicit workflow_dispatch item_number/item_numbers selections in review-plan.yml use src/repair/manual-review-enqueue.ts before review. They do not fall back to matrix publication. Admission is independently durable per item; the CLI reports failed members and continues the requested tail. Retries of the same workflow run reuse its run ID and item number, excluding attempt. Changed payloads conflict rather than borrowing an earlier delivery receipt. Admission preserves the resolved operator-selected target_branch; it discovers the repository default only when no resolved branch was supplied. The resolved codex_timeout_ms and one-off additional_prompt travel with that decision. Manual timeouts honor the requested positive duration up to the existing exact-review 45-minute cap; ordinary adaptive timeout behavior is unchanged. Ordinary item events may refresh source facts but cannot replace the selected branch, timeout, or instructions. A new explicit manual request may replace those options or clear its one-off instructions.

The queue advertises manual_publication.policy=record_comment_only and an explicit enabled bit through signed POST /internal/exact-review/admission-capabilities. The manual producer checks that contract before resolving items, independently of the aggregate public dashboard. Admission defaults off until EXACT_REVIEW_MANUAL_PUBLICATION_ENABLED=1 is configured after the consumer rollout described in repair operations. Manual decisions carry sourceAction=manual_explicit_review and immutable publicationPolicy=record_comment_only. Their completed artifacts use existing direct and batch publication capacity, fences, and retries. Source drift ends as superseded; unusable artifacts exhaust publication retries into existing dead letters instead of starting another model review. Temporary authority-service transport failures and HTTP 429/5xx responses use the existing retryable_failure / state_contention disposition and coordinator retry budget. They never authorize a write, revive a lease, or enter GitHub's inline retry loop; actual ownership/authentication rejection remains fail-closed.

Direct producers publish from the existing exact bundle's review/ directory, not raw artifacts/event output. EXACT_REVIEW_PUBLICATION_ARTIFACT_DIR selects that input relative to EXACT_REVIEW_WORK_ROOT; other callers retain the artifacts/event default. Snapshot capture and application use that same work root even when the publisher is invoked from a separate code checkout. Relative EXACT_REVIEW_BATCH_MUTATION_OUTPUT paths also resolve under the work root, including refusals that happen before publication authority is accepted. Producer selection.json, codex/, review-trees/, and sibling reports stay outside the selected bundle. The importer still rejects unexpected files, symlinks, and directories in its publication input.

Exact-item review jobs materialize the target with scripts/review-target-checkout.sh. Its cache is a blobless bare repository holding the target branch's full commit and tree history, the tags on that history, and every blob of the branch tip. Actions cache keys are <target-slug>-review-target-git-v1-<os>-<branch>-<YYYYMMDD>-<HH> (UTC). A run restores the newest entry from the same UTC day, fetches only the branch delta and the missing tip blobs, and clones the checkout locally with hardlinked objects. The first run of each hour saves the refreshed cache right after checkout, before any review input touches the workspace, and then deletes all but the two newest entries for that target branch; the first run of each day builds from GitHub again, which bounds pack and blob growth. Tip blobs fetched by id are the slow part (about 250 blobs/s from GitHub), and openclaw/openclaw changes roughly 7,800 tip blobs a day, which is why the cache is refreshed hourly rather than daily. The checkout keeps the contract of a direct git clone --filter=blob:none --single-branch: the branch at the current remote head, full non-shallow history, branch tags, and a promisor origin for lazy blob fetches. Cached tag refs are rebuilt through Git's normal tag auto-follow on every warm fetch, so deleted or moved tags cannot survive in the checkout. A non-fast-forward branch update rebuilds the cache so its old history cannot retain tags outside the current branch. A failed or partial restore is discarded, a failed cache fetch rebuilds the cache, and a failed local clone falls back to a clean clone without saving. The pinned Codex source cache is keyed by the Codex version pinned in the target checkout and is saved once per version.

The receiver workflow is .github/workflows/sweep.yml.

Important source files:

  • src/clawsweeper.ts: item selection, cadence, planning, review, dashboard, and status JSON
  • config/target-repositories.json: configured non-core target repositories and the conservative openclaw/* exact-review fallback
  • docs/target-repositories.md: target onboarding and rollout checklist
  • src/repair/workflow-utils.ts: GitHub Actions output shaping for plans
  • results/sweep-status/<repo-slug>.json: Git-backed operational state consumed by the dashboard
  • records/<repo-slug>/items/<number>.md: open item reports in the canonical Worker store
  • records/<repo-slug>/closed/<number>.md: archived item reports in the canonical Worker store

The canonical Worker owns records/**, while R2 owns ledger/v1/** and assets/**. The state branch of openclaw/clawsweeper-state retains only the operational jobs/**, results/**, notifications/**, and apply-report paths. See State storage for the ownership boundary and local hydration commands.

Broad normal and hot review workflows use run-scoped concurrency groups, so a new wave can overlap an older wave that has reached its long-tail or publish phase. Their plan jobs use a separate concurrency group per target repository with cancel-in-progress: false, serializing capacity decisions before each matrix expands. The default single-pending policy coalesces superseded planner ticks instead of retaining an unbounded history. Exact-item planners keep a run-scoped group and do not wait for broad background planning. Manual exact-item workflow_dispatch reviews use an exact-item concurrency group with the same single-pending policy, so newer revisions replace stale pending work instead of building a duplicate queue. Durable exact-review leases use lease-scoped workflow groups and remain owned by the Worker admission lane. If a successful exact-review run loses its completion callback, reconciliation uses the saved lease's accepted or deduplicated direct-publication receipt and requeue plan to preserve one owed source-drift review. The terminal requeue disposition is recorded before completion and stays on the old fenced revision. A superseded receipt cannot authorize a requeue; a newer command keeps its current decision and revision through the ordinary finishing path. The terminal-run backstop retains its */15 * * * * schedule. Event-triggered lease repair uses the reusable exact-review-reconcile-run.yml job with the exact-review-reconcile-workflow-run concurrency group and cancel-in-progress: false: one running job and one pending follow-up, with new events replacing superseded pending jobs. The lock covers both the cooldown check and reconciliation. Per-run reliability observations remain separate so coalescing does not drop telemetry; eligible events still create observer jobs.

Before event repair, scripts/exact-review-reconcile-guard.mjs reads the latest 30 runs of exact-review-reconcile.yml using gh run list --json databaseId,updatedAt. For runs updated within five minutes, gh run view --json jobs verifies that a lease-reconciliation job actually succeeded within that window, even if its observer job is still running. Observer-only runs, skipped repairs, and failed repairs do not extend the cooldown. A job-level output guard skips repair after a recent success; unavailable history or the 60-second aggregate history-read deadline allows repair to proceed. Each lookup is bounded by the remaining deadline, and a successful result is checked for freshness after the read. Scheduled and manual sweeps bypass the guard and inspect all claimed runs, so a terminal attempt whose event was coalesced or skipped is normally reconciled within about 15 minutes worst case, plus Actions scheduling and execution delay. This is a cadence backstop, not a hard wall-clock guarantee during platform outages. OpenClaw Bay is unaffected: queue schemas, telemetry observations, and public observer data contracts do not change.

Queue-completion failures remain visible separately from Codex or content failures, using the logical generation result and typed deferral rather than the review process exit alone. The workflow failure gate is unchanged. Caught Codex failures in an exact-review job also upload a separate 14-day diagnostic artifact while the runner remains alive. Its error.txt, stdout.error.txt, and stderr.tail.txt files are sanitized for repository readers and total at most 24 KiB with the readiness manifest.json. The manifest retains a bounded failure stage, reason code, and the queue's computed retryability even when unsafe raw detail is omitted. Raw reports, unstructured stdout, and non-error prompt events are omitted. It is never a publication input; cancellation, runner loss, or job termination can still prevent upload. OpenClaw Bay and queue schemas are unchanged.

Recoverable parked reviews use the nominal 5/10/20-minute retry ladder, but each item persists a schedule-time uniform jitter of 0.75-1.5x for every rung. After the third automatic recovery, operator-only HMAC-signed routes provide a bounded parked-review inventory and guarded resolution/fresh-recovery path. The five-minute dead-letter reconcile workflow inspects at most 100 parked targets, resolves terminal or repository-gone targets with an audit note, and can queue at most five fresh reviews with replay-safe recovery keys. A fresh review now requires a changed source action, head, base, draft state, body/content revision, fallback source timestamp when no content digest exists, or retry policy epoch; a queue revision by itself cannot reset the budget. Bump EXACT_REVIEW_RETRY_POLICY_EPOCH when a scanner, policy, or deployment change should reopen unchanged inputs. Manual workflow dispatches may explicitly opt into force_unchanged for a bounded operator override. Manual runs remain read-only unless execute is enabled; scheduled runs execute. Their sanitized parked inventory is uploaded beside the publication dead-letter inventory. Parked records carrying maintainer-command context remain visible with an exclusion reason but cannot be resolved or fresh-recovered by this background reconciliation path. Publication dead letters whose recorded pull-request head no longer matches the live head are never replayed or resolved from head drift alone. The reconciler may resolve them as superseded only when the existing HMAC-authenticated canonical-record read for the same target contains a complete review at that live head, then rechecks the same open GitHub node and head immediately before the guarded resolution. Each cycle checks at most ten such targets and resolves at most 20 rows per target; missing or mismatched evidence remains open as head_mismatch_unproven. tuple_protocol_invalid and workflow_cancelled rows are excluded from this path. Executed resolutions retain an audit note with the newer head and canonical endpoint and increment the publication completed and superseded totals; dry runs perform no mutation. This drains the existing queue state and does not add a dashboard health input or an OpenClaw Bay action: Bay remains observer-only. GitHub throttle deferrals use the same per-item jitter band when the queue turns the reported cooldown into its next-attempt timestamp, preventing a parked cohort from becoming eligible in lockstep; coordination and ordinary failure retries keep their existing timing. Exact publishers complete as superseded when apply verifies one trusted, complete, strictly newer durable review tuple for the same revision. The verified result travels as structured apply evidence; reason text is diagnostic only. Ambiguous or mixed results cannot terminalize the artifact, and legacy tupleless artifacts retain the existing fresh-review path.

Legacy protocol-v1 review leases without lifecycle admission rows complete without writing lifecycle facts. Once a completion commits its queue transition, alarm scheduling runs even if the subsequent lifecycle update fails, so persisted retries retain their scheduled wake-up.

Review publication and apply/comment sync use separate non-dropping queues. Queue and GitHub App requests retain their request deadline through response-body consumption. A stalled successful response becomes a typed timeout; an error response retains its HTTP status and rate-limit headers even if its body stalls. Apply treats a typed GitHub installation or abuse-rate-limit response as a bounded yield, not a failed scan. It checkpoints completed item work, records the interrupted item as skipped_runtime_budget, returns that item to the cursor, and exits successfully so a later scheduled or continuation cycle can retry it. This applies to comment-only sync and close-mode apply. Folder reconciliation also defers before mutation when its open-item scan is rate-limited; ordinary non-rate-limit failures remain fatal. The source fallback publication minimum, base, and maximum are 4, 24, and 48, but production overrides them to 8, 32, and 32, independently of the larger 80-review ceiling. The adaptive controller classifies GitHub pressure: a 403/429 or explicit rate-limit failure records a 15-minute cooldown, while GitHub 5xx failures record a 5-minute cooldown. Demand and recovery signals scale effective capacity within the production range. Production batch preparation is enabled for up to 8 concurrent size-8 batches, including 2 fresh-lane members per batch. Direct publication is also enabled and falls back to the retry/batch path when the direct result is retryable. Apply/comment sync remains per-target serialized. See docs/limits.md for effective values and docs/live-dashboard.md for the public lane telemetry. Tuple-aware state reconciliation prevents stale review snapshots from reviving closed records.

Batch admission is additionally gated by persisted GitHub credential circuits. Every batch needs the actions:openclaw/clawsweeper pool for producer artifact download, while target App circuits are owner-scoped so one exhausted installation does not stop healthy owners. A blocked pool prevents workflow dispatch, state hydration, and artifact download for each matching member until its reset-plus-deterministic-jitter recovery boundary; the alarm wakes at the earliest pending member boundary. The current batch collapses after the first pool failure, and unattempted members return without advancing their retry budget. Recovery is staggered rather than released as one cohort. An owner-scoped target App circuit also defers new exact-review admission for that owner, because review and publication share the installation quota; other owners remain admissible. Legacy dispatches normally carry the event repository's validated default branch. If an older producer omits it, intake creates a durable pre-admission branch-authority reservation instead of spending the workflow repository's Actions quota or assuming main. Resolution and direct-webhook source-head verification both consult and update the same owner-scoped target App circuit: the first quota response preserves its reset deadline, later same-owner reservations defer without another read or attempt charge, and reset recovery is bounded by the Durable Object alarm processor.

Exact publication routes only classifier-approved public reads through the repository Actions token. If that pool is exhausted after the current member's artifact is already present, the member may use the ambient target App once; later members still stop because they require the blocked Actions pool. The workflow also records a typed repository-pool observation if its final comment router dispatch encounters quota pressure. All typed observations and request counters are acknowledged with the same fenced batch completion, so a delayed cleanup is idempotent and cannot duplicate accounting.

Schedules

ClawSweeper Review Plan (review-plan.yml) owns the normal-backfill and hot-intake schedules below. sweep.yml retains exact event execution and apply; audit.yml owns audit schedules. All three keep their existing cadence.

openclaw/openclaw:

  • hot intake: */5 * * * *
  • normal backfill: 9/20 * * * * (minutes 9, 29, and 49)
  • apply: 3,18,33,48 * * * *
  • audit: 7 */6 * * *

openclaw/clawhub:

  • hot intake: 2/5 * * * *
  • normal backfill: 22 * * * *
  • apply: 8,23,38,53 * * * *
  • audit: 12 */6 * * *
  • review and apply work is gated by CLAWSWEEPER_ENABLE_CLAWHUB=1

openclaw/clawsweeper:

  • audit: 17 */6 * * *
  • self-review is primarily manual or event-driven; scheduled audit keeps the dashboard health row fresh

Failed Codex review backstop:

  • failed-review retry: 13 * * * *, in .github/workflows/failed-review-retry.yml
  • each retry sends one clawsweeper_item repository dispatch; sweep.yml routes it to the exact-review queue
  • retries remain dry-run unless CLAWSWEEPER_FAILED_REVIEW_RETRY_ENABLED=1
  • each retry is exact-item, cooldown- and attempt-bounded, and complements the immediate one-shot failed-shard recovery in the originating workflow

openclaw/fs-safe:

  • exact event review: enabled through the target repository dispatcher
  • scheduled review/apply/audit: not enabled yet
  • issues and PRs may auto-close only when already implemented on main

Generic openclaw/* and steipete/* repositories:

  • exact event/manual review: supported through configured generic fallbacks after the target dispatcher and GitHub App installation are present
  • scheduled review/audit: target fanout dispatches small cursor-based batches from target_inventory.owners
  • private and internal targets: local maintainer review only, using an operator-provided checkout
  • generic OpenClaw issues may auto-close only when already implemented on the default branch; generic OpenClaw PRs may additionally use age-gated mostly implemented there
  • generic steipete/* repositories are review/comment-only for issues and PRs

Manual review workflow_dispatch in review-plan.yml supports target_repo, target_branch, item_number, item_numbers, codex_timeout_ms, additional_prompt, and hot_intake. Explicit item selections use manual queue admission; broad runs offer due candidates to shared queue limits. Per-run batch_size, shard_count, apply_after_review, and its reason/age sub-options are retired. Existing callers must stop sending those inputs; there is no silent compatibility alias. Apply remains available through sweep.yml's separate apply_existing lane and its existing apply controls.

Target fanout dispatches review batches through repository_dispatch so each selected repository can carry its inventory default branch without consuming manual workflow inputs. Scheduled fanout runs in .github/workflows/target-fanout.yml and dispatches review plans to review-plan.yml and audit work to audit.yml. Its schedules are:

  • hot intake: 4/20 * * * *, 20 target repositories per cursor step. This 20-minute cadence is temporary containment for scheduled self-feedback; restore a faster cadence only after the loop is fixed and quota telemetry confirms it is safe. PR #959 intentionally moved this selector from every 15 minutes to every 5 minutes; this containment adjusts that current cadence without attributing the self-feedback defect to PR #959.
  • normal review: 14/20 * * * * (minutes 14, 34, and 54), 12 target repositories per cursor step
  • audit: 37 */6 * * *, 12 target repositories per cursor step

Audit fanout keeps at most 3 target audits in flight using audit.max_parallel_targets from config/automation-limits.json. It dispatches bounded waves and waits for every acknowledged run in a wave to become terminal before dispatching the next, including when a target fails or is cancelled. The dispatch API returns exact run IDs (return_run_details=true); a missing receipt, exhausted lookup retries, or 55-minute wave timeout stops further dispatch. The existing signed audit cursor stores remaining targets and outstanding run IDs after each dispatch and completion. The next invocation resumes that batch and drains the saved children before admitting another wave. Transient lookup errors retry twice with one- and two-second backoff. Audit dispatch requires available durable state; an unresolved dispatch receipt stays marked for operator investigation rather than freeing its slot. Run status is read once per minute. Coverage inventory tokens are minted again after the waves, immediately before the trailing coverage summary. The fanout job allows four hours for the twelve-target batch; ordinary review/hot fanout keeps its 30-minute timeout. The six-hour cadence, cursor selection, and twelve targets are unchanged. This bound covers the scheduled fleet fanout, not separate manual audits or the three dedicated core-repository audit schedules.

PR #1007 is directly relevant but was insufficient for the observed openclaw/libterminal#41 path. It was intended to recognize structurally proven ClawSweeper-owned comment or label activity while keeping timestamp-only or incomplete evidence eligible for conservative structural verification. Its mainline commit b83f2983da predates and is an ancestor of the ClawSweeper head used by run 31336140651. That later scheduled run still reported one structural-cache check, zero structural-cache hits, and one full hydration before publishing the durable comment again. The evidence therefore shows that #1007 did not suppress this specific execution path; it does not prove whether the receipt was missing, incomplete, stale, or bypassed at admission, and it does not make #1007 the cause of the loop. The temporary cadence reduction bounds demand while that remaining path is corrected.

There is no ClawSweeper PR #1032: the relevant record is issue #1032, which reported a post-#1007 review storm and was closed by PR #1036. PR #1036 changed the exact-review workflow and review-preparation boundary to forward sourceAction and classify scheduled_hot_intake and scheduled_normal_backfill as automatic, making them eligible for receipt and cache reuse instead of treating queued item numbers as explicit reviews. Its merge commit 138ee2f96e predates and is an ancestor of run 31336140651. The run log contains --review-source-action scheduled_hot_intake, so #1036 was effective at the classification boundary and was not bypassed there. The same run nevertheless recorded zero structural-cache hits and one full hydration, making #1036 a partial but insufficient mitigation for this case: it opened the cache-eligible path, while the available structural receipt still failed to match. This evidence does not attribute the remaining receipt mismatch to #1036 itself.

Each mode's cursor lives in the authenticated ExactReviewQueue Durable Object, not generated Git state. Reads and writes use a monotonic revision. If the canonical cursor endpoint is unavailable, normal and hot fanout warn and continue dispatch from a default cursor; persistence failure after dispatch does not fail those lanes. Audit fanout requires a batch-aware cursor store and stops on state failures so outstanding children cannot be forgotten.

Normal fanout refreshes the same signed live-open inventory consumed by GET /api/review-coverage, and both normal and hot fanout skip repositories with no open items. Normal fanout reserves a rotating slot for every selected repository, keeps the largest untracked backlog in each cycle, and apportions the remaining candidate volume by backlog share. The rotating slice is dispatched first, so one large repository can fill otherwise-idle capacity without permanently consuming smaller repositories' scheduled-feed budget.

Worker hydration also records the exact item identities present in the modern canonical tuple store. Normal fanout and each target planner use those identities for the same untracked_open boundary as the coverage endpoint. A hydrated legacy backfill report remains review context, but it does not count as coverage or yield a planner slot to a canonical re-review until a modern tuple exists.

The six-hour audit fanout also writes a GitHub Actions summary with canonical open-item reports reviewed in the trailing seven days versus batched live open issue and PR totals across the complete dynamic inventory.

Exact event review also starts Codex before generated-state hydration. The single-item review only needs the target repository and live GitHub item state; generated state is checked out afterward, just before publishing the review record, safe close result, and command-router ledger.

Default close-mode apply refreshes use that same queue-only intake. The merged apply report selects at most five distinct source-drift items in report order, with unverified-checkout holds filling spare slots. The producer resolves the repository default branch and each selected item's kind with narrow GitHub reads, then sends clawsweeper_item with source_action: source_drift_requeue and supersedes_in_progress: false. It does not send clawsweeper_target_sweep, run a broad planner, or hydrate canonical repository records before enqueue. Read or dispatch failures remain visible step failures; a dispatch notice is not a durable admission receipt. The existing intake signs the queue request.

source_drift_requeue uses the queue's existing low-priority recovery contract: existing pending or leased work, including maintainer-command context and source authority, wins; delivery deduplication, pending backpressure, and lease fencing remain queue-owned. Exact selection requests a fresh review of current source without a new force flag, stale source pin, or producer-supplied lease. Closed or missing targets still stop at the queue/executor live-state checks. General manual and broad dispatch behavior, the independent proof cursor, and close policy are unchanged. OpenClaw Bay needs no change: this producer reuses existing queue/lifecycle fields and adds no published schema, status field, or control.

Both source-drift producers, the publisher's remote-newer requeue and this apply refresh, are bounded by the queue's source-drift loop breaker: after EXACT_REVIEW_SOURCE_DRIFT_REQUEUE_LIMIT consecutive automatic generations without organic input or a command, the item parks as source_drift_loop instead of spending another review. The producer still receives a successful dedupe. A newer scheduled offer releases one review but keeps the count, because ClawSweeper's own post-review writes can feed hot intake. Separately, review_runaway_health alerts when any item claims more than EXACT_REVIEW_RUNAWAY_REVIEWS_PER_DAY reviews in a trailing day. Automation limits owns both rules.

Two deterministic no-ops are held instead of re-offered. Scheduled planning skips open-but-locked conversations, because their review run can only complete as a guarded no-op. The queue also keeps a completed locked or oversized-PR generation parked as locked_conversation or oversized_pull_request, so other automatic producers and ClawSweeper's own lease-comment writes on an oversized PR cannot reclaim it. Organic events, commands, a pushed head, and parked-review reconciliation release the hold.

Exact PR review marks ClawSweeper's own acknowledgement comment (clawsweeper-pr-ack) complete after the review snapshot and before direct publication, and GitHub moves the PR's updated_at for that edit. The producer's apply then syncs the durable review comment, may edit managed labels, and releases (deletes) its review lease comment. When direct publication is deferred, the batch publisher re-applies the same review against those writes. Apply freshness treats such an update as automation-only only when a receipt ClawSweeper recorded for the item accounts for the latest updated_at and the review's complete source, timeline, PR head, and review-activity receipt still matches the live item:

  • a live ClawSweeper-authored comment carrying a marker for the item, or a ClawSweeper-actor timeline event (a managed label edit), has exactly the item's latest updated_at; or
  • the review generation's own lease comment (review_lease_comment_id) is gone, the live durable review comment carries this generation's clawsweeper-review-version marker (item, reviewed_at, source revision, lease owner, and lease comment id) and was synced after the review snapshot, the latest update follows the generation's last recorded ClawSweeper write by at most five minutes (the shortest apply-lease hold), and no non-automation comment, review comment, or timeline event is visible after the snapshot.

Any other change in the window, including a human comment, title/body or label edit, PR review, or new head, still records skipped_changed_since_review and requeues a fresh source_drift_requeue review. Without these receipts, a close proposal kept open by its producer made every deferred re-apply drift and requeue indefinitely.

Automerge Fast Path

Automerge is an exact-item event path. A maintainer command dispatches one review for the current PR head. If review requests a repair, the adopted repair worker may push a branch fix; after a successful contributor-branch repair it immediately dispatches another exact-head review and then shepherds the repaired head for a bounded window instead of exiting immediately. That keeps the normal path to:

  1. command acknowledgement;
  2. exact-head review;
  3. optional branch repair;
  4. immediate exact-head re-review;
  5. merge after checks, review verdict, and policy gates pass.

The complete state machine is documented in docs/repair/automerge-flow.md. Keep this section as the scheduler-facing summary.

The automerge status comment is the live progress surface. It is edited in place and records review, repair, re-review, and merge events with durations, run links, and commit links.

If a no-op automerge repair finds that the PR was already the canonical fix, the worker does not stop at the observational result. It immediately continues the state machine: either queueing a fresh exact-head review, or, when the existing ClawSweeper review only asked a maintainer to land the canonical PR and the maintainer already opted into automerge, queueing the merge gate for that exact review comment.

Automerge activation also checks the OpenClaw changelog policy before spending an exact-head review pass. User-facing fix, feat, and perf PRs that touch non-doc/test files and do not already include CHANGELOG.md go straight to the adopted repair worker, so the changelog fix happens in the first loop instead of being discovered only at the final merge gate.

After live hydration, adopted automerge/autofix repairs now skip the read-only Codex planning pass entirely. The worker emits a generic structured fix artifact directly: repair the contributor branch, rebase onto current main, address comments/review findings/failing checks, add a changelog entry when required, and validate. The execute stage still owns all GitHub mutations, validation authority, push, exact-head review, checks, and merge gating.

For explicit base-sync-only repairs, the repair executor first tries a deterministic fast path: rebase onto current main, apply known mechanical conflict resolvers such as isolated CHANGELOG.md conflicts and generated config checksum three-way conflicts, push the repaired branch, then wait for exact-head review and GitHub checks. For substantive automerge repairs, Codex owns the initial rebase plus PR-comment, CI, and local-test repair loop; the executor still owns every GitHub mutation and reruns the normalized validation gate before push. If main moves during that final validation, the worker does one final base sync by default and lets the immediate exact-head review plus GitHub checks validate the pushed head; CLAWSWEEPER_FINAL_BASE_SYNC_ATTEMPTS can raise that only when extra local passes are intentionally worth the delay. Likewise, the last internal Codex /review is not a dead end: if it still finds an actionable issue, the worker can run one final review-fix pass, require changed-surface validation to pass, push the repaired branch, and leave the immediate exact-head review plus GitHub checks as the merge authority. The default shepherd wait is ten minutes with 15-second polls, controlled by CLAWSWEEPER_AUTOMERGE_SHEPHERD_WAIT_MS and CLAWSWEEPER_AUTOMERGE_SHEPHERD_POLL_MS. Terminal check failures stop the shepherd wait immediately and dispatch the router so the failed-check repair loop can start without waiting for the full timeout.

The final router gate waits up to ten minutes for transient GitHub merge state or pending required checks, polling every 15 seconds. Pending checks are wait states, not repair triggers; terminal required-check failures can still dispatch the adopted repair worker. If GitHub still reports UNSTABLE, ClawSweeper allows the merge command to try when the only visible blockers are ignored non-gating automation checks such as ClawSweeper Dispatch; GitHub branch protection still enforces required checks at merge time. If the live merge preflight reports DIRTY, BEHIND, or CONFLICTING, automerge treats that as repairable rebase work and dispatches the adopted repair worker instead of leaving the PR open with only a status comment.

Capacity

The exact-review queue owns hosted review concurrency and retry authority. Each admitted item runs one Codex session. Planners select candidates; they neither reserve a matrix of workers nor publish review results.

Repair and explicit item work retain their existing priority. Broad manual and scheduled intake share the queue's background capacity and pacing; see docs/limits.md for the configured budgets.

Current defaults:

  • exact event review: 1 shard, 1 item
  • exact manual hot intake: 1 shard, 1 item
  • scheduled hot intake and normal backfill: size candidate selection from live queue-advertised capacity, with normal fanout sharing that budget across its selected targets. If the capacity probe is unavailable, a direct single-target schedule falls back to 50 candidates; normal fanout creates a pool of 50 candidates per selected target and apportions that pool by backlog. Each selected item enters the durable exact-review queue, and every admitted item receives its own parallel workflow
  • scheduled review admission target: 220 items/hour across the fleet, with a 24-item burst. The budget meters executed reviews: every organic review that starts generation (the claimed run's startup ownership check, sent with generation_start: true) consumes the budget first. Organic admission is free, so work superseded or coalesced before claim, deduped, completed at the dispatch-time live check without a run, or claimed by a run that exits before generation costs nothing. A scheduled admission pays when admitted and prepays its own first started generation; later starts of the same item (reruns, retries, requeues) are charged when they start. Organic debt carries down to minus the burst; further debits at that floor are forgotten. Organic work remains unconditional, so this target does not cap total executions. Within the scheduled remainder, hot intake is capped at 30 items/hour by EXACT_REVIEW_HOT_INTAKE_RATE_PER_HOUR and normal backfill may use the rest (190 items/hour of lane rate), so the more frequent hot offers cannot take the remainder oldest-first backfill needs. At most 32 scheduled reviews are dispatching or leased across both lanes
  • review admission and pressure are computed independently from publication; top-level queue health describes reviews while lanes.publication retains publication backlog, retry, DLQ, and health telemetry
  • fleet fanout: 20 hot targets every 20 minutes as temporary self-feedback containment, and 12 normal targets every 20 minutes; each target cycle can offer up to 50 due items to the shared admission budget
  • broad manual runs use the same queue capacity as scheduled feeds; normal planning scans at most 250 GitHub pages and hot intake at most 10

The shared hard cap bounds each candidate offer. A target-fanout allocation may reduce that offer, but it is not a per-run worker count. Each planner uses one logical partition with no active-floor backfill, then ends after admission. Periodic schedules and the existing apply-lane backstop provide later intake; there is no matrix runtime archive, aggregate publisher, or recursive review continuation.

Broad planning does not use the active-floor backfill. It selects only due items, records each candidate's previous-review age in plan.json, and writes a run-summary funnel for selected, attempted, enqueued, deduped, shed, and deferred items. The public queue projection exposes the configured rate and replay contract under scheduled_feed in GET /api/exact-review-queue; private bucket balances are omitted. Queue telemetry distinguishes backpressure from scheduled-rate shedding so an operator can distinguish a full review queue from intentional 220/hour pacing. The 24-item burst bounds a cold-start cohort to roughly 720 GitHub requests at the observed planning average of 30 requests per completed review. Before its first enqueue, the producer reads the queue-owned contract through signed POST /internal/exact-review/admission-capabilities. Dashboard telemetry cannot block this capability check. Transport failures retain their HTTP or network diagnostic; unsupported pacing or replay contracts fail closed. Deploy the Worker before the updated producer; an older Worker returns HTTP 404. Scheduled review ingress requires scheduled_feed.enqueue_replay: scheduled_disposition_v1 before retrying transient transport or HTTP 5xx failures with the same signed delivery bytes; legacy ambiguous receipts fail closed. Publication enqueue remains single-attempt; batch lifecycle router receipts and terminal dispositions use bounded, byte-identical retries backed by durable operation identities. Operator retirement and other callers retain their single-attempt default. Legacy lifecycle payloads without replay identities remain single-attempt as well.

Normal fanout ordinarily divides one live queue-advertised candidate-capacity budget across the selected repositories; it does not grant 50 candidates to each target. If that capacity probe is unavailable, the bounded fallback for a 20-minute cycle is 50 items/target * 12 targets = 600 items/cycle. Its fallback therefore offers at most 1,800 candidates/hour before due filtering, planner capacity clamping, dedupe, and Worker admission. The direct 20-minute normal schedule also uses live advertised capacity; its fallback offers 150 items/hour before the same bounds. These paths therefore have enough candidates to keep the shared token bucket fed despite dedupe or uneven fleet distribution. The queue refills scheduled admission credit at 220/hour after bounded debits for started organic reviews. A combined load near 220 executed reviews/hour would need about 220 * 4.1 / 60 ≈ 15 concurrent review workers at a 4.1-minute mean service time and budgets roughly 6,600 GitHub requests/hour. With organic work near 110 executed reviews/hour, scheduled work receives roughly 110/hour: hot intake up to its 30/hour cap and normal backfill the rest, however many more organic items were admitted and then closed or superseded before they ran. The separate 32-slot scheduled cap also bounds old queued work and slower reviews while organic/manual requests retain admission priority. Organic overload can exceed these estimates indefinitely. Rate and burst reduce scheduled request and inference demand; the pending soft limit remains a separate queue backpressure bound and should change only when queue-memory or latency evidence requires it, not automatically with the request budget.

On saturated queues, normal planning reads the complete bounded open-item scan before selecting candidates. For the current largest repository this is about 60 REST pages per normal tick. openclaw/openclaw receives up to six normal ticks per hour (three direct, plus three fleet fanout ticks while it holds the largest untracked backlog), or roughly 360 installation-token requests per hour; that bounded cost is necessary for oldest-review fairness.

Optional planning-started and in-progress dashboard publishes in the plan job are capped at 20 seconds. They are useful telemetry, but they must not delay candidate selection or the review shard matrix; the publish job writes the final dashboard state after review artifacts land.

The plan jobs calculate live capacity from the GitHub Actions REST runs list, normalized to the same fields as gh run list. The REST endpoint is used because gh run list can miss active repository-dispatch runs in some local and Actions contexts, which would make the scheduler undercount active review workers. Every active status is paginated so fleets above 100 runs remain fully counted.

Cadence

The planner considers only open issues and PRs that pass shouldPlanItem. Protected labels and other non-reviewable items are skipped before Codex work is allocated. Scheduled scans also skip open-but-locked conversations; an unlocked webhook event admits the item again.

Review cadence:

  • items with target-side activity since the last real review: hourly
  • items created in the last 7 days without new target-side activity: daily
  • pull requests outside the hot window: daily
  • issues created in the last 30 days: daily
  • older inactive issues: weekly
  • review policy hash changes: due immediately

The activity check ignores ClawSweeper-owned GitHub mutations that are already recorded in durable report frontmatter. review_comment_synced_at covers public review comment writes, and labels_synced_at covers ClawSweeper label-only writes such as priority or advisory issue-label syncs. If GitHub updated_at is at or before either marker, the planner does not treat it as fresh reporter or maintainer activity.

Selection uses weighted buckets so hot issues cannot starve pull requests and older issue backlog forever. The normal scheduler cycles through:

  • hot issues
  • hot pull requests
  • activity-driven items
  • daily pull requests
  • recent issues
  • weekly older issues

Within each bucket, earlier due times and older reviews win before item number. The live open-item scan is compared with the canonical record index first. Items with no canonical record consume capacity before any re-review. Within that never-reviewed cohort, six-day coverage ordering and the existing weighted bucket mix still apply. Once first-review candidates are exhausted, tracked items enter the six-day coverage lane in oldest-reviewed_at order across all buckets before hot-item churn. Normal planning completes its bounded scan before applying that ordering, so a saturated item-number prefix cannot hide either untracked items or older review timestamps. The extra day is operational headroom before the seven-day freshness deadline.

Planning

The plan step runs:

pnpm run --silent plan -- \
  --target-repo "$TARGET_REPO" \
  --batch-size "$BATCH_SIZE" \
  --max-pages "$MAX_PAGES" \
  --shard-count 1 \
  --codex-model internal \
  --codex-sandbox danger-full-access \
  --min-active-shards 0 \
  --min-backfill-review-age-minutes "$MIN_BACKFILL_REVIEW_AGE_MINUTES"

pnpm run plan returns:

  • candidates: selected open items
  • shards: planner partitions (hosted queue feeds use one partition)
  • capacity: batch_size * clamped_shard_count
  • dueBacklog: due candidates found during the complete bounded scan
  • activeCodexTarget: nonempty shard count
  • oldestUnreviewedAt: oldest scanned due candidate with no existing review
  • capacityReason: why the selected count did or did not fill capacity
  • floorBackfill: selected stale current-review candidates used to fill the active floor
  • matrix: legacy CLI output, unused by hosted queue execution

pnpm run workflow -- plan-output maps that JSON to GitHub Actions outputs:

  • planned_count
  • planned_capacity
  • planned_item_numbers
  • planned_shards
  • active_codex_target
  • due_backlog
  • oldest_unreviewed_at
  • capacity_reason

Capacity reasons:

  • saturated: due backlog filled planned capacity
  • under capacity: due backlog below planned capacity
  • idle: no due candidates found
  • exact: requested item selection
  • idle: no requested open items found

Status and Dashboard

Planning and publish steps call pnpm run status, which writes structured JSON under results/sweep-status/<repo-slug>.json in generated state. Every sweep workflow status update must pass the active --target-repo so a ClawHub, ClawSweeper, or OpenClaw lane updates only its own dashboard row. The README dashboard reads that JSON and shows:

  • active Codex target
  • planned review items
  • planned review shards
  • planned review capacity
  • due backlog scanned
  • oldest unreviewed scanned
  • capacity reason

Historical matrix summaries retain their original shard fields. Current planner runs report the queue-admission funnel in their Actions summary. Use live queue and workflow telemetry for active worker counts; a selected candidate is not an active review.

Plan jobs hydrate canonical records for selection. Each admitted exact-review worker and publisher follows the existing immutable queue ownership and publication paths. Apply and audit retain their own state hydration and reconciliation.

Apply

Review is proposal-only. Apply is the only issue/PR scheduler path that mutates GitHub close state.

Apply wakes every 15 minutes for openclaw/openclaw and on offset 15-minute ticks for ClawHub. It re-fetches live GitHub state, checks labels, author association, paired issue/PR state, snapshot drift, and repository profile rules. It closes only unchanged high-confidence proposals and otherwise updates or syncs the durable ClawSweeper review comment.

Apply reconciles hydrated records before both candidate preselection and execution. Each pass immediately publishes only the item, closed, plan, and decision-packet paths for record numbers it changed. This keeps folder moves and closed-item sidecar cleanup durable even when policy filtering, an empty comment-sync batch, or an empty close queue makes the rest of the run a no-op. If another publisher updates the same tuple first, its newer tuple wins and reconciliation defers that item instead of rebuilding stale report or sidecar content.

Batch review publishers hydrate only the item tuples present in their artifacts, publish those records, and synchronize the selected durable review comments in the same job. Exact issue/PR reviews likewise synchronize their selected comments before completing. Neither path dispatches a second broad comment scan.

Automatic apply may close up to 40 items per run. Long apply runs commit checkpoints every 40 fresh closes and dispatch a continuation with a fresh GitHub App token after any checkpoint that closes at least one item. A saturated scan that closes nothing stops without chaining so the same records cannot create an unbounded runner loop.

Only automatic close-mode apply runs may queue missing hot or normal review backstops, including when no close candidates are available. Targeted apply and comments-only sync retain their requested scope, including when quota pressure ends the apply process successfully without publishing a comment.

Untargeted cursor-based close apply starts with a 600-record scan window. If the previous cursor window was a full close-mode scan, closed nothing, skipped at least 80% of processed records, and did not hit a live-fetch, runtime-budget, or missing-cursor failure, the next automatic window expands to inspect more records, capped at 1800. Each automatic checkpoint may spend up to 20 minutes in deterministic apply scanning, while the existing 55-minute App-token budget, 70-minute apply step, and six-hour coordinator job ceiling remain unchanged. This changes only the deterministic scan window: apply_limit, checkpoint size, close gates, live-state checks, and maintainer policy gates stay unchanged. The workflow logs and sweep status detail include the selected scan window and reason.

Automatic windows reserve up to two candidates for PR close-coverage proof, capped by the effective close budget. Confirmed proof-gated close proposals stay ahead of speculative promotion proofs; spare proof capacity rotates promotions on the same independent proof cursor. Reserved proofs run before deterministic closes could exhaust the checkpoint's mutation limit. An executor trace advances each cursor only through records actually examined, so a partial window preserves unexamined candidates and each pool resumes from its exact last-examined position. Coverage proof, live-state refresh, freshness checks, and close gates remain unchanged. Explicit targeted apply runs keep their requested item set and ordering policy.

Apply keeps selected report bodies in memory and loads independently reviewed paired records only when a close guard requests them. Exact-event publication does not expand its selected set. Broad apply still sorts the complete open candidate set; it no longer loads a second copy for paired lookups. Finalization reloads only requested, result, and unfinished/in-flight item records from the open and closed directories, rather than retaining the archive. This includes partial failures and runtime-budget yields and does not change cursor ordering, close eligibility, canonical baselines, or ledger identities. OpenClaw Bay needs no change: the public status, record, and ledger contracts are unchanged.

Before a close-mode apply run starts, the workflow summarizes the selected close candidate mix by quality bucket in the status detail. Buckets such as implemented-on-main, duplicate/superseded, needs PR close proof, aging/low-signal (including stalled-unproven and abandoned PRs), policy-sensitive, and retry-after-guard-skip are operator-facing telemetry only; the bucket classification does not change close limits, live-state checks, or policy gates. Stalled-unproven and abandoned PR proposals are eligible for apply selection, where the executor re-checks their PR-only age, activity, proof, status, and human-engagement gates before closing.

After a default close-mode cursor run for openclaw/openclaw, the apply job requeues up to five exact reviews for records whose close was blocked by source drift (skipped_changed_since_review) or by a stored review without verified local checkout access. Both blocks have the same cure: a fresh exact review re-verifies the close proposal at the current snapshot and writes a close-capable record, so the next apply pass can execute instead of skipping the same stale records every sweep. The per-run cap bounds review spend, and the exact-item queue's supersession semantics absorb repeat dispatches.

Apply and comment-sync Actions run titles include the target repository. Before dispatching a default cursor-based apply continuation, the workflow checks recent active or queued same-target default cursor runs and treats one of those runs as the continuation instead of adding another pending run. Custom-input and explicit-item runs have a different title and cannot suppress the default cursor lane; their own continuations still dispatch with the exact inputs. The log identifies the default cursor run that covered the continuation.

Continuation and Recovery

Broad review runs end after queue admission. Periodic schedules and the apply-lane review backstop offer later candidates under the same shared limits. The retired matrix publisher and failed-shard recovery jobs do not dispatch continuations or replay review work. The planner still offers existing canonical vision-fit OpenClaw reports (under their separate opt-in) and viable reports for other eligible targets to the bounded implementation-intake dispatcher. The dedicated strict-bug backfill and exact-publication hooks remain unchanged.

The existing failed-review retry selector remains bounded by its per-source attempts and cooldown. Both issues and PRs dispatch automatic exact queue work, never an explicit manual request, so a scanner hold still blocks them. PR retries retain their head pin. Issue retries carry their expected source revision through admission and the existing post-hydration guard; changed source ends the retry without model work or publication and is left to normal source-event/backfill intake. Repository dispatch uses the workflow repository's default branch; non-default --workflow-ref requests are rejected rather than silently ignored. Repair follow-ups also use repository dispatch only. A failed dispatch surfaces to the existing retry handling; it cannot fall back to a manual workflow and acquire explicit retry authority.

Historical shard artifacts and ledger events remain readable evidence, but are not active retry authority. The queue owns current review retries, scanner holds, publication recovery, and stale-lease handling.

Each item report also records durable review cost proxies in front matter and a Review Telemetry section: prompt characters, static prompt characters, GitHub context characters, output schema characters, additional prompt characters, context collection milliseconds, and Codex review milliseconds. These fields are intended for scheduler and prompt-budget experiments, so later throughput work can compare time and token proxies without scraping transient workflow logs.

The remaining operational state checkout uses a blobless shallow clone. Git publication is serialized by the Durable Object state-writer coordinator and uses one ordinary fetch, commit, and push.

Audit

Audit is target-read-only and runs separately from review and apply in .github/workflows/audit.yml. It refreshes results/audit/<repo-slug>.json and the README Audit Health table from live GitHub state. Scheduled audit currently covers:

  • openclaw/openclaw: 7 */6 * * *
  • openclaw/clawhub: 12 */6 * * *
  • openclaw/clawsweeper: 17 */6 * * *

Every manual dispatch of audit.yml runs an audit. Its only input is target_repo; do not send the retired audit_dashboard input to either workflow. Review planning dispatch uses review-plan.yml; exact event execution and apply remain in sweep.yml.

gh workflow run audit.yml -f target_repo=openclaw/openclaw

Audit runs retain the Audit ClawSweeper state title, the existing per-target scheduled concurrency groups, and the run-scoped manual operator concurrency group. The hosted public-target admission gate remains in place.

The audit lane first tries a ClawSweeper GitHub App read token for the target repository. If that token is unavailable, it falls back to the workflow token for public read-only API access so dashboard rows do not remain unknown just because mutating scheduled work is still gated.

Before calculating audit health, audit also runs the folder reconciler against live open GitHub state. This is target-read-only and mutates only canonical Worker records: reports for items no longer open move from items/ to closed/, reopened archived reports move back to items/, and duplicate closed copies are removed. GitHub Actions uses the fast reconciliation mode that does not fetch each closed item individually for closed_at; large cleanup runs therefore avoid hundreds of per-item GitHub API subprocesses. The local reconciler still fetches closed_at by default for operator runs; pass --skip-closed-at for fast canonical cleanup.

Review publishing applies newly generated artifacts first, then runs the same fast reconciler once before committing records. It does not run the slower artifact-apply reconciler and the explicit publish reconciler back to back.

After publishing Git-backed audit results and reconciling canonical records, audit dispatches the openclaw/clawsweeper-state dashboard renderer; that repository's 15-minute schedule remains the fallback if dispatch is delayed.

Monitoring

Useful commands:

gh api 'repos/openclaw/clawsweeper/actions/runs?per_page=100' \
  --jq '.workflow_runs[] | select(.name == "ClawSweeper") | {id,name,display_title,event,status,conclusion,created_at,head_sha,html_url}'

gh run view <run-id> --repo openclaw/clawsweeper --json jobs \
  --jq '[.jobs[] | select(.name == "Review exact event item") | select(.status=="in_progress")] | length'

gh api repos/openclaw/clawsweeper/readme --jq '.content' | base64 --decode

Read the remote generated README, not only the local checkout, when checking the live dashboard. Generated dashboard state is published from GitHub Actions and can be newer than local files.

Common Changes

To change review spend, set EXACT_REVIEW_TARGET_RATE_PER_HOUR; the Worker applies the fleet-wide rate while scheduled planners size their candidate batch to free review capacity. To rebalance scheduled work, set EXACT_REVIEW_HOT_INTAKE_RATE_PER_HOUR; normal backfill receives the remaining lane rate. Target fanout divides that capacity by untracked backlog after reserving its round-robin fairness slice. To change manual normal Codex sessions, update the worker limits and workflow defaults together.

To change review cadence, update the cadence constants and the scheduler bucket logic in src/clawsweeper.ts, then update dashboard labels and this document.

To add a new target repository, add a repository profile, wire schedule target resolution and concurrency target resolution in .github/workflows/review-plan.yml and .github/workflows/sweep.yml, then confirm the generated state paths remain flat under one repo slug.

Hosted owner fallback is limited to openclaw/* and steipete/*. To schedule another owner, add explicit repository profiles and include that owner in target_inventory.owners, then wire that owner's inventory token or explicit public-inventory fallback into the fanout workflow. Configuration alone does not activate a new owner. Fanout ignores every repository that is not admitted by the shared configured-profile-or-owner-fallback policy. Keep scheduled fanout public-only unless the generated records publish to a private state surface.