Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,7 @@ script-test:
$(call run-timed,bash scripts/pre-review-test.sh)
$(call run-timed,bash scripts/post-review-test.sh)
$(call run-timed,bash scripts/risk-tier1-test.sh)
$(call run-timed,bash scripts/pre-fix-test.sh)
$(call run-timed,bash scripts/post-fix-test.sh)
$(call run-timed,bash scripts/post-retro-test.sh)
$(call run-timed,bash scripts/pre-scribe-test.sh)
Expand Down
20 changes: 18 additions & 2 deletions agents/fix.md
Original file line number Diff line number Diff line change
Expand Up @@ -183,8 +183,24 @@ Bot-triggered runs (from the review agent) are capped at `ITERATION_CAP`
(default: 5). When the iteration count approaches this cap, the `needs-human`
label is added and the autonomous loop stops on the next attempt. A human can
then direct the agent with `/fs-fix` commands up to `ITERATION_CAP_HUMAN`
(default: 10) total iterations (bot + human combined). This ensures humans
are never locked out of the agent after a bot loop exhausts its budget.
(default: 10) total iterations (bot + human combined). This ensures humans are
never locked out of the agent after a bot loop exhausts its budget.

A maintainer can tighten the *autonomous* loop for a single PR with a
`fullsend-fix-budget/N` label (N a positive integer). If present, the smallest
valid label lowers the bot cap to N, so the review→fix loop escalates to a human
sooner on a risky or expensive PR. It applies to the bot cap only: the human
`/fs-fix` cap is never tightened, so the label can never lock a human out. The
label can only *tighten* the bot cap, never raise it: a value at or above the
bot cap has no effect, and malformed values (non-integer, zero, negative, or
absurdly large) are ignored so a bad label cannot silently block or widen the
loop. pre-fix enforces the tightened cap and post-fix reports against the same
effective cap (summary and the `needs-human` warning).

> **Not yet active.** The label is recognized and enforced, but the review→fix
> workflow does not yet forward PR labels to the fix agent (`PR_LABELS` is unset
> in the delivery path), so applying the label currently has no effect. See
> `docs/fix.md` and the wiring follow-up before relying on it.

## Validation retry behavior

Expand Down
9 changes: 9 additions & 0 deletions docs/fix.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,7 @@ Remove the label or use `/fs-fix` to re-engage.
|-------|---------|
| `fullsend-no-fix` | Prevents automatic fix runs on this PR. Applied by `/fs-fix-stop`. Manual `/fs-fix` commands are unaffected. |
| `needs-human` | The fix agent is approaching its iteration cap and needs human direction. Applied automatically when an automatic fix iteration reaches the warning threshold. |
| `fullsend-fix-budget/N` | **Reserved — not yet active** (requires `PR_LABELS` wiring in the review→fix workflow; see [Iteration limits](#iteration-limits)). Once active: tightens the *autonomous* review→fix loop for this PR to `N` iterations (`N` a positive integer), so the bot escalates to a human sooner. Applied by a maintainer. Lowers the bot cap only (never the human `/fs-fix` cap) and can only tighten it, never raise it; malformed values are ignored. |

## Configuration

Expand Down Expand Up @@ -156,6 +157,14 @@ The fix agent enforces iteration caps to prevent infinite review-fix loops:
`needs-human` label.
- Each `/fs-fix` comment cancels any in-flight fix run for the same PR and
starts a new one.
- **Per-PR override (reserved, not yet active):** a `fullsend-fix-budget/N`
label is intended to lower the *automatic* cap for a single PR (bot only; the
manual `/fs-fix` cap is never tightened, so a human is never locked out). The
parser and enforcement ship in pre-fix/post-fix, but the review→fix workflow
does not yet forward PR labels to the fix agent (`PR_LABELS` is unset in the
delivery path), so the label currently has no effect. Wiring it requires the
workflow to pass the PR's labels through `PR_LABELS`; until then the caps
behave exactly as above.

## Multi-forge support

Expand Down
28 changes: 28 additions & 0 deletions harness/fix.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,34 @@ env:
TRIGGER_SOURCE: "${TRIGGER_SOURCE}"
HUMAN_INSTRUCTION: "${HUMAN_INSTRUCTION}"
FIX_ITERATION: "${FIX_ITERATION}"
# NOTE: a `fullsend-fix-budget/N` PR label can tighten the autonomous (bot)
# fix cap; pre-fix/post-fix consume PR_LABELS to enforce it (the human
# /fs-fix cap is never tightened). No entry is needed in this shared block:
# childScriptEnv() builds the pre-/post-script env from os.Environ() first,
# so a step-level env var already reaches both consumers.
#
# Do not activate from github.event.pull_request.labels: reusable-fix.yml is
# workflow_call'd from workflow_dispatch (github.event.pull_request is
# empty), and the dispatcher strips .labels from event_payload. Resolve
# current labels in the "Extract PR number and context" step, the same
# pattern reusable-fix.yml already uses for HEAD_REF on issue_comment:
#
# gh api "repos/${SOURCE_REPO}/pulls/${PR_NUM}" \
# --jq '[.labels[].name] | join(",")'
#
# Write that to GITHUB_OUTPUT (e.g. pr_labels) and set
# PR_LABELS: ${{ steps.context.outputs.pr_labels }} on "Run fix agent" in
# both reusable-fix.yml and the inline fix job in reusable-dispatch.yml.
# That returns live labels, not an event-time snapshot, and works for
# issue_comment and pull_request_review. Extending the trimmed
# event_payload is an alternative that needs the injection review the
# trim comment demands. The workflow wiring is a fullsend follow-up; this
# repo only consumes PR_LABELS when the step env sets it.
#
# If an env.runner entry is ever wanted it belongs in
# forge.github.env.runner — and in forge.gitlab.env.runner only once the
# GitLab agent template exports a set-possibly-empty PR_LABELS. Never in
# this shared block (fail-closed).
REVIEW_BODY_FILE: "${REVIEW_BODY_FILE}"
PRE_AGENT_HEAD: "${PRE_AGENT_HEAD}"
PUSH_TOKEN: "${PUSH_TOKEN}"
Expand Down
52 changes: 52 additions & 0 deletions scripts/lib/fix-budget.lib.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
#!/usr/bin/env bash

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

6. Protected scripts require human approval 📜 Skill insight § Compliance

The PR modifies multiple files under the protected scripts/ path, so it must receive human review
and must not be auto-approved. The feature rationale provides context, but there is no linked issue
authorizing these governance/infrastructure changes.
Agent Prompt
## Issue description
This PR changes protected `scripts/` infrastructure and cannot be auto-approved.

## Issue Context
Route the PR for human approval and link the authorizing issue for the protected-path changes before merge.

## Fix Focus Areas
- scripts/lib/fix-budget.lib.sh[1-42]
- scripts/pre-fix.src.sh[20-29]
- scripts/pre-fix.src.sh[114-119]
- scripts/pre-fix-test.sh[1-59]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

# shellcheck shell=bash
# fix-budget.lib.sh — parse a per-PR fix-loop budget from PR labels.
#
# A label of the form `fullsend-fix-budget/N` (N a positive integer) lets a
# maintainer cap the review->fix loop for a single PR below the global
# iteration cap. The label can only TIGHTEN the cap, never raise it:
# enforcement lives in pre-fix, which applies min(label_budget, cap).
#
# Bundled into pre-fix.sh via bundle-sh.sh.
#
# Expected env vars (optional):
# PR_LABELS — PR label names separated by commas and/or newlines. Absent/empty
# is fine: parse_fix_budget then returns nothing and the cap is
# unchanged. (The upstream dispatcher comma-joins labels; a
# newline-joined value is also accepted.)

[[ -n "${FIX_BUDGET_SH_LOADED:-}" ]] && return 0
FIX_BUDGET_SH_LOADED=1

FIX_BUDGET_LABEL_PREFIX="fullsend-fix-budget/"

# parse_fix_budget [labels]
# Reads label names (arg 1, or PR_LABELS env when omitted) separated by commas
# and/or newlines. Echoes the smallest valid budget found, or nothing when no
# valid label is present. A malformed value (non-integer, zero, negative) is
# ignored, not fatal — a bad label must not silently drop the existing cap.
parse_fix_budget() {
local labels="${1-${PR_LABELS:-}}"
local best="" label n
# Accept comma-joined labels (the upstream dispatcher format) as well as
# newline-joined: normalize commas to newlines before splitting.
labels="${labels//,/$'\n'}"
while IFS= read -r label; do
Comment on lines +28 to +34

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remediation recommended

5. Feature lacks linked authorization 📜 Skill insight § Compliance

This PR adds a new parser, runtime guard, generated bundle changes, and tests well beyond the rule's
20-line threshold, but the PR metadata contains no linked authorizing issue. The non-trivial feature
therefore lacks the required explicit authorization.
Agent Prompt
## Issue description
The non-trivial feature change has no linked issue authorizing its scope.

## Issue Context
Link an issue that explicitly authorizes the per-PR fix-budget feature and confirms the intended producer wiring and enforcement scope.

## Fix Focus Areas
- scripts/lib/fix-budget.lib.sh[1-42]
- scripts/pre-fix.src.sh[114-119]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

# Trim surrounding whitespace so " fullsend-fix-budget/3 " still matches.
label="${label#"${label%%[![:space:]]*}"}"
label="${label%"${label##*[![:space:]]}"}"
[[ "${label}" == "${FIX_BUDGET_LABEL_PREFIX}"* ]] || continue
n="${label#"${FIX_BUDGET_LABEL_PREFIX}"}"
# Bound the digit count. An arbitrarily long value would overflow Bash's
# signed 64-bit arithmetic in the `-lt` comparison (e.g. 2^64 evaluates as
# 0), which would look "tighter" than any cap and block every fix run.
# A budget above 99999 is meaningless next to caps of 5/10, so treat an
# over-long value as malformed and ignore it.
[[ "${n}" =~ ^[1-9][0-9]{0,4}$ ]] || continue
if [[ -z "${best}" || "${n}" -lt "${best}" ]]; then
best="${n}"
fi
done <<< "${labels}"
[[ -n "${best}" ]] && printf '%s\n' "${best}"
return 0
}
40 changes: 40 additions & 0 deletions scripts/lib/review-labels.lib.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
#!/usr/bin/env bash
# shellcheck shell=bash
# review-labels.lib.sh — recognize pipeline-managed "control" labels.
#
# Control labels are set by the review pipeline (or a maintainer, for the
# fix-budget knob), not by the review agent. post-review refuses to add or
# remove them via agent-recommended label_actions. Kept in a sourceable lib so
# the same definition is exercised by both production and the unit test — a
# duplicated copy in the test would pass even if the production branch drifted.
#
# Bundled into post-review.sh via bundle-sh.sh.

[[ -n "${REVIEW_LABELS_SH_LOADED:-}" ]] && return 0
REVIEW_LABELS_SH_LOADED=1

REVIEW_CONTROL_LABELS=(
"ready-for-merge" "requires-manual-review" "rejected"
"ready-for-review" "fullsend-no-fix" "fullsend-fix"
)

# is_control_label LABEL — return 0 if LABEL is pipeline-managed, 1 otherwise.
is_control_label() {
local label="$1"
local cl
for cl in "${REVIEW_CONTROL_LABELS[@]}"; do
if [[ "${cl}" == "${label}" ]]; then
return 0
fi
done
# Pipeline-managed label prefixes.
if [[ "${label}" == risk/* ]]; then
return 0
fi
# Maintainer-set fix-loop budget (fullsend-fix-budget/N); pipeline-managed so
# the review agent preserves it rather than treating it as a contextual label.
if [[ "${label}" == fullsend-fix-budget/* ]]; then
return 0
fi
return 1
}
118 changes: 118 additions & 0 deletions scripts/post-fix-test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -1022,6 +1022,124 @@ run_prefix_github_validation_test "prefix-github-dotdot" \

rm -rf "${PRE_TMPDIR}"

# ---------------------------------------------------------------------------
# Budget-aware needs-human mirror (post-fix section 6)
# ---------------------------------------------------------------------------
# On main the script sets NO_PUSH=true and still reaches the iteration-cap
# warning. Leave FULLSEND_VALIDATED_ITERATION_DIR unset so a missing
# agent-result.json is a warning, not a fail-closed exit.

# post-fix prepends $HOME/.local/bin to PATH (pre-commit deps). Point HOME
# at an empty temp dir so a developer/CI gh there cannot hide this mock.
BUDGET_TMPDIR="$(mktemp -d)"
BUDGET_MOCK="${BUDGET_TMPDIR}/bin"
mkdir -p "${BUDGET_MOCK}"
cat > "${BUDGET_MOCK}/gh" <<'MOCKEOF'
#!/usr/bin/env bash
echo "$@" >> "${GH_LABEL_LOG}"
exit 0
MOCKEOF
chmod +x "${BUDGET_MOCK}/gh"

run_postfix_budget_test() {
local test_name="$1"
local trigger="$2"
local iteration="$3"
local bot_cap="$4"
local human_cap="$5"
local labels="$6"
local expect_summary="$7"
local expect_needs_human="$8" # "yes" or "no"

local run_dir="${BUDGET_TMPDIR}/run-${test_name}"
local repo_dir="${run_dir}/repo"
mkdir -p "${repo_dir}"
git init -q -b main "${repo_dir}"
git -C "${repo_dir}" config user.email "test@example.com"
git -C "${repo_dir}" config user.name "Test"
git -C "${repo_dir}" commit --allow-empty -m "init" -q

local log="${BUDGET_TMPDIR}/gh-${test_name}.log"
: > "${log}"

local exit_code=0
(
cd "${run_dir}"
export PATH="${BUDGET_MOCK}:${PATH}"
export HOME="${BUDGET_TMPDIR}/home"
mkdir -p "${HOME}"
export GH_LABEL_LOG="${log}"
export PUSH_TOKEN="fake-token"
export REPO_FULL_NAME="test-org/test-repo"
export PR_NUMBER="99"
export TRIGGER_SOURCE="${trigger}"
export REPO_DIR="repo"
export FULLSEND_FORGE="github"
export FIX_ITERATION="${iteration}"
export ITERATION_CAP="${bot_cap}"
export ITERATION_CAP_HUMAN="${human_cap}"
export PR_LABELS="${labels}"
bash "${POST_SCRIPT}"
) > "${BUDGET_TMPDIR}/stdout-${test_name}.log" 2>&1 || exit_code=$?

if [[ ${exit_code} -ne 0 ]]; then
echo "FAIL: ${test_name} — exit ${exit_code}"
cat "${BUDGET_TMPDIR}/stdout-${test_name}.log"
FAILURES=$((FAILURES + 1))
return
fi
if ! grep -qF "${expect_summary}" "${BUDGET_TMPDIR}/stdout-${test_name}.log"; then
echo "FAIL: ${test_name} — missing summary '${expect_summary}'"
cat "${BUDGET_TMPDIR}/stdout-${test_name}.log"
FAILURES=$((FAILURES + 1))
return
fi
if [[ "${expect_needs_human}" == "yes" ]]; then
if ! grep -q -- '--add-label needs-human' "${log}"; then
echo "FAIL: ${test_name} — expected needs-human label"
echo "gh log:"; cat "${log}"
FAILURES=$((FAILURES + 1))
return
fi
else
if grep -q -- '--add-label needs-human' "${log}"; then
echo "FAIL: ${test_name} — unexpected needs-human label"
echo "gh log:"; cat "${log}"
FAILURES=$((FAILURES + 1))
return
fi
fi
echo "PASS: ${test_name}"
}

BOT="fixbot[bot]"

run_postfix_budget_test "budget-tightens-needs-human-and-summary" \
"${BOT}" 2 5 10 $'fullsend-fix-budget/2' \
"Iteration: 2 of 2 (bot cap)" "yes"

run_postfix_budget_test "budget-above-global-cap-has-no-effect" \
"${BOT}" 3 5 10 $'fullsend-fix-budget/9' \
"Iteration: 3 of 5 (bot cap)" "no"

run_postfix_budget_test "budget-1-escalates-on-iteration-1" \
"${BOT}" 1 5 10 $'fullsend-fix-budget/1' \
"Iteration: 1 of 1 (bot cap)" "yes"

run_postfix_budget_test "human-trigger-keeps-full-human-cap" \
"alice" 3 5 10 $'fullsend-fix-budget/2' \
"Iteration: 3 of 10 (human cap, total across bot+human)" "no"

run_postfix_budget_test "malformed-label-leaves-global-cap" \
"${BOT}" 3 5 10 $'fullsend-fix-budget/abc' \
"Iteration: 3 of 5 (bot cap)" "no"

run_postfix_budget_test "absent-label-leaves-global-cap" \
"${BOT}" 3 5 10 "" \
"Iteration: 3 of 5 (bot cap)" "no"

rm -rf "${BUDGET_TMPDIR}"

# --- Summary ---

echo ""
Expand Down
69 changes: 68 additions & 1 deletion scripts/post-fix.sh
Original file line number Diff line number Diff line change
Expand Up @@ -1206,6 +1206,60 @@ classify_branch_vs_pr_head() {
fi
}
# END bundled: lib/branch-guard.lib.sh
# shellcheck source=lib/fix-budget.lib.sh
# BEGIN bundled: lib/fix-budget.lib.sh
# shellcheck shell=bash
# fix-budget.lib.sh — parse a per-PR fix-loop budget from PR labels.
#
# A label of the form `fullsend-fix-budget/N` (N a positive integer) lets a
# maintainer cap the review->fix loop for a single PR below the global
# iteration cap. The label can only TIGHTEN the cap, never raise it:
# enforcement lives in pre-fix, which applies min(label_budget, cap).
#
# Bundled into pre-fix.sh via bundle-sh.sh.
#
# Expected env vars (optional):
# PR_LABELS — PR label names separated by commas and/or newlines. Absent/empty
# is fine: parse_fix_budget then returns nothing and the cap is
# unchanged. (The upstream dispatcher comma-joins labels; a
# newline-joined value is also accepted.)

[[ -n "${FIX_BUDGET_SH_LOADED:-}" ]] && return 0
FIX_BUDGET_SH_LOADED=1

FIX_BUDGET_LABEL_PREFIX="fullsend-fix-budget/"

# parse_fix_budget [labels]
# Reads label names (arg 1, or PR_LABELS env when omitted) separated by commas
# and/or newlines. Echoes the smallest valid budget found, or nothing when no
# valid label is present. A malformed value (non-integer, zero, negative) is
# ignored, not fatal — a bad label must not silently drop the existing cap.
parse_fix_budget() {
local labels="${1-${PR_LABELS:-}}"
local best="" label n
# Accept comma-joined labels (the upstream dispatcher format) as well as
# newline-joined: normalize commas to newlines before splitting.
labels="${labels//,/$'\n'}"
while IFS= read -r label; do
# Trim surrounding whitespace so " fullsend-fix-budget/3 " still matches.
label="${label#"${label%%[![:space:]]*}"}"
label="${label%"${label##*[![:space:]]}"}"
[[ "${label}" == "${FIX_BUDGET_LABEL_PREFIX}"* ]] || continue
n="${label#"${FIX_BUDGET_LABEL_PREFIX}"}"
# Bound the digit count. An arbitrarily long value would overflow Bash's
# signed 64-bit arithmetic in the `-lt` comparison (e.g. 2^64 evaluates as
# 0), which would look "tighter" than any cap and block every fix run.
# A budget above 99999 is meaningless next to caps of 5/10, so treat an
# over-long value as malformed and ignore it.
[[ "${n}" =~ ^[1-9][0-9]{0,4}$ ]] || continue
if [[ -z "${best}" || "${n}" -lt "${best}" ]]; then
best="${n}"
fi
done <<< "${labels}"
[[ -n "${best}" ]] && printf '%s\n' "${best}"
return 0
}
# END bundled: lib/fix-budget.lib.sh


# ---------------------------------------------------------------------------
Expand Down Expand Up @@ -1544,6 +1598,16 @@ fi
# ---------------------------------------------------------------------------
ITERATION="${FIX_ITERATION:-1}"
BOT_CAP="${ITERATION_CAP:-5}"

# A per-PR `fullsend-fix-budget/N` label may tighten the BOT cap only (never
# raise it, never touch the human cap) — matching pre-fix. Mirror it here so the
# needs-human warning and the iteration summary reflect the cap pre-fix actually
# enforces. Without this, a budget of 2 under a global cap of 5 would report
# "2 of 5" and never add needs-human, even though pre-fix rejects the next cycle.
FIX_BUDGET="$(parse_fix_budget "${PR_LABELS:-}")"
if [[ -n "${FIX_BUDGET}" && "${FIX_BUDGET}" -lt "${BOT_CAP}" ]]; then
BOT_CAP="${FIX_BUDGET}"
fi
WARN_THRESHOLD=$(( BOT_CAP - 1 ))

# The needs-human label is based on the bot cap — it signals that the
Expand All @@ -1568,5 +1632,8 @@ echo " Trigger: ${TRIGGER_SOURCE}"
if is_bot_user "${TRIGGER_SOURCE}"; then
echo " Iteration: ${ITERATION} of ${BOT_CAP} (bot cap)"
else
echo " Iteration: ${ITERATION} of ${ITERATION_CAP_HUMAN:-10} (human cap, total across bot+human)"
# The fix-budget label never tightens the human cap, so the human escape hatch
# always reports its full ITERATION_CAP_HUMAN budget.
HUMAN_CAP="${ITERATION_CAP_HUMAN:-10}"
echo " Iteration: ${ITERATION} of ${HUMAN_CAP} (human cap, total across bot+human)"
fi
Loading
Loading