mirror of
https://github.com/NousResearch/hermes-agent.git
synced 2026-07-29 18:46:59 +00:00
Replace the static comment-pending + comment-results two-job pattern
with a live-updating comment system that polls the GitHub Actions API
every 15s, re-assembles the review comment from whatever results are
available, and upserts it via the <!-- hermes-ci-review-bot --> marker.
The comment updates in real time as each job finishes — no waiting for
the full pipeline.
Every CI job that wants to appear in the review comment emits a
review_status output — a JSON array of objects, each with a source
and a results array:
[
{
"source": "review-label-gate",
"results": [
{"kind": "action_required", "title": "...", "summary": "...",
"how_to_fix": "..."},
{"kind": "info", "title": "...", "summary": "..."}
]
},
{
"source": "ci timing",
"results": [
{"kind": "warning", "title": "CI timings", "summary": "...",
"detail": "...", "link": "..."}
]
}
]
One job can emit multiple results of different kinds. The source field
is used to exclude the corresponding job from the synthesized error
list (case-insensitive, hyphen-normalized matching against GitHub
Actions job display names).
| job | source | kind (on failure) | section |
|----------------------------|--------------------------|---------------------------|----------------------|
| review-labels | review label gate | action_required / info | Action required |
| lockfile-diff | lockfile-diff | action_required | Action required |
| ci-timings | ci timing | warning / info | Warnings |
| supply-chain scan | supply chain | error / (none) | Job failures |
| supply-chain dep-bounds | supply chain | action_required / (none) | Action required |
| osv-scanner | osv scan | warning / (none) | Warnings |
| uv-lockfile-check | uv.lock check | action_required / (none) | Action required |
| history-check | unrelated histories | action_required | Action required |
| contributor-check | contributor attribution | action_required | Action required |
Jobs that find nothing emit [] (empty array) — no noise info items.
A single comment-live job polls the GitHub Actions API every 15s,
classifies jobs into (completed, pending), assembles the comment, and
upserts it. Merges review_status outputs from all needs jobs via
toJSON(needs.*.outputs.review_status), and downloads the ci-timings
artifact when it becomes available. Shows commit SHA + message below
the header.
The assembler has ZERO job-specific knowledge. It just:
1. collect_from_statuses() — flattens all nested status objects into ReviewItems
2. collect_failed_jobs() — synthesizes errors for failed jobs with no declared status
3. _attach_job_urls() — fills in per-job log links for ALL items
4. render_comment() — groups by severity, renders with group headers
Each item shows links inline next to the title: View report (job-emitted
URL) and View job (auto-attached logs link). Each info item is its own
collapsible <details> block.
# ૮ >ﻌ< ა ci review
running on abc1234 — commit message first line
## ❌ Job failures
### {title} · [View job](url)
{summary}
## ⚠️ Action required
### {title} · [View job](url)
{summary}
**How to fix:**
{how_to_fix}
## ⚠️ Warnings
### {title} · [View report](url) · [View job](url)
{summary}
{detail}
<details><summary>{title}</summary>
{content}
</details>
Still running 3 jobs: ci-timings, docker
- test_assemble_review_comment.py (48 tests): collect_from_statuses,
collect_failed_jobs with exclude_sources, _attach_job_urls,
render_comment (group headers, inline links, commit info, per-item
details, pending footer), assemble integration
- test_live_comment.py (16 tests): classify_jobs pure function
- test_timings_report.py (10 tests): generate_review_status nested format
- test_lockfile_diff.py (6 tests)
- test_classify_changes.py (32 tests, pre-existing)
382 lines
16 KiB
YAML
382 lines
16 KiB
YAML
name: CI
|
|
|
|
# Orchestrator workflow. Runs ``detect-changes`` once, then conditionally
|
|
# calls the sub-workflows that a PR can actually affect. A final
|
|
# ``all-checks-pass`` gate job aggregates results so branch protection only
|
|
# needs to require a single check.
|
|
#
|
|
# Sub-workflows are triggered via ``workflow_call`` and keep their own job
|
|
# definitions, matrices, and concurrency settings. They no longer have
|
|
# ``push:`` / ``pull_request:`` triggers of their own — everything flows
|
|
# through this file.
|
|
|
|
on:
|
|
pull_request:
|
|
push:
|
|
branches: [main]
|
|
|
|
permissions:
|
|
contents: read
|
|
pull-requests: write # needed by lint (PR comment) + supply-chain review_status
|
|
actions: read # needed by osv-scanner (SARIF upload)
|
|
security-events: write # needed by osv-scanner (SARIF upload)
|
|
packages: write # needed by docker build
|
|
|
|
concurrency:
|
|
group: ci-${{ github.ref }}
|
|
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
|
|
|
|
jobs:
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
# detect: run the classifier once. Every downstream job reads its outputs
|
|
# to decide whether to run. On push/dispatch the classifier fails open
|
|
# (all lanes true) so post-merge validation is never weakened.
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
detect:
|
|
name: Detect affected areas
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
outputs:
|
|
python: ${{ steps.classify.outputs.python }}
|
|
frontend: ${{ steps.classify.outputs.frontend }}
|
|
site: ${{ steps.classify.outputs.site }}
|
|
scan: ${{ steps.classify.outputs.scan }}
|
|
deps: ${{ steps.classify.outputs.deps }}
|
|
npm_lock: ${{ steps.classify.outputs.npm_lock }}
|
|
docker_meta: ${{ steps.classify.outputs.docker_meta }}
|
|
mcp_catalog: ${{ steps.classify.outputs.mcp_catalog }}
|
|
ci_review: ${{ steps.classify.outputs.ci_review }}
|
|
event_name: ${{ github.event_name }}
|
|
steps:
|
|
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
- name: Detect affected areas
|
|
id: classify
|
|
uses: ./.github/actions/detect-changes
|
|
with:
|
|
# Forks get no repo secrets (AUTOFIX_BOT_PAT is empty); fall back to
|
|
# the built-in read-only token so classification still works there.
|
|
github-token: ${{ secrets.AUTOFIX_BOT_PAT || github.token }}
|
|
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
# Lane-gated sub-workflows. Each runs in parallel after detect finishes.
|
|
# Skipped workflows (if condition is false) don't spin up runners.
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
tests:
|
|
name: Python tests
|
|
needs: detect
|
|
if: needs.detect.outputs.python == 'true'
|
|
uses: ./.github/workflows/tests.yml
|
|
with:
|
|
slice_count: 8
|
|
secrets: inherit
|
|
|
|
lint:
|
|
name: Python lints
|
|
needs: detect
|
|
if: needs.detect.outputs.python == 'true'
|
|
uses: ./.github/workflows/lint.yml
|
|
with:
|
|
event_name: ${{ needs.detect.outputs.event_name }}
|
|
secrets: inherit
|
|
|
|
js-tests:
|
|
name: JS & TS checks
|
|
needs: detect
|
|
if: needs.detect.outputs.frontend == 'true'
|
|
uses: ./.github/workflows/js-tests.yml
|
|
secrets: inherit
|
|
|
|
e2e-desktop:
|
|
name: Desktop E2E
|
|
needs: detect
|
|
if: needs.detect.outputs.python == 'true' || needs.detect.outputs.frontend == 'true'
|
|
uses: ./.github/workflows/e2e-desktop.yml
|
|
|
|
docs-site:
|
|
name: Docs Site
|
|
needs: detect
|
|
if: needs.detect.outputs.site == 'true'
|
|
uses: ./.github/workflows/docs-site-checks.yml
|
|
secrets: inherit
|
|
|
|
history-check:
|
|
name: Deny unrelated histories
|
|
needs: detect
|
|
if: needs.detect.outputs.event_name == 'pull_request'
|
|
uses: ./.github/workflows/history-check.yml
|
|
secrets: inherit
|
|
|
|
contributor-check:
|
|
name: Check contributors
|
|
needs: detect
|
|
if: needs.detect.outputs.python == 'true'
|
|
uses: ./.github/workflows/contributor-check.yml
|
|
secrets: inherit
|
|
|
|
uv-lockfile:
|
|
name: Check uv.lock
|
|
needs: detect
|
|
uses: ./.github/workflows/uv-lockfile-check.yml
|
|
secrets: inherit
|
|
|
|
lockfile-diff:
|
|
name: package-lock.json diff
|
|
needs: detect
|
|
if: needs.detect.outputs.event_name == 'pull_request' && needs.detect.outputs.npm_lock == 'true'
|
|
uses: ./.github/workflows/lockfile-diff.yml
|
|
secrets: inherit
|
|
|
|
docker-lint:
|
|
name: Lint Docker scripts
|
|
needs: detect
|
|
if: needs.detect.outputs.docker_meta == 'true'
|
|
uses: ./.github/workflows/docker-lint.yml
|
|
secrets: inherit
|
|
|
|
docker:
|
|
name: Build&Test Docker image
|
|
needs: detect
|
|
if: needs.detect.outputs.python == 'true' || needs.detect.outputs.frontend == 'true' || needs.detect.outputs.docker_meta == 'true'
|
|
uses: ./.github/workflows/docker.yml
|
|
secrets: inherit
|
|
|
|
supply-chain:
|
|
name: Supply-chain scan
|
|
needs: detect
|
|
if: needs.detect.outputs.event_name == 'pull_request' && (needs.detect.outputs.scan == 'true' || needs.detect.outputs.deps == 'true')
|
|
uses: ./.github/workflows/supply-chain-audit.yml
|
|
with:
|
|
event_name: ${{ needs.detect.outputs.event_name }}
|
|
scan: ${{ needs.detect.outputs.scan == 'true' }}
|
|
deps: ${{ needs.detect.outputs.deps == 'true' }}
|
|
|
|
review-labels:
|
|
name: Review label gate
|
|
needs: detect
|
|
if: needs.detect.outputs.event_name == 'pull_request' && (needs.detect.outputs.ci_review == 'true' || needs.detect.outputs.mcp_catalog == 'true')
|
|
uses: ./.github/workflows/review-labels.yml
|
|
with:
|
|
ci_review: ${{ needs.detect.outputs.ci_review == 'true' }}
|
|
mcp_catalog: ${{ needs.detect.outputs.mcp_catalog == 'true' }}
|
|
secrets: inherit
|
|
|
|
osv-scanner:
|
|
name: OSV scan
|
|
uses: ./.github/workflows/osv-scanner.yml
|
|
secrets: inherit
|
|
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
# Live-updating PR review comment.
|
|
#
|
|
# A single ``comment-live`` job polls the GitHub Actions API every 15s
|
|
# for job statuses in this run, re-assembles the review comment from
|
|
# whatever results are available, and upserts it via the
|
|
# ``<!-- hermes-ci-review-bot -->`` marker.
|
|
#
|
|
# The poller exits when all non-infra jobs are completed (or on
|
|
# timeout). ci-timings' review_status is picked up automatically when
|
|
# its artifact becomes available — the poller downloads and merges it.
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
comment-live:
|
|
name: CI review comment (live)
|
|
needs: [detect, review-labels, lockfile-diff, supply-chain, osv-scanner, uv-lockfile, history-check, contributor-check]
|
|
if: always() && github.event_name == 'pull_request' && github.event.pull_request.head.repo.fork != true
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 40
|
|
steps:
|
|
- name: Checkout code
|
|
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
|
|
- name: Run live comment poller
|
|
env:
|
|
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
|
GITHUB_REPOSITORY: ${{ github.repository }}
|
|
GITHUB_RUN_ID: ${{ github.run_id }}
|
|
PR_NUMBER: ${{ github.event.pull_request.number }}
|
|
RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}
|
|
# Commit info for the review comment header.
|
|
COMMIT_SHA: ${{ github.event.pull_request.head.sha }}
|
|
COMMIT_MESSAGE: ${{ github.event.pull_request.head.commit.message }}
|
|
COMMIT_URL: ${{ github.server_url }}/${{ github.repository }}/pull/${{ github.event.pull_request.number }}/commits/${{ github.event.pull_request.head.sha }}
|
|
# Structured review statuses from workflow_call jobs.
|
|
# Each job outputs a JSON array of {source, results: [...]} objects
|
|
# that the assembler renders directly — no hardcoded job-name
|
|
# matching. We merge all available outputs into one array.
|
|
REVIEW_STATUSES: ${{ toJSON(needs.*.outputs.review_status) }}
|
|
run: |
|
|
set -uo pipefail
|
|
|
|
# REVIEW_STATUSES is a JSON array of strings (some may be empty
|
|
# when a job was skipped). Parse each string and merge into one
|
|
# flat array for the assembler.
|
|
python3 - <<'PYEOF'
|
|
import json, os, sys
|
|
|
|
raw = os.environ.get("REVIEW_STATUSES", "")
|
|
merged = []
|
|
if raw:
|
|
try:
|
|
arr = json.loads(raw)
|
|
except (json.JSONDecodeError, TypeError):
|
|
arr = []
|
|
for item in arr:
|
|
if not item:
|
|
continue
|
|
try:
|
|
statuses = json.loads(item)
|
|
except (json.JSONDecodeError, TypeError):
|
|
continue
|
|
if isinstance(statuses, list):
|
|
merged.extend(statuses)
|
|
|
|
# Write merged array to a temp file the poller reads.
|
|
with open("/tmp/review_statuses.json", "w") as f:
|
|
json.dump(merged, f)
|
|
print(f"Merged {len(merged)} review status entries")
|
|
PYEOF
|
|
|
|
python3 scripts/ci/live_comment.py \
|
|
--interval 15 \
|
|
--timeout 2100 \
|
|
--review-statuses-file /tmp/review_statuses.json
|
|
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
# Gate: runs after everything. ``if: always()`` ensures it reports a
|
|
# status even when some deps were skipped. Only actual ``failure``
|
|
# results cause it to fail; ``skipped`` is treated as success.
|
|
#
|
|
# Branch protection should require ONLY this check.
|
|
#
|
|
# Outputs ``needs-json`` — a compact ``{job_name: result}`` dict — so
|
|
# the live comment poller can list failed jobs in the PR comment.
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
all-checks-pass:
|
|
name: All required checks pass
|
|
needs:
|
|
- tests
|
|
- lint
|
|
- js-tests
|
|
- e2e-desktop
|
|
- docs-site
|
|
- history-check
|
|
- contributor-check
|
|
- uv-lockfile
|
|
- lockfile-diff
|
|
- docker-lint
|
|
- supply-chain
|
|
- review-labels
|
|
- osv-scanner
|
|
# comment-live is a polling job — it doesn't block the gate.
|
|
# we don't require docker to pass rn because it's so slow lol
|
|
# - docker
|
|
if: always()
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
outputs:
|
|
needs-json: ${{ steps.evaluate.outputs.needs-json }}
|
|
steps:
|
|
- name: Evaluate job results
|
|
id: evaluate
|
|
env:
|
|
NEEDS: ${{ toJSON(needs) }}
|
|
run: |
|
|
echo "$NEEDS" | python3 -c "
|
|
import json, sys
|
|
needs = json.load(sys.stdin)
|
|
# Emit compact {job_name: result} for the comment assembler.
|
|
compact = {name: info['result'] for name, info in needs.items()}
|
|
print(f'needs-json={json.dumps(compact)}')
|
|
with open('$GITHUB_OUTPUT', 'a') as f:
|
|
f.write(f'needs-json={json.dumps(compact)}\n')
|
|
failed = [name for name, info in needs.items() if info['result'] == 'failure']
|
|
for name, info in sorted(needs.items()):
|
|
result = info['result']
|
|
icon = '✅' if result in ('success', 'skipped') else '❌'
|
|
print(f'{icon} {name}: {result}')
|
|
if failed:
|
|
print(f'::error::{len(failed)} job(s) failed: {\", \".join(failed)}')
|
|
sys.exit(1)
|
|
print('All checks passed (or were skipped)')
|
|
"
|
|
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
# CI timing report: collect per-job/step durations from the GitHub API,
|
|
# cache them on main (as a baseline), and on PRs generate an HTML diff
|
|
# report with a gantt chart + per-step breakdown. The report is uploaded
|
|
# as an artifact and a markdown summary is written to $GITHUB_STEP_SUMMARY.
|
|
#
|
|
# The live comment poller picks up ci-timings' completion automatically —
|
|
# it reads review-status.json from the artifact when the job finishes.
|
|
# ─────────────────────────────────────────────────────────────────────
|
|
ci-timings:
|
|
name: CI timing report
|
|
needs: [all-checks-pass, docker]
|
|
if: always()
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 10
|
|
steps:
|
|
- name: Checkout code
|
|
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
|
|
- name: Restore baseline cache (PR only)
|
|
if: github.event_name == 'pull_request'
|
|
uses: actions/cache/restore@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5.0.5
|
|
with:
|
|
path: ci-timings-baseline.json
|
|
# Prefix-match: exact key will never hit (run_id differs), so
|
|
# restore-keys finds the most recent baseline from main.
|
|
key: ci-timings-baseline-never-exact
|
|
restore-keys: |
|
|
ci-timings-baseline-
|
|
|
|
- name: Collect timings and generate report
|
|
env:
|
|
# Forks get no repo secrets (AUTOFIX_BOT_PAT is empty); fall back to
|
|
# the built-in read-only token so the timings API read still works
|
|
# there instead of hard-failing this advisory job on every fork PR.
|
|
GITHUB_TOKEN: ${{ secrets.AUTOFIX_BOT_PAT || github.token }}
|
|
run: |
|
|
python3 scripts/ci/timings_report.py \
|
|
--baseline ci-timings-baseline.json \
|
|
--output ci-timings-report.html \
|
|
--json-out ci-timings.json \
|
|
--summary-out ci-timings-summary.md \
|
|
--review-status-out review-status.json
|
|
|
|
- name: Upload HTML report + review status
|
|
# Advisory report — artifact-service blips must not fail the job.
|
|
continue-on-error: true
|
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
|
|
id: ci-timings-artifact
|
|
with:
|
|
name: ci-timings-report
|
|
path: |
|
|
ci-timings-report.html
|
|
review-status.json
|
|
retention-days: 14
|
|
|
|
- name: Output summary
|
|
env:
|
|
REPORT_URL: ${{ steps.ci-timings-artifact.outputs.artifact-url}}
|
|
run: |
|
|
echo "# CI Timing report" >> "$GITHUB_STEP_SUMMARY"
|
|
echo "[View the full interactive report]($REPORT_URL)" >> "$GITHUB_STEP_SUMMARY"
|
|
cat ci-timings-summary.md >> "$GITHUB_STEP_SUMMARY"
|
|
|
|
- name: Save baseline cache (main only)
|
|
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
|
|
run: |
|
|
# Degraded runs (API rate-limited) produce no ci-timings.json —
|
|
# skip rather than fail, and never cache an empty baseline.
|
|
if [ -f ci-timings.json ]; then
|
|
cp ci-timings.json ci-timings-baseline.json
|
|
else
|
|
echo "No timings JSON this run — skipping baseline update"
|
|
fi
|
|
|
|
- name: Upload baseline to cache (main only)
|
|
if: github.event_name == 'push' && github.ref == 'refs/heads/main' && hashFiles('ci-timings-baseline.json') != ''
|
|
uses: actions/cache/save@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5.0.5
|
|
with:
|
|
path: ci-timings-baseline.json
|
|
key: ci-timings-baseline-${{ github.run_id }}
|