You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Needs improvement: Label Closed PRs (0% success, 85% AR), Content Moderation (11% success, 89% AR), PR Description Updater (16% success, 84% AR), AI Moderator (2% success, 64% AR), Q (0% success, 58% AR), Agentic Commands (36% success, 64% AR — but 112 runs, largest volume)
Note on data limitations:metrics/latest.json in shared memory is stale (Jan 2026, "no GitHub API access during collection" per Metrics Collector — tracked in persistent issue #43292). This report uses fresh gh api repos/.../actions/runs data (last ~3h window) instead, since gh run list intermittently errored ("error connecting to api.github.com") while gh api succeeded reliably.
Performance Rankings (this window)
Top Performing Agents 🏆
Agent
Runs
Success
AR
Notes
Auto-Close Parent Issues
16
100%
0%
Fully stable
PR Data Prefetch
11
100%
0%
Fully stable
Running Copilot Code Review
22
100%
0%
Fully stable
CJS
23
74%
13%
Healthy, minor noise (2 cancelled, 1 failure)
Matt Pocock Skills Reviewer
11
64%
0%
CLI-hang fix appears to have landed for this workflow — 0% AR now vs. historical 100% AR (see #49577 caveat below)
Impeccable Skills Reviewer
11
73%
0%
CLI-hang fix (#44254/#48499) holding in this window
Persistent AR — tracked in closed #47925/#48894 cluster, appears to have regressed again
Content Moderation
9
11%
89%
89% AR this window — was reported "recovered" in prior run (Jul 8 memory said 5/5 success); regression since
PR Description Updater
19
16%
84%
Persistent — open issue #49747 "reported incomplete result"
AI Moderator
52
2%
64%
Highest run volume of the underperformers; recurring — closed issues #49440, #49190, #48970 (rate-limit/pagination fix) suggest still not fully resolved
Q
43
0%
58%
58% AR + 42% skipped; no open blocking issue found this window (prior blocker PR #43527 status unconfirmed)
Agentic Commands
112
36%
64%
Largest run volume in repo (112 runs/~3h) — 64% AR is a major aggregate drag on ecosystem-wide success rate
CGO
34
41%
26%
Mixed — 6 failures, 4 cancelled
CI (plain CI, not agentic)
18
6% success, 61% failure/cancelled
0%
Not agentic — do not treat as AR; likely genuine CI regression, separate from agent quality
CI Approval-Pending (not agent failures)
No action_required runs observed this window for CI/CGO/CWI/CJS/CPI that map to the "PR approval gate" pattern specifically — CI shows failure/cancelled (18 runs), which is a genuine build issue, not an approval gate. CGO/CWI combine action_required (approval-gate, non-agentic) with real failure — flagging for maintainers to distinguish, not filing as agent defect.
Cross-checked against existing tracking (no re-filing)
Q workflow / Agentic Commands — both persistently tracked ([aw] Impeccable Skills Reviewer failed #43079 for Agentic Commands per historical memory). Q's AR rate (58%) is lower than the ~79% reported Jul 8, suggesting partial improvement, but still a leading contributor to ecosystem-wide low success rate.
Volume concentration risk: Agentic Commands (112 runs) and AI Moderator (52 runs) together account for over 27% of all runs in this window, and both have failure/AR rates above 60% — meaning a large share of total ecosystem compute is currently being spent on runs that don't complete successfully.
Re-open regression tracking for Content Moderation / Label Closed PRs / PR Description Updater — these were marked recovered/closed but show 84–89% AR now. Confirm whether the underlying CLI-hang-on-exit fix regressed or a new trigger reintroduced it.
Fix stale Metrics Collector ([aw] Metrics Collector failed #43292) — shared metrics memory (metrics/latest.json) is 7+ months stale, forcing this analysis to fall back to live gh api queries every run. Restoring the collector would improve trend analysis (week-over-week deltas) for future performance reports.
Analysis window: 2026-08-02T09:54Z – 2026-08-02T13:03Z (~3h, 600 runs sampled)
Data source: live gh api actions/runs queries (shared metrics/latest.json is stale — see #43292)
Next report: next scheduled run
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Agent Performance Report — 2026-08-02T13:06Z
Executive Summary
Performance Rankings (this window)
Top Performing Agents 🏆
Agents Needing Improvement 📉
CI Approval-Pending (not agent failures)
No
action_requiredruns observed this window forCI/CGO/CWI/CJS/CPIthat map to the "PR approval gate" pattern specifically —CIshowsfailure/cancelled(18 runs), which is a genuine build issue, not an approval gate.CGO/CWIcombineaction_required(approval-gate, non-agentic) with realfailure— flagging for maintainers to distinguish, not filing as agent defect.Cross-checked against existing tracking (no re-filing)
Behavioral Patterns
Recommendations
High Priority
Medium Priority
metrics/latest.json) is 7+ months stale, forcing this analysis to fall back to livegh apiqueries every run. Restoring the collector would improve trend analysis (week-over-week deltas) for future performance reports.Next Steps
All reactions