feat(models): upgrade Kimi defaults to K3 - #9348
Conversation
|
OpenAPI Spec Update The OpenAPI specification has changed. Please review the generated spec in the workflow artifacts. |
|
OpenAPI Spec Update The OpenAPI specification has changed. Please review the generated spec in the workflow artifacts. |
Aragora Code ReviewAdvisory-only review. No issues found. |
Co-authored-by: codex[bot] <codex[bot]@users.noreply.github.com>
|
OpenAPI Spec Update The OpenAPI specification has changed. Please review the generated spec in the workflow artifacts. |
OpenAI independent model reviewReviewer: openai (openai) — independent adversarial model review via Codex CLI OpenAI harness, grounded on the exact PR head. Verdict: PASS
dogfood: yes |
Co-authored-by: codex[bot] <codex[bot]@users.noreply.github.com>
|
OpenAPI Spec Update The OpenAPI specification has changed. Please review the generated spec in the workflow artifacts. |
…nce record Quorum round-2 findings at 416fd85: - P2 (claude): drop the kimi K3 lane flip — revert to moonshotai/kimi-k2.6. The reliability record itself defers K3 until live verification + a catalog entry exist (catalog work #9348 is unmerged). Test pin updated; doc notes the deferral explicitly. - P2 (openai): the Tier-4 evidence artifact was gitignored. Committed verbatim as docs/governance/records/20260716T2200Z-gemini-reviewer- reliability-record.md with a provenance header; all citations (quorum_evidence.py, REVIEW_AUTHORITY_PRINCIPLES.md, governance test) repointed to the committed path, plus a test pinning its presence. - P3 (openai): removed Google from the stale WESTERN_FAMILIES description comment. - P3 (claude): doc now states Chinese-routed lanes fully count at Tier 0-2 per the tier table (model-pin changes do not alter counting authority), and explicitly frames the reliability record + this PR's Tier-4 settlement as an operator-approved substitute for the docs/specs/ design-doc requirement, scoped to this demotion only. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…-20260716 # Conflicts: # README.md
|
OpenAPI Spec Update The OpenAPI specification has changed. Please review the generated spec in the workflow artifacts. |
|
OpenAPI Spec Update The OpenAPI specification has changed. Please review the generated spec in the workflow artifacts. |
Claude independent model reviewReviewer: claude (anthropic) — independent adversarial model review via the Aragora Claude reviewer, grounded on the exact PR head. Verdict: PASS Reviewed the complete 42-file diff. This is a coordinated, well-tested
Findings:
dogfood: yes |
OpenAI independent model reviewReviewer: openai (openai) — independent adversarial model review via Codex CLI OpenAI harness, grounded on the exact PR head. Verdict: CHANGES-REQUESTED
dogfood: yes |
Exact-head soak and catalog dispositionPR: #9348 Do not collect or apply new evidence on this head. The current OpenAI P2 dissent is substantive, and the prior Tier-4 settlement authorization at Live classification:
Optimal dispositionPark this PR until the K3 soak boundary rather than adding another ad hoc price table or settling current-head dissent. On or after
If K3's live ID, price, context, or availability changes before the boundary, keep the PR parked and update the catalog evidence rather than preserving stale values. No branch, CI, evidence, settlement, merge, label, or protection state was changed by this packet. |
…y record [Tier 4] (#9363) * fix(swarm): remove gemini from the counting quorum set; kimi lane to K3 2026-07-16 founder directive following a repeat fabricated-claim pattern in gemini merge-quorum reviews (invented model release dates asserted twice after refutation, false METRICS-drift claims, demands for nonexistent route ids — full record in operator-context 20260716T2200Z). Gemini reviews still post and remain readable as advisory evidence; they no longer count toward Tier 3-4 quorums or satisfy the Tier-2 Western condition. The kimi advisory lane upgrades moonshotai/kimi-k2.6 -> kimi-k3 (live-verified 2026-07-16: $3/$15 per MTok, 1M context). Counting set retains claude/openai/grok/mistral/hermes. Promoting kimi (moonshot) into the counting set would amend the jurisdiction principles doc and is left as an explicit founder decision, not folded in here. TIER 4: quorum eligibility surface — requires founder exact-head settlement. 259+3 quorum-evidence tests green. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(governance): align roster pin test + authority principles doc with gemini demotion Quorum round-1 findings at 7659c55: - P1: test_western_families_match_spec now pins the post-directive roster (gemini removed); no other tests assert the old counting set (recognizer/ retrigger gemini references are recognition-layer, intentionally kept). - P2: REVIEW_AUTHORITY_PRINCIPLES.md drops Google from the Western counting list, documents the gemini advisory-only demotion citing the 2026-07-16 reviewer-reliability record, and notes the demotion is itself a Tier-4 change made with operator preapproval and settled through the Tier-4 chain. - P3 (note): documented the ARAGORA_OPENROUTER_REVIEWER_MODELS JSON override for advisory-lane model pins (kimi-k3) and its contained failure mode. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(governance): narrow PR to gemini demotion; commit auditable evidence record Quorum round-2 findings at 416fd85: - P2 (claude): drop the kimi K3 lane flip — revert to moonshotai/kimi-k2.6. The reliability record itself defers K3 until live verification + a catalog entry exist (catalog work #9348 is unmerged). Test pin updated; doc notes the deferral explicitly. - P2 (openai): the Tier-4 evidence artifact was gitignored. Committed verbatim as docs/governance/records/20260716T2200Z-gemini-reviewer- reliability-record.md with a provenance header; all citations (quorum_evidence.py, REVIEW_AUTHORITY_PRINCIPLES.md, governance test) repointed to the committed path, plus a test pinning its presence. - P3 (openai): removed Google from the stale WESTERN_FAMILIES description comment. - P3 (claude): doc now states Chinese-routed lanes fully count at Tier 0-2 per the tier table (model-pin changes do not alter counting authority), and explicitly frames the reliability record + this PR's Tier-4 settlement as an operator-approved substitute for the docs/specs/ design-doc requirement, scoped to this demotion only. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(swarm): advisory-only is a first-class family classification — gemini never counts for or against, at any tier Quorum round-3 findings at 975bda0 (both reviewers, same P2): dropping gemini from WESTERN_FAMILIES only stopped its supportive signal at Western-keyed tiers; it still counted at Tier 0-1 and its CHANGES-REQUESTED still promoted blocking dissent. The record's mandate is 'gemini dissent is NOT to be counted anywhere'. - ADVISORY_ONLY_FAMILIES = {gemini} (cites the committed record), plus CHINESE_ROUTED_FAMILIES so the classification is explicit and TOTAL: every FAMILY_PROVIDERS key belongs to exactly one of western / chinese-routed / advisory-only (partition pinned by governance tests). - TierQuorumRule.counted_families drops advisory-only families at EVERY tier — covers all three rule-derived surfaces (auto-settle, review-queue signal_count, reconcile diagnostic) even for raw reviewer-id lists. - EvidenceItem: advisory-only would_count demoted at the shared __post_init__ choke point (prepared artifacts cannot smuggle it back); .dissenting is False for advisory-only families. - review_queue: _dissenting_views_from_comments never promotes advisory-only blocking dissent; _get_validated_review_classification skips advisory-only reviews before the blocking scan and heard/dissent accounting (an advisory-only P0/P1 cannot veto advisory_settle either). Reviews still post/parse/lint — advisory visibility preserved everywhere. - [P3] taxonomy: doc + code comment now classify western / chinese-routed / advisory-only explicitly; payload-jurisdiction maps gemini explicitly to the Western (Google/US) column — demotion removes counting authority, not payload eligibility. Stale Google mention in the WESTERN_FAMILIES comment fixed. - [P3] design-doc-waiver text moved out of REVIEW_AUTHORITY_PRINCIPLES.md into a clearly-marked in-repo addendum on the committed record; the principles doc keeps the roster fact + citation only. - Tests: partition totality/disjointness; gemini PASS never counts (tiers 0-4 x both gate regimes); gemini CHANGES-REQUESTED (P1-backed) never blocks while the same review from claude still does; outcome-level dissenting/counting exclusion. Note: tests/swarm/test_pr_review_protocol.py:: test_resolve_provider_slots_prefers_available_candidates fails pre-existing on this branch (fails with this diff stashed too); unrelated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(swarm): close advisory-only enforcement gaps — dogfood leg, protocol dissent, packet visibility, alias ids Quorum round-4 findings at d55528c (claude CR, openai PASS): - P2: the required adversarial-dogfood leg no longer accepts an advisory-only-attributed item (a gemini dogfood comment could satisfy the Tier 1+ requirement via _known_model_reviewer_id). It stays visible in dogfood_evidence but does not satisfy the leg. - P2: protocol-payload dissenting_views are now filtered for advisory-only identities like the comments path (a {agent: gemini} view from a merge-protocol payload or stale prepared artifact no longer blocks); shared helper _view_is_advisory_only checks model_family + agent (family:role form), canonicalized. - P3: a blocked-severity ([P0]/[P1]) advisory-only CHANGES-REQUESTED no longer vanishes from the merge packet — recorded as an advisory view, with the reasons note naming the roster demotion (advisory-only family) rather than the severity gate as the non-blocking cause. - P3: counted_families canonicalizes via canonical_family before the advisory-only subtraction, and _FAMILY_ALIASES gains google -> gemini (mirrors FAMILY_PROVIDERS / the recognizer markers), so raw alias/provider ids cannot dodge the exclusion at any filter site. - Tests, one per gap: gemini dogfood fails the required leg (deepseek contrast satisfies); protocol-payload gemini dissent does not block; P1-backed gemini review appears in advisory_views with the advisory-only-family reason; alias forms (google/Google/' GEMINI ') excluded across tiers, EvidenceItem counting, and dissent. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(cli): extract review-queue renderers to review_queue_render.py review_queue.py sits at its 6000-LOC bridge ceiling (#8553: extract, do not grow) and this PR's advisory-only roster logic pushed it to 6054. Move the pure presentation helpers (_render_packet, _render_merge_authorization_packet, _render_active_auto_handle_alerts) to a satellite module, re-exported under their historical names so call sites and tests are unchanged. 5783 LOC; ratchet PASS. * fix(quorum): advisory-only families excluded from emitted counted ids _counted_model_reviewer_ids now drops advisory-only families at the source, so merge packets (counted_reviewer_ids/counted_model_families) and evidence-lint would_count match the never-count roster contract — downstream automation reads those fields as counts-toward-quorum (#9363 round-5 openai [P2]x2). Reviews stay visible as reviewer_signals/advisory_views. Regression tests for both surfaces. * fix(quorum): canonicalize gemini-cli registry id to the gemini family Live protocol payloads carry AgentRegistry names ('gemini-cli:role'); without the alias, demoted gemini dissent re-enters through protocol dissenting_views and dodges the advisory-only exclusion (#9363 round-5 openai [P2]). Regression tests on both surfaces. * docs(governance): align REVIEW_AUTHORITY_PRINCIPLES with the counted-family code Tier 3-4 counts only the frontier-grade Western subset (claude, openai, grok); the doc still said all Western families count at every tier, an authoritative spec/code mismatch about merge authority (#9363 round-6 openai [P2]). Kimi lane example refreshed to the live catalogued pin and the lifted K3 deferral replaced with the surviving requirement — a reviewer pin must name a catalogued model ([P3]). * fix(quorum): close the antigravity gemini-demotion escape antigravity is the current primary Gemini surface (agy CLI, default_model=gemini-3.5-flash) but its registry id did not canonicalize to gemini, so a protocol dissenting_view attributed to 'antigravity:<role>' passed _view_is_advisory_only and BLOCKED merges — exactly the 'gemini dissent counted somewhere' this PR exists to prevent (#9363 round-6 claude [P2]). Rather than hand-adding a third alias and waiting for the next surface, test_every_gemini_registry_surface_is_demoted walks AgentRegistry and fails CI on any Gemini-family agent whose id escapes the demotion. Verified the guard catches the leak with the alias removed. Also recognizes 'antigravity' as a gemini marker on both recognizer tables so a self-labelled Antigravity review is preserved as an advisory view instead of dropped ([P3]). --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Summary
moonshotai/kimi-k3OpenRouter pin and migrate active Kimi defaults, reviewer selection, PDB invocation/pricing, model metadata, UI labels, seed data, benchmarks, scripts, and current documentationKimiK3Agentwhile retainingKimiK2Agentas an import compatibility aliaskimi-thinkingto K3 mandatory reasoning and make credential detection accurately require OpenRouter for current Kimi agentskimi-legacyAPI contract, and Factory Droid K2.5 because the installed Droid catalog does not expose K3Availability proof
moonshotai/kimi-k3endpoint, 1,048,576-token context, text/image input, reasoning/tools support, and current pricingValidation
603 passedfocused agent, credential, selector, PDB, model-pin, and quorum-evidence teststsc --noEmitpython3 scripts/check_portability.pygit diff --checkbash scripts/automation_pr_preflight.sh origin/main HEADGovernance
Draft because the change includes the Kimi model used by quorum-evidence review. No evidence, settlement, workflow rerun, or merge action is included.