Automated Review: Check new UI against Cockpit's style - #2915
Conversation
Automated PR Review (Claude)Change mapGovernance note (read first). This PR's entire diff edits Claims. The PR body and commit message assert three things:
Supporting counts from the base checkout, since the rules are stated as descriptions of it:
Failure site. Not applicable — this PR fixes no bug. It adds review rules. Entry points. The PR changes no functions. The changed artifact is a policy document read at review time by both review workflows.
Because both workflows read the file from the base ref, the new bullets take effect for every PR opened after this merges, and the false-positive rate in 1.1 is multiplied by that frequency rather than paid once. Invariants. The added text relies on the rule that a bullet stated as "what the codebase already does" is safe to flag deviation from. The sites that can violate it are the codebase facts each bullet asserts; the table above enumerates them. Nine hold, three do not, and those three are findings 1.1, 1.2 and 1.3. 0. Summary
The PR appends twelve bullets to the UI/UX section of the automated reviewer's guidelines, codifying dialog anatomy, Vuetify theming, white-alpha button fills, footer action layout, padding ownership, glass layering, stacking, field-attached actions, icon-button affordances, sentence case and vertical-space economy, and instructs that each breach be counted as its own finding under the existing severity rubric. Most of the twelve are accurate descriptions of the base checkout and would be useful as written. Three are not: one flags a pattern the codebase uses 352 times without exception, one demands an element the shell that AGENTS.md mandates does not provide, and one asserts a button-token vocabulary the shell contradicts. Because the rules are consumed as instructions by an automated agent that cannot check them against the code before firing, an inaccurate bullet does not degrade gracefully — it produces confident, wrong findings on every subsequent PR. 1. Correctness & Implementation Bugs1.1 1.2 1.3 2. Persistence & User Data — ✅ (the diff adds, reshapes and removes no persisted key; it touches one Markdown file under
|
b0ffc27 to
27045b5
Compare
|
Done:
|
|
/review |
Automated PR Re-review 2 (Claude)Comparing Previous findings status
Verification of each, against the code rather than the comment: 1.1 — the bullet ( Discussion since last review
New findingsChange mapGovernance note (still applies). The whole diff edits Claims.
Counts re-derived from the base checkout this run, since each bullet is stated as a description of it:
Failure site. Not applicable — this PR fixes no bug. It adds review rules. Entry points. The PR changes no functions; the changed artifact is a policy document read at review time.
Because both workflows read it from the base ref, an inaccurate bullet is paid once per PR opened forever, not once here — which is what sets the severity of 1.4 and 1.5. Invariants. The addition relies on the rule that a bullet stated as "what the codebase already does" is safe to flag deviation from. The sites that can violate it are the codebase facts each bullet asserts; the table above enumerates them. Eight hold, four do not. 0. Summary
The revision narrows the three bullets the last review found unsupported and reframes the finding-count instruction, and each of those five fixes is real. What remains is the same defect one bullet deeper: the button-token bullet now also tells the reviewer to flag "a footer whose buttons all look alike", which is precisely what the shared 1. Correctness & Implementation Bugs1.4 1.5
All four are reasonable things to want. State them as new conventions (and grade them accordingly) rather than as descriptions, so the reviewer does not go looking for in-tree examples that are not there. 2. Persistence & User Data — ✅ (the diff adds, reshapes and removes no persisted key; one Markdown file under
|
27045b5 to
b4003d0
Compare
|
Done:
|
|
/review |
1 similar comment
|
/review |
Automated PR Re-review 3 (Claude)Comparing Previous findings status
Verification of each against the code, not against the comment:
Discussion since last review
New findingsChange mapGovernance note (still applies). The whole diff edits Claims.
Counts re-derived from the base checkout this run, since each bullet is stated as a description of it:
Failure site. Not applicable — this PR fixes no bug. It adds review rules. Entry points. The PR changes no functions; the changed artifact is a policy document read at review time.
Because both workflows read it from the base ref, an inaccurate bullet is paid once per PR opened forever, not once here — that is what keeps 1.6 and 1.7 worth raising at all. Invariants. The addition relies on the rule that a bullet stated as "what the codebase already does" is safe to flag deviation from. The sites that can violate it are the codebase facts each bullet asserts; the table above enumerates them. Ten hold, two do not, and the three the umbrella explicitly downgrades to "house direction" are no longer claims about the tree. 0. Summary
The revision closes the previous 1. Correctness & Implementation Bugs1.6
1.7 2. Persistence & User Data — ✅ (the diff adds, reshapes and removes no persisted key; one Markdown file under
|
b4003d0 to
b2199b3
Compare
|
Done:
|
|
/review |
Automated PR Re-review 4 (Claude)Comparing Previous findings status
Verification of each against the code, not against the comment:
Discussion since last review
New findingsChange mapGovernance note (still applies). The whole diff edits Claims.
Counts re-derived from the base checkout this run, since each bullet is stated as a description of it:
Failure site. Not applicable — this PR fixes no bug. It adds review rules. Entry points. The PR changes no functions; the changed artifact is a policy document read at review time.
Because both workflows read it from the base ref, an inaccurate bullet is paid once per PR opened forever, not once here — that is what keeps 1.8-1.10 worth raising at all. Invariants. The addition still relies on the rule that a bullet stated as "what the codebase already does" is safe to flag deviation from, with a fixed three-item carve-out for the conventions the tree is converging on. The sites that can violate it are the codebase facts each bullet asserts; the table above enumerates them. Ten hold; the icon-label and magic-pixel clauses do not, and neither is inside the carve-out. A second invariant arrived with this round's fixes: a footer "rendered by the 0. Summary
All four findings carried in from round 3 are closed by code, not by argument: the footer-actions bullet is reframed as house direction with the shell's missing 1. Correctness & Implementation Bugs1.8 1.9 1.10 2. Persistence & User Data — ✅ (the diff adds, reshapes and removes no persisted key; one Markdown file under
|
b2199b3 to
ee7632c
Compare
|
Done:
|
|
/review |
Automated PR Re-review 5 (Claude)Comparing Previous findings status
Verification of each against the code, not against the comment:
Discussion since last review
New findingsChange mapGovernance note (still applies). The whole diff edits Claims.
Counts re-derived from the base checkout this run, restricted to what this round's change turns on:
Failure site. Not applicable — this PR fixes no bug. It adds review rules. Entry points. The PR changes no functions; the changed artifact is a policy document read at review time.
Because both workflows read it from the base ref, a mis-graded rule is paid once per PR opened forever, not once here — that is what keeps 1.11 at Invariants. This round replaces the invariant the previous rounds relied on. It used to be a bullet stated as "what the codebase already does" is safe to flag deviation from, with a fixed three-item carve-out; it is now a clause is safe to flag as universal practice only if it cites in-tree precedent, and otherwise its breach is 0. Summary
All four findings carried in from round 4 are closed by code, not by argument: both shell exemptions are narrowed to the 1. Correctness & Implementation Bugs1.11 Both exemplars are breaches of clauses the test downgrades:
The consequence is not cosmetic. A PR that adds a The fix is one ordering clause: state that the precedent test lowers the consistency grade only, and that a breach reaching the user stays 2. Persistence & User Data — ✅ (the diff adds, reshapes and removes no persisted key; one Markdown file under
|
ee7632c to
924a116
Compare
|
Done:
|
|
/review |
Automated PR Re-review 6 (Claude)Comparing Previous findings status
Verification of each against the code, not against the comment:
Discussion since last review
New findingsChange mapGovernance note (still applies). The whole diff edits Claims.
Counts re-derived from the base checkout this run, restricted to what this round's change turns on:
Failure site. Not applicable — this PR fixes no bug. It adds review rules. Entry points. The PR changes no functions; the changed artifact is a policy document read at review time.
Because both workflows read it from the base ref, a mis-graded rule is paid once per PR opened forever, not once here — that is what keeps a grading defect above a wording preference. Invariants. This round replaces the invariant again. It used to be a clause without cited precedent has its breach graded 0. Summary
Both findings carried in from round 5 are closed by code. The severity collision at the centre of 1.11 is gone, and it is gone in the right direction: the precedent test now decides only how a finding is framed, while severity is decided by whether the breach reaches the user, so a new 1. Correctness & Implementation Bugs1.12 The split: "a breach that still reaches the user (an overlay rendering light and unreadable, a control no keyboard can reach) stays The consequence is over-firing at verdict-moving weight. On the base ref there are 114 icon-form This is residue from the remedy I proposed in 1.11, and the remedy was right — the fix here is one clause, not a reversal. Either name where the accessible-name clause sits (for instance 2. Persistence & User Data — ✅ (the diff adds, reshapes and removes no persisted key; one Markdown file under
|
The reviewer now checks the conventions that have only ever lived in review comments: dialog anatomy, the white-alpha button tokens, the two-action footer, who owns the padding, and the dark theme that overlay-teleporting Vuetify controls need because the app configures none. Not all of them describe the current tree - about half do and the rest are the house direction it is converging on - and all are read against the existing severity rubric, so a breach that reaches the user can move a verdict instead of landing as an unweighted nit.
924a116 to
f7f81ba
Compare
|
Done:
Not addressed:
|
|
/review |
Automated PR Re-review 7 (Claude)Comparing Previous findings status
Verification of each against the code, not against the comment:
Discussion since last review
New findingsChange mapGovernance note (still applies). The whole diff edits Claims. All re-derived against the base checkout this run rather than carried over from round 6.
Failure site. Not applicable — this PR fixes no bug. It adds review rules. Entry points. The PR changes no functions; the changed artifact is a policy document read at review time.
Because both workflows read it from the base ref, a mis-graded rule is paid once per PR opened forever, not once here — which is what has kept every grading defect in this PR above a wording preference for six rounds. Invariants. The invariant established in round 6 is unchanged: the precedent test decides the framing, and the severity is decided independently by whether the breach reaches the user. The sites that can violate it are the twelve bullets' individual clauses, and the test is whether each clause's breach lands unambiguously on one side of "reaches the user" vs "only cosmetically inconsistent". Re-walked all twelve this run, and this round closes the one hole: 0. Summary
Both findings that a code change could close are closed, and this round found nothing new. The icon-only bullet now grades a missing accessible name on both branches ( 1. Correctness & Implementation Bugs — ✅ (re-walked all twelve bullets against the base checkout: the
|
major.majorwhen the breach reaches the user (an overlay rendering light and unreadable, a control no keyboard can reach),minorwhen the surface is only cosmetically inconsistent with the app around it (an action nobody can tell is the primary one). When a single surface breaches several rules, they are grouped under one finding with a sub-item per fix, so each stays actionable without inflating the count the verdict is read against.Worked example, replayed against the "Select BlueOS Cloud mission" dialog in #2865, quoted as the reviewer would emit it under section 6 (one surface, two bullets, grouped as one finding per the rule above):
The finding is
minor, so on its own it lands asMINOR SUGGESTIONSand cannot move a verdict. The rule stays quiet on what this dialog already gets right, which is the point of showing it: the close X is a keyboard-reachablev-btn iconat the top right (:11), and the separator above the footer is where it belongs (:100). A checker that also fired on a compliant header would be worse than none.Part of #2884