UI/UX round F — before/after vision scores
Scores use a 1–10 scale; each cell is before→after.
The same P-visual (including severity definitions) and P-first5 prompts were sent verbatim to read ?q= for all 18 named after shots: 36 responses, no retries, no parse errors. Raw response wrappers and text use the same shot/key/visual/first5 schema as before. Parsed findings use the exact scores/worst/first5/findings schema at the round root, as requested by Main. Vision scores and issues are judgments [INFERENCE], not new live interaction assertions.
Per-shot scores
| Shot | Hierarchy | Density | Alignment | Contrast | Discoverability | Polish |
|---|---|---|---|---|---|---|
| J1-desktop | 6→5 | 6→6 | 8→7 | 7→6 | 6→6 | 6→6 |
| J1-mobile | 5→5 | 5→4 | 8→7 | 6→6 | 5→5 | 6→6 |
| J2-desktop | 6→6 | 5→6 | 8→8 | 7→7 | 6→6 | 6→6 |
| J2-mobile | 5→4 | 5→4 | 7→7 | 7→7 | 5→5 | 6→6 |
| J3-desktop | 5→5 | 3→4 | 6→5 | 6→6 | 4→4 | 5→5 |
| J3-mobile | 3→3 | 4→3 | 5→5 | 6→6 | 4→4 | 4→4 |
| J4-desktop | 5→5 | 4→4 | 6→6 | 6→5 | 4→4 | 5→5 |
| J4-mobile | 4→4 | 4→4 | 6→5 | 6→6 | 4→4 | 5→5 |
| J5-desktop | 7→7 | 6→6 | 8→8 | 7→7 | 4→4 | 7→6 |
| J5-mobile | 4→4 | 3→3 | 7→7 | 5→5 | 2→2 | 4→4 |
| Q1-desktop | 5→5 | 5→5 | 7→7 | 6→6 | 4→4 | 5→5 |
| Q1-mobile | 3→3 | 2→3 | 5→5 | 4→4 | 3→3 | 3→3 |
| L1-desktop | 4→4 | 4→4 | 7→6 | 5→6 | 4→4 | 5→5 |
| L1-mobile | 3→3 | 3→4 | 4→5 | 5→5 | 4→4 | 3→4 |
| D1-desktop | 3→3 | 4→4 | 6→6 | 5→5 | 4→4 | 4→5 |
| D1-mobile | 3→3 | 2→3 | 3→5 | 4→4 | 4→4 | 3→4 |
| S1-desktop | 4→4 | 4→4 | 6→6 | 4→4 | 4→4 | 5→5 |
| S1-mobile | 4→4 | 4→4 | 6→7 | 4→5 | 4→4 | 5→5 |
Severity totals
| Severity | Before | After | Delta |
|---|---|---|---|
| P1 | 6 | 7 | +1 |
| P2 | 90 | 87 | -3 |
| P3 | 76 | 76 | +0 |
| Total | 172 | 170 | -2 |
Matching rule and outcome
[INFERENCE] Findings are compared only within the same shot (surface + viewport), using element/function AND underlying observable problem. Rewording, coordinates, rank, severity, and broader/narrower descriptions of the same symptom do not make a finding new. Every baseline severity, including P3, is eligible. Cross-shot matches and vocabulary overlap alone are ineligible. A compound finding is matched when an explicitly described core element/symptom carries forward; that does not claim every added clause was previously reported. NEW means absent from that same-shot baseline critique, not introduced by the fix.
All 94 after P1/P2 findings were compared: 88 matched, 6 NEW (0 P1, 6 P2). Two of the 88 matches are borderline; a stricter toolbar-specific rule would make those two additional NEW P2s (86 matched / 8 new overall). Complete ID mapping and reasons: finding-matches.json.
All seven after P1s have a baseline symptom match, including two severity regradings.
| After P1 | Baseline match | Baseline severity |
|---|---|---|
| F-J3-desktop-1 | F-J3-desktop-1 | P1 |
| F-J3-mobile-1 | F-J3-mobile-1 | P2 |
| F-J4-desktop-1 | F-J4-desktop-1 | P1 |
| F-J4-mobile-1 | F-J4-mobile-1 | P1 |
| F-J5-mobile-1 | F-J5-mobile-1 | P1 |
| F-Q1-mobile-1 | F-Q1-mobile-1 | P1 |
| F-Q1-mobile-2 | F-Q1-mobile-7, F-Q1-mobile-4 | P3, P2 |
Remaining after P1 findings — verbatim
F-J3-desktop-1 — P1
Element: Bottom action bar and quiz viewport
Problem: At this 1280×900 viewport, the bar spans approximately x=123–1143 and y=805–884, covering document content. The second question’s toolbar runs behind it, leaving controls partially clipped. Persistent chrome competes with the content users are editing.
F-J3-mobile-1 — P1
Element: Bottom “YOUR QUIZ” action panel
Problem: The roughly 125px-tall panel overlays the quiz rather than occupying a clearly separate region. Answer content is visibly hidden behind it. Between the settings above and this panel below, only the quiz heading and first question stem remain unobstructed. Whether scrolling can fully expose every answer cannot be established from the screenshot.
F-J4-desktop-1 — P1
Element: Bottom action bar overlapping quiz content
Problem: The bar occupies approximately x123–1142, y805–884 and paints over the document. The second question’s toolbar begins around y765, so its lower portion and subsequent answer content are obscured. Users must reposition content to work around the overlay.
F-J4-mobile-1 — P1
Element: Bottom “YOUR QUIZ” action panel
Problem: At approximately x30–345, y759–885, the panel overlays the quiz document. Answer content is partially visible beneath it near the bottom edge, while a roughly 126px-high block of chrome interrupts reading immediately after the first question. The screenshot cannot establish whether scrolling can fully reveal the covered controls.
F-J5-mobile-1 — P1
Element: Question text, answer cells, and metadata
Problem: Most quiz text is approximately 6–8 px tall; metadata and the footer are smaller still. The title is readable, but the content needed to complete the quiz is not comfortably legible at this scale.
F-Q1-mobile-1 — P1
Element: Sticky “Move to Trash” / “Update” bar
Problem: At approximately y=482–521, the full-width bar overlays the question list. Question 4 is partially hidden, and a delete icon is visible only as a fragment below the bar. This interrupts reading and makes the covered controls unusable in the visible state.
F-Q1-mobile-2 — P1
Element: Question text, helper text, and row controls
Problem: At the supplied 217 px image width, much of the interface renders at roughly 6–8 px text height. Several icons occupy only about 10–14 px. Reading long questions and distinguishing adjacent reorder, edit, and delete actions is impractical at this rendered size.
NEW after P1/P2 findings — verbatim
F-J1-desktop-2 — P2
Element: Category selector, approximately x318–963, y382–758
Problem: Twenty category tiles compete at equal visual weight, with no visible indication of whether selection is single-choice or multiple-choice. The full-width ‘Surprise me’ tile uses a dashed outline unlike every other tile, making its role ambiguous: random selection, a separate mode, or a special action.
F-J1-mobile-5 — P2
Element: Surprise me tile
Problem: The full-width ‘Surprise me’ tile is placed ahead of every category, but its dashed border gives it a different interaction language without explaining the difference. It is unclear whether it selects a random category, generates a mixed quiz, or immediately starts something.
F-J1-mobile-9 — P2
Element: Collapsed Customize row
Problem: ‘Customize’ and a chevron communicate expansion, but not what settings are available. Nothing visible tells a first-time user whether this contains question count, difficulty, format, or something else.
F-Q1-mobile-6 — P2
Element: “Answers” switches
Problem: “Answer review for players” appears pale blue while “Answer sheet in PDF” appears gray, but neither has an explicit On/Off label. The policy notes imply site-level constraints without clearly saying whether either switch is locked. Enabled, disabled, and policy-controlled states are difficult to distinguish.
F-D1-mobile-5 — P2
Element: “Subjects,” “Attention,” and “Next content ideas”
Problem: The same subject shortages are surfaced repeatedly: subject rows show question counts, “Attention” links to subject problems, and “Next content ideas” repeats subjects with explanatory sentences and individual “Add questions” links. The repetition extends the page without adding equivalent decision value.
F-S1-mobile-4 — P2
Element: Question Imports workflow, approximately y=1060–1373
Problem: File selection, a gray “Import draft quiz” control, explanatory copy, and a separate draft-generation form are stacked inside one card. The disabled-looking import action has no adjacent, explicit requirement explaining how to enable it. “Create draft” is visually dominant, but its relationship to importing a JSON file is unclear.
Borderline matched P2 findings — verbatim for review
F-J3-desktop-2 — P2
Element: Question editing toolbars
Problem: Each question has six faint, icon-only controls. The plus and paired-arrow symbols do not clearly explain their effects or scope. These controls look disabled even where they may be actionable; the first question’s up arrow is even lighter, but the distinction is subtle.
Continuity judgment [INFERENCE]: Matched to baseline F-J3-desktop-3 for the same undiscoverable per-question editing actions. Toolbar-specific wording/state ambiguity is additional detail; a stricter toolbar-only rule would classify this as NEW.
F-J4-desktop-2 — P2
Element: Question editing toolbars
Problem: Each question has six small, pale, icon-only actions. The plus, paired arrows, and overlapping rectangles do not explain whether they affect questions or answers. These are essential editing controls, but the screenshot provides no visible labels or instructions.
Continuity judgment [INFERENCE]: Matched to baseline F-J4-desktop-2 for the same undiscoverable per-question editing actions. Toolbar-specific wording/state ambiguity is additional detail; a stricter toolbar-only rule would classify this as NEW.
Matching caveats
Uncertainties and granularity
- F-J3-desktop-2 → F-J3-desktop-3 and F-J4-desktop-2 → F-J4-desktop-2 are the two borderline matches. Their shared functional symptom is undiscoverable per-question editing operations. The baseline wording emphasizes missing clear affordances; AFTER emphasizes visible but unintelligible icons. These are not matched to the unrelated top question-count icons merely because both use plus symbols. A stricter requirement for an explicitly baseline-described per-question toolbar would instead make these two new, yielding 42 matched / 5 new.
- F-J3-desktop-5 is a compound finding: its checkbox complaint is directly baseline-reported, but its button-like page-count facet is not independently explicit in the J3-desktop baseline. The matched status does not claim every clause is old. Likewise, the two uncertain toolbar matches do not establish baseline reporting of every icon-state detail.
- The three new classifications mean absent from the same-surface baseline report. They do not establish newly introduced runtime defects or causation by P1/P2 changes.
Uncertainty / compound-scope caveats
Binary decisions above follow finding-level continuity: NEW requires no same-element, same-symptom baseline. A matched compound finding does not mean every clause was previously reported.
- F-S1-desktop-5 — partial coverage: the unexplained disabled import prerequisite directly matches F-S1-desktop-6. The AFTER-specific JSON-versus-library workflow distinction is additional detail, not separately established by that baseline problem.
- F-S1-desktop-6 — partial coverage: unlabeled toolbar/diagnostics controls directly match; the appended bare “7”/missing question-unit complaint has no identified baseline counterpart.
- F-S1-mobile-2 — changed-state caveat: header crowding directly matches F-S1-mobile-3. BEFORE F-S1-mobile-2 said no save/autosave confirmation; AFTER explicitly sees “Saved” but questions whether autosave is explained. Do not describe the former absence assertion as unchanged. The crowding match determines this finding’s matched status.
- F-S1-mobile-5 — wider-element scope: unlabeled header search/eye controls directly match F-S1-mobile-3; diagnostics controls were not identified in any mobile baseline problem.
- F-S1-mobile-6 — wider comparison scope: the experience-card symptom directly matches F-S1-mobile-4; the extension comparing language, Feedback, and ordinary-button outlines is not independently covered by that baseline.
- Interpretation-sensitive NEW exclusions: F-D1-mobile-5 and F-S1-mobile-4 are deliberately not merged into generic “fragmented guidance” or “disabled button” complaints. Their observable symptoms are respectively repeated subject-health information and missing workflow/prerequisite explanation—not merely related vocabulary or shared controls. No binary match is left undecided.
Measurement and interpretation limits
- The 25 worse metric rows contain 19 single-run timing differences and 6 non-timing differences. tDom is one navigation timing and measureMs includes screenshot/measurement/tab-walk work; neither is a repeated matched performance benchmark.
- visibleIndicators is a count across the driver’s 40 Tab steps, not a count of distinct controls. Its 40→11 change does not by itself prove 29 controls lost focus styling. All recorded consoleErrors/pageErrors on the 16 DOM-measured shots were empty; J5 has no such collector.
- Full-page screenshots have different dimensions on 9 of 19 PNG pairs; no resizing/cropping was used to manufacture ratios. J5 contains freshly generated random quiz content, so PDF pixel differences cannot be attributed solely to styling.
- F-Q1-mobile-2 refers to a supplied 217 px image width, but the actual after PNG IHDR is 375×2712. That pixel-size wording is retained verbatim as requested, not endorsed as a CSS-pixel measurement. [INFERENCE] Full-page image presentation/downscaling can affect apparent text size in the vision critique.
- The desktop journey manifest contains only a draft entry; the mobile manifest contains draft, artifact, artifact, preview-meta. Both J5 PDFs returned status 200 and rendered with 3 pages; desktop page count came from pdf.js rather than preview metadata. The incomplete desktop capability recording is reported without modifying the identical driver.