The Rams Design Score is a single number from 0 to 100 that summarizes the design quality of a codebase's UI layer. It's applied the same way every time, against the same ruleset, by the same reviewer.
It reads code the way a senior designer reviewing a pull request would. It does not run the app, take screenshots, or evaluate visual aesthetics.
Rams reviews the UI code itself: the .tsx, .jsx, .vue, .svelte, SwiftUI .swift, and styling files that produce the rendered interface. Issues are evaluated against 348 rules across 9 design categories, each catching a specific failure mode.
Every rule lives in one of these. Each specialist reviews the diff with its own rules, the way a real design team splits the work.
Semantic HTML, keyboard navigation, screen-reader compatibility, focus management, ARIA correctness, color contrast on text. Things WCAG 2.2 AA cares about and that real users with disabilities run into.
Hardcoded hex values that bypass design tokens, contrast ratios, semantic color usage, hover and focus states.
Heading hierarchy correctness, font scale adherence, line height, font weight, text truncation patterns.
Whether layout values come from a scale (4/8/16) or are arbitrary one-offs. Padding and margin consistency. Gap usage.
Component composition, prop API design, primitive reuse, design-system token leaks, in-line style overrides.
Loading states, empty states, error handling, destructive action confirmations, form patterns, scroll behavior.
Animation duration and easing, prefers-reduced-motion support, infinite loops, layout shifts caused by animations.
Patterns specific to AI-generated UI code: arbitrary radius values, magic numbers, overly defensive null checks, unused props, copy-paste duplication.
SwiftUI and Apple-platform conformance: SF Symbols over emoji and raster icons, Dynamic Type, safe areas, semantic system colors and Dark Mode, VoiceOver traits, Reduce Motion, right-to-left layout.
Each issue carries a severity. The score deducts points by severity, and criticals cap the score — one at 59, two at 49, three or more at 39 — so a 60 or above always means zero critical issues. The number answers one question: is this ready to merge?
Critical: breaks core functionality, blocks accessibility, ships visible bugs. Heaviest deduction.
Serious: degrades quality, leaks design-system intent, will cause user-facing regressions. Moderate deduction.
Moderate: minor polish or convention issue. Light deduction. (Excluded from the public score view to focus attention.)
Brighter is safer to ship. A high score means the UI clears the bar with no critical issues left standing.
The score has a boundary on purpose. These belong to other tools, and reading them into a design number would only blur it.
The homepage widget scores a representative slice. For continuous review on every pull request, the GitHub App reads exactly the files that changed.
When you score a public repo via the homepage widget, Rams pulls up to 30 UI files prioritized by location (app/, pages/, components/) and reviews them as a representative sample. A score of 100 is reserved for a repo small enough to read end to end. Anything larger is sampled, and sampling caps the score at 98. The number never claims more coverage than the review had. For continuous review on every pull request, install the GitHub App. It reviews exactly the files that changed.
The rubric is intentionally stable, and scores are versioned: every score names the engine that produced it. We don't loosen rules to flatter scores. When the engine changes meaningfully, we rescore the public set on the new version and keep the history in the open — a score never moves silently. Every change is logged below.
When rules are added or weights revised, it lands here. The counts and areas are public. The rule text stays in the engine.
Thirty-five rules added across seven categories, from forms that lose your work and navigation faked with click handlers to exits that never run and pages that scroll sideways on a phone. Twenty-nine existing rules sharpened. Benched before shipping: recall on the planted-defect fixtures went up on two defects and held on the rest.313348 rules
A deterministic core: a first set of checks is now measured from the code rather than judged — the same finding every run, cited exactly. More graduate as they prove out.
Four rules added: motion and typography.309313 rules
The engine grew eyes: changed components are rendered and the pixels reviewed, with before-and-after renders shipping in reviews as proof. Three rules added from production reviews.306309 rules
Fifteen rules added across seven categories, and a sharpening pass over the existing set.291306 rules
The critical cap is now a ladder: one confirmed critical still caps a score at 59, two cap at 49, three or more at 39. Sixty and above still means zero criticals. The public catalog was re-capped in place.
Thirty-three rules added across eight categories — motion, input, and icon craft drawn from a study of the agent skills the design-engineering community ships. Four rules sharpened; touch targets raised to Apple’s 44px floor.258291 rules
The engine measures before it judges: contrast ratios and token usage are computed from the code and cited exactly, and findings the code cannot evidence are dropped instead of published. Reviews run about 20% faster. The catalog was rescored — scores moved in both directions, almost always at the critical cap.
Largest batch yet, and a ninth category: Native, for SwiftUI and Apple-platform work. New rule families for agent-built UI and cross-app coherence. Review restraint recut: zero findings on clean work is a valid verdict, and every critical is re-verified before it can cap a score.194258 rules
7 rules added from production reviews. Reviewer precision improved: the engine reads a codebase’s own design system before judging it.167174 rules
1 craft rule added, from a pattern seen repeatedly in the wild.166167 rules
47 rules added across five categories, and 25 existing rules sharpened with tighter detection.119166 rules
Scoring: any critical issue now caps the score at 59, so a score of 60 or above always means zero critical issues.
Re-reviews now verify fixes: resolved findings are named and the score recovers. 10 rules added, 5 sharpened.109119 rules
Initial public methodology: 109 rules across 8 categories, severity-weighted scoring.
Free on public repos. Rams reviews every pull request against these 348 rules and posts inline fixes.
Frameworks