Season 0. The designer agreement read is on Nov 6.

How Rams measures design

Rams puts one number on how well an interface is designed. This page is the method behind that number: what we look at, how it is computed, how sure we are, and where it can be wrong. It is our opinion, held to a fixed method and checked against designers.

Method
1
Updated
Oct 10
Judge
Claude Sonnet 4.6
Cross-judge
GPT-5.6 Terra
Viewports
1440 · 390 px
Rules
348
Designer check pending · read Nov 6

Three measures

One judge and one method, applied to three kinds of work.

  • Models

    Rams Index

    Every model builds the same pages from the same written briefs. We render each page and grade it. Scored 0 to 100.

  • Live sites

    Rams Index · the web

    The public home pages of well-known products, rendered as a signed-out visitor sees them. The first cut scored 140 sites on the same scale as the models.

  • Pull requests

    Rams Score

    The UI code in a pull request, read against 348 public rules. Scored 0 to 100 and posted on the pull request with fixes.

What we look at

Rendered pages for the two indexes. Code for the pull request score.

Seven grades
Hierarchy, typography, spacing, color, components, composure and craft. The judge grades each from 1 to 5, once on the desktop render and once on the phone render.
Two widths
Every page is rendered in Chrome at 1440 px and at 390 px, and both screenshots are judged. A page that only works at one width shows it.
Rendered checks
Measured, not judged: text contrast, text under 12 px, tap targets too small for a finger, content wider than the phone, and overlapping or clipped elements. The same page gives the same result every time.
The rules
348 public rules across 9 categories, from accessibility to motion. The pull request score reads code against them. Every rule is listed on the rules page.

How the number is computed

Plain arithmetic on the judge's grades and the review's findings.

Rams Index

(grade sum − 7) / 28 × 100

The grade sum runs from 7 (every grade a 1) to 35 (every grade a 5). We average desktop and phone, then the judge's passes, then the briefs, so every brief counts the same. The rescale is a straight line, so ranks and ties are the grade sum's own.

The web
The same formula on the same seven grades, so a site and a model sit on one scale. The first cut is one judge pass per site. Beside it we print the page score, the number a Rams page review returns.
Defects stay separate
On the Rams Index the rendered checks are counted beside the score and never blended into it. A page can look great and still ship small text.
Rams Score
Each issue in the code carries a severity, and the score deducts points by severity. Critical issues cap it: one at 59, two at 49, three or more at 39. A 60 or above always means no critical issue.

How sure we are

Every Index score carries an interval, and ranks come from the intervals.

Intervals
Each model gets a 95% interval. We resample the briefs, the builds of each brief and the judge passes on each page 4,000 times, paired across models so every model faces the same draws.
Ranks
A model’s rank is 1 plus the number of models clearly better than it. Clearly better means the 95% interval on the paired difference sits above zero. Models inside the noise share a rank, and the board says so.

Judges

A model grades the pages, so we check the model.

Primary
Claude Sonnet 4.6 grades every page. On 37 real home pages and 21 iPhone apps we also tried Claude Opus 5.5, GPT-6.1 Sol and OpenAI’s Decisions API as judges. None did better on both the web pages and the app screenshots, so the judge stayed.
Cross-judge
GPT-5.6 Terra judges the same pages through the same steps, and its score is printed beside ours. A Claude judge ranks Claude models, so we measure how much kinder it is to Claude pages than the second judge is. On the current board that difference is −0.2 points (95% interval −3.7 to +3.6): no sign that either judge favors its own lab’s models.
Next
Both judges’ orders are printed on the Index. Next is an open-weight judge calibrated only on designer ratings, so the measure belongs to no lab. Planned, not built.

Repeatability

Run it twice, get the same answer. ICC is how closely two judge passes on the same page agree, where 1 means identical.

Real home pages
0.97

Grade sum ICC across two passes on 37 sites.

App Store screenshots
0.94

Grade sum ICC across three passes on 21 apps.

Generated pages
0.76

Grade sum ICC on 15 model-built pages judged twice. They sit close together, so passes swap them more often.

Pull request score
~2

Average move in points when the same review runs again, from 96 reruns on real pull requests. About 9 in 10 land within 2 points.

Agreement with designers

The test that decides whether the number means anything.

Status
Pending. Outside designers, paid and working alone, compare pairs of pages without seeing any score: 37 home pages, 21 iPhone apps and 45 model-built pages.
The rule
Written down on Oct 6 and amended once on Oct 7, before any answer came back. Rams passes on a surface if its agreement with each designer is clearly above zero and within 0.10 of how well the designers agree with each other (Kendall tau). If the designers barely agree with each other, the read is void.
When
The read is Nov 6. We publish the result whether it passes or fails, with the numbers.
So far
Against one person’s quality tiers for the 37 home pages (design-led, typical and rough, set before scoring), the grade sum agrees at Kendall tau 0.71. One rater is a sanity check, not validation.

Keeping the test fair

A number people aim at gets gamed. These are the defenses.

Private briefs
From season 1, 25 private briefs carry the rank. They are never published, and pages built from them are never shown. The current board uses the 5 public briefs only.
Anchors
The 5 public briefs, on rams.ai since July, run beside the private set. A model that does better on them than its private score predicts, by more than the field does, gets a note on its row.
Rotation
Each quarter the five oldest private briefs retire into the public sample, and five new ones replace them.
Leak checks
Before every run and on the first of each month, we search for each brief online. A brief found word for word is retired the same day, the board is recomputed without it, and the changelog says so.
We run every model
Through its public API, at the provider’s default settings. Labs don’t send outputs, can’t pick their best page, and can’t pay to be listed, run again or removed. If a listed lab is a Rams customer, its row says so.

What we don't do

Lines the measure never crosses.

Customer code
Never used to build or tune the measure. Code sent for review is analyzed and discarded, never stored and never used for training. Details are on the privacy page.
Screenshots
We don’t publish screenshots of other companies’ sites, and we don’t license or sell them as training data.
Blocks
We don’t get around robots.txt, logins or bot checks. A site that blocks us is skipped and named, never scored.
The judge’s prompts
Stay private, with the rubric wording and the check thresholds. The method, the scores and the intervals are public.

Known weaknesses

Where the number can be wrong today, and what we do about it.

Brand halo
The judge often knows whose page it is: it named the site in 45 of 140 summaries in the first web cut. So we tested it. On 25 sites we hid the name and logo and judged again, beside a control that made the same edit to ordinary words. Hiding the brand cost no more than the control (+0.5 points, 95% interval −2.9 to +4.3), so we see no brand halo bigger than about 4 points. We repeat the test each season.
Banners
Cookie banners and sign-up sheets cover the page. Since 10/10 the renderer hides consent banners from 28 consent platforms and floating cookie, region and sign-up prompts it recognizes by their words, without clicking anything, and the row says what it hid. Prompts that sit in the page flow stay, because a visitor sees them too.
Bot walls
A challenge page can be graded as if it were the site. Since 10/10 the renderer refuses bot walls, access-denied and maintenance pages before the judge runs (it caught all five that got past the old check). A person still looks at every screenshot before an edition goes out.
Generated pages
Frontier models all build competent pages, and the scores sit close together. The judge repeats less well on them (ICC 0.76) than on real sites, so small gaps between models are noise and show as shared ranks.
Our renderer
Pages that need WebGL can fail in headless Chrome. A page that didn’t render is withdrawn, not scored.

Corrections

How a site, a lab or a team can dispute a score.

Tell us
If you think a site, a model or a pull request is scored wrong, email rams@rams.ai with the page and what looks off. A person reads every one.
What we do
We render the page again and judge it on the current method. If our render was wrong (a bot wall, a banner we should have hidden, a page that didn’t load), we fix or withdraw the score.
If we were right
The score stands and we tell you why. Nobody can pay to change a score. Any score we change goes in the changelog with the reason.

Versions

A score never moves without a line in a changelog.

Boards
Every board names its method version, judge model and engine commit. When the judge, the renderer or the checks change, we judge every saved page again on the new version, bump the version and keep the old board with its date.
Pull request scores
Each review names the engine that produced it. Rule changes are logged in the rules changelog.
This page
  • 2026-10-10Method 1Published with the Rams Index (season 0): the cross-judge result, the brand-halo test, and the renderer’s bot-wall and banner handling.
  • 2026-10-09Method 1First version of this page, in draft.
  • 2026-10-07Index 0The Rams Index headline became the grade sum rescaled to 0 to 100. The finding-based visual score moved to a secondary column.
Method 1 · updated 2026-10-10Questions or a disagreement: rams@rams.ai