roast-my-design-system

Find where your AI agent will invent UI.

Your AI can write the UI. This makes sure it writes your UI.

An agent builds UI the way a new hire does on their first day: it looks around the repo and copies what it finds. We measured that on 10 real products, 259 agent sessions, with and without this tool. On everyday work the agent reused the components and tokens that were there and stayed on-system. It went off-system in the places where the repo had no answer to copy: a chart in a codebase with no chart palette, a theme in a codebase with no named surfaces. In each case it made the values up and hardcoded them.

npx roast-my-design-system@latest

Run it at the root of a UI repo. About a second later a self-contained HTML report opens. No account, no network, no telemetry, nothing in your repo is changed. Free, on npm, MIT.

Roast. Teach. Guard.

The mess an agent adds is a map of the gaps in your system

An agent copies what it finds. Where the repo has no answer, it makes one up, and the next agent copies that. This tool draws that map before the gap becomes a layer. It finds the places with no answer, measures the mess already there, writes the rules for your agent, and checks every edit the agent makes, so nobody has to remember to ask.

A script does the counting. Claude writes the explanation. Every number in the report comes from a deterministic read of your files, the same numbers every run. Where an AI reads the scan for you, its text is labelled as written by AI and kept apart from the measurements.

What you get

What the agent runs showed

Claude Code, headless, on 10 public products pinned to one commit each (cal.com, Dub, Metabase, Plausible, SigNoz, trigger.dev and four more). We counted the design-system findings each session added, by this tool's own rules. 355 sessions, Sonnet 5 and Haiku 4.5, September and October 2026. Routine work stayed on-system with or without the tool: 7 findings without and 9 with, over 80 sessions. Work that needed something new did not.

The agent does not always call a tool when it should: Sonnet in 7 sessions of 24, Haiku in 2. A check on every edit does not have that problem. Nothing was rendered, so zero findings means the code follows these rules. It does not mean the design was reviewed. Method, tables and limits are in the research write-up, which will be published separately.

Three ways to use it

1. Roast the repo
npx roast-my-design-system@latest

The scan and the report. Every flag is in the table below.

2. Put the system inside the agent
claude mcp add roast -- npx roast-my-design-system@latest --mcp

The same engine as a local MCP server: five read-only tools the agent calls while it writes UI, from "is there a Button already?" to "review my changes". The Claude Code plugin bundles it, so plugin users skip this step. Verified in Claude Code, Cursor and Windsurf (now Devin Desktop); any MCP client can register the same command.

3. Check what the agent changed
npx roast-my-design-system@latest --check

Three doors to one check. --check reads your git diff in the terminal and exits 1 on findings. The plugin's review skill does it in chat. And the plugin's edit hook runs it after every file the agent edits, writes or produces with a shell command, handing back only the findings that edit added, so the agent does not have to remember to ask. For pull requests, the sister package guard-my-design-system runs the same rules as a GitHub Action and fails the check when new mess arrives.

Install the Claude Code plugin
/plugin marketplace add gregkozakiewicz/roast-my-design-system
/plugin install roast-my-design-system@roast-my-design-system

What it installs: two skills, one local MCP server and two checks that run on their own, nothing else. roast (/roast-my-design-system) scans the whole repo, writes the report with Claude's read of the numbers inside it, then walks the fixes with you. review (/roast-my-design-system:review) checks only what changed, in about a second. Older Claude Code without the marketplace: clone the repo and copy skills/roast-my-design-system into ~/.claude/skills/ (or ~/.codex/skills/ for OpenAI Codex).

Read only, no network, no telemetry, one shareable HTML report. Every release is provenance signed on npm and passes a 408-check test suite before it ships.

What it works on

Supported
Recognised, not supported yet

Every command

One scan powers all of it; the flags decide what lands on disk. Combine freely. The full documentation lives at github.com/gregkozakiewicz/roast-my-design-system.

CommandWhat you get
npx roast-my-design-system@latestThe scan and design-system-roast.html, opened in your browser
npx roast-my-design-system@latest <path>Scan a different repo than the current directory
... --applyThe generated agent rules injected into every agent file you have: CLAUDE.md, AGENTS.md, .cursorrules, .cursor/rules/, .windsurfrules and .github/copilot-instructions.md, inside a marked block. Re-running replaces only that block, never your own text
... --rulesThe same rules written to design-system-rules.md instead, for pasting by hand
... --cardroast-card.svg: a shareable 1200x630 card with the score and worst findings. Pure SVG, embeds in a README
... --sarifdesign-system-roast.sarif for GitHub code scanning: upload it in CI and findings appear in the Security tab, annotated on files
... --mcpThe scan as a local MCP server: five tools your agent calls while writing UI, from "is there a Button already?" to "review my changes", plus the roast-fix prompt serving the top fix from a fresh scan. Nothing leaves your machine
... --checkThe working tree's changed files checked against the design system, in the terminal. Exits 1 on findings, so it slots into scripts
... --by "Dwayne Hicks"A requester credit in the report header, next to the scan date
... --notes <file.md>An agent-written analysis embedded in the report as "What the repo teaches the agent": labelled as written by AI, kept apart from the measured numbers. The Claude Code skill writes and passes this automatically; the flag is here so any agent can
... --section "Title" <file.md>An agent-written chapter appended after the notes, same styling, same written-by-AI label, with sub-headings allowed. Repeatable, so analysis that outgrows the notes still lives inside the report instead of a hand-built page
... --exclude lab/Leave a folder out of the scan, or list folders in a .roastignore file at the repo root. The report prints every exclusion in the header with its file count, so a scoped score always says it is scoped
... --jsonThe scan summary as JSON on stdout, for scripts and pipelines
... --theme light / --out <file> / --no-openLight report, custom report path, don't open the browser
/roast-my-design-system:review (in Claude Code)The plugin's second skill: the same check as --check, in chat. Each changed file's findings with the fix named, the fixes applied on request, the check re-run. About a second, no score
/roast-my-design-system (in Claude Code)The full experience: the roast in chat and embedded in the report as "What the repo teaches the agent", the rules offer, and the fix loop with Claude on your own numbers

One scan writes rules for every agent: Claude, Cursor, GitHub Copilot, and Windsurf. Every scan also checks the agent rules you already have and flags stale references, no flag needed.

What your agent gets over MCP

One real exchange against Unleash, an MUI product that scores 60/100. The agent has a grey in hand and a padding in mind. Every answer is the server's own text, unedited.

roast_find_token #6b7280:
Nearest token: #607d8b, 11 channel steps from #6b7280. Unless the difference is a deliberate decision, use the token.
It writes the first draft anyway and runs roast_validate before saving:
2 findings:
✕ L5 Colour #6b7280 written onto an MUI component, and the theme has no such colour.
   Fix: Add it to the theme once (frontend/src/themes/dark-theme.ts), then read it there: color: 'text.secondary' in sx.
✕ L5 Pixel size p: 12px on an MUI component.
   Fix: 12px is between steps 1 (8px) and 2 (16px) on the default theme. Keep it with a comment, or use the nearest step in sx.
Second draft: color: 'text.secondary' and p: 1.5. roast_validate again:
No measured violations found. Checked: hardcoded colours vs the token set, near-identical colour twins, off-scale spacing, off-scale radii, font sizes and shadows, typefaces outside the system, arbitrary bracket values, static inline style blocks, !important, duplicate component definitions, chart colours against the chart palette, colours and pixel sizes written onto kit components where the theme has a value.

Four calls, under 800 tokens, and the new component reads the theme instead of adding colour number 44. Where a repo has charts but no palette, a new chart that paints by hand gets one warning naming the existing chart that does the same and asking for the palette once. That is the gap report, live.

Example roasts

Real reports from public repos, hosted exactly as the skill generates them. Every number is deterministic; every path is real.

npx shadcn create, freshfactory install, all 61 componentsread as a fresh install: "the score is the kit's, not yours"; 13 colours, every theme variable in place, shadcn's own 24 bracket values named and not counted
100/100
UnleashMUIread as an MUI product: 1,109 files import the kit and the theme is read 6,166 times; 4 colours and 2 spacings per 100 kit files are written onto components, against a target of 4
60/100
MetabaseMantineits own wrapper over Mantine counts as the kit, so 2,679 files are read instead of 392; nothing written onto components, 2 spacings per 100 kit files
47/100
SigNozAnt Designread as an Ant Design product: 2 colours and 7 spacings per 100 kit files written in style objects where a token exists
51/100
Apache AirflowChakra UIread as a Chakra product: no colours written onto components, 4 spacings per 100 kit files as pixel strings where a space step exists
43/100
vercel/ai-chatbotshadcn install71 values like [13px] written outside the Tailwind scale, and 66 palette colours per 100 files where a theme variable exists
80/100
excalidraw/excalidraw78 off-scale spacing values and 90 !important declarations
55/100
dubinc/dub642 arbitrary bracket values and 21 duplicated components
20/100
telekom/scaleStencil95 Stencil components read by tag; 66 spacing values outside the scale where about 12 would do
55/100
magicuidesign/magicuiregistryread as a shadcn registry and counted on the components it publishes: 52 off-theme colours per 100 files in the code it ships, its docs site kept out and named
78/100
adobe/spectrum-web-componentsLit740 colour tokens with 8 hardcoded colours beside them, and 37 !important declarations
83/100

Your AI can write the UI. This makes sure it writes your UI.