Find where your AI agent will invent UI.
Your AI can write the UI. This makes sure it writes your UI.
An agent builds UI the way a new hire does on their first day: it looks around the repo and copies what it finds. We measured that on 10 real products, 259 agent sessions, with and without this tool. On everyday work the agent reused the components and tokens that were there and stayed on-system. It went off-system in the places where the repo had no answer to copy: a chart in a codebase with no chart palette, a theme in a codebase with no named surfaces. In each case it made the values up and hardcoded them.
Run it at the root of a UI repo. About a second later a self-contained HTML report opens. No account, no network, no telemetry, nothing in your repo is changed. Free, on npm, MIT.
Roast. Teach. Guard.
- Roast the repo to find its real design system, the mess in it, and the gaps.
- Teach the agent, through generated rules and a local MCP server.
- Guard every edit, so the agent does not have to remember to ask. When the agent says it is done, a last review looks at everything the session changed and sends it back once if a problem it added is still there.
The mess an agent adds is a map of the gaps in your system
An agent copies what it finds. Where the repo has no answer, it makes one up, and the next agent copies that. This tool draws that map before the gap becomes a layer. It finds the places with no answer, measures the mess already there, writes the rules for your agent, and checks every edit the agent makes, so nobody has to remember to ask.
A script does the counting. Claude writes the explanation. Every number in the report comes from a deterministic read of your files, the same numbers every run. Where an AI reads the scan for you, its text is labelled as written by AI and kept apart from the measurements.
What you get
- The gaps, named. The report opens with one question, where will your agent have to guess, and answers it first: the places where the repo has no answer yet, with the files that prove each one and the single move that closes it, or "No gaps found". Today it knows one gap for certain, because the agent runs found it: charts that paint their series colours by hand in a repo with no chart palette. Across 126 public repos, 43 look like that. The list grows only when a gap has been measured.
- A health score you can defend in a meeting. 0 to 100, the same number every run, measured against 3 yardsticks: the ideal norms of a design system, the median of the 34 product repos at the core of a 119-repo benchmark, and 10 reputable systems (Primer, Polaris, Carbon, shadcn/ui and others). Monorepos get a score per package, so the tidy UI library stops hiding the messy app.
- Where the colour comes from. The colour usage bar counts, of every 100 colour uses, how many read a theme colour by name, how many use Tailwind's palette and how many are strays written by hand. It follows CSS variables, Tailwind themes, Sass and Less variables and JavaScript theme objects.
- Every finding with its file path. Every colour and its near-identical twin, every off-scale spacing value, typeface, duplicated or never-imported component, inline style block and !important, each with a real file path. The adoption map draws your real system to scale, every component's tile sized by its import count, the never-imported grouped by the year git last saw them.
- What to fix first. A "Fixes you can make right now to increase the health score" list derived from your own numbers, with what each move is worth and where the score lands if you do them all. Every move carries a copy button with a ready-made fix prompt for your agent. Fix, rescan, press the next button.
- Rules that stop the mess coming back. A generated design-system-rules.md with the canonical components, your token file and the known duplicates to avoid. One flag writes it into every agent file you have: Claude, Cursor, GitHub Copilot and Windsurf. Every scan also checks the rules you already have for stale references.
What the agent runs showed
Claude Code, headless, on 10 public products pinned to one commit each (cal.com, Dub, Metabase, Plausible, SigNoz, trigger.dev and four more). We counted the design-system findings each session added, by this tool's own rules. 355 sessions, Sonnet 5 and Haiku 4.5, September and October 2026. Routine work stayed on-system with or without the tool: 7 findings without and 9 with, over 80 sessions. Work that needed something new did not.
- 72 in 24 runs, to 1 in 24. Haiku 4.5 asked for a chart and a new component on 4 products, without the tool and with the plugin on. October 2026, roast 10.1.
- 21 in 24 runs, to 7 in 24. Sonnet 5, same tasks. 5 of the 7 are chart colours in a repo with no chart palette. The agent kept them and wrote a comment saying why.
- 28 to 2. Findings a Christmas theme added across 10 products, without and with this tool. September 2026.
The agent does not always call a tool when it should: Sonnet in 7 sessions of 24, Haiku in 2. A check on every edit does not have that problem. Nothing was rendered, so zero findings means the code follows these rules. It does not mean the design was reviewed. Method, tables and limits are in the research write-up, which will be published separately.
Three ways to use it
The scan and the report. Every flag is in the table below.
The same engine as a local MCP server: five read-only tools the agent calls while it writes UI, from "is there a Button already?" to "review my changes". The Claude Code plugin bundles it, so plugin users skip this step. Verified in Claude Code, Cursor and Windsurf (now Devin Desktop); any MCP client can register the same command.
Three doors to one check. --check reads your git diff in the terminal and exits 1 on findings. The plugin's review skill does it in chat. And the plugin's edit hook runs it after every file the agent edits, writes or produces with a shell command, handing back only the findings that edit added, so the agent does not have to remember to ask. For pull requests, the sister package guard-my-design-system runs the same rules as a GitHub Action and fails the check when new mess arrives.
/plugin install roast-my-design-system@roast-my-design-system
What it installs: two skills, one local MCP server and two checks that run on their own, nothing else. roast (/roast-my-design-system) scans the whole repo, writes the report with Claude's read of the numbers inside it, then walks the fixes with you. review (/roast-my-design-system:review) checks only what changed, in about a second. Older Claude Code without the marketplace: clone the repo and copy skills/roast-my-design-system into ~/.claude/skills/ (or ~/.codex/skills/ for OpenAI Codex).
Read only, no network, no telemetry, one shareable HTML report. Every release is provenance signed on npm and passes a 408-check test suite before it ships.
What it works on
- React repos: Next, Remix, Vite and plain React.
- Web-component repos: Stencil and Lit.
- Every repo compared with repos built the same way: shadcn installs, products on MUI, Mantine, Chakra or Ant Design, Tailwind themes, web-component systems, shadcn registries. A repo that uses two kits is told so.
- 5 kit profiles: Tailwind theme, MUI, Mantine, Chakra, Ant Design. The header names the kit, and both kits when there are two.
- Any styling on top: Tailwind, styled-components, Emotion, Sass, Less, vanilla-extract, Stitches, CVA, CSS Modules.
- Vue, Angular and Svelte: named in the header, colours and spacing still counted, but components are not measured and the report says so.
- HeroUI, NextUI, Radix Themes, Fluent UI, React Bootstrap and Grommet: named in the header, no kit rules.
Every command
One scan powers all of it; the flags decide what lands on disk. Combine freely. The full documentation lives at github.com/gregkozakiewicz/roast-my-design-system.
| Command | What you get |
|---|---|
npx roast-my-design-system@latest | The scan and design-system-roast.html, opened in your browser |
npx roast-my-design-system@latest <path> | Scan a different repo than the current directory |
... --apply | The generated agent rules injected into every agent file you have: CLAUDE.md, AGENTS.md, .cursorrules, .cursor/rules/, .windsurfrules and .github/copilot-instructions.md, inside a marked block. Re-running replaces only that block, never your own text |
... --rules | The same rules written to design-system-rules.md instead, for pasting by hand |
... --card | roast-card.svg: a shareable 1200x630 card with the score and worst findings. Pure SVG, embeds in a README |
... --sarif | design-system-roast.sarif for GitHub code scanning: upload it in CI and findings appear in the Security tab, annotated on files |
... --mcp | The scan as a local MCP server: five tools your agent calls while writing UI, from "is there a Button already?" to "review my changes", plus the roast-fix prompt serving the top fix from a fresh scan. Nothing leaves your machine |
... --check | The working tree's changed files checked against the design system, in the terminal. Exits 1 on findings, so it slots into scripts |
... --by "Dwayne Hicks" | A requester credit in the report header, next to the scan date |
... --notes <file.md> | An agent-written analysis embedded in the report as "What the repo teaches the agent": labelled as written by AI, kept apart from the measured numbers. The Claude Code skill writes and passes this automatically; the flag is here so any agent can |
... --section "Title" <file.md> | An agent-written chapter appended after the notes, same styling, same written-by-AI label, with sub-headings allowed. Repeatable, so analysis that outgrows the notes still lives inside the report instead of a hand-built page |
... --exclude lab/ | Leave a folder out of the scan, or list folders in a .roastignore file at the repo root. The report prints every exclusion in the header with its file count, so a scoped score always says it is scoped |
... --json | The scan summary as JSON on stdout, for scripts and pipelines |
... --theme light / --out <file> / --no-open | Light report, custom report path, don't open the browser |
/roast-my-design-system:review (in Claude Code) | The plugin's second skill: the same check as --check, in chat. Each changed file's findings with the fix named, the fixes applied on request, the check re-run. About a second, no score |
/roast-my-design-system (in Claude Code) | The full experience: the roast in chat and embedded in the report as "What the repo teaches the agent", the rules offer, and the fix loop with Claude on your own numbers |
One scan writes rules for every agent: Claude, Cursor, GitHub Copilot, and Windsurf. Every scan also checks the agent rules you already have and flags stale references, no flag needed.
What your agent gets over MCP
One real exchange against Unleash, an MUI product that scores 60/100. The agent has a grey in hand and a padding in mind. Every answer is the server's own text, unedited.
Nearest token: #607d8b, 11 channel steps from #6b7280. Unless the difference is a deliberate decision, use the token.
2 findings: ✕ L5 Colour #6b7280 written onto an MUI component, and the theme has no such colour. Fix: Add it to the theme once (frontend/src/themes/dark-theme.ts), then read it there: color: 'text.secondary' in sx. ✕ L5 Pixel size p: 12px on an MUI component. Fix: 12px is between steps 1 (8px) and 2 (16px) on the default theme. Keep it with a comment, or use the nearest step in sx.
color: 'text.secondary' and p: 1.5. roast_validate again:No measured violations found. Checked: hardcoded colours vs the token set, near-identical colour twins, off-scale spacing, off-scale radii, font sizes and shadows, typefaces outside the system, arbitrary bracket values, static inline style blocks, !important, duplicate component definitions, chart colours against the chart palette, colours and pixel sizes written onto kit components where the theme has a value.
Four calls, under 800 tokens, and the new component reads the theme instead of adding colour number 44. Where a repo has charts but no palette, a new chart that paints by hand gets one warning naming the existing chart that does the same and asking for the palette once. That is the gap report, live.
Example roasts
Real reports from public repos, hosted exactly as the skill generates them. Every number is deterministic; every path is real.
Your AI can write the UI. This makes sure it writes your UI.