Design systems · AI-assisted development

A design system written for AI to build from

I created and lead a next-generation design system for an enterprise software company. Product teams don't get a component library or a Figma handoff. They get a versioned specification that their AI coding agent builds from, and an automated gate that checks the result.

HIGHLIGHTHow three AI test builds caught a defect that review had missed
Role
Creator and lead
Timeline
Jul 2026 – ongoing
Scope
Spec, gate, assets, releases
Tools
Claude Code, a 15-agent harness

At a glance

The specification is the product

THE BET

If an AI agent generates every component, the design system doesn't need a component library. It needs rules precise enough for a machine to follow and to check.

WHAT I OWNED
  • The spec and its visual language
  • The validation gate
  • The icon and pictogram pack
  • The agent harness
  • Releases and governance
  • The adoption guide
WHERE IT STANDS

The foundations are shipped. Pilots with other products showed how easy the system is to implement, and wider rollout now follows each team's roadmap.

0component libraries to build or maintain
15specialist AI agents in the harness I built
AAWCAG 2.2 in both themes. A failure blocks the release.
3/3AI test builds caught a defect that review had missed

The problem and the approach

Remove the library and keep the contract

In a traditional design system, product teams wait for a central team to design, build and release each component, then adapt it until products drift apart. AI coding agents can now generate screens faster than any central team can supply them, but nothing tells the agent what on-brand or accessible means. So I rebuilt the workflow around the agent.

Before

1

Designer draws the componentA Figma file per component

2

Handoff to engineering WAITSpecs, redlines, questions

3

Central team builds and releases WAITOne library, one framework

4

Product team adapts itDrift starts here

5

Accessibility auditIssues found after release

After

1

Team pins a release of the specOne line in their project

2

Their AI agent reads the contractTokens, component rules, definition of done

3

Agent generates the UIIn whatever framework the product uses

4

The gate checks it in CITokens, contrast, keyboard, ARIA, both themes

5

People review what machines can't judgeIs it the right screen? Does it read clearly?

01

One spec, any framework

Tokens, component rules and the definition of done live in three plain-text files that work with React, Vue or plain HTML.

02

Accessibility lives in the tokens

Every colour pairing the system allows meets WCAG 2.2 AA in both themes, so the agent can't pick a failing pair.

03

A release is a contract

Teams pin a tagged release, and a hash proves they built against exactly the spec they asked for.

Getting buy-in

Winning over teams who had just migrated

The Head of UX and I made the case together. The hardest audience was teams who had only just moved onto the previous version of the design system.

THE OBJECTION
Why redo work we've only just finished?

Those teams had spent real time on the last migration and didn't want to spend it again.

WHAT CHANGED THEIR MINDS
01

Speed

Teams saw UI built from the spec far faster than through the old handoff.

02

No cost to them

My agent harness does the building, so I could take on the front-end implementation myself. Teams didn't have to find the time.

03

Proof from pilots

Pilots with other products showed how easy the system was to implement.

THE OUTCOME
"We love it, but we can't do it right now."

That's a question of timing, not a rejection. Rollout follows each team's roadmap, so nobody is forced into a migration mid-delivery.

The harness

How one person covers the front end

I own the harness end to end. It's a team of 15 specialist agents, each with one job and a written standard, plus guards that stop unsafe changes. The agents do the work. I make the decisions.

QUEUERanked backlog

Every session opens with the ranked work queue.

MEOne decision

The driver turns the queue into a single choice for me to make.

AGENTSSpecialists do the work

Each task goes to the agent built for it.

Design

  • System architectIs this token the right shape?
  • Agent-experience designerCan a rule be read two ways?
  • Adoption designerIs the route in easy for teams?
  • Prior-art researcherWhat do standards say?

Quality

  • Accessibility gatekeeperDoes it meet WCAG 2.2 AA?
  • Generation proberDoes the spec build correctly?
  • Spec auditorDoes it match the spec?
  • VerifierDoes the fix fix the defect?

Delivery

  • Check authorWrites the failing test first
  • Reproduction writerProves a bug before anyone fixes it
  • Scope wardenIs this the smallest fix?
  • Cross-repo porterKeeps both repos in step

Release

  • Release stewardWas the procedure followed?
  • Consumer-impact assessorWhat changes for teams?
  • Dependency stewardAre dependencies safe?

Decisions stay with me. Agents can land work that can be measured. A change to the spec or the release process waits for me.

Brand values are locked. Any edit that moves a brand value stops and must be explained.

No fix without proof. A bug must be reproduced before an agent may fix it, and a separate agent checks the fix.

Where designers work now

Designers build in the same session as the agent

Without Figma handoff, designers work directly with an AI agent against the pinned spec. I designed the contribution flow around how designers actually work, so nobody has to memorise a rulebook.

Designer

Builds a screen

The designer describes the screen in their own words. The agent generates it from the spec.

Agent

Spots a missing asset

No pictogram in the set fits. The agent offers to draft one without leaving the session.

Agent

Applies the rules

It checks for duplicates, follows the style guide, writes the catalogue entry and opens a review.

Gate

Flags the draft

The gate lists the draft without failing the build, so the screen ships on time and the draft can't slip through unseen.

Guardrails belong in the environment, not in people's memory. A pictogram can look right and still break a rule, such as poor legibility in dark mode or a silent duplicate. So the agent enforces the style guide, and designers spend their attention on whether the asset means the right thing.

Visual language

Every visual choice is a rule the agent can follow

An agent can't interpret "make it feel premium". So I turned each aesthetic decision into a named rule with a token behind it. These are recreations of five of them.

Two themes

Brand energy where it counts

Dark is the signature theme, with deep navy, luminous teal and depth expressed through light. Light is the working surface, restrained and precise for dense enterprise workflows. The same tokens drive both.

Sprint overviewOn track
Research synthesis
Prototypes
Usability tests
modal · 20
Assign
In review surface · 16 · control · 10 · pill

Radius by role

Components reference a role, never a raw value. Nested shapes get tighter as they get smaller, and changing every card's corner is a single edit.

shadow
glow + edge

Depth through light in dark mode

Shadows vanish on navy, so dark mode lifts surfaces with a soft glow and a luminous edge. One elevation token resolves to a shadow in light and a glow in dark.

Hover or tab to each button

Colour-matched hover glow

In dark mode, hover glows in the button's own colour. Teal marks everyday actions, pink is the accent and red warns of something destructive. In light mode the same token gives a subtle lift.

Suggested categoryMaintenanceAI-generated
Suggested categoryMaintenanceAI-generated

AI disclosure, designed for regulation

The one component unique to this system marks content or actions produced by AI. It's built for the EU AI Act's transparency duty, which applies from August 2026.

  • Always visible and never dismissible
  • A text label is mandatory, so it never relies on colour
  • The gradient is decoration only and never sits under text

The gate

Trust, but verify every change

A spec only works if someone checks it was followed. I built a validation gate with Claude Code that blocks any change that breaks the contract.

STATIC

Reads the source on every commit. It flags raw colours, off-scale spacing, unknown tokens, broken type pairings and CSS that would break right-to-left layouts.

RENDERED

Opens a real browser before merge. It measures contrast on the actual pixels in both themes, tabs through every control and checks for a visible focus ring.

HUMAN

Says where automation stops. A screen can pass the gate and still be the wrong screen. The adoption guide says so plainly, so reviewers spend their time there.


    

The moment that proved the model

Three AI builds found what review had missed

Until then, the spec had only been checked by reading it. So I tested it the way teams would use it, by building from it.

01 · BUILD

Three fresh sessions

Each AI session got only the pinned spec and built the same components from scratch.

02 · MEASURE

Same failure, three times

All three failed the rendered gate on one button state. Every token was valid, so no static check could see it.

03 · FIX

Rule fixed, process changed

I corrected the rule. Any release that changes a rule now gets the same test.

Ghost buttonActive state, as specified–FAILS AA
Ghost buttonActive state, corrected–PASSES AA

Recreation. Ratios are calculated live from the colours shown. AA needs 4.5:1.

CAUGHT BY THE GATE

A rule that broke its own standard

The button rule sent the label to a surface at 3.34:1. The same paragraph cited that ratio when it banned the pairing elsewhere. Reading missed it, but measuring the pixels didn't.

MISSED BY THE GATE

A rule that contradicted itself

A table rule asked for two things that can't both be true. The agent picked one and wrote clean code, and noted the conflict. So the release process now reads the agent's notes as well as the result.

Testing a spec by building from it finds defects that reviewing it never will.

Trade-offs

What this approach costs, and how I contained it

The decision log records what each major choice costs, not only why I made it. These are three of them.

Cost

AI output varies from run to run

Three runs from an identical brief used the same tokens and values but structured the code differently, from 407 to 930 lines.

How I contained it

Govern the contract, not the code

The gate checks tokens, contrast and behaviour, which held steady. I rejected snapshot tests because they would need rewriting every time the spec or the model changed.

Cost

Draft assets can look finished

Letting designers place a draft pictogram keeps them moving, but a draft could ship by accident.

How I contained it

Make drafts impossible to miss

Every draft is marked in the code and listed by the gate on every run, with a link to its review.

Cost

Existing products weren't built for this

Most early adopters have a shipping product, its own CSS and a designer with expectations.

How I contained it

Write the guide for existing products first

The adoption guide has a section for existing products. It names the four problems they'll hit, in the order they'll hit them.

Running it as a product

Shipping a design system like software

A spec that teams build from has to be as dependable as code. Before 1.0 I ship small and often, with 90 releases in 14 weeks. Teams pin a release, so nothing changes in their product until they choose to upgrade. Every decision goes in a dated log of more than 150 entries, and a mistake gets corrected in a new entry, never erased.

Where it stands, October 2026

SHIPPED
  • Spec covering 21 components and 6 layout primitives
  • Static and rendered gate
  • Icon and pictogram pack, with one meaning per icon
  • Adoption guide and one-command setup
  • Pilots with other products
IN PROGRESS
  • Rollout to product teams, timed to their roadmaps
  • Every known adoption blocker is a tracked issue
  • The weekly report states the adoption position

Reflection

What I'd tell the next team

Test a spec by building from it

Careful reading missed a defect that three AI builds found on the first try. Now every release that changes a rule is tested that way.

Put guardrails where people work

Asking designers to memorise a style guide fails quietly. Building the rules into their AI session made the right path the easy one.

Remove the cost of saying yes

Teams liked the idea from the start. What won them over was seeing the speed and knowing the implementation work wouldn't land on them.

Let's talk about your design system

I'm looking for a senior role leading design systems and design operations.