Catching design–code drift before PR review

I built a lightweight method to compare our design system to code on a pull request, route real issues to the right owner, and keep Figma honest when design should get updated. The hard part was teaching the tool what useful signals look like, and a laundry list isn't it.

01 The challenge

Designers were spending ~45 minutes per pass and roughly a full day each month catching visual bugs that should have been caught long before review. Figma and code both acted like source of truth, but nothing kept them in sync, so mismatches kept surfacing late in Storybook sweeps and product review.

02 The results

In an audit of our Alert component, the script evaluated 22 style properties, caught 4 legitimate visual bugs (including a focus ring mismatch), and generated almost no noise. It took roughly 25–35 hours to move from proof of concept to trustworthy output, with the team estimating ~150 hours of manual parity checks saved at scale.

35 hours

About 25–35 hours from first proof of concept to output I would trust in a real review

4

Legitimate visual issues caught in one alert crit demo, with almost no noise

03 The solution

I focused on creating a shared filter for what is worth a ticket, what is noise, and who owns the fix, combining severity checks with human review so handoffs between design and engineering stay predictable.

Project details

Company

WP Engine

Premium web hosting for WordPress sites.

Role

Product Designer

I owned the workflow end to end and spent most of the time shaping what "good" output looks like for real reviewers.

Team

Myself

Design systems, Engineering

Timeline

3 weeks

Approximately 25–35 hours

Tools

Figma, Storybook, Cursor, Github and Jira. Models: Composer 2 and Composer 2.5

Type

Internal workflow

Internal workflow for design–code parity during PR review.

Full case study

I built this in Cursor by feeding the Agent our design system and codebase. My goal was to name what "drift" means for us, keep humans in the loop where judgment matters, and make handoffs feel predictable instead of surprising.

The challenge

A branch can look fine in Storybook while Figma and code have quietly diverged. The design problem is not finding every delta. It is protecting reviewer attention so people still trust the report.

Solution

Calibrating taste was where the real design work occurred. Letter-spacing and focus rings are worth a conversation, but icon opacity usually isn't something worth fixing. I review curated output before anything becomes a ticket or a Figma change, and we re-check after fixes so we're not trading one mismatch for another.

Constraints and trade-offs

I still confirm before ticketing or applying changes. Hidden layers (nested containers without meaning) in Figma taught us to be careful about exclusions and practicing better file hygiene.

Impact and results (what is true today)

Early demos surfaced issues engineers actually engaged with. We're still rolling the shared toolkit out to the wider team.

The filter is the product, not the amount of differences found. Next are publishing the skills in the team repo so designers can compare Figma to a pull request without living inside a local dev environment.