← All skills

ux-craft

v2.5.3MITMaking something

The other half of design-craft: how a thing behaves rather than how it looks. Flows, forms, error states, the wording on buttons, and whether a person can really tap the thing on a phone. Journeys get a named default sequence so two onboardings are not the same three-step wizard.

Install
/plugin install ux-craft@fledgeling-plugins

Needs the marketplace added first — how to do that.

Reach for it when

Book-grounded UX engine for implementing, mocking, and reviewing web and mobile UIs, layouts, and user flows, plus marketing and transactional emails — grounded in the UX canon (Norman, Nielsen, Krug, Yablonski, Refactoring UI) and the psychology behind it, with a deterministic gate that refuses what prose only asks.

Not for

NOT for visual artifact production (use design-craft:design-craft), pixel-matching an implementation (use mockup-fidelity:mockup-fidelity), Figma email graphics (use email-mockups), or research strategy (use intent-layer or discovery-sentinel).

Uses multiple models

Uses multiple modelsThis skill may ask a different AI for a second opinion. Usually to check its own work, because a reviewer from the same family tends to agree with it. The defer skill picks which one, from OpenAI, Google, xAI or another Claude, based on what the job is and which account has room left. Nothing leaves your machine unless a skill you ran asks for it.Read about defer →

What ships with it

Say any of this

  • why do users drop off
  • is this intuitive
  • review the UX
  • audit this flow
  • make this easier to use
  • will users understand this

Taken from the skill’s own trigger description — these are the phrases it listens for. You do not have to match them exactly.

Easily confused with

A Claude Code plugin for the UX half of building interfaces: forms, flows, states, error recovery, interface copy, and the review that catches what a screenshot hides. It pairs with design-craft: that one is the visual hands, this one is the UX brain.

Generated journeys used to ship the same skeleton every time (hero CTA, three features, three-step how-it-works). Build mode now names that default and varies the sequence on one axis, or keeps it with a reason.

This README is the functional version. It gets its voice pass, its icon and its banner in the brand phase; the content below is accurate now.

What it does

Three modes, picked from the shape of what you ask:

You haveModeYou get
Something that exists (code, a URL, a screenshot, an email, a flow description)ReviewA prioritised report with pasteable fixes, severity calibrated to user impact, and an honest list of what could not be checked
Something to build or mock (a screen, a flow, a form, an email)BuildThe goal sentence, the existing system matched, the flow shaped and settled, a counted state grid, the real words, then the gate
A question about behaviour ("why do users drop off", "modal or inline")AdviseThe answer in the first sentence, the mechanism chain behind it, both options argued honestly, and a rating of how strong the evidence actually is

The gate

skills/ux-craft/scripts/ux-lint.py: stdlib-only Python, no dependencies, two modes.

ux-lint.py --static src/checkout        # walk HTML/JSX/TSX/Vue/Svelte/CSS
ux-lint.py --probe http://localhost:3000/checkout   # measure a rendered page

It refuses the failures that ship silently: a <div onclick> carrying navigation with no role and no tabindex, outline: none with no focus style anywhere in the file, a placeholder standing in for a label, motion with no reduced-motion guard, a <form novalidate> with no per-field error states, lorem ipsum in the artifact, and any surface asserting its own verification. Warnings cover contrast where both colours resolve, competing primary actions, a destructive action whose only gate is a toast, a live region inserted with its text already inside it, and target sizes under the WCAG floor.

Three properties are the point:

  • Every finding names three things: what you did, what the user silently gets, and the fix. Not "invalid".
  • Only exit 0 is a pass. A run that examined zero files exits 2, because a clean sheet over nothing is a lie. A check that raised exits 4: unknown, not clean.
  • Every run prints a never-empty "Not checked" list. A check that cannot measure says so rather than reporting zero. Screen-reader output, real keyboard traversal, colours behind a var(), and everything the render engine cannot see all appear there by name.

The accessibility floor, resolved

The three touch-target numbers in circulation are not interchangeable, and mixing them produces a finding a client can disprove from the spec:

NumberStandardLevel
24 × 24 CSS pxWCAG 2.2 SC 2.5.8 Target Size (Minimum)AA, the only one a WCAG failure may cite
44 × 44 CSS pxWCAG 2.2 SC 2.5.5 Target Size (Enhanced)AAA, a craft target; a miss is not an AA failure
44 × 44 ptApple Human Interface Guidelinesnot WCAG; pt is density-independent
48 × 48 dpAndroid / Materialnot WCAG; dp is density-independent

Both WCAG numbers were read at w3.org with their exceptions. Neither vendor page could be read at source in the session that wrote this, and the evidence file says so rather than smoothing it over.

Install

/plugin marketplace add fledgeling-co/fledgeling-plugins
/plugin install ux-craft@fledgeling-plugins

What it knows about its own limits

skills/ux-craft/references/evidence.md rates the replication status of every behavioural law the skill cites, including the ones it argues against: nudge effects near zero after publication-bias correction, choice overload context-dependent rather than universal, Hick's Law largely failing to transfer to a structured interface, Miller's 7±2 misapplied twice over, type-to-confirm universally adopted and never measured. Every measured claim in the skill carries its run and date, and the two whose run was never recorded are marked as such.

The skill also carries a Known limits section: it cannot substitute for usability testing, cannot measure drop-off, and cannot verify assistive-technology behaviour. It is told not to promise any of those.

References

Thirteen, each justified by the failure it prevents rather than by a contents list.

FileWithout it
review-playbook.mdA review becomes a framework dump, and a render that failed gets reported as a pass
flows-and-forms.mdA form ships with one reachable state, a live region that never announces, and a list of 387 items with no way to find one
flow-shape-variety.mdA journey defaults to the same 3-step how-it-works skeleton; names one axis of variation and the project ledger
persuade-conversion.mdA Persuade surface has a CTA but no offer/audience/action contract or proof placement; loaded only for landing and campaign surfaces
psychology-laws.mdFindings become taste and citations become name-drops
evidence.mdThe skill holds its own claims to a lower standard than it holds its citations
mobile-ux.mdA desktop layout gets shrunk instead of prioritised, and hover carries something load-bearing onto a device with no hover
email-ux.mdThe footer clips past Gmail's 102 KB limit and takes the unsubscribe link with it
ai-product-ux.mdAn AI surface overwrites user work silently and treats retrieved content as trusted
data-provenance.mdA figure with no provenance renders as the strongest claim available
ux-writing.mdThe copy reads as machine-written and the empty states say "No data"
checklists.mdThe closing sweep runs from memory, which is where the legally-required links go missing
model-calibration.mdA model family that needs a cell to fill gets a paragraph to read

Browsers

Obscura only. Playwright, Puppeteer, chrome-devtools-mcp, Playwright MCP and browser-use are not used and not recommended. The playbook carries that engine's measured blind spots as a table, because a reviewer who does not know them files engine artifacts as product defects, and the sharpest one lands exactly on this skill's subject: a native radio input renders as nothing, which looks precisely like a missing affordance.

Credit

Rebuilt from the ux-craft:ux-craft skill in the diolog-plugins marketplace, which supplied the canon, the review playbook, the psychology reference and the measured live-site defects that most of these rules are built on. The predecessor's two-tier evidence taxonomy with a declared n=1, its find-wide-then-filter-hard rule, its section arguing against its own citations, and its provenance reference are kept here largely as they were written, because they were better than anything a rewrite would have produced.