Two surfaces.
One engine.
A dashboard where product and design teams follow UX quality release by release. An MCP server that brings the same findings into the developer's IDE before commit. Same engine, same lenses, same fix-ready output.
The report that keeps up with your product.
Not a PDF you download once. Every release is scored against the last one, regressions surface on their own, and every finding comes with the fix your team can ship today.
The UX Score, release by release
Five pillars, weighted into one score. Behaviour is only scored when a source measured it; without one it drops out and the other four carry the score.
Web weights. Mobile leans on usability and access.
Prioritised findings
Ranked by what they cost, each with its evidence: the contrast ratio, the pixel size, the screen and the step where it happened.
- CTA contrast 2.8:1, below WCAG AAHigh
- Touch target 32px on the mobile navMedium
- Form label not tied to its fieldLow
axe-core verified · desktop and mobile · per screen
Fix-ready code
Real CSS, HTML or JS you paste into your IDE, written in your design tokens instead of hard-coded values.
/* contrast 2.8:1 → 4.6:1 */ .cta-button { color: var(--canvas-white); background: var(--brand-deep); min-height: 44px; }
The neurodiversity lens
15 to 20% of people think differently. The lens reads each screen for what WCAG alone does not: ADHD friendliness, dyslexia readability, autism predictability, sensory load and colour-vision safety. It comes from the same run, with nothing to set up.
10 tools inside your IDE.
The agent that wrote the UI asks Corexi to review it, pulls the UX rules, starts a run on the live product and reads what it changed. No browser tab, no copy and paste.
{
"mcpServers": {
"corexi": {
"url": "https://corexi.ai/api/mcp",
"headers": { "Authorization": "Bearer crxi_live_..." }
}
}
}- Cursor
- Claude Code
- VS Code
- Windsurf
- Replit
- Claude Desktop
One-click install for Cursor, VS Code and Claude Desktop, a copied config for the rest. Listed on cursor.directory.
- get_ux_rulesResearch-backed UX patterns before writing code.
- get_ux_checklistPre-flight checklist for 12 surface types.
- review_codeFindings with severity, evidence and fix code.
- review_screenshotA mockup or Figma export, reviewed by vision AI.
- trigger_scanStart an Autopilot run on the live product.
- get_scan_jobFollow a run's progress until it is scored.
- get_run_diffWhat the run fixed, brought back and found new.
- get_findingsPrioritised findings with fix code, filterable.
- get_latest_scanThe latest run's result for any product.
- list_sitesYour products and their latest UX Score.
Connects to what you already use.
Analytics bring the behavioural side into the score. IDEs and agents take the fix to where the code is written.
- Clarity
- GA4
- PostHog
- Mixpanel
- Amplitude
- Firebase
- Corexi snippet
- Hotjar (connection only)
- Cursor
- Claude Code
- VS Code
- Windsurf
- Replit
- Claude Desktop
It keeps going after the first run.
Each run is compared with the last: what was fixed, what came back, what is new. And when something slips, you hear it before your users do.
Autopilot runs on your schedule and after releases: signs in, walks the product, scores every screen. A run cut short by a restart is scored afterwards, not lost.
A digest of score changes, screens that regressed and what to fix first. In your inbox, Slack or Teams.
Context builds up: your stack, your tokens, behavioural signals, earlier findings. Each run starts from what the last one learned.
Three rules, checked after every run.
A run scores the product under the number you set. It fires on a partial run too: it is about what was actually measured.
PX 63, below your threshold of 75
The product scored clearly lower than the run before, with the new high findings named in the message.
PX fell from 69 to 58 · 4 new high findings
The accessibility pillar fell since the last run: WCAG and the rules an automated engine can verify.
Accessibility fell from 74 to 61
- In the app
- Slack
- Microsoft Teams
- Signed webhook (alert.fired)
You set the threshold and where it lands; nothing fires twice for one run. A run that reached far fewer screens than the last is a smaller sample, not a fall, and the message says so.
Is your product ready for agents?
Your product is no longer used only by people. Assistants and buying agents walk it on someone's behalf, and where an agent stalls, a person stalls too.
A customer's assistant fills your forms, places the order, changes the setting. An unlabelled button, a click that says nothing, a flow that ends short: the agent stops there, and so would a person.
The question moves from "does my screen look right" to "can an agent use my product safely". Products agents cannot use are products agents will not recommend.
Flows reached without help, steps against the fewest needed, actions that did what they said, hand-offs to a person. Read from the same run that scores your product, shown beside PX.
Autopilot is itself an agent, so every run already holds the evidence. The score is arithmetic over those records, not a model's guess, and it reads in lines you can check: “6 of 8 flows reached their end without help”, “Critical flow ‘Pay’ stopped at /checkout”.
Quick scan or Autopilot
The quick scan is the taster. The run is the product.
A quick scan opens one public page and reads what it sees, free and without an account. Autopilot signs in to your product, uses it the way a new user would, maps the flows and scores every screen it reaches — then does it again on the next release and tells you what changed.
| Quick scan | Autopilot run | |
|---|---|---|
| Where it looks | One page, from outside | Inside the product, signed in, every screen it reaches |
| Flows | None | Sign-up, onboarding, checkout and the rest, with steps and detours measured |
| Evidence | The page's visuals | Visuals, the screen's controls and your real user data (GA4, Clarity, Firebase) |
| Cadence | Once | Every release, with a diff of what got better and what regressed |
| Findings | A sample, the rest locked | Every finding with evidence, priority and fix-ready code |
The whole method is published: how a run is scored and how it signs in to your product.
From finding the fix to shipping it.
Today it shows what stands between people and your goal, and hands your team the fix. Next, it ships the fix itself, checks it with real users and keeps what works.
- v1.0Live now
Finds it, prices it, hands it over
- Signs in and walks every screen, at desktop and mobile width
- Reads each screen through five lenses and weighs it against your analytics
- Prices every finding against the goal you set
- Hands the fix to Jira, GitHub, Slack and your editor over MCP
- v1.xNext
Goes further on its own
- Walks native iOS and Android apps, not only the screens you upload
- Opens the fix pull request itself, for your team to review
- Scores AI readiness: whether AI agents and assistants can use your product, from finding the button to finishing the checkout
- v2.0Where it is going
Self-driving UX
- Ships the fix behind a feature flag, to a slice of your users
- Measures it against your goal with real people
- Keeps it or rolls it back, and tells you which and why
LED-214 · Send invoiceflag · 10%Invoices sent, first weekWithout the fix61%With the fix68%Kept, and rolled out to everyone
v1.0 is what runs today. The later versions are our plan, shown without dates.
Then it goes round again.
Give it your product's URL. We set up the login with you once; from then on it goes round on its own, after every release.