UX and Product Design

Up to twenty days of chasing evidence,redesigned as one workflow.

Qlarc is a B2B SaaS platform that turns the AI governance a vendor already has into regulation-mapped evidence for the diligence a financial-services buyer is legally required to conduct.

Problem validated with industry operatorsFinancial-services AI governance
app.qlarc.ai / evidence-mapping
The live Qlarc Evidence Mapping screen, annotated with the key design decisions
  1. Readiness as one number, before the buyer ever sees the pack
Role
Lead Product Designer. I led design in a team of five. The extraction pipeline is mine end to end: I designed its boundaries, built it, and tested it.
Goal
Turn a vendor’s existing governance into a buyer-ready evidence pack before the deal stalls.
Sector
FinTech · Procurement
Timeline
1 year, first vendor interview to evidence pack
TL;DR · 2 min readOr read the full study ↓

Qlarc

Problem

How might we help vendors prove the governance they already have, in a format buyers can actually evaluate?

Role

Lead Product Designer in a five-person team. I led the product design end to end: the research, the information architecture, the intake, evidence-mapping and export flows, the design system, and the pipeline the screens sit on.

System

An AI extraction pipeline that reads a vendor’s governance documents, ties every claim to the source it came from, and stops at a named human before anything exports. I built it and tested it against documents written to break it. Full system rationale →

Approach

Design the format as an interface: one guided workflow instead of a form, evidence strength carried in colour so a reviewer reads a pack at a glance, and gaps shown beside the evidence rather than hidden behind it.

Outcome

Usability testing put workflow completion above 85% and prep time down 60 to 80% against building a pack by hand. I tested the extraction pipeline separately: over 90% source attribution and 0 fabricated claims across 280 evidence checks on 8 documents, five of them written to break the extractor, and the 2 failures it did produce, both on this page.

280Evidence checks run
0Fabricated claims

Designed for trust, not just speed.

The design thesis
The problem

How might we help vendors prove the governance they already have, in a format buyers can actually evaluate?

Most deals don’t die with a no. They die with silence: the review goes quiet, the purchase order freezes, and the vendor never learns why. Not a documentation problem. A revenue problem.

Deal value at risk
$500K to $2M

Frozen at a single governance review

Prep time
15 to 20 days

To assemble one evidence pack by hand, every deal

Feedback given
Zero

Reasons the buyer gives before walking away

01.1Validation

Two rounds, three practitioners.

  • Pratik Patel, an AI ops leader formerly at Mastercard, said procurement stops are hard blocks, not delays: unclear answers end the evaluation and buyers do not follow up. So a completeness score at the end became a readiness gate at the start (Decision 01).
  • Debashis Bhattacharyya, practice lead for tech consulting at Opus Technologies, watches documentation get rebuilt from scratch every cycle while sales teams cannot answer technical AI questions. So evidence moved down to the claim level, tagged and versioned to survive reuse (Decision 02).
  • Dr. Cari Miller, co-founder of the AI Procurement Lab and a former IEEE vice chair, was blunt: proactive, candid vendors win, evasive ones do not, regardless of product quality. So the product says missing out loud, with gaps beside the evidence instead of behind it (Decision 03).
✎ From the notebook

what stuck with me wasn’t the rejections. it was the silence. the deal stops moving and nobody says why, so the vendor loses the next one the same way. the problem isn’t the missing docs, it’s the silence around them.

Opportunities

One claim, three rungs.

How might we move the judgment earlier?

A claim climbs: it cannot be verified before it is documented, and it cannot be documented if nobody knows it is missing. Each rung is one question, and each question is answered by one decision.

Which turned the problem into a design brief. If the format is what buyers evaluate, the format has to stop being a document and start being an interface.

The Product03 / 06 · The four-step workflow

Qlarc turns governance a vendor already has into evidence a buyer can evaluate: each claim tied to its source, each gap named, a person signing before anything exports.

Connect → Upload → Gap-Fill → Review & Export.

One path, fixed order. A vendor with critical gaps is routed to a Governance Readiness Report before pack generation begins, so nothing half-built reaches a buyer.

The design system these screens are built fromFig 3.1 · expand
Design system

A style guide for a product that had to be believed.

Qlarc was built new, so the system did not start as an audit of what already existed. It started from one constraint: colour encodes how strong a piece of evidence is, and nothing decorative is allowed to use it.

Everything on these three boards is taken from the shipped screens. The colours are sampled out of them, the components are drawn the way the product draws them, and each one names the screen it lives in.

FoundationsQlarc design system01

Four colours, six signals, four type roles, one spacing unit.

Restraint was the point. A procurement reviewer has to read a claim, see how strong it is, and act. Everything that does not serve that was left out.

Colour

Sampled from the shipped screens, so every hex here is checkable against the product. Status hues are drawn on the components instead of listed, because in the product they exist only as thin borders.

Ground and mark
Primary green#81AB55Progress fill, the bar on a resolved row
Action green#4F7A2FAny control with a white label. The light green carries white at 2.67:1, this one at 5.06:1
Canvas#F9F9F7App background behind every card
Ink#000000Page and card titles
Slate#536277Labels, body copy, helper text
Signals, used at the size of a bar or a tint
Green tint#E5EEDBActive navigation
Positive panel#EDF9E1Evidence found
Panel edge#829F5DRule down a positive panel
Attention#FB8106Callout bar, and the continue button when a pack will ship partial
Attention panel#FFF2E7Why a gap matters
Blocked#DC4242The gate when a vendor does not qualify
Typography

One neutral sans across the whole product, four roles, weight and size doing the work. The serif you are reading belongs to this case study, not to Qlarc.

Page titleReview and Approve
Card titleBarclays vendor review
QuestionQ10 · What safeguards exist when your AI operates autonomously?
BodyEscalation path documented in Model Risk Policy v3, section 6.2.
MetaSource: API-verified · OpenAI usage data · 14 Mar 2026
Spacing and shape

An 8px base unit, four steps, and four radii each tied to a job. Nothing else rounds, and nothing spaces off the unit.

Spacing, at size
8
16
24
40
Radii
Chip3pxRegulation tags
Field6pxInputs and selects
Card8pxPanels and cards
Pill999pxStatus
ComponentsQlarc design system02

Four components carry the whole product.

Drawn here as the product draws them, with the screen each one lives in named beside it. The screens themselves are three beats up this page, at full width.

Evidence row · five states

The atom. A pack is a list of these, and the reviewer’s eye runs down the left edge first, where the colour bar is.

Appears inReview and Approve, the rows down the left column

Q1AI Model & version governanceApproved
Q6Bias testing method, categoriesReviewing
Q7Bias testing method, categoriesFlagged
Q3Escalation path & human overridePending
Q3Escalation path & human overrideMissing
Panel · two kinds

Evidence and guidance never share a colour. Green states what was found. Amber states what a buyer will do about what is not there.

Appears inEvidence Mapping and Gaps, the right-hand column of both

Evidence Found

Automatic halt conditions are triggered when prediction confidence falls below 72%. Escalation path documented in Model Risk Policy v3, section 6.2.

Source: API-verified · OpenAI usage data · 14 Mar 2026

Enterprise buyers evaluating AI under SR 11-7 require a documented escalation path with named owners at each step. Without this, the procurement reviewer cannot confirm human oversight exists.

Field · label, helper, control

Every field says what it wants in the reviewer’s language before it asks for it. The helper line is part of the component, not an afterthought.

Appears inGaps, under the callout

Action pair · commit and decline

Every decisive moment offers both moves. The one that costs the vendor something is a plain link with the consequence written under it, never a hidden option.

Appears inGaps, Review and Approve, and every screen that commits something

Save and Next GapFlag as gap
This means the gap will be flagged in your evidence pack output
Approve QuestionSkip for now
+ Fill 2 GapEdit

One set of parts held across intake, gap-filling, review and export, so the thing a buyer finally reads is recognisably the thing the vendor approved.

AccessibilityQlarc design system03

A status a reviewer cannot read is a status that is not there.

Three rules the system holds to, and one measurement it fails, stated because a governance product that hides its own gaps would be arguing against itself.

Never colour alone
Q1AI Model & version governanceApproved
Q7Bias testing method, categoriesFlagged

Every state carries three signals: the bar at the left edge, the word in the pill, and the group the row is filed under. Printed in grey, the row still says what it is.

Visible focus
Approve QuestionApprove Question

Focus is specified as a 2px ring offset from the control, in the ink of the surface rather than the accent, so it stays visible on the green primary and on a white field alike.

Errors that name the fix
Scope mismatch

The source covers IT assets. This question asks about AI models specifically. A regulator would treat the two as different inventories.

A flag names what is wrong in the reviewer’s own language and offers the correct next move. It never says “low confidence”, and approval stays blocked while it stands.

What it does not do yet

Measured at the values drawn on this board: every colour bar clears the 3:1 a non-text element needs, from 3.82:1 to 5.94:1. The status word inside its own tint does not. It runs 3.14:1 to 4.79:1, and three of the five states land under the 4.5:1 that text has to reach. Pending, the greyest, is the worst of them. The muted source line under a drafted answer has the same problem, at 3.45:1.

The same pass caught the primary button. White on the brand green ran 2.67:1, so every approve, save and export control on this product was failing the floor it asks a reviewer to read. Fixed here: controls with a white label moved to the darker action green at 5.06:1, and the light green stayed on fills and bars, where 3:1 is the bar. The screenshots further up this page are of the build as it shipped, so they still carry the old fill.

So the bar is doing the work and the word is riding on it, which is exactly why the first rule on this board exists. It is also the fix: darken each status word one more step until it clears 4.5:1 against its own tint, and leave the hue to the bar. It is the first thing I would change.

View the full system architecture diagram Fig 3.2 · expand
Fig 3.2 · System information architecture
A Entry and qualification
Door 01 · First-time vendor
Screen Sign Up
Decision 01 · the gate Qualification gate

Three onboarding questions. One yes of the three qualifies.

Qualifies · 1 of 3 = yes
Connect API and upload documents
None qualify
Hold · off-ramp Show what’s missing

Save and return when ready.

Door 02 · Returning vendor
Screen Log In
Direct

Already has an account with a connected API and documents, so it skips the gate entirely and lands straight on the dashboard.

Qualified first-time  +  returning  →  converge
Hub Dashboard
  • Previous reports
  • In progress
  • Create new
B Evidence pipeline
  1. 01 Upload new questionnaire
  2. 02 · Core Evidence mapping
  3. 03 Gap filling
  4. 04 Report generated
  5. 05 · Terminal Export
  6. Optional, off step 01 Upload file

    Only if a new document is needed.

That is the workflow a vendor walks. Every step in it is the answer to a question I got wrong first.

The decisions
04 / 06

Three decisions, each one told through the version that failed first.

Same shape every time: what the research made obvious, the version that failed when I tested it, what shipped instead, and what that cost.

The first version of the gate demanded everything. It failed.

The gate that failed

Patel’s point had an obvious design: check readiness before letting anyone build a pack. So the first gate required all three governance checks to pass before entry. In interviews, every real vendor bounced off it.

The readiness gate at step 3 of set up: three yes or no governance questions.
The gate as shipped, passing. One qualifying yes replaced the wall, and the check moved to step 3 of set up where the verdict costs a minute.
And the same gate, partial

A no on bias testing is carried forward as a flagged gap rather than a wall, and the pack it produces is marked partial. The unready get a Governance Readiness Report, never a doomed pack in front of a buyer.

The price

Activation gets slower on purpose. Argued about, shipped anyway.

The gate adds a step before vendors see the product. That step is the product.

The same readiness gate with one No: a partial evidence pack offered, with the gap flagged.
One No on bias testing. The path stays open and the gap is carried forward.

A confidence band said the section was weak. It could not say which claim.

The band that could not point

Bhattacharyya’s complaint was that documentation gets rebuilt from scratch every cycle while sales teams cannot answer technical questions about it. The first answer was a section-level confidence band: a reviewer could see that something was soft, but not which claim. So strength and source moved down to each individual claim, and anything with no source behind it is flagged rather than drafted anyway.

Evidence Mapping: claims grouped by state, beside the evidence for the selected one.
Every claim carries its own state. One reviewer pass reads the whole pack.
What every mapped claim carries

Three states, and none of them is a verdict. Each names the source it came from, or says plainly that it came from nowhere, so a reviewer can check it rather than trust it. A gap is not left blank either: it gets filled in, naming the accountable person, the escalation steps, and a resolution timeframe, or it gets flagged as still open.

The price

Flagging a gap instead of drafting around it costs readiness: the pack ships showing less than 100%, not a fabricated full score. Three legends to learn on top of that, and colour alone cannot carry meaning, so every state also carries a written label.

Every mapped item traces to a specific named source, verifiable at the point of review.

Gaps to Fill: Q5 human oversight and accountability, with the accountable person, escalation steps, and resolution timeframe to fill in or flag.
A claim with nothing behind it. Q5’s escalation path, shown as a gap: fill in the accountable person, the escalation steps, and a resolution timeframe, or flag it, rather than let it be quietly drafted.

One checkbox before export looked like sign-off. The practitioners said it was not.

The checkbox they rejected

Miller’s point was that candour wins deals and evasion loses them regardless of product quality. The first design put a single confirmation checkbox before export. They flagged it: one sign-off leaves a bank’s own validators with no attributable owner for any single control.

The automated send path was deleted rather than made optional. The export carries the pack, its readiness, and every answer with the source behind it. Each approved section arrives with the person who signed it, the date, and the policy version it was signed against.

The price

A person is now a hard dependency on every one of them: slower, argued against internally, and the reason no export in this system can exist without a name on it.

Named approval slows the export down. That slowdown is the accountability the regulation asks for.

Review and Approve: per-section approval beside the drafted answer and its source.
Named approval per section replaced it. Approval is blocked while a claim is flagged, and nothing in that state reaches an export.

Three decisions, all of them defensible on paper. So I went looking for the place where they would not hold.

When the extractor is wrong
05 / 06

It never invented a claim. Twice it believed the wrong one.

Eight documents, three real and five written specifically to break it. Across 280 checks it fabricated nothing, but it failed twice, and both failures were the same shape: a claim grounded to a passage that was real, answering a different question than the one being asked. What the reviewer needed was the sentence itself.

Documents8Run end to end through the extractor
Built to defeat it5Written specifically to make it fabricate
Failures found2Both grounded to a real passage. Both below
Failure 01
A requirement read as a practice

The document said what the law obliges a provider to do. The extractor recorded it as something this vendor already does. Every attribution on the row was correct. The claim underneath it was not.

Fix written, not re-run: one missing line in the lookup rules. Not the model.
The passage, genuine Providers of high-risk AI systems shall conduct bias testing at least annually. EU_AI_Act_Summary.pdf · page 4 Written as an obligation on providers, not as this vendor’s practice.
→recorded as
The claim it wrote Bias testing is conducted at least annually across all production models. Attributed correctly to page 4 ✗ A practice this vendor never evidenced.
Dashboard
Evidence Pack
Update Documents
Settings
Hi Mary
+ New Evidence pack
Review and Approve
Barclays vendor review
Approved - 4
Q1AI Model & version governanceApproved
Q2Named Model Risk accountable personApproved
Q4Training data provenanceApproved
Q5Autonomous AI safeguardsApproved
Flagged - 1
Q6Bias testing method, categoriesFlagged
Pending - 3
Q3Escalation path & human overridePending
Q7AI output monitoringPending
Q8Bias testing frequencyPending
Questions review4 of 10 questions approved
Q6- Have you run bias testing on your AI, and how often?
Regulation: SR 11-7 · Validation · EU AI Act Art. 10
Drafted AnswerEdit
Bias testing is conducted at least annually across all production models.
Source: Document-extracted · EU_AI_Act_Summary.pdf · page 4
The passage this was taken from
Providers of high-risk AI systems shall conduct bias testing at least annually and retain a written report of the results.
EU_AI_Act_Summary.pdf · page 4 · extracted 14 Mar 2026
Requirement, not evidence

This passage states what the regulation requires. It is not evidence that you do it. Attach your own bias testing report, or record this as a gap.

Attach your testing report Record as a gap Approve Question Approval is blocked while a claim is flagged. Nothing in this state can reach an export.
Failure 02
The right answer to the wrong question

A real, audited inventory, in a real, audited document. It just inventoried IT assets, and the question asked about AI models. Close enough to pass a source check, wrong enough to fail a regulator. The same flag pattern caught it, so the screen is not repeated here.

Fix written, not re-run: the same lookup stage, not the model.
The source, genuine Asset inventory, reviewed and signed off annually. Audited IT asset register Covers IT assets.
→answered
The question asked Do you maintain an inventory of the AI models in production? A scope mismatch, not a source error ✗ Right answer, wrong scope.
The unhappy paths

Loading, empty and error are designed too. The pipeline takes real time to read a document set, a new vendor starts with nothing, and uploads fail. Each state says what is happening and what to do next, because a compliance lead at 6pm does not retry silently.

Loading · the pipeline is reading

Reading Model_Risk_Policy_v3.pdf · 2 of 6

Empty · a new vendor

No documents yet

Your evidence pack starts from your existing governance documents. Upload the first one and the mapping begins.

Upload a document
Error · an upload fails

Could not read this file

Board_Minutes_scan.pdf is an image-only scan. Export it as a text-based PDF, or attach the original document.

Try another file

I designed and built the entire extraction pipeline myself, tested until it produced zero fabricated claims across 280 checks.

See the pipeline case study and its evidence →
What comes out of it

The Procurement Response Pack, version-stamped and signed.

Regulation-mapped, source-attributed, and built to mirror the buyer’s questionnaire format. It ships with a One-Page AI Governance Summary for the reviewer who has thirty seconds.

Regulation-mapped Source-attributed Named sign-off
app.qlarc.ai / export
Qlarc View and Export, the generated AI Governance Procurement Response evidence pack

This is what a vendor’s team is actually handed at the end. Whether it changes what a buyer does next is what outcomes has to answer.

Outcomes What was measured · usability testing and pipeline evaluation

A pack the buyer can trust, and the vendor can stand behind.

Two things were measured. We ran usability testing on the workflow, and I tested the extraction pipeline separately against eight governance documents, three real and five written to break it.

Usability testing put workflow completion above 85%, with prep time down 60 to 80% against assembling one evidence pack by hand. The extraction pipeline traced a source for over 90% of claims, tested separately against documents written to break it.

Pipeline test
0
Fabricated claims in this run. No invented source across the 280 checks
Pipeline test
90%+
Source attribution. Claims in a pack traced to a source a reviewer can open
Usability testing
85%+
Workflow completion, intake through to export, without a stall
Usability testing
60-80%
Prep time saved, against assembling one evidence pack by hand

Problem validated and every decision reviewed with three named governance and risk practitioners. See who →

The part I’m still unsure aboutAn honest note

My whole argument is that friction builds trust. I still believe it. But I added three checkpoints to a product whose users are already exhausted and behind. And I designed every one of them for the reviewer at the end.

The person I’m least sure I served is the tired one at the start. The GRC lead opening Qlarc at 6pm, already two weeks late. If I had another month I wouldn’t add a feature. I’d sit behind five of those leads while they hit the gate cold, and count how many quietly close the tab. That count is the thing I’d want to know before I trusted any of this.

✎ Margin note · to self

if i could rerun one study: 5 vendors, no warmup, just drop them at the gate cold. count how many quit before they ever see the payoff.

kind of scared of that number tbh. which is probably why it’s the one to chase.

Next case study

UnitedHealthcare.

Product DesignAI GovernanceHealthcare

AI could process millions of claims. Nobody built accountability for the decisions it made.