UX case study · AIMSLO
Designed to protect the 20% of monthly revenue at risk on every SLA breachfor a B2B SaaS provider.
Swaayata fixes SLA breaches autonomously. I designed the product and its human-in-the-loop.

- 01The agent doesn’t just alert. It acts, and shows its working in the live feed.
Swaayata
An SLA breach is billed automatically: a 20% credit clause on the monthly invoice. Human response took 41 minutes of scramble.
Lead Product Designer, owning the console, the human-in-the-loop, and the reasoning surfaces.
I designed the three moments a human is actually watching: the breach, the pause for override, and the resolution, with confidence thresholds that stop the agent when it should not act alone.
Breach to corrective action designed down to under 5 minutes, with 70% of repeat incidents resolved without a human.
The agent takes the heavy, predictable work. The team keeps the judgment.
The design thesis
Two terms run through the whole product. An SLA is the promise a company signs. An SLO is the measurable target it has to become. Swaayata reads the first into the second.
Take one. “The payment system will respond in under two seconds, 99.9% of the time” is a sentence in a contract, signed by people who will never get paged. Swaayata reads it into a live target, checked every second and something an agent can act on: p99_latency < 2000ms, success_rate ≥ 99.9%.
Then, on every target it tracks, the agent
When performance slips, it takes corrective action, not just another alert.
AutonomyEach human override becomes a training signal. The model sharpens through RLHF.
FeedbackEvery action is visible and reversible. You step in only when you want to, oversight by choice, not a gate.
AccountabilityAn SLA is a financial contract, not a dashboard reading. Every target carries a credit clause the company has already agreed to pay, so a missed target isn’t counted in downtime. It’s counted in dollars:
lost
And it fires on every breach. Miss enough months and the account that took a nine-month sale to win starts shopping.
So I didn’t design most of it. I designed the three places a human is watching, where an engineer decides whether to let the agent run at all. The rest, it could keep.
“Are we actually meeting our commitments?”
The system answers in color, before a single number is read.
“What is it doing to production, right now?”
Every decision the agent makes is visible, and one click from reversed.
“Why did it make that call?”
The reasoning chain is the artifact, not just the conclusion.
- 01See it.
“Are we actually meeting our commitments?”
The system answers in color, before a single number is read.
- 02Stop it.
“What is it doing to production, right now?”
Every decision the agent makes is visible, and one click from reversed.
- 03Understand it.
“Why did it make that call?”
The reasoning chain is the artifact, not just the conclusion.
“How do you show the health of dozens of SLA promises at once, without overwhelming an engineer?”
The system answers in color, before a single number is read.

Status in color, before any number is read.
Two levels: headline, then drill-down.
Engineers already have every metric they could want. What they lack is a single answer to the question their manager asks every morning: are we meeting our commitments?
So the top of the screen answers that in color, before a number is read. Detail waits below, never competing with the headline.
“How do you make an autonomous AI feel safe enough that an engineer lets it touch production?”
Every decision the agent makes is visible, and one click from reversed.

Every decision logged, with the why.
Override, one click, always there.
The agent restarts containers, runs failover, reallocates memory, on its own. Terrifying to hand a VP unless every decision is visible, reviewable and reversible.
So the agent shows its work, before and after it acts, with override one click away. Every correction is training data, and I made that visible so the work feels worth doing.
rlhf only works when engineers engage carefully. at 3am with something on fire they hit override and move on, not thinking about teaching the model. so the loop only improves in calm conditions, which is exactly when you need it least. i designed for the ideal. the hard edge case is still open.
“How do you show the reasoning behind an AI decision, so an engineer can judge whether it was right?”
The reasoning chain is the artifact, not just the conclusion.

Red marks where the chain breaks.
Promise to target to component to metric.
“Agent restarted Container X” tells you what happened, not why. So I made the reasoning itself the artifact: a promise connects to a target, to a component, to a metric.
Red nodes show exactly where the chain is breaking, and an engineer reads the diagnosis without opening a log or writing a query.
Color carries the answer before any label is read. That consistency is the design language of the product.
The agent takes the heavy, predictable work. The team keeps the judgment.
The whole product is built on the agent being sure. The case I never designed for is the one it has never seen before.
Today it does one of two things, both wrong: it acts confidently on a situation it doesn’t understand, or it escalates with no context and leaves someone to reconstruct it at 3am. The honest version needs an explicit “I’m not confident here” state, where the agent surfaces its own uncertainty and hands off cleanly, with enough context for a person to actually act.
every screen i designed assumes the agent knows what it’s doing, the dashboard, the override, the reasoning graph, all of it. none of them have a state for “i’m not sure.”
that absence is the next thing i’d design.
Qlarc.
Vendors had the AI governance. Buyers required proof of it. The gap between them was killing deals.