UX and product design · Swaayata · 2024

Protecting 20% of a monthly invoiceby making every action an AI agent takes visible and reversible.

Swaayata is an SLO platform for SRE teams. Fixing a breach by hand took 41 minutes, and the clause lands on the month’s invoice either way. I designed the console, the override and the record that let the team allow the agent to act in seconds instead.

swaayata / human-in-the-loop
Swaayata Human in the Loop overview: decision log with confidence levels and the override panel.
  1. The agent doesn’t just alert. It acts, and shows its working in the live feed.
Role
Lead Product Designer
Goal
Make every autonomous action visible, reversible, and explainable.
Sector
Enterprise SaaS · 2024
Timeline
6 months · concept to handoff
TL;DR · 2 min readOr read the full study ↓

Swaayata

Problem

The agent could fix an outage in seconds. Nobody could watch it, stop it, or explain it afterwards.

Role

Lead Product Designer. I owned the console and its four routes, the human-in-the-loop layer, the reasoning graph, and the loading and failure states the build never had.

Approach

Permission, not autonomy. So I designed the three moments a person is watching: a console before the agent acts, an override during, a record for the morning after. No approval queue, because a queue costs the minutes it protects.

Outcome

41 minutes from alert to fix, timed on the team’s own path. Designed down to under five, with 70% of repeat incidents resolved without a human.

41 minAlert to fix, existing pathTimed
<5 minOnce the agent carries itDesign target

The agent takes the heavy, predictable work. The team keeps the judgment.

The design thesis
The Swaayata dashboard on a laptop screen: SLO status, incidents, service health and live metrics in one view.
The problem

The invoice moves faster than the engineers do.

The agent could fix an outage in seconds. But who would watch it, stop it, or explain it afterward?

A service level agreement is the uptime a company signs for. The SLO is the tighter internal target that protects it. Miss it and a credit lands on the next invoice, with no approval and no appeal. The engineers were not slow; they were doing the only thing they were allowed to do, and it took forty-one minutes. Not a monitoring problem. A billing problem.

Time cost
41 min
Her own timed run of the team's alert-to-fix path, before any of this was designed
Measured
Revenue cost
$53K/mo
The credit clause: 20% of a $267K monthly fee, landed on the next invoice
Model

One measured herself, one modelled from the contract terms.

By the time we know which promise it touched, the month is already lost.
Opportunities

The brief was autonomy. The blocker was permission.

The agent worked. Every conversation still ended in the same place. Most of an autonomous system runs with no one watching, and it should. What needs designing is the small number of moments where a person is watching, and deciding whether to keep letting it run.

Which turned one vague problem into three specific ones, in the order an engineer asks them.

See it

“Are we actually meeting our commitments?”

Asked every morning, by someone who is not on call

A console that answers before a number is read, then lets you go down two levels without ever losing the answer. It also has to answer for an engineer who cannot see the colours.

Answered in Decision 01The console
02.1The solution

Swaayata is an SLO platform, not a bigger dashboard.

It sits on the monitoring a team already runs, turning a signed uptime promise into a live target the agent can act inside before a breach reaches an engineer. The three questions above become three coordinated surfaces: a console that answers are we fine, an override that lets the agent act without a queue, and a reasoning graph that survives being asked why. One decision path, built as three screens, because a person needs to trust each moment on its own.

Decision 01 · the console

The first console showed everything. It answered nothing.

Engineers already have every metric they could want. The first build gave them all of it at equal weight, which is a data problem solved and a judgment problem untouched. Status moved to the top, and the detail underneath it opens in place: nobody has to leave the glance to read the history behind it.

A 99% target allows 1% of failure, the error budget: a monthly allowance for downtime. What matters is not how much is spent, but how fast. 60% gone by day 27 of 30 is fine, the month is nearly over. 60% gone by day 2 is not, since at that pace it runs out three weeks early. Past 100% the promise is already broken, and some services below read 150% instead of 100%.

Rung 01DashboardExport · 803px, shown at source
Swaayata dashboard: SLO met, not met and incident counts, SLO compliance, service health and live metrics.
Counts are build data, not production telemetry.
Insight 01

Status first, detail on demand

What it does

Answers the one question a manager asks each morning, in colour, before a number is read. It counts: 1 met, 10 not met, 15 incidents.

What it does not do

Name a single service. A count is not something you can act on.

The decision

The first build put forty metrics here at equal weight and answered nothing. Counting became the only job this screen has.

Rung 02Components overviewRebuilt · a route in the build, no export survives
1Sorted worst first, so the answer is at the top before a colour is read. Take the colour away and the ranking, the words and the markers all still work.
2checkout-api is the selected row, and it is what sets the scope line on the next screen. Without this rung that scope line arrives from nowhere.
Insight 02

Position is the answer

What it does

Ranks. The only surface holding names and targets beside each other, sorted by error budget spent.

What it does not do

Show anything inside a service. No metrics, no history, no reasoning.

The decision

The sort order is the answer. Position carries severity before any colour is read, which is also what makes this screen work for someone who cannot see the difference between the red and the green.

Decision 01, continued · inside the service

A name is not yet a cause.

Now the breach has a name. A name is not yet a cause.

Rung 02 ends with checkout-api at the top of the list and a scope line the next screen inherits. It does not say what inside that service moved, when it moved, or whether anybody acted. The last two rungs answer that, and they are the same route: 03 is the screen as it was built, 04 is that screen with one row opened. The states after them are new work, and labelled as new work.

Rung 03Components detailsExport · shown at the Performance block
Components details, the Performance block: Response Time, Throughput and P99 Latency as closed rows, each with a chevron.Performance block · 3 of 6 rows
Insight 03

A delta is not a history

What it does

Six metrics for one service, each with its current value and its movement over seven days. p95 Latency 24 ms, down 5%.

What it does not do

Tell you when it moved, or by how much, or whether it crossed anything.

The decision

Every row with a target opens. p99 Latency has no target set in Metric, so it does not open, and the row says that rather than failing silently.

Rung 04Components details, openedRebuilt · designed, never exported
Insight 04

Detail arrives in place

What it does

Draws the metric against its own SLO target, in place, with the 03:41 crossing on it and the seven day register underneath.

What it does not do

Navigate anywhere. The row header, the value and the delta all stay on screen.

The decision

Nobody ever loses the glance. I watched someone lose the top of the screen to read a graph, so the graph came to them instead.

The same rowWhen the data is not thereNew for this case study · never designed

Holds the row open when the value has not arrived, or cannot. Same component, same layout, two states it never had. A gap in collection is never drawn as a zero, and an unread metric is never drawn as a healthy one.

Loading · designed for this case study
›Response Time
›
›

Skeleton, not a spinner. Nothing moves, because a pulse that never changes state is decoration. The row name holds position so the page does not jump.

Error · designed for this case study
›Response Time23 ms
Throughput could not be read

The scrape for this target last succeeded at 02:14, 17 minutes ago. This is not a reading of zero, and the agent has not acted on it.

An error has to say what is not known, not just that something broke.Missing data must never read as a healthy zero, because that is the reading that lets a breach pass unnoticed.

Which answers what the screen has to show. It says nothing about who is allowed to act on it without asking first, and that was the harder call.

Decision 02 · the human in the loop

The safe answer costs the thing it protects.

The safe answer was an approval gate. It fails on its own terms.

Human in the Loop: KPI cards, component selector, decision performance chart, the Override panel, and the decisions table with confidence and AI Assisted or Human Assisted pills.
Full screen, as exported. SLO’s, decisions and improvement, human-assisted share and conversion rate as KPI cards; a component selector; the decision performance chart; the Override panel; and the decisions log.
Insight 05

Oversight by choice, not a gate

The obvious answer is a human who approves each action. I designed it, then timed it: nineteen minutes to acknowledge, while the credit clause keeps burning. The gate protects nobody from the thing that actually costs money.

Close-up of the Override panel: the AI Assisted Decision reasoning text and the Add Human Input control.
Cropped from the screen on the left. It states the reason before anyone asks, verbatim from the export, typos included: “The route fall back module was choose for api gate way based on the pervious route choose,” 92% confidence, with the override control one tap away.
A gate that pages someone at 03:41 is a forty-one minute response with a sign-off screen in front of it.

So there is no gate. The agent acts, states its reason before it moves, and logs the decision with its confidence. Every action stays one click from rolled back for twenty-four hours, and the log says so on every entry.

The price: the agent can be wrong in production before anyone sees it. That is the honest cost of a system that acts.

Which means the agent now acts alone, on its own confidence, with no one in the room. The record it leaves behind has to survive being questioned the next morning.

Decision 03 · the reasoning

The derivation is the artefact, not the conclusion.

A log is a record. It is not an explanation.

Entity reasoning graph: SLO circles, component hexagons, goal nodes and metric nodes, with breaching nodes in pink.
Node shape carries the entity type and colour carries the breach, so the diagnosis reads without opening a log.
Insight 06

The chain, with the break marked

A log line tells an engineer what happened, never why. “Agent restarted container X” is a record, not an explanation. They were reconstructing the reasoning by hand at exactly the hour they had least capacity for it.

So the surface became the derivation instead of the log line: the signed clause, the target it became, the component it binds to, and the number that broke it. Here that chain is 99.9% monthly availability, signed, down to p95 under 30 ms, the SLO that protects it, down to p95 = 34 ms, the reading that breached it.

The price: every entity has to be modelled and named before a single node can be drawn. A wrong parse still draws a clean graph, and nothing on the screen says so.

See it, stop it, understand it: three decisions, one thread. What that thread actually shipped, in six months, is below.

Outcomes

Six months, concept to handoff.

The same path, agent carried
<5 min
the design target, down from 41 minutes timed
Design target
Resolved without a human
70%
of standard, repeating incidents, designed for
Design target
At risk on a missed month
20%
of the monthly invoice, caught before the clause fires
Contract term

The agent takes the heavy, predictable work. The team keeps the judgment. I designed the part that lets an engineer believe that at 03:41, and stop it if they do not.

07 / What I'd still closeAn honest note

Every screen assumes the agent has a confident read. None of them has a state for the one it does not.

When the agent hits a pattern it has never seen, it either acts confidently on something it does not understand, or it escalates with no context, leaving someone to reconstruct the incident at 3am. And the override that is meant to become a training signal only fires when nothing is burning: at 03:41, with a real breach live, an engineer hits override and moves on, so the loop sharpens in exactly the calm conditions where it matters least.

✎ Margin note · to self

i can describe both gaps precisely. i have not designed the answer to either.

Next case study

Qlarc.

Systems DesignStrategyAI GovernanceB2B SaaS

Vendors had the AI governance. Buyers required proof of it. The gap between them was killing deals.