← Selected Work

Decant AI Guidance · Fortna / Freelance · 2026

AI Guidance for the Moment of Doubt

My earlier project, Decant Product Design, made the everyday operator workflow faster, but warehouse work rarely goes as planned; counts don't match, labels won't scan, items arrive damaged. In those moments, operators stop, find a supervisor, and wait. I began research on a context-aware AI guidance layer that would help operators make those decisions at their station. I was part of a reduction in workforce during the research phase, so I never saw the final product. This case study shares the service strategy and measurement plan I would have implemented.

Company
Fortna / Freelance
Role
Senior UX Designer
Duration
2 months
Methods
Service Design · Enterprise AI · Service Blueprint · Journey Map · Human-AI Interaction · Metrics Strategy
Fortna decant operator screen with a Level 3 escalation notifying a supervisor of a short
AI Guidance Level 3: Escalation. Operator clicked End Bin, system detected a short, and notified a supervisor

Strategic Recommendation

Place context-aware AI guidance at the exact moments operators hesitate, so exceptions are resolved at the station instead of escalated to a supervisor. Define success metrics before designing screens, and validate in an onsite pilot before scaling.

Target Outcome

Fewer supervisor escalations, less idle time at the station, and faster ramp-up for new operators, without trading away inventory accuracy.

My Role

Senior UX Designer and Researcher

  • Initiated user research into how operators handle decant exceptions.
  • Framed the problem as a service design challenge spanning operators, supervisors, and warehouse systems.
  • Completed the service blueprint, AI guidance principles, and measurement plan using secondary research after my departure.

Define

The “Happy Path” was solved, but doubt remained

My earlier decant redesign cut clicks by 52% and processing time by 36% but those gains lived on the happy path flow. When something unexpected happens, the operator has to leave the flow, find a supervisor, explain the problem, and wait for a decision. The system knows about the item, the order, and the site's rules, but none of that context reaches the operator when they need it.

Research Questions

  • —Which exceptions happen most often, and which cost the most amount of time?
  • —What does an operator do today when they're unsure, and how long does it take?
  • —What would an operator need to see to feel confident deciding on their own?
  • —When should a decision still go to a supervisor?

Secondary Research

Why the moment of doubt is expensive

Warehouses have some of the highest turnover in any industry. In 2021 there was a 49% annual turnover rate reported for the sector.2 A large share of operators are always new, it takes them several weeks to get up to speed, and new operators encounter more frequent moments of doubt during routine tasks. Every exception case that needs a supervisor slows down two people (the operator and supervisor) when completing a warehouse task.4

Errors are also costly. Industry estimates put the cost of a single inventory error between $10 and $250,3 due to downstream effects like shipping, dissatisfied customers, broken service-level agreements (SLA’s).

There is hope on the horizon however; the most encouraging evidence comes from outside of the warehouse industry. In a large study of over 5,000 Customer Support Agents, AI assistance raised productivity by 15% on average, and less experienced employees improved in speed and quality.5 This is the type of enhancement that a warehouse could benefit from.

Annual Turnover

49%

Industry Average (2021)

Error Cost

$10–$250

Single Inventory Error

AI Productivity Gain

15%

For new/less experienced hires

Service Design

Mapping the doubt

I mapped a quantity mismatch, one of the most common decant exceptions, as a service blueprint. Today, the operator's path splits off at the moment of doubt: they stop, locate a lead, and wait. The proposed path keeps them at the station. The system detects the mismatch, pulls the relevant context (SKU data, bin configuration, site procedure), and offers a recommendation the operator can accept or override.

Service blueprint comparing the operator's path today versus with AI guidance across start task, scan item, doubt, get help, and resolve stages, with frontstage, backstage, support, and metrics rows
Service Blueprint: Before and After AI Guidance

Measurement

Designing the scorecard

Because I didn't get to see this project through, I focused on how I would prove it works. The north-star metric is: time to resolve an exception.

  • Efficiency: supervisor escalations per shift, operator idle time at their workstation, units decanted per hour
  • Quality: inventory accuracy, downstream errors traced to decant
  • Trust: guidance acceptance rate, override rate, and whether overrides were right
  • People: time to proficiency for new hires
KPI tree diagram with the north star metric time to resolve an exception and Speed, Quality, Trust, and People branches
KPI Tree: a way to decide if AI guidance is helpful

To judge the pilot, I'd use a weighted scorecard (see KPI Scorecard visual). Each KPI gets a weight from 1 to 3 based on how central it is to the problem. 1 means the KPI is supportive to determining if AI is helpful, 2 means the KPI is important, and 3 means critical.

In addition to the weight, each KPI will get a result value from 0 to 2 based on how it moved against its baseline. 0 means it was worse than the baseline, 1 means no change, and 2 means it improved.

The score is calculated by multiplying the weight (1-3) by the result (0-2). The highest score possible is 36. A team would decide what score is needed to “pass” this assessment. I used 70% as an example, with the caveat that inventory accuracy cannot decrease, no matter the score, prioritizing quality over quantity. I'd set the success bar before launch to eliminate confirmation bias.

Weighted KPI scorecard table listing KPIs by area with weights, blank result and score columns, a row scoring guide, an inventory accuracy pass or fail guardrail, and a 26 of 36 points success bar
KPI Scorecard: a tool to eliminate confirmation bias during pilot evaluation

AI Guidance Principles

Guidance that keeps operators in charge

The biggest risk in AI decision support is that users trust it blindly. We’ve all heard stories of students using AI to write book reports, or executives using AI to create presentations, but they fail to do a sanity check before sharing their work, and get found out in front of their peers. People accepting an AI's incorrect answer without checking, is the most common error in AI-assisted decision-making tools. The article “Guidelines for Human AI-Interactions” outlines 18 AI Design Guidelines1 for user interface design; 4 of them stuck out to me, and I adopted these principles to help design this agent workflow.

G2

Make clear how well the system can do what it can do

Know when to hand off to a user. Example: Low confidence or high-value warehouse items route to a supervisor, with context already attached.

G3

Time services based on context

Notify the user only at the moment of doubt. Avoid pop-up fatigue; guidance appears only when an exception is detected or requested.

G8

Support efficient dismissal

Make recommendations, don't independently decide. The operator always confirms, can override with a reason, or can ignore the recommendation.

G11

Make clear why the system did what it did

Show your work, AI. Every recommendation cites its source, such as the vendor history or site rules.

Fortna decant operator screen with a Level 1 tip about scanning a pack label once instead of each unit
AI Guidance Level 1: Tip. System detected scanning inefficiency
Fortna decant operator screen with a Level 2 recommendation to record a short of two units
AI Guidance Level 2: Recommendation. Operator clicked End Bin and system detected a short
Fortna decant operator screen with a Level 3 escalation notifying a supervisor of a short
AI Guidance Level 3: Escalation. Operator clicked End Bin, system detected a short, and notified a supervisor

What Success Would Look Like

Validating in three phases

1. Baseline (Weeks 1–4)

Track existing decant flows to learn frequency, resolution time, and escalation rates.

2. Pilot (Weeks 5–12)

Turn on guidance at a subset of stations and compare performance against a control group.

3. Scale

Iterate on guidance prompts and roll out across client sites through the configurable product.

Success means operators resolve more exceptions on their own, supervisors spend less time fighting fires, and accuracy holds steady or improves.

Sources

  1. Amershi et al., "Guidelines for Human-AI Interaction," CHI 2019
  2. Eclipse Advantage, “Reduce Warehouse Employee Turnover Before Summer Hits”
  3. Powerhouse AI, “Why inventory inaccuracies can cost hundreds of thousands of dollars per warehouse”
  4. Supply Chain Dive, “Pay is only one piece of the warehouse worker retention puzzle”
  5. World Economic Forum, "Beyond the desk: How AI is transforming the frontline workforce"