Skip to main content
Not Robbie Clark
← All AI work

Research case study · T-Mobile

NORA. Turning root-cause analysis into answers a care agent can use.

NORA delivers findings from T-Mobile’s root-cause analysis (RCA) engine to care agents. It checks a customer’s account, device, and network and reports what’s likely wrong. The findings were accurate — but written for higher-tier care agents, not a frontline agent with a customer on the line.

We’re adding a persona-aware layer that turns NORA’s technical output into fast, easy-to-understand diagnostics and troubleshooting steps.

Role
Senior Experience Designer
Team
Project manager (my side) · 2 software engineers (T-Mobile)
Timeline
May 2026 – present · Phase 2 completed September 2026
Methods
Generative interviews · Contextual inquiry · Personas · Journey mapping · Concept validation · AI output evaluation
Tools
Dovetail · Figma · Figma Make · Claude · OpenAI API
Where it stands
The consumer care agent prompt is built, tuned against the Playbook, and ready for testing in a controlled NORA environment.

The outcome in one line

Before

“The exact root cause of the customer’s issue is that their tethering data usage has exceeded the 5GB allocation available within their subscribed plan (SOC: [code] among others)…”

After

“The customer’s tethering data has reached the plan limit, triggering speed restrictions.”

Same diagnosis, rewritten so an agent can read it and say it mid-call. See how the Playbook rules got it there →

Part 1

User research

Understand the agents, their calls, and what a first response has to do — then write the rules.

01 · The problem

Accurate answers, hard to use.

T-Mobile’s engineers came to us with a clear signal: care agents found NORA’s output too complicated. We wanted to understand why before changing anything.

“The summarization and the next steps and action in the output, it’s still way too technical.”

T-Mobile stakeholder
Original NORA response to a slow data query: five sections of technical findings, signal metrics, site IDs, and engineering-level next steps.
The original output for “customer is having slow data on the device.” Tap to enlarge.
  • Five sectionsof technical findings to read while the customer waits.
  • Signal metrics and site IDsRSRP in dBm, RSRQ, a cell site code.
  • Next steps written for engineers“Check localized coverage… in site monitoring tools.”
  • No customer-ready wordingThe agent still has to translate all of it, live.

02 · Research plan

Start with the decision, not the answer.

The engineers’ feedback told us that agents struggled. The research had to tell us why — and what a better first response would need to contain. One constraint shaped everything: no new UI, no builds. Every improvement had to come from what NORA says, not from changes to the screen.

  1. 01

    Desk research

    Frontline care work, agent cognitive load, and how AI assistance is used in care.

  2. 02

    Stakeholder discovery

    Care SMEs, engineers, and consumer care leads on how NORA is used today.

  3. 03

    Agent interviews

    10 remote 1:1 sessions with frontline agents.

  4. 04

    Observation

    5 sessions watching agents use NORA as they worked.

03 · Interview protocol

Writing questions that don’t lead.

Each session followed the same path — starting from real work, moving toward NORA.

  1. 1

    Walk me through a recent call

    What was difficult? What was easy?

  2. 2

    Expectations

    When you submit a query to NORA, what do you expect to see?

  3. 3

    Translation

    How do you explain the issue to the customer on the phone?

  4. 4

    Escalation

    How do you handle calls where a ticket needs to be created?

  5. 5

    NORA output

    What’s difficult to interpret? What’s easy?

The bias

SMEs had already told us agents found the output hard to read. My first draft took that as fact — some questions were written to confirm it.

AI-assisted vetting

I ran the draft through Claude to check for leading and compound questions, bias, and gaps. It flagged the assumption.

The fix

Balanced questions — what’s difficult and what’s easy — so agents could tell us if NORA worked. T-Mobile SMEs then reviewed the protocol for relevance.

04 · Interviews & observation

Listening first, then watching.

10

remote 1:1 interviews with T0/T1 agents

45–60

minutes per interview

5

one-hour observation sessions

1

tool to record, transcribe, and tag: Dovetail

What observation showed that interviews didn’t

What we saw

Agents asked NORA only once — no follow-up questions.

What it meant

The first response has to do all the work: diagnosis, next steps, and talk track in one pass.

What we saw

After entering caller details, agents typed very basic prompts.

What it meant

NORA can’t rely on a well-written prompt. The output has to be useful no matter how the agent asks.

What we saw

Ticket creation made agents anxious.

What it meant

Escalation needs a clear, confident path: when to file, what to say, what happens next.

05 · Synthesis

From transcripts to principles.

I tagged moments across all 10 transcripts in Dovetail to find recurring themes. AI sped up the clustering, with one rule: for every theme it proposed, I asked for the supporting quotes and timestamps and checked them against the transcripts. No evidence, no theme.

Translation is the work

Agents have to interpret technical signals before they can explain them with confidence.

Live calls demand scannability

A short headline, a clear next step, and speakable wording beat comprehensive detail.

Confidence needs visible limits

Agents need to tell a reliable diagnosis from an informed guess.

Consistency protects the experience

The same root cause should produce the same explanation and action, whichever agent takes the call.

Empathy must survive the handoff

Guidance should acknowledge the customer’s situation, not just report system status.

Seven guiding principles

I wrote seven principles that became the foundation of the NORA Playbook.

Translation is the work

01Speak Plainly

Live calls demand scannability

02Surface Only What Moves Things Forward

Agents asked NORA only once (observation)

03Always Land on a Next Action

Ticket anxiety (observation) + empathy

04Make the Handoff Clean

Consistency protects the experience

05Use Consistent Structure

Live calls demand scannability

06Balance Skimming and Reading

Confidence needs visible limits

07Communicate System Status

Design implication

Every first response pairs a short diagnosis with confidence cues, customer-ready language, and one next best action the agent is allowed to take.

06 · Persona

One persona: the Frontline Consumer Care Agent.

Interviews and observation showed T0 and T1 agents share nearly all the same goals and frustrations. T1 agents take messier, higher-stakes calls, but the core need is the same — so we built one persona, validated by T-Mobile’s SMEs.

Persona card: Frontline Consumer Care Agent, with profile, goals and motivations, general pain points, and tool pain points.

Who they are

Early-career, often in their first 6–12 months. Comfortable with consumer tech, not RF engineers. Ping peers before escalating.

What they want

Resolve on first contact. Keep calls steady. Give clear, confident, consistent answers — without guessing.

What gets in the way

Output that reads like engineering signal. 2–5 possible actions instead of one, so agents freeze or escalate.

The principle we designed by: treat the care agent like a customer.

07 · The Playbook

Turning research into rules an AI can follow.

The Playbook is the single source of truth for everything NORA says to a consumer care agent. I built it as an interactive reference in Figma Make, on top of T-Mobile’s existing Core Communication Standards — Clear and Accurate · Empathetic and Protective · Respectful of Time. Additions and edits went to T-Mobile care SMEs in bi-weekly review meetings, so the rules were checked against real care work as they took shape.

Every NORA output should feel like a briefing from a knowledgeable colleague who knows what you need and respects your time — not a system log.

NORA Playbook interactive reference, Section 4.2 Findings Summary, showing the Conditions tab with a state/condition/content rules table and labeled good and bad examples.
The Findings Summary rules, live in the Playbook. Tap to enlarge, or browse the full NORA Playbook.

Capabilities & guardrails

NORA informs the agent. It doesn’t decide, act, or replace judgment — and it says when it has gaps.

Target persona

The Frontline Consumer Care Agent, as the lens for every rule.

7 core principles

The foundation every section rule builds on.

Audience styles

Headline style for the agent (“Reset network settings”); script style for the customer (“Go ahead and restart your gateway…”).

Controlled vocabulary

Preferred and avoided terms, kept as a living list.

Section rules

Structure, length, tone, and hedging rules for each part of the output, with do/don’t examples.

Tone & personalization

Five tones matched to the customer’s situation.

One structure for every response

1 · Findings

One-line summary, then Account, Device, and Network diagnostics — only what’s relevant.

2 · Next steps

Ranked actions the agent can take, with optional wording for the customer.

3 · Last resort

If nothing works: exactly what ticket to file and what to tell the customer.

One finding, before and after

Before

“The exact root cause of the customer’s issue is that their tethering data usage has exceeded the 5GB allocation available within their subscribed plan (SOC: [code] among others). In alignment with the provisioning policy and SOC restrictions attached to this account, throttling is applied after the tethering threshold is breached…”

Three sentences of policy before the diagnosis. Internal codes.

After

“The customer’s tethering data has reached the plan limit, triggering speed restrictions.”

Leads with the finding. Plain language. One sentence.

From Playbook to prompt. NORA runs on GPT models, so the rules had to become instructions a model could follow. I translated the Playbook into a structured markdown system prompt — principles, section rules, suppression lists, and tone logic.

08 · The pushback

Should agents label a caller’s sentiment by hand?

The proposal

Agents identify the caller’s sentiment and add it to their prompt, so NORA responds in the right tone.

  • Bias. Two agents would label the same caller differently.
  • Discoverability. Nothing in the interface would tell agents the option existed — and under a no-build constraint, nothing could be added.
  • Research quality. Free-form sentiments can’t be compared, tested, or improved.

My solution

Signal-based tone triggers — no manual flagging.

NORA reads signals it already has, like repeat callers, new customers, and time of day…

…maps them to five predefined tones…

BaseEmpatheticDirectReassuringAccountable

…and writes tone-appropriate scripts automatically.

No extra work for agents, no labeling bias, and fixed categories we can test. It became the Playbook’s tone and personalization framework.

Part 2

AI research

Turn the Playbook into a prompt, then find out whether the model actually follows it.

09 · Validating the new outputs

Showing people the new NORA before anything ships.

With nothing to build, I used the system prompt and screenshots of real NORA outputs to generate new example responses in Claude, then shared them with T-Mobile care SMEs and frontline agents — in meetings, by email, and in shared documents.

The reaction was overwhelmingly positive. One issue surfaced: agents were confused when a category came back “Incomplete” even though one of its checks had failed.

The fix: a status precedence rule

✕ Bad → – Incomplete → ✓ Good

If any check fails, the category is Bad — even if others didn’t complete. If nothing failed but something couldn’t be checked, it’s Incomplete. Only a category where everything completed and passed is Good. An agent never sees “Incomplete” hiding a real problem. The rule went into the Playbook, the prompt, and the evaluation harness.

Revised NORA output for a repeat home internet contact: Findings with Good, Bad, and Good category statuses and a one-line summary; eight ranked next steps with empathetic customer scripts; and a Last Resort ticket section.
A revised output generated from the system prompt — home internet, repeat contact, Empathetic tone. Scroll to read.

10 · Evaluation

A rule the model follows three times out of five isn’t a rule yet.

You can’t see that from a single output. So I built an evaluation harness with Claude Code that runs the system prompt through the OpenAI API at scale.

  1. 1

    Scenarios

    Real NORA diagnostic outputs, plus synthetic cases for situations the real ones don’t reach.

  2. 2

    Repeat runs

    Each scenario runs several times through the model.

  3. 3

    Automated checks

    39 machine-verifiable rules: section structure and order, one status per category, summary length, approved scripts only, and leak detection for every account number, device ID, site code, and internal system name in the source data.

  4. 4

    Dashboard

    A plain-language page showing which rules broke, where, and whether they broke every time or only sometimes.

  5. 5

    Tune & re-run

    Adjust the prompt where rules break unpredictably, then run again.

Evaluation dashboard: 20 outputs generated, 19 of 20 followed every rule, 1 of 39 automated checks broken; a grid of runs per scenario with one failure.
Harness dashboard from a recent run. Tuning is ongoing.

What it can’t check

Judgment calls — whether a diagnosis is right, or whether a suggested step makes sense. A clean run means no automated rule broke, not that the output is correct. Those questions still need SME review. Tone scenarios are being added so all five tones are validated.

11 · Learnings & next

Standards only matter if you can measure them.

  1. 1

    Design for the moment of use

    A rep mid-call, with the customer waiting, set the length, order, and tone of every output.

  2. 2

    Check your own assumptions first

    Stakeholders told me the output was hard to read. My protocol had to leave room for agents to tell me otherwise.

  3. 3

    Observation reveals what interviews miss

    “Agents ask only once” never came up in an interview — and it shaped the entire output structure.

  4. 4

    With no build budget, language is the design

    Every improvement came from what NORA says, not what the screen shows.

  5. 5

    Test AI like a system, not a demo

    Consistency across repeated runs shows which rules the model actually follows.

What’s next

Keep refining

Continue tuning the consumer prompt after testing in a controlled NORA environment.

Scale the approach

Apply what we learned — and the materials we built — to a Playbook and prompt for three user groups in T-Mobile for Business.