top of page

The Model Dignity Check in Practice: Who Becomes Invisible When We Ship

Jeff Abbott
2 days ago
3 min read

Arjun’s launch was ten days out. His company, a Bangalore health-tech startup, had built an AI intake assistant for a network of clinics — a triage bot that would greet patients through the clinic app, sort symptoms, and route appointments, trimming the morning crush at the front desk. The demo was smooth, the clinic directors were eager, and the engineering burn-down chart pointed cleanly at the ship date.

Then the company’s newest clinical advisor asked, in a review meeting, a question that wasn’t on the chart: “Who did we test this on?” The answer — younger urban patients with smartphones, comfortable in English or Hindi — was defensible and, Arjun realized as he said it aloud, damning. Chapter 9 of the book gives that realization a discipline: the Model Dignity Check — five questions answered in writing before any AI system goes live. Arjun took the launch doc and did something product leads rarely have the standing to do ten days from a date: he made the answers a launch requirement.

Who becomes invisible when we optimize?

The book insists on specific people, not categories, so the team named them:

Mrs. Lakshmi — the composite the clinics know well — seventy-two, feature phone, arrives at 6 a.m. and trusts the queue because the queue has never needed her to read anything. The migrant construction workers whose phone numbers change with each job site. The families where one smartphone serves five patients.

Writing actual profiles, Arjun says, changed the meeting’s temperature; you cannot wave away a person you have just described.

What “normal” is baked into the training data?

Our data assumes one patient per phone, continuous connectivity, and text literacy. It learned from the people who already found the app easy. Every dataset tells a story about who matters, and ours had a very specific protagonist.

How does this perform for our most vulnerable users?

The edge-case tests the checklist demanded — shared phones, interrupted connections, voice-only interaction — produced the finding that reshaped the sprint: the bot handled a dropped connection by restarting triage from the beginning, which for a patient borrowing a neighbor’s phone meant never finishing at all. A failure invisible in every demo, because demos don’t borrow phones.

Can affected humans understand and contest decisions?

The original design buried “speak to a person” three menus deep, on the theory that surfacing it would depress adoption metrics. It now appears at every step, in every supported language, because — as the checklist’s logic runs — opacity breeds distrust, and a triage bot that cannot be questioned will be routed around, then abandoned.

Does this strengthen or erode human agency?

The team’s written answer became a design principle taped to the wall:

The bot suggests; the front desk decides; the queue never closes.

The walk-in path — the 6 a.m. queue that Mrs. Lakshmi trusts — survives untouched, not as a legacy concession but as a designed equal citizen of the system.

The launch slipped three weeks. Arjun expected that conversation with his CEO to be the hard part; instead, the five written answers did the arguing for him — a page of named people and tested failures is a different object than a product manager’s unease. The bot shipped in the new form, adoption climbed anyway, and the five questions are now a standing section of the company’s launch template, run again — as the book prescribes — at every retraining.

The check’s real product, Arjun says, wasn’t the delay or even the fixes. It was the discovery of who had been standing outside the demo all along.

Arjun is a composite — drawn from conversations with product leaders shipping AI into healthcare and public-facing services, with details changed. Mrs. Lakshmi stands for someone in every launch.

Try it yourself

Before your next AI feature goes live, write answers to all five questions — naming specific people, not categories. If any answer troubles you, redesign before deploying. Vague answers hide real problems; that’s not a warning, it’s the test.

 
 
 

Comments


bottom of page