No. 03DisciplineDistinguished Designer, AI Design2016–2025

Design for AI

The discipline didn’t exist. The job was to write it.

0
Playbooks to inherit
5
Ethics focus areas
200+
Guild members

In 2016 the answer didn’t exist at IBM. The job was to write it.

Challenge. IBM was putting AI into products with no design discipline behind it. No shared principles. No ethics practice a product team could actually use. No methods built for systems that answer differently every time you ask.

Solution. Build the discipline in layers, each one because the last hit its limit. Principles, then ethics, then people and method, then craft deep enough to become product. The last layer, ambient intelligence, is still open.

Impact. IBM’s first practical AI ethics guide, published free. A guild of 200+ practitioners that still meets. Enterprise Design Thinking extended for data and AI. The foundation Carbon for AI stands on, and the ambient intelligence blueprint that points past the chat box.

Three gaps, stacked

Fair warning: every other case in this set ends in a thing you can point at. A product, a system, a room. This one ends in a discipline. The things came later, and this version of the case is about the things.

In 2016, five years after Watson won Jeopardy, “cognitive” was on every IBM roadmap. What wasn’t on any roadmap: what these systems owed the people using them. Three gaps, stacked.

01

No language

Cognitive, AI, machine learning. Three names for the same thing in the same meeting, and nobody could say what any of them should feel like to sit across from.

02

Broken habits

Enterprise software design had spent forty years on a quiet promise: same input, same output. Probabilistic systems ended that promise. Our methods didn’t notice.

03

No practice

The ethics conversation lived in research papers and keynotes. Nothing existed for the designer starting a project on a Tuesday morning.

A title, a sentence, and nothing to inherit

IBM appointed its first three Distinguished Designers in 2016, and I was one of them. My brief fit in a sentence: define what it means to design human experiences for AI. There was no field to join. No canon. No case law. A title, a sentence, and nothing to inherit.

Principles · 2016–2017

I started by writing down what didn’t exist. Defining Cognitive Design, October 2016, version 0.4. Twenty-six pages, released deliberately unfinished, with an open invitation to break it and send back the pieces. It even shipped with my notes-to-self still in the margins, in caps, because waiting for a polished document meant waiting another year the company didn’t have.

What was actually in it. A definition to stand on: design is the purpose, planning, and intention behind an action, and cognitive design is that intent applied to simulated thought. The bet underneath everything: if a system can understand, reason, and learn, computing stops being transactional and becomes a relationship, so the discipline should be built the way relationships are. I still compress that bet into one sentence on stages: we’re entering the relationship era of computing.

A working taxonomy of the roles an AI can play, which I called contexts: spotlight, assistant, coach, companion, ambient. A set of traits a system had to earn: adaptive, stateful, natural, transparent, with humility written into “natural” on purpose, because a system that makes a person feel stupid has already failed.

A borrowed standard of conduct: the machine as a good host, taken from Eames, who took it from Saarinen, all of whose energy goes into anticipating the needs of guests. And a structure for the relationship itself, mapped onto Knapp’s model of how strangers become intimates: initiating, experimenting, intensifying, integrating, bonding.

By 2017 that document had hardened into teachable material. Three principles IBM still recites: augment human intelligence rather than replace it. Earn confidence through transparency. Know enough about people to hold up your end of a relationship. Purpose, value, trust as the foundation sequence, in that order, because trust is earned last. Each word carried a working definition, because a sequence you can’t define is a slogan. Purpose: the reason for a person to engage with the system at all, and it shifts as the user and system learn from each other. Value: what the system provides that tangibly improves their life. Trust: the willingness of a user to place their confidence in the efficacy of the system.

The Knapp stages each got a design move: at initiating, establish the system’s tone, personality, and presence. At experimenting, display provable authenticity. At intensifying, deliver multi-step, context-aware interactions. At integrating, build a fully realized profile shared between system and user. At bonding, the slide just said three question marks, because I refused to pretend we knew.

And one diagram that did years of work, the AI/Human Context Model: intent feeding data, data feeding understanding, understanding feeding reasoning and memory, reasoning feeding output, output wearing an expression, expression provoking a user reaction, reaction feeding learning, learning feeding the outcome. Business on one end, human on the other, machine in the middle. Teams used it to find out which box they’d forgotten to staff.

The fundamentals went public at ibm.com/design/ai, where they still live.

The limit: principles on a page don’t change what ships. Posters don’t argue with roadmaps.

Ethics · 2018–2019

Principles said what good looked like. Nothing said what to do about it on that Tuesday morning. Around then, Stack Overflow surveyed developers and nearly half said the humans building AI are responsible for its ramifications. Not the bosses. Not the middle managers. The people with their hands on it. Those people had no tools.

So with Milena Pribić and Lawrence Humphrey I co-wrote Everyday Ethics for AI, IBM’s first practical AI ethics guidance for product teams. Five focus areas, each with a one-line spine.

  • Accountability: every person involved in the creation of AI, at any step, is accountable for considering its impact in the world.
  • Value alignment: design to the norms and values of your user group, not your own.
  • Explainability: humans should be able to perceive, detect, and understand how the system decides.
  • User data rights: protect the data and preserve the user’s control over it.
  • Fairness: minimize bias and design for inclusive representation.

The structure was the innovation as much as the content. Each area came in four parts: what it means, recommended actions to take, secondary research to consider, and questions to ask your team, so a lead could run any area as a working session instead of a lecture.

One running example threaded the whole guide, a hotel chain building an AI concierge into its rooms, and it kept the abstractions honest: the team interviews real guests face to face, implements a feedback learning loop when the assistant misses, and gives every guest the ability to shut the AI off for their entire stay. The guide pointed forward too, to AI FactSheets out of IBM Research, the nutrition-label idea that later became the factsheet link inside every Carbon for AI explainability popover.

Then I carried it to every stage that would have me, Think in 2019, AI Day the same year, with two companion arguments. More data was coming with less trust attached, so trustworthiness and trust are different jobs: one is engineering, the other is earned. And trust itself has anatomy, four parts borrowed from Dietz and Hartog: ability, empathy, integrity, predictability. Machines could demonstrate the first and the last. The middle two were on us. IBM published the guide free at ibm.biz/everydayethics, and it’s still up.

The limit: a document and its authors don’t scale. Three people can’t sit in every kickoff.

People & method · 2020–2022

If the practice was going to outlive its authors, it needed members. The scaling came in three shapes: a workshop, a guild, and a method.

The workshop. By 2020 the basics were packaged as Team Essentials for AI, a facilitated workshop with a badge behind it. It taught whole product teams the ecosystem model and the purpose-value-trust sequence, but its real cargo was working norms: write before you talk, there are no bad ideas until you converge, laptops away, yes-and. And three tips that did more work than any framework. Start small and be bold. Don’t think of AI as magic. Consider if you should even if you could.

The prep itself eventually compressed into a kitchen metaphor: mise en place. Understand the business need before you start, because “the boss is hot for AI” isn’t a good intent. Align the whole team on a well-articulated intent. Get acquainted with your data and what it actually represents. Get comfortable with hyper-personalization right away. And pre-arrange the regular ethics discussions instead of hoping they happen.

The workshop leaned on one table I’d been drawing for years: humans are good at intuition, emotional judgment, common sense, and creativity. Machines are good at probabilities, volume, repetition, and staying objective. Each is bad at the other’s list. Design the seam accordingly.

The guild. In 2021, Hal Wuertz and I founded the Design for AI Guild inside a single business unit. Around two dozen volunteers. Four hours a week. No budget line. The mechanics were deliberately boring and deliberately durable: a roster on Airtable, a cross-workspace Slack channel, plenaries every other week with a fixed anatomy.

Each plenary opened with signals, a technique borrowed from strategic foresight where a member brings something observed in the world that might mean big change, then project updates, then two fifteen-minute critiques where teams put unfinished work in front of the whole guild. Attendance earned YourLearning credit, so the hours counted somewhere official. The guild even grew a service arm, Design for Good, six-week IBM Service Corps deployments where guild members ran the AI methods with nonprofits.

That first year the guild shipped six foundational projects: an AI skills ladder so managers could hire and grow AI design competence on purpose, an org-wide AI glossary so two teams stopped meaning different things by “model,” a set of AI quality heuristics for evaluating AI experiences the way Nielsen’s heuristics evaluate interfaces, a competitive analysis framework tuned for AI products, a question-based explainability workshop built with IBM Research, and the modules that became Enterprise Design Thinking for Data & AI.

The method. That last one deserves the deepest look, because it’s where the discipline got installed into how IBM already worked. With Mara Pometti and Milena Pribić, I took the method I’d helped scale once before and taught it to ask about data. It runs in two halves. A strategy workshop, two to three days with executives and product owners, sets the intent. A technical workshop about a week later, with the data scientists, designers, and developers, turns that intent into something buildable.

The toolbox holds the activities in four sections plus ethics, and the sequence tells its own story. Set intents: business opportunity statements, a fill-in-the-blank I’ve called the offspring of MadLibs and design thinking. We struggle today because. We need ways to. This will achieve our strategic objectives by. Our worked example, the Clearbridge hotel chain, filled it in like this: we struggle today because we lack the means to proactively address our guests’ needs; we need ways to give guests what they need before they think of it themselves.

Then the team picks from six AI intents, the only six things this technology reliably does for a business: accelerate research and discovery, enrich interactions, anticipate and preempt disruptions, recommend with confidence, scale expertise and learning, detect liabilities and mitigate risk. Then hypotheses, written as bets with a price on waiting: if we knew X, we could Y, which moves this metric by this much, and every week we delay costs us Z.

Identify: information types, a to-be use case map that should end up looking like a pine tree tipped on its side, then paths and components so the map decomposes into things a team can actually schedule. Evaluate: data source identification, where every information type gets traced to a named source and a named steward and colored by whether the data is proprietary, public, or the user’s own. This is the moment most AI ambitions meet their actual inventory. Plan: Hills written as who, what, and wow, an assumptions grid that sorts the guesses by risk, and a 30-60-90 to-do wall with an owner on every sticky.

And riding along the whole way, the ethics exercises, ready whenever the team needs them: layers of effect for forecasting primary, secondary, and tertiary consequences, an inverted behavior model that reverse-engineers BJ Fogg to predict misuse, dichotomy mapping, stakeholder maps asking explainable to whom, consideration cards, a motivation matrix, a 360 review, power mapping. Ethical decision-making moved from a review at the end to exercises inside the workshop, which was the whole point of layer two.

We trained coaches on it through 2022, and three years later the same template was still running inside the AI Accelerator, aligning watsonx teams before they touched a component. The projects were good enough that in 2022 the guild went company-wide. Christopher Noessel stepped up to run it, and roughly 200 new members joined from across IBM’s disciplines and geographies. Members call that era Guild 2.0. Claudia Elbourn runs it today.

The limit: a community spreads literacy. Literacy isn’t craft. Someone still has to know how the thing should behave, sentence by sentence, pixel by pixel.

Craft · 2017–2023

The craft never waited its turn. It ran under every layer above, and by 2023 it had produced a shelf.

Conversation. In 2017, with Zack Causey and Lawrence Humphrey, I wrote Building Conversational Agents, the first hands-on guide for the teams building on Watson Conversation. It was unglamorous on purpose: dialog architecture explained through intents (understanding what the user means) and entities (catching what they said), evaluation, jump-to logic, placeholder syntax, slots, context-aware dialog, and a closing argument called intents and entities the right way. It moved through a Slack channel and a Box folder, which is how craft actually travels inside a company.

By 2019 the conversation work had grown a design theory. Designing for Dialogue opened with a real text thread between Milena and Lawrence, then the same thread with one side deleted: this is what your bot experiences. Conversation got an anatomy: topics provide context, exchanges communicate information, utterances are the individual statements. Exchanges got a taxonomy borrowed from linguistics, greeting and greeting back, question and answer, request and grant or refusal or challenge, complaint and apology or excuse or justification, farewell and farewell.

If that sounds academic, its use was blunt: most chatbots fail because their creators don’t define their purpose, set goals that are too ambitious, and launch before they’re ready. The same three failures the 2020 workshop would later teach whole teams to avoid. Eventually the practice matured into a full conversation design framework led by Bob Moore of IBM Research, built on his Natural Conversation Framework.

Explainability. The 2018 focus area grew into the guild’s most complete craft suite. In-product guidance: the AI label as the single path in, a call-out card with a summarized local explanation one click deep, then two tear-sheet variants for the full story, with a rule to use one or the other, never both. A levels model separating local explanations (why this output), global explanations (how the model works), and XAI documentation (what it was trained on and how it performs).

A self-evaluation on Airtable that scored a product’s explainability across six dimensions and returned an explainability card. An examples library, tagged and searchable by persona. A mental models activity built with IBM Research, plus a lite version, for surfacing what users actually believe the AI is doing. A UX research toolkit with a project plan template, an interview discussion guide, and a survey template tuned for measuring trust and understanding.

And writing guidance developed under IBM Legal’s eye, down to patterned language: say the system can or is designed to, never that it ensures. Say it mainly or likely used these types of data. Never sell. My favorite artifact in the suite explains explainability itself with a BLT: the high-level summary of the sandwich, the mid-level ingredients, the detailed sourcing of the bacon. Progressive disclosure you can taste.

Data literacy. In February 2023, the guild’s literacy thread produced The Friendly Guide to Data Literacy, built with Maria Sánchez Domènech and Andrea Barbarin for everyone who nods along when the word data comes up. It starts by admitting there is no common definition of the word, then hands designers the working anatomy.

Six characteristics: nature (quantitative or qualitative, and how each converts to the other, the way 13 degrees with wind becomes “brisk and chilly”), format (numeric, textual, rich media), shape (structured, unstructured, and the semi-structured middle where email lives), origin (observed, with the observer effect named; volunteered, with social desirability named; inferred, dark clouds meaning rain; and metadata), provenance (the audit trail of where data has been and who touched it), and quality (accuracy, completeness, timeliness, consistency, interconnectedness).

Then five principles to keep everyone honest. Data is an artifact of human behavior. Data codifies the past, so it struggles anywhere history has no precedent. Data is a means to an end. Representative beats big. And data gets collected, used, and disposed of ethically, because we’re all accountable for that one. The 2016 document had called data the fuel and insights the destination. The guide finishes the thought: it’s not the map, it’s not the vehicle, and it can’t be an answer without the right question.

The codification. Five months later, in July 2023, ten of us from across IBM Software compiled the accumulated craft into UX for AI, thirty-five pages of baseline guidance with two jobs: make designers competent AI designers, and give everyone else on the team enough feel for the craft to collaborate without having to master it.

It opens where the guide left off. The systems are probabilistic now, generative variability is the design material, and data literacy is officially part of the designer’s job. Bias gets the plainest metaphor in the book: train on a single teacher with a strong accent and you’ll speak with one too.

The paper defines the AI designer’s responsibilities as a triangle: people, data, systems, with a warning attached to the third, know the technology well enough to stop proposing sci-fi solutions that can’t be built. It insists on prototyping early, Wizard of Oz if needed, LLMs making functional fakes cheap, and it tells you what to validate: purpose, value, trust, usability, interpretability, task efficiency, and the possibility of misuse. It widens who counts as a user to four archetypes: influencers who commission the system, makers who build it, users who act on it, and affected people who never touch it but live with its decisions.

Its core is a six-layer model of AI experience: AI intent, human contexts, AI contexts, AI modalities, AI UX components, all of it wrapped in ethics. Human contexts get spectrums a team can actually locate a user on: happy to sad, indoor to outdoor, alone to with others, flow to frustrated. AI modalities get named and bounded: sidebar, call-out, ambient (display-only, think speedometer), spotlight, and avatar, with a rule I’ll defend forever: an avatar should never attempt humanoid form.

And the ten principles get stated plainly: AI must be human-centered. Simple yet powerful. Design for human perception. Assistive automation. Earn trust with transparency. Always build for explainability. Design for a human in the loop. Feedback is essential, not an afterthought. Informative insights from data. Design for resilience, because the machine will be wrong and the experience has to survive it.

Some of it was new craft. Some of it was the 2016 document, vindicated. The AI contexts (assistant, coach, teacher, partner) descend directly from the ones I’d sketched seven years earlier, and the assistant’s description survives nearly word for word. Ambient is in there too, filed as a modality, waiting for its turn. The paper even asks a question I first typed in October 2016, almost verbatim: if we’ve moved past typing into a field and pressing submit, what does it mean to design this? Both times it landed on page four. Seven years, same question, better answers.

The four contexts earned one-line definitions along the way. An assistant assists, proactively or when invoked, and keeps the bulk of the interaction in the user’s hands. A coach guides people within the actions they take, toward a better outcome. A teacher educates and shows how to get something done properly. A partner learns with the user and highlights what it considers relevant, as if it were co-experiencing the solution.

Seven years, same question, better answers.

The limit: guidance is still prose. Product teams ship components.

Where it landed · 2023–2024

When watsonx shipped and every IBM software team was told to embed generative AI, the discipline was sitting there, waiting. The principles became recognition standards: one universal AI label, localized into every supported language, so a person anywhere on earth can tell what the machine made. The ethics practice became an explainability popover, one click from any AI output, with the factsheet idea from the 2018 guide living at the bottom of it.

The methods became the front half of the AI Accelerator, a five-phase onramp that meets teams where they are, two weeks or six months: kickoff, then alignment through Enterprise Design Thinking for Data & AI, then learning through Carbon for AI and explainability and system prompt writing, then experimentation, then rapid co-design and prep for development. A team entering the Accelerator in 2025 was walking through the layers of this case in order, without knowing it.

The practice became patterns, and the patterns became product. That product is Carbon for AI, and it has its own case. Agentic AI is stress-testing all of it now. That story belongs to Carbon too.

The unfinished layer · 2025

The last layer was the oldest idea in the stack. The 2016 document had named the contexts an AI could occupy: spotlight, assistant, coach, companion, ambient. Ambient meant intelligence the way a room has temperature, present everywhere, pointed at the task instead of at itself. The document guessed ambient would end up the most common form. It took nine years for the technology to catch up to the sentence.

My last work at IBM was the blueprint for that layer. Ambient Intelligence in Enterprise Software, April 2025, with a primer for the portfolio that summer. The cover line carried the thesis: we’ve taught people to adapt to software, and it’s time software returned the favor. A chat box is a vending machine for text. You put language in, it spits product out. The pocket-calculator phase of AI. Ambient is the bet on what comes after: an intelligence layer inside the work rather than a panel beside it.

The blueprint is specific about how. A six-layer stack, sensor to learning, rebuilt for the enterprise, where the sensors aren’t hardware but the digital exhaust already flowing: activity logs, system events, chat and task traffic, presence signals, performance metrics. A context layer that turns that exhaust into answers to real questions: who’s working on what, is this task critical, is a bottleneck forming.

A fusion layer that reconciles what no single system can see, like a ticket marked done with no activity behind it. A decision layer choosing the lowest-friction help. An actuation layer that nudges in Slack, re-routes tickets, reprioritizes a sprint board, holds notifications while someone is heads-down. And a learning layer watching which interventions people ignore, because ignored help is a design verdict.

The governance isn’t an appendix; every layer carries its own requirements, and the 2018 ethics practice is load-bearing all the way down. Consent to observation, with opt-ins split by data type: data about the task, data about how you do the task, data about how everyone does it together. Contestability, so a person can see and correct what the system inferred about them.

Provenance tracking on every synthesized signal. The Carbon for AI frameworks applied to every suggestion and automation. And user agency as a hard floor: accept, modify, reject, or reverse anything the system does. Frustration itself becomes telemetry, read from undo patterns and repeated navigation, which is the ethics practice grown teeth.

Underneath the whole thing sits a humane interaction grammar borrowed from Raskin: a small, universal set of verbs (find, summarize, create, adjust, delegate, notify, schedule, monitor, undo) with modifiers for time, entity, format, priority, condition, and privacy. The LLM interprets whatever a person says; the grammar normalizes it into something predictable, composable, and reversible.

Chain it and it reads like work actually sounds: find last quarter’s invoices, summarize for execs, notify the CFO. Domains get dialects, a datalake speaks ingest, catalog, validate, transform. Undo is a first-class verb everywhere, so a person can predict the machine even when the machine is busy predicting them. The paper closes on Weiser, the most profound technologies are the ones that disappear, and on a day-in-the-life of an analyst whose tools quietly triage, draft, shield her focus, and hand off her shift summary before she asks.

The limit: this layer never hit one. I left before it could. The question came with me.

What changed

Everyday Ethics is still free and still cited. The fundamentals are still public. The method still onboards teams that will never know my name. That was the point.

Credits

Seven years of this belonged to many hands.

  • Milena Pribić · Everyday Ethics co-author; ethical AI practices, start to finish
  • Lawrence Humphrey · Everyday Ethics co-author; conversational agents guides
  • Zack Causey · conversational agents guides
  • Hal Wuertz · Design for AI Guild co-founder
  • Christopher Noessel · guild lead, Guild 2.0
  • Claudia Elbourn · guild lead, 2024 on; explainability content
  • Mara Pometti · Enterprise Design Thinking for Data & AI
  • Vera Liao · question-based explainability workshop, IBM Research
  • Bob Moore · conversation design framework and the Natural Conversation Framework, IBM Research
  • Maria Sánchez Domènech and Andrea Barbarin · The Friendly Guide to Data Literacy
  • UX for AI co-authors · Dawn Ahukanna, Katrina Alcorn, Jillian Byra, Emily DiCesaro, Stefanie Lauria, Mayan Murray, Jason Telner, Adi Veerubhotla
  • The Design for AI Guild · 200+ volunteers who did all of this on top of their day jobs

What still matters

The warning came from this era, and it was never about the machines. The 2016 document bet everything on a machine one day holding up its end of a relationship: understand, reason, learn. The 2017 conversation guides said designing those relationships requires us to consciously understand ourselves before anything else. By Think 2019 it was a slide: to design for this relationship with AI, we need to know ourselves first. On the TED stage I asked it as a question about friendship.

There’s a distinction that keeps this case honest: designing for AI and designing with AI are different jobs. This case is the first one. The rest of my life became the second.

The thing I designed against back then, handing over the reasoning for the convenience, I now live inside, ten hours at a stretch. I wrote about what the warning means from in there: Know Yourself First.