Xerox PARC (Palo Alto Research Center) researchers invented the personal computer (the Alto), the GUI, Ethernet, object-oriented programming, and ubiquitous computing—foundational technologies the world runs on. Recent breakthroughs span human-machine collaboration, distributed autonomy, advanced sensing, and bio-inspired devices.
My team was building a virtual customer care agent that could understand and respond to natural language beyond keyword matching. To train it, we needed to annotate thousands of real customer-agent conversations at the semantic level — labeling speech acts, identifying turn structure, tracking sentiment across time, aligning audio to transcript. No tool existed for this. Most annotation environments were built for simple text tagging, not for the multi-layered conversational analysis our team required.
Working with a linguist and a research engineer, I designed CAAP from scratch — a developer-facing NLP tooling environment that became the core research infrastructure for the AI team's model development work.
- Audio/transcript synchronization interface with waveform visualization and cursor-linked navigation — analysts could select a turn in the audio and see the corresponding text, or vice versa
- Annotation layer system with per-turn speech act labeling (greeting, instruction, clarification, question, acknowledgement, apology, escalation, and more) — taxonomy developed alongside linguists through iterative field research
- Multi-annotator workflow supporting parallel annotation of the same sessions for inter-rater reliability measurement, essential for training data quality
- Sentiment and emotion tracking across conversation timelines — emotional valence visualized as a signal layer beneath the transcript
- Model iteration workflow connecting annotation output to the training pipeline, so researchers could close the loop between annotated data and model behavior
The core question wasn't interface design — it was: what does a conversational AI system need to understand about human interaction in order to respond appropriately? I co-led the field research that answered this: ethnographic observation at two contact centers, analysis of 1,000 audio files and 200+ chat sessions, and formal conversation analysis grounded in CA methodology.
The key finding was that chat interaction and spoken conversation are not the same behavioral substrate. Turn-taking, repair, and emotion surface differently across mediums. The implications for an AI conversational system were significant — a model trained on chat data would not transfer cleanly to voice. This research shaped both the annotation taxonomy and what the engineering team decided to build first. It was presented at the MOOD-Y academic conference in 2014 and covered by the Wall Street Journal.
The annotation taxonomy, the speech act labeling scheme, and the decision to treat chat and voice as fundamentally different systems — all came directly from my empirical analysis of 1,000 calls and 200+ chat sessions.
Designing how a model should interpret a turn, when it should ask for clarification rather than guess, what constitutes a complete vs. incomplete utterance — these are active questions in LLM-era product design. I was working on structurally identical problems in 2013, grounded in empirical data from real human conversations, before conventions existed.
WDS, a Xerox company, wanted to build a virtual customer care agent that could handle inbound mobile device support — not by scripting responses to anticipated questions, but by genuinely understanding what the customer was asking and reasoning toward a solution. This was 2013. The commercial landscape was IVR trees and keyword bots. The goal was an AI system that could learn from human agents, adapt to individual customers, and know when to hand off.
My role spanned ethnographic research, behavioral architecture, and interaction design — defining both the system's conversational logic and the interfaces that allowed human agents to monitor and intervene.
- Conversation regulation system: The logic for how Xi moved between automated response, human assistance, and human fallback. The Customer Experience Regulator I defined monitored whether the customer had asked for a human, whether too many turns had occurred without resolution, and whether the emotion analyzer detected sustained frustration
- Human monitoring interface: Dashboard allowing agents to observe multiple live Xi conversations simultaneously — current state, sentiment trajectory, problem status — and to inject a response or take over without disrupting the interaction
- Adaptive dialog management: With the NLP team, defined how the dialog manager would decide next steps: invoke NLG for more information, route to the solution finder, or escalate. This decision architecture became the engineering specification
- Customer modeling layer: Defined what data the system would hold about each customer (interaction history, device, brand relationship, language-pattern indicators) and how that model shaped Xi's individualized responses
- Personalized response generation: Worked on how NLG would produce responses that felt natural rather than canned — informed by field research on what made real human agent responses feel genuine vs. stilted
The interesting challenge wasn't the happy path — it was edge cases. What does the system do when the customer's utterance is ambiguous? When they express frustration not at Xi but at the organization? When they want a human but haven't said so explicitly?
Field research at two contact centers gave me the empirical basis for these decisions. Watching real agents navigate exactly these situations — the micro-decisions about when to redirect, when to clarify, when to acknowledge emotion before solving the problem — those patterns became the specification for Xi's conversational logic.
My white paper that launched the Xi project — defining the system's behavioral architecture, the ethnographic research agenda, and the design constraints that shaped every subsequent product decision.
My field research behind Xi's escalation logic, emotion detection thresholds, and the decision to treat chat and voice channels as requiring separate behavioral architectures.
This is work most directly analogous to what a Model UX Designer does: defining the behavioral envelope of a conversational AI system — when it acts, when it defers, how it handles ambiguity, how it signals uncertainty — and translating that into both a product users experience and a specification engineers build from.
P&G's Olay brand wanted to build a skin consultation product powered by computer vision and machine learning — a system that could analyze a selfie, estimate skin age relative to chronological age, identify specific concerns by facial region, and recommend a personalized product regimen. The technology was real. The design problem was trust.
Consumers would share facial images, receive a numerical skin age, and get recommendations from an algorithm. Every one of those interactions had the potential to feel cold, inaccurate, or invasive — or useful, credible, and empowering. The difference was in how the system communicated.
The underlying consumer problem was clutter. Faced with too many products and little scientific skin care knowledge, women experimented, wasted money, and got results they didn't want. The question PARC and Olay set out to answer was how to take the knowledge behind a lab bench or a beauty counter and apply it to each woman's own face.
I led design and ethnography on a team of computer vision and machine learning scientists — as the non-technical contributor, responsible for concept development and validation, interaction design, usability testing, and product and project management.
We ran a two-week field study in Shanghai with ten frequent buyers of beauty products, P&G's SK-II audience: in-home interviews, participants demonstrating their facial regimens at the sink, shop-alongs at department store counters, and follow-up interviews. The probe was a Wizard-of-Oz prototype — a realistic app with a hidden human simulating the system's analysis — so we could test how the AI should behave with real people before committing to what the model had to do.
We went in with written hypotheses about how a personalized advisor should act. Several were wrong, and the wrong ones shaped the product most.
- Where the output comes from. We assumed questions asked up front wouldn't feel intrusive if they guided the image analysis. They did. Participants doubted the technology and suspected the recommendations came from their answers rather than their face. Analysis had to come first and visibly drive the result. One participant put it plainly: start from my face, with dots or links on it, so I know what you're suggesting for me.
- How the output is organized. We planned to structure results by part of the face. Participants preferred results organized by concern — by what the system could actually sense.
- Consistency. If her skin didn't change over a week, identical results felt like nothing was happening. Results that changed anyway felt untrustworthy, and raised the suspicion that the app wanted her buying new products. Output had to hold steady where the signal was steady and still give her something new to learn.
- Familiarity. The relationship had to develop from acquaintance to friend to confidante. Immediate familiarity read as insincere. It should remember and recognize her between visits, and earn its way to using what it knew.
- Control. Women balancing work and family didn't want an app dictating a daily selfie. Evening "me time" felt like the app caring; a rushed-morning prompt felt like pressure. She decides what to share and with whom, and didn't want to be compared with her friends.
- Whose interests. Regimens that mixed in products she already owned, including other brands, were more credible than Olay-only regimens. Ignoring her budget, or treating a purchase as the end of the relationship, cost trust immediately.
These findings became the behavioral specification: image analysis first and shown on her face, output organized around sensed concerns, a soft sell, and concrete value she could walk away with whether or not she bought anything.
- Capture flow: Guided selfie experience with framing guidance, lighting cues, face positioning feedback, and quality scoring surfaced as a percentage. PARC's computer vision scientists built the software controlling lighting, camera distance, and facial expression; the design goal was an image good for the model rather than a flattering selfie, without the interaction feeling like a test
- Annotation requirements: Worked with Olay and PARC's ML scientists on the annotation requirements for the image datasets being built. The models learned to detect target skin features from human-graded images used as ground truth, so what got labeled determined what the system could recognize as valuable to consumers
- AI output presentation: How the system communicated skin age diagnosis, concern identification by facial region (forehead, crowsfeet, cheeks, chin), severity indicators, and explanations. Core challenge: presenting a machine's assessment of someone's face in a way that feels informative rather than judgmental
- Trust architecture: Field research surfaced a consistent finding — users wanted to feel in control, not evaluated. This shaped framing results as "here's what we see" rather than "here's what you are," showing the evidence behind conclusions, giving context rather than just a number
- Personalization logic: Translated skin analysis into product recommendations, explaining the connection between concerns and ingredients so users understood why they were seeing specific products
- Adaptive learning signal: When users confirmed their real age, that feedback directly improved subsequent predictions — a loop where using the product made it better
- Prototypes of increasing fidelity: From the Wizard-of-Oz probe to "looks-like" and "works-like" prototypes, shown to Olay customers throughout, so validation continued while the platform was being built
The launch version combined three kinds of AI — computer vision, deep learning, and adaptive recommendations — in a web platform that diagnosed skin from a selfie, recommended product and regimen changes, and was designed so repeated use improved the results and deepened the relationship.
A personalized AI product at 5M+ scale surfaces every edge case eventually. What does the system say when image quality is insufficient for reliable analysis? How does it communicate uncertainty without undermining credibility? How does it handle skin conditions outside its training distribution? How does it avoid outputs that could feel discriminatory around age, skin tone, or beauty norms?
The age-confirmation loop had its own version of the problem. It only improved predictions if people answered honestly, and a result that reads as "you look five years older" is exactly when someone might report a younger age instead. Collecting feedback from users was a design problem in its own right, not a form field.
These questions were worked through over three years of iteration with the P&G team and PARC researchers. The answers were embedded in the language, framing, and confidence calibration of every screen. Forbes' coverage focused on the combination of privacy and personalization, and gave a section to trust running both ways: consumers trusting the system with their photos, and the system relying on what consumers told it.
- About 5 million visitors globally, with each visit improving the system for subsequent users. More than 2 million unique regimen combinations; 94% of women who tried it agreed the recommended products were right for them.
- Engagement and conversion. Compared with regular Olay.com visitors, Skin Advisor users converted at twice the rate, checked out with 40% larger baskets, bounced three times less often, and spent four times as long on the site. The platform launched with local versions in ten countries.
- Behavioral data changed the product roadmap. Usage showed a large share of consumers wanted fragrance-free products, which weren't planned. The fragrance-free versions P&G then released performed as well as the originals.
- Built on existing research. Repurposing and refining computer vision and ML work PARC had already developed kept development cost down and delivered under budget.
- Data practices. The work is credited with addressing bias in the training data and with establishing a Consented Data Promise for how consumers' images and data were used.
"Working with PARC, we were able to utilize machine learning to develop a platform that can both inform and delight Olay customers." — Dr. Frauke Neuser, Associate Director Scientific Communications, Olay, Procter & Gamble
My ethnographic field research that surfaced the trust gap between what users expected a personalized AI to know and what it could actually deliver — and directly shaped the confidence calibration, uncertainty language, and evidence framing built into Skin Advisor.
The questions from Shanghai are the questions personalized assistants face now: how quickly a system should act on what it knows about someone, whether people can see where an output came from, what stays stable across sessions, and what the user controls about their own context and memory. The method carries over too — simulate the behavior, put it in front of real people, and write down what they expect before specifying what the model has to do. I did this for 5M+ users, with real consequences for how people felt and what they bought.
TIBCO's ActiveMatrix platform managed SOA infrastructure for large enterprises — thousands of interconnected web services exchanging data across organizations. When something went wrong, the consequences were significant and the cause was hard to find. TIBCO wanted to add a machine learning layer that could identify anomalous patterns, predict failures before they occurred, and trigger automated responses — without requiring operations teams to manually monitor thousands of service metrics.
My initial brief on day one: a spreadsheet listing hundreds of web services and their performance metrics. No design brief. No existing product to reference. A technical domain I knew nothing about; Voilà! The challenge I'd been looking for.
- Role-based dashboard architecture: Separate views for Business Analysts (SLA compliance and business impact), System Architects (service relationships and dependency mapping), and IT Administrators (operational health and alerts) — each surfacing the same ML data through a lens appropriate to that user's decisions
- Anomaly and alert visualization: Visual language for distinguishing normal variation from statistically anomalous behavior — response time trends, throughput charts, and hit-rate tracking that gave users enough context to act without requiring them to understand the model
- Rule and objective architecture: Interaction model for a nested policy engine (Rule Packages, Objective Groups, Rules) defining how the system should respond to different conditions — designing the UI for non-ML-experts to configure and modify automated behavioral rules
- Trigger and condition design: Visualization of when rules fired, what conditions caused them, and what automated responses they invoked — so users understood what the system had done on their behalf and why
The central challenge was the human-automation relationship. The system was making decisions autonomously — re-provisioning resources, triggering recovery actions, flagging services for attention. Users needed to trust those decisions enough to let the system act, but retain enough oversight to catch errors. Too much automation and users lost confidence; too much manual intervention and the product had no value.
I designed the continuum between automated action and human intervention — what the system could do on its own, what required confirmation, what surfaced as an alert for human judgment. The same design problem that appears in every agentic AI system.
Designing adaptive, data-driven behavioral systems for technical users, in a domain learned from scratch, with a small team and no precedent — that's the shape of early-stage AI product work. The design principles developed here (role-based information architecture, calibrated confidence communication, explicit automation/oversight boundaries) apply directly to model UX design.
Inspired by a predictive sentiment product I designed for Xerox call centers, I've been experimenting with ML models that monitor and analyze speech and language dynamics, identify behavior and mood patterns, and surface meaningful observations that support self-reflection. The goal is a digital tool that gently helps people with BPD, ADHD, and PTSD develop awareness and self-regulation. A family member is the first user.
I started by researching mood, behavior, thought patterns, and daily functioning, and by approximating those processes with speech and biometric data captured from connected devices. Prototyping followed, with a gradual commitment to the technologies, architecture, and interfaces needed for a platform that integrates language, physiological signals, contextual memory, and longitudinal analysis.
- Live speech capture and Apple Watch biometric integration
- Structured, modular observation architecture
- Longitudinal memory
- Explainable AI reasoning with confidence modeling
- Human feedback loops
- AI-generated reflections
Each prototype answered a different product question:
- What should AI observe automatically?
- What information should remain invisible?
- When does AI provide insight instead of simply analysis?
- Which observations build trust?
- What explanations increase confidence?
- How should uncertainty be communicated?
Technical feasibility initially appeared within reach, and the system signals offered promising insights. Reliable, repeatable capture, analysis, and diagnosis proved elusive: open-source ASR tools and device OS limitations required substantial technical investment and expertise. What came out of the work is an architecture that deliberately separates raw observations from higher-level interpretation, so future models and workflows can evolve without redesigning the underlying system.
Prototype and user testing raised a larger question: can real-world sensor data and clinical data be combined to support predictive modeling and earlier diagnosis, helping mitigate disease and optimize personalized medicine?
Before the PARC work, and before "ML product design" was a job category, the same underlying problems appeared at a different scale. At Netscape and AOL I designed personalization systems — My Netscape, AOL's personalization services, recommendation and collaborative filtering systems — that served more than 100 million users globally.
Collaborative filtering in 2001 raised the same design questions that show up in every adaptive AI product today: how do you surface recommendations that feel personally relevant without feeling surveilled? How do you communicate that the system is learning from behavior without making that learning feel intrusive? How do you handle the cold-start problem — what does the system show someone it knows nothing about yet? How do you design for the gap between what users say they want and what their behavior reveals they actually engage with?
The Liberty Alliance federated identity work from the same period was early permission and consent architecture — defining how users controlled what systems knew about them and what those systems could share across platforms. The same design problem that shows up today in every agentic AI system that needs to act on a user's behalf.
The design problems in collaborative filtering (2001), conversational AI (2013), and today's LLM products are structurally the same: how does a system that learns from behavior communicate what it knows, express appropriate uncertainty, and maintain user trust across that loop? The context changes. The questions don't.
The two Xerox PARC papers are primary research, written during that work and grounding the design decisions in the case studies above. The Liberty Alliance specification is an industry standard I contributed to at AOL, and the personalization work that follows it is where the same questions played out in product. Shorter essays are on the Thoughts page.
Drawing on conversation analysis of 1,000 audio calls and 200+ live chat sessions collected at two mobile device support contact centers, this paper argues that chat-based AI systems require fundamentally different behavioral architecture than spoken dialogue systems — not simply the addition of speech I/O components.
The analysis covers four structural differences that have direct implications for AI system design: asynchronous turn-taking breaks the real-time "proof procedure" that allows interlocutors to verify mutual understanding; repair mechanisms that work in speech arrive too late or misfire in chat; emotion accumulates silently on one side of a chat interaction in ways that have no spoken analog; and the reduced moral accountability of text produces interaction dynamics that spoken conversation's social norms prevent.
The conclusion: moving from a chat AI system to a spoken dialogue system is not an engineering problem — it's an architectural one. This work directly informed the behavioral specification for the Xi conversational agent.
The white paper that launched the Xi research project at Xerox PARC (2013–2016). Written to define the system's scope, technical architecture, and research agenda for a team of 10–16 researchers across PARC's centers.
The document defines Xi's full behavioral architecture: Natural Language Understanding (utterance parsing, slot-filling, knowledge frame population), Dialog Management (next-step decision logic, solution finding, disambiguation), a Customer Experience Regulator (escalation thresholds based on turn count, explicit human requests, and detected frustration), an Emotion Analyzer, Natural Language Generation, a Customer Modeler (individualized interaction based on stored customer profiles), and the human assistance and human fallback systems that govern when Xi passes control to an agent.
The paper also defines the ethnographic research agenda that drove the design: field work at contact centers to understand human agent workflows, conversational analysis to understand what "normal" interaction looks like so Xi could approximate it, and customer motivation research to understand why people call rather than self-serving — and what a system responsive to those motivations would look like.
This document represents the earliest stage of what became a deployed conversational AI system — the point where research questions, design constraints, and system architecture were first made explicit.
Best-practice guidance from the Liberty Alliance Project, a consortium of more than 150 companies writing open specifications for federated network identity: single sign-on across services and permissions-based sharing of personal attributes, work that later fed into SAML 2.0. I contributed from AOL, defining the user experience for authentication, identity, and privacy and bringing a human-interaction perspective to a largely legal and technical working group.
The document maps Liberty's specifications onto fair information principles and national privacy law, then sets recommendations for implementing companies: clear notice of who collects what and why; real choice and consent before attributes are shared; access and correction for the individual; use limited to the purpose consented to; retention only as long as needed; and a way to resolve complaints. It closes with security guidance on browser and protocol vulnerabilities.
The same questions (what a system may know about a person, who decides, and how that consent is shown) now sit at the center of AI agents that act on a user's behalf.
Personalization, recommendation, and collaborative-filtering systems at Netscape and AOL, including My Netscape, serving more than 100 million users.
The consent principles in the Liberty specification became practical design questions here: how to surface recommendations that feel relevant without feeling surveilled, how to show that a system is learning from behavior, what to show someone the system knows nothing about yet, and how to handle the gap between what people say they want and what they actually engage with.
Most designers who work on conversational AI products don't have primary research publications grounding that work. These papers show that the design decisions in the Xi case study weren't intuitions — they were conclusions from empirical study. The Chat vs. Talk analysis in particular is directly applicable to current LLM product design: questions about turn structure, when a model should wait vs. respond, how uncertainty surfaces differently across interaction modes, and what "repair" looks like in AI-mediated conversation are all live problems in 2026 that this research addressed in 2014.
Three findings map directly onto model responses. A customer asking a simple question got a stock "sorry to hear this is giving you trouble" reply, and the exchange immediately read as a bot — the same failure as boilerplate sympathy that doesn't fit the turn. When one side packed several points into a single turn, the other had to guess which to answer, and misunderstandings surfaced twelve minutes later, with nothing to repair them earlier; long, multi-part responses carry the same risk. And voice needs machinery chat never did: tracking a turn as it unfolds, coming in at the first possible completion, giving continuers while someone tells a story, and recognizing when the other person is starting to close.