The instrument comes first
What two decades of platform research taught me about trust in automated systems
In an experiment late in my Pinterest years, a number moved the wrong way. A navigation change had produced a dip in campaign creation for one segment. On most teams that number triggers a fight: one camp reads it as proof the change failed, another reads it as noise, and the loudest reader wins.
We did not have that fight. Before the experiment ran, we had designed a three-tier evaluation model that assigned a meaning to every signal in advance.
Tier one was the direct signal: did navigation opens drop.
Tier two was the guardrail: did page visits hold.
Tier three was the real question: did downstream engagement move.
The dip landed in a place the model had already mapped. It read as a discoverability problem, and it produced a mitigation instead of a rollback.
The lesson was not that the mitigation worked, although it did. The decision had been made months earlier, when we designed the measurement. By the time the data arrived, there was nothing left to argue about.
I have spent most of two decades on platforms where a professional has to interpret what a system tells them and act on it under pressure. Advertising platforms at Meta and Pinterest. An operations platform at Opendoor. Now an AI platform where analysts make consequential decisions against a clock. Different domains, same finding every time: the instrument you design before the work determines the quality of every decision after it.
Here is the failure mode the instrument exists to prevent.
Automated systems are good at looking confident. A recommendation arrives with a percentage attached. A detection lands on a map with a pin. A summary reads clean and complete. Polish implies rigor, and people calibrate their trust to the polish.
But polish can lie. A pin with no citation. A confident summary drawn from a window of time the person did not intend. A result set that is silently thin because coverage was missing, presented with the same visual certainty as a result set that is rich. I call this certainty theater: the interface performs a confidence the system has not earned.
The cost is not that people get fooled once. The cost is that professionals, who are not fools, get burned once and then discount everything the system says afterward. Trust in automated systems does not degrade gracefully. It cliffs.
So the design question underneath every human-AI collaboration problem I have worked on is less how to get people to trust the system, and more how to make the system show its work honestly enough that trust, once given, survives contact with reality.
What that looks like in practice, six ways.
Verification designed in, not bolted on. Inside an ads platform, I helped build a framework for automated recommendations. The design decision that mattered most was not what advice to show. It was separating commercial prompts from genuine performance advice, tying each recommendation to the predictive model that generated it, and instrumenting the window after a person adopted a recommendation so they could verify it actually worked, against an industry baseline rather than our own marketing. The system asked to be checked. Adoption followed.
Meaning assigned before data arrives. The three-tier model above. The general principle: any metric that gets its meaning assigned after the result is in will mean whatever the most senior person in the room needs it to mean. Pre-committing to interpretation is the difference between an experiment and a Rorschach test.
Percentages replaced with honest tiers. A result arrives wearing a confidence score. 87. What does 87 mean, measured how, against what. I have yet to meet the person under deadline who can calibrate that number, and a number a person cannot calibrate is worse than no number at all. So I design toward semantic tiers traced to corroborating sources: enough to act on, nothing to perform with. The product is decision support, not the decision, and the interface has to know the difference.
Research that outlives the readout. The standard failure of research in product organizations is that a study gets read once, one actionable item survives, and the rest evaporates. On my current engagement I took the studies and converted them into archetype-level checkpoints: gates a design proposal gets tested against before anything is built. The research keeps working after the meeting ends.
Segments kept separate. In two information architecture studies, we segmented enterprise, small business, and creators, and analyzed their mental models separately: card sorts, clustering, then tree tests to validate the drafted structure. The finding that mattered was that the segments disagree. Averaging them produces an information architecture that belongs to nobody. The same is true of any consolidation effort on any platform: unify the chrome without a job lens and you quietly break the thing one segment depended on. The polished, unified interface is often the certainty theater version of design itself.
The instrument for what people will not say. Behavioral analysis across thousands of enterprise advertiser accounts told us what people did. It could not tell us what they were trying to do. And the platform had no qualitative channel for the people actually doing the work, so we built its first in-feed sentiment instrument, placed where people worked instead of where surveys live. Instruments measure outcomes, and they also hear the things your dashboards are structurally deaf to.
Every one of these is the same move. Before the build, before the launch, before the consolidation, someone has to design the thing that will tell you the truth afterward. The evaluation model. The verification window. The honest tier. The checkpoint. The segment boundary. The listening instrument.
That work is invisible when it succeeds, which is why organizations underinvest in it. Nobody celebrates the argument that did not happen because the measurement had already settled it. But it compounds. A team that pre-commits to interpretation stops relitigating results. A system that asks to be verified earns adoption that survives its first mistake. A research program that converts into standing checkpoints keeps making decisions long after the researcher has moved on.
The field is starting to call this discipline AI evaluation, or human-AI collaboration design, or trust calibration. The names are newer than the practice. It is the oldest rule in empirical work, applied to interfaces: decide how you will know before you look.
The first product I ever helped ship was CNN’s first iPhone application. It launched in 2009 and won Macworld’s best news application award. And I cannot tell you what it did. Not what it did to readership, not what it did to revenue, not whether the things we believed about it were true. Nothing was designed to know.
Nobody flagged it, including me. The discipline barely existed in newsrooms then, so the absence never registered as an absence. The award is the only measurement the launch ever got, and awards measure juries.
It stayed that way for three years. Then the same app shipped an iPhone 5 redesign with real numbers attached, a 7 percent lift in ad revenue, a 25 percent increase in article read-through. Somebody had designed to know. The first launch succeeded, as far as anyone can say. That phrase, as far as anyone can say, is the cost.
Twenty years in, I have never once regretted building the instrument first. I have regretted every launch where we skipped it.



