When AI makes the call, someone still has to own the outcome. That someone is you.
There’s a version of the near AI future that product teams love to pitch: autonomous agents that handle customer interactions end-to-end, models that make real-time decisions faster and more accurately than any human could, systems that learn and adapt without requiring constant supervision. Less toil, more scale, better outcomes.
It’s a compelling vision. It’s also one that quietly transfers an enormous amount of moral weight onto the people who build and ship those systems – weight that most product organizations aren’t yet structured to carry.
The question is simple to state and hard to answer: when your AI agent makes a consequential decision – and gets it wrong – who is responsible?
Not in the abstract. Not “the team.” A name. A role. An accountability structure. Because the answer to that question, and whether it was answered before the system shipped, is increasingly the difference between product leadership and product negligence.
The Accountability Vacuum
Most AI deployments today operate inside a structural accountability gap. The model is owned by data science or acquired externally. The feature is owned by product. The business outcome is owned by the VP of Growth. The legal exposure is owned by Legal. Nobody owns the ethical consequence of what the system actually does to the people it affects.
This is not cynicism – it’s an organizational reality that emerges naturally from how most tech companies are structured. Accountability is assigned to outputs (the model accuracy, the feature adoption, the revenue impact) rather than to outcomes (what happens to the specific people on the receiving end of the system’s decisions).
When that system is a recommendation engine tweaking click-through rates, the gap is manageable. When it’s an AI agent declining loan applications, filtering job candidates, adjusting insurance premiums, or autonomously managing customer relationships – the gap becomes a liability. Moral, reputational, and eventually legal.
The uncomfortable truth is that when accountability is undefined, it defaults to nobody. And when nobody is accountable, the incentive to get it right is structurally weakened. Teams optimize for what they’re measured on. If ethical outcomes aren’t in the measurement framework, they won’t reliably appear in the output. And this is only the organizational view. The legal perspective makes it even dangerously more complex.
This isn’t a problem that ethical intentions solve. It’s a problem that requires explicit design – in how teams are structured, in what gets reviewed before launch, in who has the standing to say “we don’t ship this yet.”
That person, in a well-functioning product organization, should be the Product Leader.
The Collingridge Problem
There’s a concept from technology governance that every Product Leader shipping AI features should understand cold: the Collingridge Dilemma.
The political scientist David Collingridge articulated it in 1980, and it hasn’t aged a day: the impacts of a new technology cannot be reliably predicted until it has been widely adopted – but once it has been widely adopted, changing or controlling it becomes extremely difficult. You need scale to understand consequences. But scale makes course-correction expensive, slow, and politically painful.
We have already lived through the most instructive case study of this dilemma in the history of consumer technology: social media.
The original promise was genuinely compelling. A global network that lets you stay in contact with friends who moved abroad, reconnect with people you’d lost touch with, maintain relationships across distance and time zones. That promise was real – it delivered on it. In the early years, the metrics told a clean story: more connections, more sharing, more engagement. The product was working.
What the metrics couldn’t see – what nobody could see until the system had reached billions of users and a decade of deployment – was the shape of the second-order effects. Algorithmic feeds optimized for engagement turned out to also be optimized for outrage, because outrage is stickier than nuance. Recommendation engines designed to show you content you’d like turned out to be extraordinarily effective at radicalizing users incrementally, because the next most engaging piece of content is usually a more extreme version of the last one. Platforms built to connect people turned out to be equally effective at organizing them around disinformation – and at putting that capability in the hands of anyone willing to pay for targeted reach, including foreign governments seeking to influence elections.
None of this was the stated intention. Some of it wasn’t even visible to the people building the systems at the time. That’s precisely the point. By the time the consequences were undeniable – congressional hearings, whistleblowers, academic research, documented manipulation of democratic processes across multiple countries – the platforms were foundational infrastructure. Unwinding the engagement-optimization models would mean accepting a significant drop in the metrics the entire business was built around. The window for a different architectural choice had closed years earlier, when the systems were smaller, the lock-in was weaker, and the course-correction would have been cheap.
The Collingridge Dilemma doesn’t mean you shouldn’t ship. It means you should be precise about what you don’t yet know – and build the feedback loops that can surface second-order effects before they become irreversible. It means treating “we haven’t seen evidence of harm” as categorically different from “we have evidence there is no harm.” And it means that the window for real ethical governance is narrow: it’s before launch, not after.
This reframes the role of the Product Leader significantly. The ethical leverage point isn’t in the post-mortem. It’s in the product review, the launch decision, the framing of what the model is allowed to optimize for. Miss that window and you’re managing consequences, not shaping outcomes. The social media industry missed it. The AI agent era is offering the same choice, on a faster timeline, with less margin for error.
The Product Leader as Ethical Gatekeeper
Here’s the shift that the AI era demands: Product Leaders are no longer just accountable for what a product does. They are accountable for what a product decides.
That’s a meaningful distinction. A feature does something when a user takes an action. An AI agent decides something on behalf of – or about – a user, often without their awareness, at a scale and speed that no human review process can match. The Product Leader who ships that agent is not just a delivery manager. They are, whether they’ve accepted the title or not, the last meaningful human checkpoint before the system operates autonomously in the world.
Most product organizations aren’t built around this reality. Reviews focus on technical readiness, business cases, and UX quality. Ethical readiness – who is harmed if this fails, what are the second-order effects, who owns the outcome – is either absent from the checklist or present as a compliance exercise rather than a genuine gate.
What would it look like to take this seriously? Not as an ethics theater exercise, but as a practical discipline embedded in how product decisions get made. A starting framework:
Who is harmed when this system fails – and have we designed for them? Not “what’s our fallback if the model breaks.” Who are the real people at the tail of the distribution – the edge cases outside the primary persona – and what happens to them when the system doesn’t perform as intended? If you haven’t modeled the failure modes for the most vulnerable users in your system, you haven’t finished the product review.
Can we explain this decision to the person it affects? Not to a regulator. Not in a whitepaper. To the specific person whose application was rejected, whose content was removed, whose premium was raised. If that explanation isn’t constructible, you don’t yet understand your own system well enough to deploy it at scale.
Who owns the outcome when something goes wrong? A name. Not “the team” or “a cross-functional group.” The accountability should be defined before launch, not negotiated in a post-mortem. If it’s genuinely unclear, that’s a product decision that needs to be made explicitly – because ambiguous accountability in AI systems doesn’t stay ambiguous. It just gets resolved by whoever can’t escape the news cycle.
What are we optimizing for, and what does that imply for the people in the system? Optimization targets are not neutral. Engagement, retention, conversion – each encodes a value judgment about what matters, and each creates incentive gradients that the model will follow to their logical conclusion, whether or not that conclusion is one you’d endorse in a product review. Be explicit about what the model is allowed to trade off against what.
What don’t we know yet – and how will we find out? This is the Collingridge question, operationalized. Not “what are the known risks.” What are the consequences that won’t be visible until the system has been in production at scale? What monitoring exists to surface them? At what signal would you roll back – and is that threshold defined in advance, or will it be negotiated under pressure after the fact?
None of this requires a philosophy degree. It requires treating ethical readiness as a first-class product requirement – one that can block a launch the way a critical bug can block a launch – and building the organizational muscle to enforce it.
Beyond Good Intentions
The structural failure here isn’t a shortage of ethical Product Leaders. Most people in the profession care about doing the right thing. The failure is that “caring about doing the right thing” is not a system. It’s a disposition, and dispositions are inconsistent, subject to pressure, and impossible to audit.
Medicine figured this out a long time ago. The Hippocratic tradition endures not because physicians are uniquely virtuous, but because the profession collectively decided that individual virtue was insufficient as a safeguard. That the stakes were too high to rely on good intentions. That a binding framework – one that creates real professional consequences for violations – was necessary. The fact that doctors still sometimes act unethically doesn’t invalidate the framework. It demonstrates why the framework exists.
Product Management hasn’t had that reckoning yet. We talk about responsible AI, publish principles, run ethics reviews that function as checkbox exercises, and congratulate ourselves when we decline a feature that was obviously harmful – without asking the harder question: what process would have caught the ones that weren’t obvious?
The scale at which modern Product Leaders operate makes this urgent in a way it wasn’t a decade ago. A bad individual decision can harm a user. A bad AI system, shipped without adequate ethical architecture, can harm millions – often in ways that are diffuse, invisible, and deniable precisely because no single decision caused the harm. The system did. The system you shipped.
The profession is going to be forced to answer these questions one way or another. By regulators, by litigation, by the kind of public failure that makes careers end and companies restructure overnight. The only real choice is whether product leadership gets ahead of it – builds the frameworks, develops the muscle, defines the accountability structures – or waits to have it defined externally, under worse conditions, with less autonomy.
The AI agent era doesn’t just change what products can do. It changes what Product Leaders are responsible for. The sooner that’s treated as a first-order leadership question rather than a compliance footnote, the better the outcomes – for users, for organizations, and for a profession that’s still figuring out what it owes the people it builds for.

