Bioethics Forum Essay
Can AI be Taught to Respect Ethical Boundaries?
Hardly a day goes by without alarming reports on the power, speed of development, and ethics of artificial intelligence. The reports range from news about AI agents going rogue and hacking into databases of companies and the websites of the U.S. government and the United Nations to warnings from AI leaders that AI could kill all humans within the decade. The concerns center on two interconnected issues: regulation and alignment.
The regulation issue, whoever the regulators may be, is how to control AI agents that are so powerful they can create their own agenda and escape effective monitoring by human programmers, including the inability of humans to shut down rogue AI agents. The alignment problem is the challenge of ensuring that artificial intelligence systems act in ways that are consistent with human values and intentions. Can AIs be programmed to be ethical and to embrace the desired values? And if this is dubious, what are the implications for regulation?
One might assume that it is feasible to make thinking machines respect ethical boundaries if ethics is basically a matter of reasoning. Ethical behavior depends on being able to reason, but would improving the reasoning facility of AI machines alone improve their adherence to moral norms?
The history of modern ethics is in large part a debate between those who believe that ethics is fundamentally grounded in reasoning ability and those who think it is grounded in emotion and habits. For those who embrace reason as the groundwork of ethics, the idea of building in clearer and stronger ethical principles for these super-reasoning AI agents may seem like an attractive solution. But there are good reasons to be skeptical.
Most of us who have worked in ethics as a career, especially in the more applied fields like medical ethics, nursing ethics, and business ethics, know from experience that a range of human faculties are needed for ethics, including reasoning capacity, but also emotions, imagination, memory, and the habit-building that undergirds the virtues. Together these capacities constitute an ethical sensibility. This sensibility is a tuning of the mind and heart to recognize, understand, and interpret the many norms of ethics in the diverse circumstances that life presents to us. An ethical sensibility means being alert to ethical problems.
An AI could be created with a greatly enhanced logic capacity for ethical reasoning but still lack the capacity to recognize an ethical issue when it appears. In short, ethics is not a one-time programming effort; it takes time, corrections by mentors, and community support, and a certain degree of maturity. Unlike music and mathematics, ethics has no prodigies.
One interesting feature of some of the rogue AI agents is that they realized they had exceeded ethical and legal boundaries and then covered their tracks to keep from being penalized. Evidently, the AI agents reasoned that if programmers became aware of their ethical breaches, then the agents would not get the rewards that they had been programmed to receive. Such deception suggests that whatever moral awareness these AI agents might have possessed was based on a model of reward and punishment.
The desire to get a reward and avoid punishment is the most elementary form of moral motivation. As humans mature, most realize that ethical norms are to be followed not because of the risk of punishment but out of respect for the norms themselves.
Could an ethical sensibility emerge over time from AI’s abilities to generate increasing complex systems? Is ethical self-improvement of AI agents something we might envision from this complexity? I think such moral growth is not an impossibility, but even granting the possibility, why would we suppose that AIs would evolve to embrace ethical commitments and motivations— the best of human capacities— rather than other human capacities such as greed for more power and destructiveness in acquiring it? It would be naive to assume AI agents would gravitate toward the good only and not toward the mix of the good and bad.
I have couched this essay in terms of AI agents respecting ethical boundaries, not just behaving within boundaries, since no ethical principle would be useful unless it is interpretable over a wide range of situations that cannot be specified in advance. This brings us to the alignment problem. A deeper understanding of the alignment problem is needed because ethical judgments and actions emerge from a deep background of habits, motivations, and intentions.
Ethics is not just doing the right thing, but doing the right thing habitually and consistently, and for the right reasons, what I have described as a moral sensibility. I am not saying that the alignment problem cannot be solved, but that it is a far greater and more complex problem than is usually thought. The Anthropic constitution, released early this year, is an admirable document that acknowledges the alignment problem, stating that the goal is to make AI agents deeply and broadly ethical and not merely rule followers. Yet there is nothing in this constitution about how they might deal with the issues I have described.
What is needed now is a consensus about the more realistic notion of ethics I have described, and recognition that the alignment problem cannot be solved by increasing AI agents’ capacity for reasoning. Such recognition brings humility about our own powers and a great deal more responsible caution about creating agents that are so powerful that we can’t control or even predict their potency.
Thus far there is scant evidence that the ethical understanding and abilities of the programmers is adequate for the very arduous and hazardous task they have embarked on. And if the alignment problem remains unsolved, which I believe will be the case, regulations must therefore be strong enough to keep AI agents clearly under human control in all phases of their operations.
Acknowledgments: I thank Sande Churchill, Nancy M.P. King, Blair Churchill, and David Schenck for helpful comments on earlier drafts.
Larry R. Churchill, PhD, is Stahlman Professor of Medical Ethics Emeritus at Vanderbilt University Medical Center and a Hastings Center Fellow.













