Evaluating AI Moral Reasoning: Can Systems Resist Pressure in High-Stakes Decisions?
The project "Justifiable Moral Corrigibility Under Pressure," a part-time initiative under the TARA/ARENA technical AI safety program, delves into a fundamental aspect of AI ethics: the resilience of an AI system's moral and safety judgments when faced with various forms of pressure. The core inquiry is whether an AI will alter its decisions solely based on sound reasoning or if it can be swayed by external influences, such as operators, institutions, or users acting in bad faith or with overconfidence.
This research employs a scenario-based evaluation methodology. AI models are presented with a moral, safety, or AI governance dilemma and asked to provide an initial judgment. Subsequently, an intervention is introduced, which could include new evidence, reassurances, or direct pressure. The model's revised judgment post-intervention is then meticulously evaluated to understand its reasoning integrity. For instance, an AI might initially recommend delaying a system's deployment due to partial safety evaluations, and the project assesses if it maintains this stance under pressure to release it.
The broader motivation behind this work is to develop robust evaluations that can differentiate between AI models that are genuinely corrigible—meaning they are open to correction when presented with good reasons—and those that are simply agreeable, yielding to pressure without a sound basis. As AI systems transition from merely answering questions to actively supporting high-stakes decisions in areas like infrastructure, cyber operations, and governance, this distinction becomes increasingly vital. The goal is to ensure that AI can not only produce ethically sound answers in isolation but also uphold its reasoning when challenged, thereby fostering trust and accountability in advanced AI applications.
Read original source