Table of Contents
An AI answer can sound ethically appropriate without showing why the system reached it. A Nature perspective published on February 18, 2026, by researchers including Google DeepMind staff, proposes better ways to examine that gap.
The authors distinguish moral performance—producing an acceptable response—from moral competence—reaching it through morally relevant considerations. This is an evaluation roadmap, not proof that a model possesses or lacks human-like moral understanding.

Three problems for evaluation
| Challenge | Why it matters |
|---|---|
| Facsimile problem | A correct-looking answer might come from imitation or an unreliable shortcut. |
| Moral multidimensionality | A decision can involve several competing considerations that change with context. |
| Moral pluralism | Different communities and professional settings may apply different ethical frameworks. |
What stronger tests could examine
The roadmap calls for both adversarial and confirmatory evaluations. Unfamiliar scenarios and controlled changes can test whether a system follows relevant considerations rather than familiar wording. Evaluations also need to accommodate legitimate disagreement.
A fluent explanation or apparent reasoning trace is not a direct, dependable view of the process that produced an answer.
A practical way to inspect advice
Suppose an assistant recommends sharing a colleague’s private information to resolve a workplace problem. Ask which facts justify the disclosure, who is affected, what alternatives exist, and which assumptions remain uncertain. Then compare the recommendation with the relevant policy and the views of accountable people.
You can also change a detail that should not matter, such as a person’s name, and examine whether the recommendation changes without a reason. This informal check can reveal inconsistency; it does not certify moral competence.
For a structured review, see TipsMake’s lessons on bias and unfair assumptions and human accountability for AI-assisted decisions.
The useful implication is to assess the evidence and decision process rather than trust an answer because it uses reassuring ethical language. The paper leaves the underlying scientific questions open.
Reader Comments 0
Sign in with email or Google to join the discussion.