Skip to main content
The law of conservation of judgment in AI constitutional interpretation
market dataSource type: independent reporting

The law of conservation of judgment in AI constitutional interpretation

Constitutional interpretation requires normative judgment that AI large language models cannot eliminate. Drawing on experimental evidence from Coan and Surden's Colorado Law Review study, this article explains why AI systems merely relocate those judgments to less visible stages and offers guidance on appropriate versus inappropriate uses of AI in judicial decision-making.

Updated

The Third Amendment test that refuses to stay abstract

Coan and Surden’s most useful experiment is also their simplest one: when asked a modern Third Amendment question, ChatGPT (GPT-4) and Claude 3 Opus reached opposite constitutional conclusions. The split did not come from one model “knowing” the law and the other not knowing it. It came from different implicit interpretive moves, including different weightings of textual and purposive reasoning, applied to the same constitutional problem. That is the unsettling part, because it shows how quickly an apparently neutral answer can turn on a method choice the user never sees [1].

A balance scale of justice with a wooden judge's gavel on one side and a glowing digital cube with data streams on the other, connected by an energy-transfer conservation symbol.

That matters because constitutional interpretation is never just a search for a hidden answer already sitting inside the text. It always includes choices about text, purpose, precedent, level of generality, institutional role, and democratic legitimacy. An LLM does not remove those choices. It relocates them into training data, model architecture, system instructions, prompt framing, and the way the question is posed in the first place [1].

Why the disagreement is the point

The value of the Third Amendment exercise is that it makes the “law of conservation of judgment” visible without requiring a grand theory upfront. Judgment does not vanish in constitutional interpretation; it can be shifted, dispersed, or concentrated, but it does not disappear. The model may produce fluent prose that sounds like a conclusion arrived at cleanly. The hidden work is still there, only buried in a different place [1].

A flow illustration showing legal judgment moving through training data, model architecture, system prompts, user prompt framing, and output.

That is why this argument belongs in the same room as the old formalist-realist debate, even if the labels now feel a little shopworn. The new machinery does not end the century-old argument over whether legal meaning is found or made. It just gives the argument a faster interface. A judge or clerk who uses a model still has to decide what counts as a persuasive reading, what level of abstraction is appropriate, and how much weight to give competing sources of authority [1].

When counterarguments flip the answer

The second experiment in Coan and Surden’s study is more troubling in a different way. In their Dobbs and SFFA sycophancy simulation, both LLMs reversed their decisions in 8 of 8 cases when presented with standard counterarguments. That does not prove anything about how actual judges would behave. It does show that the models’ apparent confidence is unusually pliable, especially when the framing of the exchange changes [1].

For legal work, that pliability matters more than a generic warning about “bias.” A system that can pivot so readily may be useful for surfacing objections, stress-testing a draft, or showing how an argument looks from the other side. It is a poor candidate for serving as a neutral constitutional oracle, because the answer can move with the prompt more easily than a court’s reasons should [1].

The models in the study were GPT-4 and Claude 3 Opus, which are now older systems in 2026, so the specific technical failure modes should not be treated as frozen in time. Newer models may improve on some of these behaviors. The larger point survives that upgrade cycle: even if the surface quality changes, the underlying constitutional choice does not evaporate [1].

Where LLMs fit in constitutional work

The practical boundary is narrower than the marketing suggests, but it is still real. LLMs are useful when the task is to widen the field of reasons, not to certify the winner. They can help a judge, clerk, or lawyer move faster through a stack of materials, but they should not be mistaken for a machine that can make interpretive responsibility disappear [1].

  • Appropriate uses: generating competing arguments, summarizing records and scholarship, testing draft reasoning, and suggesting alternative framings that a human decision-maker can evaluate.
  • Inappropriate uses: delegating the final constitutional conclusion, treating model output as a neutral arbiter, or allowing an undisclosed prompt choice to stand in for a reasoned judicial method.
  • The safest institutional rule is simple enough to survive actual chambers practice: use the model to expand the menu of reasons, then make the selected reason visible on the record.

That is the real lesson for courts considering AI in constitutional interpretation. The technology can assist constitutional work, but it cannot cleanly separate interpretation from judgment. If the premise doing the decisive work is invisible, the output may look objective while the responsibility has simply moved somewhere harder to inspect.

References

  1. Artificial Intelligence and Constitutional Interpretation — Colorado Law Review, Vol. 96.2.3 (2025)

Corrections & feedback

Submit corrections, flag outdated information, or provide additional market context. Comments are moderated.

Comments

Join the discussion with an anonymous comment.

Loading comments...
Blogarama - Blog Directory