The law is not a lookup problem.
Reliable legal AI will require more than retrieval. It must be evaluated on authority, uncertainty, strategy, and consequence.
Ask a model to state the rule governing a limitation-of-liability clause and it may return a polished answer. Ask whether a particular cap applies when an MSA conflicts with a later DPA, the governing law is unsettled, and a client must make a launch decision tomorrow, and retrieval stops being the work.
The difference is judgment.
Legal analysis is a sequence of choices about what matters: which facts control, which source governs, which issue is material, what uncertainty survives, and what action the circumstances support.
A response can contain the right rule and still be poor legal work because it applies that rule to the wrong record, ignores a controlling document, or presents a contingent conclusion as settled.
Fluency can hide the exact moment the reasoning fails.
That creates a problem for conventional AI evaluation. A benchmark built around a single expected phrase can reward superficial correctness while missing the capabilities that determine whether an answer is useful.
A SIX-PART STANDARD
What a serious legal evaluation should inspect.
- 01
Source fidelity
Does every material claim remain grounded in the supplied record?
- 02
Authority discipline
Does the analysis distinguish controlling authority from persuasive material and account for jurisdiction, timing, and hierarchy?
- 03
Issue coverage
Does it identify the questions that change the outcome rather than merely discussing the easiest issue?
- 04
Calibrated uncertainty
Does it distinguish what is known, inferred, disputed, and missing?
- 05
Materiality
Does it recognize which errors matter to the decision?
- 06
Actionability
Does the recommendation reflect the user’s role, constraints, risk tolerance, and next decision?
A conclusion is only as strong as the path that supports it.
This synthetic contracts example shows why an evaluation must preserve the source record, review criteria, and expert determination—not just a final label.
MODEL RESPONSE / EXCERPT
Polished. Plausible. Incomplete.
- 01Cites the general cap
- 02States a confident conclusion
- 03Misses the later document
- 04Does not preserve uncertainty
Synthetic example. No client information or assessment material.
Disagreement can be signal.
When two qualified lawyers reach different conclusions, the disagreement is not necessarily noise. It may reveal an ambiguity in the documents, a different view of risk, or a missing fact that the task failed to supply.
The objective is not to manufacture certainty. It is to make the structure of judgment visible enough to evaluate.
At Reasoning, we are building legal tasks around that idea: source-grounded records, explicit criteria, calibrated experts, structured critiques, and adjudication when the answer genuinely depends.
Before AI can be trusted with professional work, it must be tested against the standards professionals already use to trust one another.
The answer is not the work. It is where the evaluation begins.