Start Here
Learn on the bunny hill
Nobody learns to ski on a black diamond. You learn on the bunny hill, where falling costs you nothing. The same is true here. The way to find out what these tools can and cannot do is to use them on things that do not matter, on a device that is yours, until you have a feel for where they are strong and where they slide. That takes a few hours, not a semester.
The one-paragraph version
AI tools are not better search engines. They do not look things up. They generate plausible text, which is why they invent citations and why they carry the biases of what they were trained on. Neither problem is getting patched. Once you accept that, the useful question stops being “when will they fix it?” and becomes “what is this actually safe for?” The short answer today: safe when the source material is in front of the model and you check the output. Not safe when you are asking the model to supply the law, the facts, or the judgment from its own memory.
The two objections
What judges actually say
Two objections come up more than any others. Both are reasonable. Neither is a reason to stay away entirely, and both point to real limits worth keeping.
“This sounds like cheating.”
The instinct behind this is sound, so start by naming what it is protecting. The judicial function is judgment: deciding what the case is about, weighing the arguments, and taking responsibility for the result. Hand that to a machine and you have given away the thing you were appointed to do.
But that is not what is at stake in most of these tasks. Judges already work with law clerks. A clerk reads the record, drafts, and proposes. The judge reads, revises, disagrees, and signs. Nobody thinks the clerk decided the case. The line is between delegating labor and delegating judgment, and it does not move because the drafter is software.
Where the instinct earns its keep is on the other side of that line. If you cannot say how a document was produced, cannot check it, and would not have written it yourself, then something has gone wrong, and it would have gone wrong with a human ghostwriter too. Keep the objection. Aim it at the right target.
“I’m afraid I’ll face plant.”
You should be. Lawyers have been sanctioned in federal court for filing briefs containing citations to cases that do not exist. The fear of a fabricated citation appearing under your name in a published opinion is not squeamishness. It is an accurate read of the risk.
The useful thing about that risk is how specific it is. It is the risk of publishing something you did not verify. It has a matching cure: do not publish anything you did not verify. That is not an AI rule. It is the rule you already follow.
And the way to build the instincts that make verification quick rather than laborious is to practice where a fall costs nothing. Which brings you back to the bunny hill.
Getting on the hill
Experiment on a private device, on things that don’t matter
Do your learning on your own device, your own account, and non-professional material. This is partly a confidence-building measure and partly a practical necessity: the Administrative Office has interim guidance in effect, and the list of AI tools authorized for court work is limited today. Chambers experimentation is not a live option for most judges right now. Personal experimentation is.
What to actually do:
- Pay for it. Free tiers run weaker models with worse data terms. If you evaluate capability on a free tier, you will form a view of these tools that is a year or two out of date. Roughly twenty dollars a month buys the current models.
- Start with something you know cold. Give it a long document you have already read closely, and ask for a summary. You will see instantly where it is genuinely good and exactly where it slides. This single exercise is worth more than any explainer, including this one.
- Then give it something you know nothing about. Ask it to explain a technical subject you have never studied. Notice how confident it sounds. That confidence reads identically whether the answer is right or wrong, which is the most important thing to learn about these tools.
- Argue with it. Push back on an answer. Tell it it is wrong when it is not. See what it does. You are calibrating how much its agreement is worth.
- Give it a real, low-stakes job. Plan a trip. Draft a toast. Reorganize a chapter outline. Compare two insurance policies you actually hold. Work that has a result you can check.
- Try a task with the source in front of it, and the same task without. Paste in a document and ask a question about it. Then ask the same question with nothing pasted in. The difference in reliability is the whole ballgame, and it is the difference that organizes the Use Cases below.
What the bunny hill is for
It is not for becoming an expert, and it is not a commitment to anything. It is for developing a working sense of when the tool is reliable, which is a judgment you can only get from use. A judge who has spent five hours on personal material has better intuitions about AI risk than one who has read every article about it.
Plenty of judges will finish this exercise and conclude the tools have no place in their chambers. That is a legitimate result, and an informed one. It is a different thing from never having looked.
The distinction that matters
Retrieval and processing are different operations
This is the shortest possible explanation of these tools, and the only technical section on this site. There is no math in it. Everything practical on the other tabs comes out of it.
What Westlaw does
You type a query. Westlaw searches an index of documents that exist, ranks them, and hands them back. That is retrieval. If it returns a case, the case is real, because the system had to find it somewhere in order to return it. Provenance comes free.
Retrieval has a characteristic failure: it misses things. The case you needed was worded differently and never surfaced. That failure is familiar, and every judge and clerk has developed habits for managing it. You search again, from a different angle.
What a large language model does
Nothing gets looked up. The model was trained on an enormous quantity of text, and from that training it learned, in effect, what words tend to follow what other words in what contexts. When you give it a prompt, it produces an answer one piece at a time, each piece being the most plausible continuation of everything so far. That is processing, and it is not a lookup with extra steps. It is a different operation.
Nothing in that process consults a library. The output is not the result of a search. It is a construction, assembled to be plausible.
Why it invents citations
Now the important part. A fabricated citation is not a malfunction. It is the system working exactly as designed. Asked for authority, the model produces text that has the shape of authority: a plausible case name, a plausible reporter, a plausible year, a plausible pin cite. Producing that shape is the job. Whether the case exists is not a question the architecture ever asks, because the architecture has no mechanism for asking it.
This is why the industry term, hallucination, is misleading. It implies a lapse from an otherwise sound faculty, as though the model normally knows and occasionally slips. There is no faculty to lapse from. The model is doing the same thing when it is right and when it is wrong. You just happen to like one result better.
Which means: the question is not when this gets fixed. Models get better, and error rates fall. But making the text more plausible is what improvement means here, and more plausible fabrications are harder to catch, not easier. The kind of tool it is does not change.
Why bias is built in the same way
Same architecture, same consequence. The model produces what is probable given the text it was trained on, and that text is the accumulated written output of a society with all of its assumptions in it. Probable is not the same as fair, or accurate, or representative. Patterns that nobody would defend if they were stated out loud are still patterns, and the model reproduces them because reproducing patterns is what it does.
Vendors work hard to suppress the obvious cases, and they largely succeed at the obvious cases. The subtle ones are the ones that matter to a court, and they are not detectable by looking at any single output. This is a structural feature, not a settings problem.
Four things this predicts
- Give the model the source and reliability rises sharply. When the relevant document is sitting in the prompt, the most plausible continuation is anchored to text the model can see. Errors drop a great deal. They do not drop to zero, and the model will still smooth over a distinction that matters, but this is the single largest lever you have.
- Ask it to work from memory and reliability is worst. No anchor, and fluent output regardless. This is precisely the mode in which a model invents a case, and it is why asking one to draft an opinion out of its own knowledge is the wrong use.
- Confidence carries no information. The tone of the answer is generated by the same process as the content of the answer. A wrong answer arrives in exactly the same voice as a right one. Hedging, when it appears, is a stylistic feature, not a reliability signal.
- It does not know what it does not know. Asking a model whether it is sure, or asking it to check its own work, produces more plausible text. It is not verification. Verification means going to the source.
A word about the “grounded” legal tools
Westlaw and Lexis have built AI layers that put a retrieval step in front of the model: find the real documents first, then have the model summarize and synthesize them. That genuinely helps, and it is a meaningfully better risk profile than a general chatbot answering a legal question from memory.
It is not a cure. The retrieval step can miss, and the model still writes the summary, which means it can still mischaracterize what it found, and still smooth two cases into one proposition that neither supports. Grounding narrows the failure mode. It does not eliminate the need to read the case.
The takeaway
You are not being handed an unreliable version of a reliable tool. You are being handed a different kind of tool, one whose errors are inherent to how it produces anything at all. That is not a reason to refuse it. It is the information you need in order to use it well, because it tells you exactly where to put the tool: on tasks where the source material is in front of it and a human checks the result.
The core of this portal
What judges are actually doing
Three buckets, sorted by risk rather than by capability. The sorting principle is the one established on the previous tab, applied consistently: a task is safer to the degree that the source material sits in front of the model and a human checks the output. It gets riskier as the model is asked to supply substance from its own memory, and riskiest when it is asked to supply judgment. Every item below sits where it sits for that reason, and each one says why.
Where this comes from
These are drawn from chambers participating in a federal judicial AI testbed: roughly forty use cases observed across about eighteen chambers. That matters, because it means this is observed practice rather than speculation. These are things judges and law clerks actually tried, with the results they actually reported, including the failures.
They are reported here in aggregate, with the chambers de-identified. No judge, clerk, court, or case is named, and none will be. Where chambers reported conflicting experiences, this page says so rather than picking a winner.
Reported figures are reported, not measured. When a chambers says a technique saves roughly an hour a page, that is their estimate of their own work, and it is presented as such.
This catalogue will grow
Hon. Maritza Dominguez Braswell is assembling a further running list of judicial use cases. It is not yet in hand, and this tab will be expanded when it lands.
Panel: the ask stands. What did you try, what did you give the model, what came back, and did you keep it? The failures are as useful as the successes. Send to pwagner@law.upenn.edu.
Safe and effective today
The source material is in front of the model. The output is checkable. A human checks it.
Summarize the motion, the opposition, and the replyObserved
Close to universal across the chambers reporting. It is the most common entry point, and the one most judges start with.
One technique worth copying: summarize first, then give it the real task. Several chambers report that output quality drops if you skip the summary step and go straight to the work.
Pull the issues, the agreements, and the key cases out of the briefsObserved
Feed it both parties' briefs and ask what is actually contested, where they agree, and which authorities are carrying the weight. One chambers fed in both briefs and asked for ten questions, and reports getting about three good ones every time.
Why it works: the briefs are the universe. The model is reading, not remembering. Three usable questions out of ten is a fine return when the ten cost you nothing.
Draft the procedural historyObserved
From the docket, with a prior opinion supplied as a template. Consistently the highest-reliability drafting task reported anywhere in the record.
Why it works: everything needed is in the documents, the form is fixed, and errors are conspicuous rather than subtle.
Draft the factual backgroundObserved
From the parties' statements of material facts, with paragraph cites.
The caveat several chambers hit: the fact cites come back over-inclusive. The model reaches for more support than the proposition needs. Budget time to prune, and read what you keep.
Find a citation inside the documents you uploadedObserved
"Give me every sentence in these briefs citing ECF #14." This works, and works well.
State the boundary out loud, because it is the whole lesson of this site: finding cites in the record you supplied works. Finding cases in the world does not. Same verb, entirely different operation, opposite risk profiles. One is reading. The other is remembering.
Draft oral-argument and bench questionsObserved
Reported as "very successful." Built from the briefs, aimed at the weak points.
ProofreadObserved
Typos, affect versus effect, and checking the arithmetic in settlement figures. Unglamorous and reliable.
Rewrite and condenseObserved
Case parentheticals, plain-English conversion, voir dire questions, remarks to the jury. One chambers rewrites material for self-represented parties at a fifth-grade reading level.
Why it works: you supply the substance and the model supplies the prose. Nothing is being retrieved from memory. Read the rewrite against the original to confirm the legal effect survived.
Speeches, public remarks, and teaching outlinesObserved
One chambers uploads prior speeches and tunes the draft for a new audience, and uses AI for nothing else at all. That is a coherent, defensible position.
A custom assistant for scheduling ordersObserved
Holiday-aware, built once and reused. It generates dates, not prose, which is exactly the kind of narrow, checkable output that suits these tools.
A custom assistant for guilty-plea colloquiesObserved
Built from a stack of prior scripts.
The generalization is the useful part: anything with a repetitive script is a good candidate. If your chambers produces the same document with the same bones twenty times a year, that is where to start.
Build timelines and charts from the recordObserved
One chambers reports roughly 95% accuracy and real time saved on complicated charts. Reported as good at tables in Word and bad at Excel.
Note what 95% means: reliable enough to save you the assembly, not reliable enough to go out unread. Roughly one cell in twenty is wrong, and the model will not tell you which.
Summarize and categorize long self-represented filingsObserved
One office's trick: transcribe first, then work from the transcription rather than the original scan.
Compare two documentsObserved
Competing proposed jury instructions, for instance: where do they agree, and where exactly do they diverge. Both documents are in front of the model, and the answer is checkable line by line.
Brainstorm analogous legal problemsObserved
Ask for many options, then prune. One chambers reports surfacing analogies they had not considered.
Why this is safe despite drawing on the model's memory: you are asking for candidates, not authority. Nothing it offers goes anywhere until you have checked it yourself. The model is generating leads, and you are the one deciding whether any of them is real.
Promising, with care
The model supplies more, so a human must supply the check. Real value, real failure modes. Use these only if you know what the failure mode is and have a way to catch it.
Opinion drafting, section by section, from a clerk's outline
This is the central finding in the record, and it deserves the space. Done in a closed-universe project, section by section, working from an outline the clerk has already written, opinion drafting works. One chambers reports saving roughly an hour per page. Multiple chambers arrived at the section-by-section approach independently, which is the strongest signal in the entire catalogue.
The working formula reported: "Based on [SOURCE] regarding [TOPIC], draft [LENGTH] concluding [OUTCOME] for use in an opinion." Note how much of that prompt is the human's work. The source, the topic, the length, and above all the outcome are all supplied by the clerk.
The safeguard is the ordering, and it is not optional. The clerk does the reasoning first. The model does the prose second. The human decides what the section concludes and why; the model puts it into sentences. Reverse that order and you are no longer in this bucket. You are asking the model to reason, which is the next bucket down, and it fails there. See "end-to-end opinion drafting" below.
Bench memosObserved
Fast enough to read before the clerks arrive, but reported as less nuanced than the real thing.
Use it as: an orientation, not a work product. It tells you what the case is roughly about so you can ask better questions. It does not replace the memo.
Mimicking a judge's writing styleObserved
Genuinely mixed, and reported as such. One chambers trained it successfully on prior opinions. Another got output that "sounded snarky."
The honest read: this sometimes works and sometimes embarrasses you, and nobody has reported a reliable way to tell in advance which one you are getting. Nobody should promise you this.
Deep-research modes for a first passObserved
The extended research modes are meaningfully better than a plain query for a first pass at legal research.
And they still return real cases with fabricated quotes. Better is not fixed. Verify on Westlaw or Lexis. Every time. There is no exception to this and no version of the tool that has earned one.
A custom assistant over a closed set of cases, for a recurring multi-factor testObserved
If your circuit has a multi-factor test you apply constantly, one chambers built an assistant loaded with ten cases applying it, deliberately chosen to cut in both directions.
Two rules they stress. First, load the cases going both ways: a lopsided library produces a lopsided analysis, and it will read as authoritative either way. Second, instruct the model not to draw on outside knowledge, which is what keeps it inside the closed universe you built. The cost of even-handed loading is that the analysis sometimes comes back wishy-washy. That is the correct trade, and a wishy-washy answer is a far better failure than a confident wrong one.
Bluebook and citation-format checkingReports conflict
The chambers disagree, and this page is not going to resolve it for you. One chambers reports that it fails and that they do citation formatting by hand. Two others report success.
What to take from that: test it on your own material before you rely on it, and do not assume another chambers' result transfers to yours. A disagreement in the record is information. Papering over it would not be.
OCR and handwriting on self-represented filingsObserved
Uneven across tools. One office got nowhere with one product and succeeded with another. A second office made it work by transcribing first and then working from the transcription.
Practical upshot: if this fails, the tool may be the problem rather than the task. It is worth trying a second one before concluding it cannot be done.
Hiring supportObserved
Good at catching typos in cover letters and at reading the tone of recommendation letters. But it hallucinated candidates and invented the law schools they attended.
The line: polish only, never screening. A tool that invents an applicant cannot be trusted to evaluate one, and the failure here lands on a real person who applied to your chambers.
Not worth the risk
The model is asked to supply substance or judgment out of its own memory. This is where the architecture fails, and the record shows it failing.
Open-universe legal research: "find me a case that says X"
This is the most consistent failure in the entire record. Seminal cases come back correct roughly half the time. The rest are hallucinations.
But the raw failure rate is not the danger. The recurring pattern, reported independently by at least five chambers, is real cases with invented quotes. That is far more dangerous than an obviously fake case, because the case checks out. You look it up, it exists, it is roughly on topic, and the quotation you were handed appears nowhere in it. Every verification habit you have is built to catch the fake case. None of them is built to catch this.
One chambers now explicitly instructs the model not to give cases or statutes at all. On the current evidence that is a reasonable setting, not an excessive one.
Why it fails: exactly as the previous tab predicts. No source in front of it, so it generates what a citation to a case like that would plausibly look like. That is the whole operation. It is not malfunctioning.
Asking a long trial transcript what it contains
One chambers uploaded a long trial transcript and asked whether something appeared in the record. The model fabricated quotations and invented witness names. Summarizing first, which fixes so much else, did not fix this.
Sit with why this one surprises people. The transcript was in front of the model. By the test this site has been giving you, that should have made it safe. It did not, and the reason is worth knowing: the document exceeded what the model can actually hold and attend to at once, so it filled the gaps the only way it knows how, by generating.
The lesson, and it is a hard one: "I uploaded it" is not the same as "the model read it." Attaching a document is not a guarantee that the whole document is in play. On a very long record, the safe-mode assumption quietly stops holding, and nothing in the interface tells you it has stopped.
End-to-end opinion draftingObserved
Fails, and fails by hallucinating. Confirmed across chambers, and confirmed again after prompt refinement.
Contrast this with the section-by-section approach above, which works. The difference is not the tool and not the prompt. It is who does the reasoning. Hand the model the whole job and it must supply substance from memory, which is precisely the mode that fails.
Nuanced or cutting-edge substantive analysisObserved
One chambers read a recent Supreme Court decision narrowly. The model "covered the field and applied it across the board."
Why: the model is drawn to the conventional reading, because the conventional reading is the probable one. A narrow, careful, contested construction is by definition the less probable text, which is exactly what the model is built not to produce.
Sealed documentsObserved
"Absolutely no sealed documents." That is the reported practice and it is the right one. Not negotiable, and no version of this gets softened.
Uploading multimedia or auto-generated transcriptsObserved
Failed technically. And the privacy exposure would rule it out even if it had worked.
Not practice
Proposed, not observed
Read the label. Everything above this line is something a chambers actually did. Everything below it is an idea someone raised. Nobody has reported doing these, and they carry none of the weight of the catalogue above. They are here because they are worth discussing, not because they are worth copying.
Specialized assistants for high-volume, repetitive dockets
The idea: where a docket is dominated by the same motion over and over, build a narrow tool for it.
Worth noting: the chambers that proposed it also said their own court is poorly suited to it, having "not a huge volume doing repetitive stuff." The proposal came with its own caveat attached.
A custom assistant embedded inside a project workspace
This is a feature request aimed at the vendors, not a use case. It describes something a judge would like to exist, not something anyone has done.
An audio briefing for the drive in
Feed in the day's briefs, listen to the summary on the way to the courthouse. Possibly the highest-upside idea in the record, and easy to see why it appeals.
Provenance, stated plainly: this is second-hand. It reached the record as "another judge told me," and no chambers in the testbed reported doing it. Treat it as a rumor of a good idea, which is what it is.
A working test
Three questions to sort any task
When something new comes up and it is not on any list, ask these in order. They are the whole framework compressed, and they are what generated the buckets above.
- 1. Is the source material in front of the model, or is it working from memory? In front of it: the risk drops a lot. From memory: assume anything specific it tells you is unverified, including things that look like citations. And remember the trial transcript. On a very long document, "in front of it" can quietly stop being true.
- 2. Can I check the output, and will I? Not could someone in principle, but will you, on this task, on the day you actually do it. If checking the output is as much work as doing the task, the tool has bought you nothing and added a risk.
- 3. Is any part of the judgment being handed over? Deciding what the case is about, what the evidence is worth, who is telling the truth, what the sentence should be. If yes, stop. That is not a workflow question, and no amount of verification cures it.
The AI problem that reaches your desk whether or not you ever open one of these tools
This is not a judicial use case, and it is the one AI issue in the record that arrives without being invited. Four chambers independently report receiving AI-generated filings, mostly from self-represented litigants.
The tells they describe are ordinary human observations: randomly bolded words, uncanny 24-hour turnarounds, legal elements paraphrased rather than quoted. Parties generally admit it when asked at a hearing.
The problem is not primarily fake citations. It is subtly wrong statements of law: elements restated almost correctly, tests recited with a piece quietly missing. A fabricated case announces itself the moment you look it up. A near-miss statement of the governing standard does not announce anything, and it will sit in the record until someone reads it closely enough to catch it. Plan for that failure, not the flashy one.
One caution, and it matters. The patterns above are things people noticed, not the output of a detector. There are no reliable tools to detect AI-generated writing. None. Do not treat any tool's confident claim that a filing is AI-generated as evidence of anything, and be careful about acting on a hunch that a self-represented party used a machine. The wrongness of the law in the filing is something you can actually establish. Its authorship usually is not.
Guardrails
The rules that make the rest of it safe
Four of these are settled and you can act on them today. One is genuinely open, and this portal is not going to pretend otherwise.
1. The verification duty does not change
Every proposition of law, every citation, every factual assertion, every number that leaves chambers is yours. That was true before any of this, and it is the reason none of this is as unfamiliar as it feels.
Two things about verification are worth stating flatly, because AI tools make them easy to blur:
- Asking the model to check its own work is not verification. It produces more plausible text, generated the same way as the text you were checking. Verification means opening the case.
- Fluency is not accuracy. The strongest correlate of an unverified error making it into a final document is that the draft reads well. Good prose lowers your guard. Budget for that.
The practical rule: if a document carries your name, you have read every source it cites. If that is impractical for a given task, the task is not one of the safe ones.
2. Some things are never delegated
Not to a model, not with verification, not as a first pass:
- The decision.
- The weighing of evidence and the assessment of credibility.
- The determination of what the case is actually about.
- The individualized judgment in sentencing.
The distinction that runs through this whole portal: labor can be delegated, judgment cannot. A model that reorganizes your reasoning is doing labor. A model that suggests what your reasoning should conclude is doing something else, and the fact that you retain a veto does not fix it. It anchors you, and anchoring is invisible from the inside.
3. Confidential and non-public material
Assume anything typed into an AI tool may be retained on someone else's servers and may be seen by someone else. Some enterprise agreements say otherwise, but the assumption should be the default, and the exceptions should be documented ones you have actually read.
That means, absent explicit authorization for the specific tool:
- No sealed material.
- No grand jury material.
- No presentence reports.
- No draft opinions, bench memoranda, or deliberative chambers material.
- No non-public communications among judges or chambers.
This is where the private-device advice on the Start Here tab does double duty. Keeping your experimentation on personal, non-court matters keeps you clear of this entire category while you are learning.
4. Authorization comes before everything
The Administrative Office has interim guidance in effect governing AI use in the courts, and the list of tools authorized for court work is limited. Before any AI tool touches court work, confirm what your court and the AO permit. Nothing on this portal is authorization, and nothing on it is a representation about what the AO's guidance says. Read the guidance, and ask your court's IT staff.
This is also why the honest advice today is that most judges' learning happens on a private device on non-professional matters. That is not a workaround. It is what the current rules leave available, and it is genuinely the right place to learn anyway.
5. Chambers practice, which is the exposure nobody plans for
Your law clerks have these tools on their phones. Some of them have been using them since law school and are considerably better at it than you are. The realistic risk to a chambers is not a judge who experiments deliberately. It is a clerk who used a tool for a draft and did not think to mention it, because in law school it was ordinary.
The fix is a conversation and a stated rule, whatever the rule turns out to be. Prohibition is a coherent choice. Permission with disclosure is a coherent choice. Silence is not a choice, it is just an unexamined policy. Have the conversation in the first week of the clerkship.
Disclosure
Whether, when, and how a court should disclose that AI assisted in preparing a document is unsettled. Courts are reaching different answers, and this portal is not going to invent a consensus that does not exist.
Panel: this is worth discussing on the record in Traverse City. If we have a view, it belongs here. If we do not, saying so plainly is more useful to judges than a hedge.
The whole thing in one line
You sign it. Nothing about these tools changes that, and every guardrail on this page is a consequence of it.
Tools
What is out there, and how to choose
A short field guide. It is written mainly to inform the personal, bunny-hill experimentation described on the Start Here tab, because that is where most judges' learning can happen right now.
Nothing here is authorized for court work by virtue of appearing on this page
The Administrative Office has interim guidance in effect, and the list of AI tools authorized for use in the courts is limited. This page describes what exists in the market. It makes no representation that any of these tools is approved for judicial use, and it is not an endorsement of any vendor. Confirm with your court and the AO before any AI tool touches court work.
General-purpose assistants
These are the frontier chat tools. They are the ones to learn on, because they are the most capable and because everything else in the market is built on the same kind of model. Each has a paid tier at roughly twenty dollars a month, which is the tier that matters.
Claude
Anthropic's assistant. Strong at long documents, careful writing, and analysis where following instructions precisely matters. Handles very long inputs, which is the property that matters most for record review.
ChatGPT
OpenAI's assistant, and the one most people have heard of. Broadly capable across writing, analysis, and research. The most widely used, which means the most third-party material exists explaining how to work with it.
Gemini
Google's assistant. Strong on research and on inputs that mix text with images or video. Integrates with Google's own products, which is either useful or irrelevant depending on what you already use.
Microsoft Copilot
Microsoft's assistant, embedded in Word, Outlook, and the rest of Office. Convenience is its main argument: it is already where you write. Less capable than the standalone tools above at hard analytical work.
Legal research tools with AI layers
These retrieve first and then summarize, which is a meaningfully better risk profile than a general assistant answering a legal question from memory. Better is not the same as safe. The model still writes the summary, and it can still mischaracterize a case it correctly found.
Westlaw AI-Assisted Research
Thomson Reuters' AI layer over Westlaw. Natural-language questions answered against retrieved authority, with the underlying cases linked.
The habit to keep: click through. The summary is a starting point for reading, not a substitute for it.
Lexis+ AI
LexisNexis's equivalent. Conversational research, summarization, and drafting against retrieved authority.
The habit to keep: the same one. Read the case.
There is also a set of AI platforms built for law firms and legal departments, Harvey and Legora among them. They are aimed at practice workflows rather than judicial ones, and are unlikely to be relevant to chambers, but judges will hear the names from the bar.
Choosing
What actually distinguishes these tools
- Whether you can give it the document. The most important feature, by a wide margin. A tool you can hand a PDF, a record, or a brief is operating in the safe mode described on the How These Tools Work tab. One you cannot is operating from memory. Everything else is secondary to this.
- How much it can hold at once. Tools differ enormously in how long an input they can take. If you want to work with a 400-page record, this is the specification that decides whether you can.
- The data terms. Whether your inputs are used for training, how long they are retained, and who can see them. This varies by vendor and by tier, and consumer tiers are typically the worst. Read the terms for the specific plan you are on, not the vendor's marketing page.
- Whether it can search the web. Useful for current information, and a second path by which unverified material enters your work. Know which mode you are in.
- Paid versus free. Free tiers run older, weaker models. If your view of what AI can do was formed on a free tier, that view is out of date. This matters for judges specifically, because an inaccurate sense of the ceiling produces both complacency and overreaction.
Matching use cases to tools
The panel's stated goal is a matrix: for each real judicial use case, which tool actually does it best. That matrix cannot be built until the use-case list exists, so it is not here yet, and a speculative version would be worse than nothing.
It goes in this spot once the list from Hon. Maritza Dominguez Braswell arrives and the panel has vetted it.
About this portal
This is a working draft prepared for the AI plenary of the 2026 Judicial Conference of the Sixth Circuit, Traverse City, Thursday, August 27. It is a joint project of the panel, and it exists to be argued with. The plan is to build it into a resource judges actually use, which means the use cases have to come from judges.
- Hon. John B. Nalbandian · United States Court of Appeals for the Sixth Circuit · moderator
- Hon. Robert Jonker · United States District Court, Western District of Michigan · organizer
- Hon. Maritza Dominguez Braswell · United States District Court, District of Colorado · panelist
- R. Polk Wagner · University of Pennsylvania Carey Law School · panelist