A practical resource for federal judges and chambers staff, maintained with support from the Penn Carey Law AI Project. Visit pennai.law →

Judicial AI Field Guide

Evaluate a task, identify the risks, and preserve human judgment.

The core idea

Search retrieves. An AI model generates.

Many legal AI products combine retrieval and generation: the product retrieves authority or documents, then an AI model generates a response. That source boundary matters, but generated prose remains fallible. The practical questions are what the system received, what it may invent or omit, and how a human will verify the result.

Start Here

How These Tools Work

Guardrails

Tools

What Courts See

Use it where being wrong costs nothing

You cannot calibrate a tool you have never used. The way to find out what these tools can and cannot do is to use them on things that do not matter — on a device that is yours, on material that is personal — until you have a feel for where they are strong and where they give way. That takes a few hours, not a semester.

The one-paragraph version

Search retrieves stored material. An AI model generates a response from learned patterns. Many legal products combine the two: the product retrieves authority, then an AI model generates an answer from it. A bounded source lowers risk, but generated prose still requires verification. The useful question is not whether the tool is generally safe. It is whether this task is authorized, source-grounded, checkable, and free of delegated judgment.

What judges actually say

Two objections come up more than any others. Both are reasonable. Neither is a reason to stay away entirely, and both point to real limits worth keeping.

“This sounds like cheating.”

The instinct behind this is sound, so start by naming what it is protecting. The judicial function is judgment: deciding what the case is about, weighing the arguments, and taking responsibility for the result. Hand that to a machine and you have given away the thing you were appointed to do.

But that is not what is at stake in most of these tasks. Judges already work with law clerks. A clerk reads the record, drafts, and proposes. The judge reads, revises, disagrees, and signs. Nobody thinks the clerk decided the case. The line is between delegating labor and delegating judgment, and it does not move because the drafter is software.

Where the instinct earns its keep is on the other side of that line. If you cannot say how a document was produced, cannot check it, and would not have written it yourself, then something has gone wrong, and it would have gone wrong with a human ghostwriter too. Keep the objection. Aim it at the right target.

“I’m afraid I’ll face plant.”

You should be. Lawyers have been sanctioned in federal court for filing briefs containing citations to cases that do not exist. The fear of a fabricated citation appearing under your name in a published opinion is not squeamishness. It is an accurate read of the risk.

The useful thing about that risk is how specific it is. It is the risk of publishing something you did not verify. It has a matching cure: do not publish anything you did not verify. That is not an AI rule. It is the rule you already follow.

And the way to build the instincts that make verification quick rather than laborious is to practice where being wrong costs nothing. Which is the whole point of starting on personal material.

Experiment on a private device, on things that don’t matter

Do your learning on your own device, your own account, and non-professional material. This is partly a confidence-building measure and partly a practical necessity: the Administrative Office has interim guidance in effect, and the list of AI tools authorized for court work is limited today. Chambers experimentation is not a live option for most judges right now. Personal experimentation is.

What to actually do:

What the first few hours are for

It is not for becoming an expert, and it is not a commitment to anything. It is for developing a working sense of when the tool is reliable, which is a judgment you can only get from use. A judge who has spent five hours on personal material has better intuitions about AI risk than one who has read every article about it.

Plenty of judges will finish this exercise and conclude the tools have no place in their chambers. That is a legitimate result, and an informed one. It is a different thing from never having looked.

Retrieval and processing are different operations

This is the shortest possible explanation of these tools, and the only technical section on this site. There is no math in it. Everything practical on the other tabs comes out of it.

What Westlaw does

You type a query. Westlaw searches an index of documents that exist, ranks them, and hands them back. That is retrieval. If it returns a case, the case is real, because the system had to find it somewhere in order to return it. Provenance comes free.

Retrieval has a characteristic failure: it misses things. The case you needed was worded differently and never surfaced. That failure is familiar, and every judge and clerk has developed habits for managing it. You search again, from a different angle.

What a large language model does

Nothing gets looked up. The model was trained on an enormous quantity of text, and from that training it learned, in effect, what words tend to follow what other words in what contexts. When you give it a prompt, it produces an answer one piece at a time, each piece being the most plausible continuation of everything so far. That is processing, and it is not a lookup with extra steps. It is a different operation.

Nothing in that process consults a library. The output is not the result of a search. It is a construction, assembled to be plausible.

Why it invents citations

Now the important part. A fabricated citation is not a malfunction. It is the system working exactly as designed. Asked for authority, the model produces text that has the shape of authority: a plausible case name, a plausible reporter, a plausible year, a plausible pin cite. Producing that shape is the job. Whether the case exists is not a question the architecture ever asks, because the architecture has no mechanism for asking it.

This is why the industry term, hallucination, is misleading. It implies a lapse from an otherwise sound faculty, as though the model normally knows and occasionally slips. There is no faculty to lapse from. The model is doing the same thing when it is right and when it is wrong. You just happen to like one result better.

Which means: the question is not when this gets fixed. Models get better, and error rates fall. But making the text more plausible is what improvement means here, and more plausible fabrications are harder to catch, not easier. The kind of tool it is does not change.

Why bias is built in the same way

Same architecture, same consequence. The model produces what is probable given the text it was trained on, and that text is the accumulated written output of a society with all of its assumptions in it. Probable is not the same as fair, or accurate, or representative. Patterns that nobody would defend if they were stated out loud are still patterns, and the model reproduces them because reproducing patterns is what it does.

Vendors work hard to suppress the obvious cases, and they largely succeed at the obvious cases. The subtle ones are the ones that matter to a court, and they are not detectable by looking at any single output. This is a structural feature, not a settings problem.

Four things this predicts

  • Give the model the source and reliability rises sharply. When the relevant document is sitting in the prompt, the most plausible continuation is anchored to text the model can see. Errors drop a great deal. They do not drop to zero, and the model will still smooth over a distinction that matters, but this is the single largest lever you have.
  • Ask it to work from memory and reliability is worst. No anchor, and fluent output regardless. This is precisely the mode in which a model invents a case, and it is why asking one to draft an opinion out of its own knowledge is the wrong use.
  • Confidence carries no information. The tone of the answer is generated by the same process as the content of the answer. A wrong answer arrives in exactly the same voice as a right one. Hedging, when it appears, is a stylistic feature, not a reliability signal.
  • It does not know what it does not know. Asking a model whether it is sure, or asking it to check its own work, produces more plausible text. It is not verification. Verification means going to the source.

A word about the “grounded” legal tools

Westlaw and Lexis have built AI layers that put a retrieval step in front of the model: find the real documents first, then have the model summarize and synthesize them. That genuinely helps, and it is a meaningfully better risk profile than a general chatbot answering a legal question from memory.

It is not a cure. The retrieval step can miss, and the model still writes the summary, which means it can still mischaracterize what it found, and still smooth two cases into one proposition that neither supports. Grounding narrows the failure mode. It does not eliminate the need to read the case.

The takeaway

You are not being handed an unreliable version of a reliable tool. You are being handed a different kind of tool, one whose errors are inherent to how it produces anything at all. That is not a reason to refuse it. It is the information you need in order to use it well, because it tells you exactly where to put the tool: on tasks where the source material is in front of it and a human checks the result.

What judges are actually doing

Three buckets, sorted by risk rather than by capability. The sorting principle is the one established on the previous tab, applied consistently: a task is safer to the degree that the source material sits in front of the model and a human checks the output. It gets riskier as the model is asked to supply substance from its own memory, and riskiest when it is asked to supply judgment. Every item below sits where it sits for that reason, and each one says why.

Where this comes from

These are drawn from chambers participating in a federal judicial AI testbed: roughly forty use cases observed across about eighteen chambers. That matters, because it means this is observed practice rather than speculation. These are things judges and law clerks actually tried, with the results they actually reported, including the failures.

They are reported here in aggregate, with the chambers de-identified. No judge, clerk, court, or case is named, and none will be. Where chambers reported conflicting experiences, this page says so rather than picking a winner.

Reported figures are reported, not measured. When a chambers says a technique saves roughly an hour a page, that is their estimate of their own work, and it is presented as such.

Lower risk if authorized

The source material is in front of the model. The output is checkable. A human checks it.

Summarize the motion, the opposition, and the replyObserved

Close to universal across the chambers reporting. It is the most common entry point, and the one most judges start with.

One technique worth copying: summarize first, then give it the real task. Several chambers report that output quality drops if you skip the summary step and go straight to the work.

Pull the issues, the agreements, and the key cases out of the briefsObserved

Feed it both parties' briefs and ask what is actually contested, where they agree, and which authorities are carrying the weight. One chambers fed in both briefs and asked for ten questions, and reports getting about three good ones every time.

Why it works: the briefs are the universe. The model is reading, not remembering. Three usable questions out of ten is a fine return when the ten cost you nothing.

Draft the procedural historyObserved

From the docket, with a prior opinion supplied as a template. Consistently the highest-reliability drafting task reported anywhere in the record.

Why it works: everything needed is in the documents, the form is fixed, and errors are conspicuous rather than subtle.

Draft the factual backgroundObserved

From the parties' statements of material facts, with paragraph cites.

The caveat several chambers hit: the fact cites come back over-inclusive. The model reaches for more support than the proposition needs. Budget time to prune, and read what you keep.

Find a citation inside the documents you uploadedObserved

"Give me every sentence in these briefs citing ECF #14." This works, and works well.

State the boundary out loud, because it is the whole lesson of this site: finding cites in the record you supplied works. Finding cases in the world does not. Same verb, entirely different operation, opposite risk profiles. One is reading. The other is remembering.

Draft oral-argument and bench questionsObserved

Reported as "very successful." Built from the briefs, aimed at the weak points.

ProofreadObserved

Typos, affect versus effect, and checking the arithmetic in settlement figures. Unglamorous and reliable.

Rewrite and condenseObserved

Case parentheticals, plain-English conversion, voir dire questions, remarks to the jury. One chambers rewrites material for self-represented parties at a fifth-grade reading level.

Why it works: you supply the substance and the model supplies the prose. Nothing is being retrieved from memory. Read the rewrite against the original to confirm the legal effect survived.

Speeches, public remarks, and teaching outlinesObserved

One chambers uploads prior speeches and tunes the draft for a new audience, and uses AI for nothing else at all. That is a coherent, defensible position.

A custom assistant for scheduling ordersObserved

Holiday-aware, built once and reused. It generates dates, not prose, which is exactly the kind of narrow, checkable output that suits these tools.

A custom assistant for guilty-plea colloquiesObserved

Built from a stack of prior scripts.

The generalization is the useful part: anything with a repetitive script is a good candidate. If your chambers produces the same document with the same bones twenty times a year, that is where to start.

Build timelines and charts from the recordObserved

One chambers reports roughly 95% accuracy and real time saved on complicated charts. Reported as good at tables in Word and bad at Excel.

Note what 95% means: reliable enough to save you the assembly, not reliable enough to go out unread. Roughly one cell in twenty is wrong, and the model will not tell you which.

Summarize and categorize long self-represented filingsObserved

One office's trick: transcribe first, then work from the transcription rather than the original scan.

Compare two documentsObserved

Competing proposed jury instructions, for instance: where do they agree, and where exactly do they diverge. Both documents are in front of the model, and the answer is checkable line by line.

Brainstorm analogous legal problemsObserved

Ask for many options, then prune. One chambers reports surfacing analogies they had not considered.

Why this is lower risk despite drawing on the model's memory: you are asking for candidates, not authority. Nothing it offers goes anywhere until you have checked it yourself. The model is generating leads, and you are the one deciding whether any of them is real.

Promising, with care

The model supplies more, so a human must supply the check. Real value, real failure modes. Use these only if you know what the failure mode is and have a way to catch it.

Opinion drafting, section by section, from a clerk's outline

This is the central finding in the record, and it deserves the space. Done in a closed-universe project, section by section, working from an outline the clerk has already written, opinion drafting works. One chambers reports saving roughly an hour per page. Multiple chambers arrived at the section-by-section approach independently, which is the strongest signal in the entire catalogue.

The working formula reported: "Based on [SOURCE] regarding [TOPIC], draft [LENGTH] concluding [OUTCOME] for use in an opinion." Note how much of that prompt is the human's work. The source, the topic, the length, and above all the outcome are all supplied by the clerk.

The safeguard is the ordering, and it is not optional. The clerk does the reasoning first. The model does the prose second. The human decides what the section concludes and why; the model puts it into sentences. Reverse that order and you are no longer in this bucket. You are asking the model to reason, which is the next bucket down, and it fails there. See "end-to-end opinion drafting" below.

Bench memosObserved

Fast enough to read before the clerks arrive, but reported as less nuanced than the real thing.

Use it as: an orientation, not a work product. It tells you what the case is roughly about so you can ask better questions. It does not replace the memo.

Mimicking a judge's writing styleObserved

Genuinely mixed, and reported as such. One chambers trained it successfully on prior opinions. Another got output that "sounded snarky."

The honest read: this sometimes works and sometimes embarrasses you, and nobody has reported a reliable way to tell in advance which one you are getting. Nobody should promise you this.

Deep-research modes for a first passObserved

The extended research modes are meaningfully better than a plain query for a first pass at legal research.

And they still return real cases with fabricated quotes. Better is not fixed. Verify on Westlaw or Lexis. Every time. There is no exception to this and no version of the tool that has earned one.

A custom assistant over a closed set of cases, for a recurring multi-factor testObserved

If your circuit has a multi-factor test you apply constantly, one chambers built an assistant loaded with ten cases applying it, deliberately chosen to cut in both directions.

Two rules they stress. First, load the cases going both ways: a lopsided library produces a lopsided analysis, and it will read as authoritative either way. Second, instruct the model not to draw on outside knowledge, which is what keeps it inside the closed universe you built. The cost of even-handed loading is that the analysis sometimes comes back wishy-washy. That is the correct trade, and a wishy-washy answer is a far better failure than a confident wrong one.

Bluebook and citation-format checkingReports conflict

The chambers disagree, and this page is not going to resolve it for you. One chambers reports that it fails and that they do citation formatting by hand. Two others report success.

What to take from that: test it on your own material before you rely on it, and do not assume another chambers' result transfers to yours. A disagreement in the record is information. Papering over it would not be.

OCR and handwriting on self-represented filingsObserved

Uneven across tools. One office got nowhere with one product and succeeded with another. A second office made it work by transcribing first and then working from the transcription.

Practical upshot: if this fails, the tool may be the problem rather than the task. It is worth trying a second one before concluding it cannot be done.

Hiring supportObserved

Good at catching typos in cover letters and at reading the tone of recommendation letters. But it hallucinated candidates and invented the law schools they attended.

The line: polish only, never screening. A tool that invents an applicant cannot be trusted to evaluate one, and the failure here lands on a real person who applied to your chambers.

Not worth the risk

The model is asked to supply substance or judgment out of its own memory. This is where the architecture fails, and the record shows it failing.

Open-universe legal research: "find me a case that says X"

This is the most consistent failure in the entire record. Seminal cases come back correct roughly half the time. The rest are hallucinations.

But the raw failure rate is not the danger. The recurring pattern, reported independently by at least five chambers, is real cases with invented quotes. That is far more dangerous than an obviously fake case, because the case checks out. You look it up, it exists, it is roughly on topic, and the quotation you were handed appears nowhere in it. Every verification habit you have is built to catch the fake case. None of them is built to catch this.

One chambers now explicitly instructs the model not to give cases or statutes at all. On the current evidence that is a reasonable setting, not an excessive one.

Why it fails: exactly as the previous tab predicts. No source in front of it, so it generates what a citation to a case like that would plausibly look like. That is the whole operation. It is not malfunctioning.

Asking a long trial transcript what it contains

One chambers uploaded a long trial transcript and asked whether something appeared in the record. The model fabricated quotations and invented witness names. Summarizing first, which fixes so much else, did not fix this.

Sit with why this one surprises people. The transcript was in front of the model. By the test this site has been giving you, that should have lowered the risk. It did not, and the reason is worth knowing: the document exceeded what the model can actually hold and attend to at once, so it filled the gaps the only way it knows how, by generating.

The lesson, and it is a hard one: "I uploaded it" is not the same as "the model read it." Attaching a document is not a guarantee that the whole document is in play. On a very long record, the source-grounded assumption quietly stops holding, and nothing in the interface tells you it has stopped.

End-to-end opinion draftingObserved

Fails, and fails by hallucinating. Confirmed across chambers, and confirmed again after prompt refinement.

Contrast this with the section-by-section approach above, which works. The difference is not the tool and not the prompt. It is who does the reasoning. Hand the model the whole job and it must supply substance from memory, which is precisely the mode that fails.

Nuanced or cutting-edge substantive analysisObserved

One chambers read a recent Supreme Court decision narrowly. The model "covered the field and applied it across the board."

Why: the model is drawn to the conventional reading, because the conventional reading is the probable one. A narrow, careful, contested construction is by definition the less probable text, which is exactly what the model is built not to produce.

Sealed documentsObserved

"Absolutely no sealed documents." That is the reported practice and it is the right one. Not negotiable, and no version of this gets softened.

Uploading multimedia or auto-generated transcriptsObserved

Failed technically. And the privacy exposure would rule it out even if it had worked.

Proposed, not observed

Read the label. Everything above this line is something a chambers actually did. Everything below it is an idea someone raised. Nobody has reported doing these, and they carry none of the weight of the catalogue above. They are here because they are worth discussing, not because they are worth copying.

Proposed

Specialized assistants for high-volume, repetitive dockets

The idea: where a docket is dominated by the same motion over and over, build a narrow tool for it.

Worth noting: the chambers that proposed it also said their own court is poorly suited to it, having "not a huge volume doing repetitive stuff." The proposal came with its own caveat attached.

Proposed

A custom assistant embedded inside a project workspace

This is a feature request aimed at the vendors, not a use case. It describes something a judge would like to exist, not something anyone has done.

Proposed

An audio briefing for the drive in

Feed in the day's briefs, listen to the summary on the way to the courthouse. Possibly the highest-upside idea in the record, and easy to see why it appeals.

Provenance, stated plainly: this is second-hand. It reached the record as "another judge told me," and no chambers in the testbed reported doing it. Treat it as a rumor of a good idea, which is what it is.

Judge Dominguez Braswell's inventory

A subject-matter primer before an expert hearing

Use the output to identify vocabulary and questions, then verify every substantive claim against authoritative sources.

Judge Dominguez Braswell's inventory

Public procedural guidance

Potentially useful for authorized, public-facing self-help material. It should not become case-specific assistance or an ex parte channel.

Judge Dominguez Braswell's inventory · high sensitivity

Settlement-conference preparation

Consider only with specific authorization, secure tooling, careful treatment of confidential information, independent review, and appropriate party consent.

Three questions to sort any task

When something new comes up and it is not on any list, ask these in order. They are the whole framework compressed, and they are what generated the buckets above.

The AI problem that reaches your desk whether or not you ever open one of these tools

This is not a judicial use case, and it is the one AI issue in the record that arrives without being invited. Four chambers independently report receiving AI-generated filings, mostly from self-represented litigants.

The tells they describe are ordinary human observations: randomly bolded words, uncanny 24-hour turnarounds, legal elements paraphrased rather than quoted. Parties generally admit it when asked at a hearing.

The problem is not primarily fake citations. It is subtly wrong statements of law: elements restated almost correctly, tests recited with a piece quietly missing. A fabricated case announces itself the moment you look it up. A near-miss statement of the governing standard does not announce anything, and it will sit in the record until someone reads it closely enough to catch it. Plan for that failure, not the flashy one.

One caution, and it matters. The patterns above are things people noticed, not the output of a detector. There are no reliable tools to detect AI-generated writing. None. Do not treat any tool's confident claim that a filing is AI-generated as evidence of anything, and be careful about acting on a hunch that a self-represented party used a machine. The wrongness of the law in the filing is something you can actually establish. Its authorship usually is not.

The rules that make lower-risk use possible

Four of these are settled and you can act on them today. One is genuinely open, and this portal is not going to pretend otherwise.

1. The verification duty does not change

Every proposition of law, every citation, every factual assertion, every number that leaves chambers is yours. That was true before any of this, and it is the reason none of this is as unfamiliar as it feels.

Two things about verification are worth stating flatly, because AI tools make them easy to blur:

  • Asking the model to check its own work is not verification. It produces more plausible text, generated the same way as the text you were checking. Verification means opening the case.
  • Fluency is not accuracy. The strongest correlate of an unverified error making it into a final document is that the draft reads well. Good prose lowers your guard. Budget for that.

The practical rule: if a document carries your name, you have read every source it cites. If that is impractical for a given task, the task is not one of the lower-risk tasks.

2. Some things are never delegated

Not to a model, not with verification, not as a first pass:

  • The decision.
  • The weighing of evidence and the assessment of credibility.
  • The determination of what the case is actually about.
  • The individualized judgment in sentencing.

The distinction that runs through this whole portal: labor can be delegated, judgment cannot. A model that reorganizes your reasoning is doing labor. A model that suggests what your reasoning should conclude is doing something else, and the fact that you retain a veto does not fix it. It anchors you, and anchoring is invisible from the inside.

3. Confidential and non-public material

Assume anything typed into an AI tool may be retained on someone else's servers and may be seen by someone else. Some enterprise agreements say otherwise, but the assumption should be the default, and the exceptions should be documented ones you have actually read.

That means, absent explicit authorization for the specific tool:

  • No sealed material.
  • No grand jury material.
  • No presentence reports.
  • No draft opinions, bench memoranda, or deliberative chambers material.
  • No non-public communications among judges or chambers.

This is where the private-device advice on the Start Here tab does double duty. Keeping your experimentation on personal, non-court matters keeps you clear of this entire category while you are learning.

4. Authorization comes before everything

A public October 21, 2025 letter from AO Director Robert J. Conrad Jr. describes interim guidance distributed Judiciary-wide on July 31, 2025. It addresses accountability, confidentiality, security, education, independent verification, core judicial functions, and possible disclosure. Before any AI tool touches court work, confirm what your court and the AO permit.

This is also why the honest advice today is that most judges' learning happens on a private device on non-professional matters. That is not a workaround. It is what the current rules leave available, and it is genuinely the right place to learn anyway.

5. Chambers practice, which is the exposure nobody plans for

Your law clerks have these tools on their phones. Some of them have been using them since law school and are considerably better at it than you are. The realistic risk to a chambers is not a judge who experiments deliberately. It is a clerk who used a tool for a draft and did not think to mention it, because in law school it was ordinary.

The fix is a conversation and a stated rule, whatever the rule turns out to be. Prohibition is a coherent choice. Permission with disclosure is a coherent choice. Silence is not a choice, it is just an unexamined policy. Have the conversation in the first week of the clerkship.

Disclosure

Whether, when, and how a court should disclose that AI assisted in preparing a document is unsettled. Courts are reaching different answers, and this portal is not going to invent a consensus that does not exist.

A chambers self-assessment

Use it as a conversation guide. Answers remain on this device and are not saved.

The whole thing in one line

You sign it. Nothing about these tools changes that, and every guardrail on this page is a consequence of it.

What is out there, and how to choose

A short field guide. It is written mainly to inform the personal, low-stakes experimentation described on the Start Here tab, because that is where most judges' learning can happen right now.

Read this before anything below

Nothing here is authorized for court work by virtue of appearing on this page

The Administrative Office has interim guidance in effect, and the list of AI tools authorized for use in the courts is limited. This page describes what exists in the market. It makes no representation that any of these tools is approved for judicial use, and it is not an endorsement of any vendor. Confirm with your court and the AO before any AI tool touches court work.

General-purpose assistants

These are the frontier chat tools. They are the ones to learn on, because they are the most capable and because everything else in the market is built on the same kind of model. Each has a paid tier at roughly twenty dollars a month, which is the tier that matters.

Claude

Anthropic's assistant. Strong at long documents, careful writing, and analysis where following instructions precisely matters. Handles very long inputs, which is the property that matters most for record review.

ChatGPT

OpenAI's assistant, and the one most people have heard of. Broadly capable across writing, analysis, and research. The most widely used, which means the most third-party material exists explaining how to work with it.

Gemini

Google's assistant. Strong on research and on inputs that mix text with images or video. Integrates with Google's own products, which is either useful or irrelevant depending on what you already use.

Microsoft Copilot

Microsoft's assistant, embedded in Word, Outlook, and the rest of Office. Convenience is its main argument: it is already where you write. Less capable than the standalone tools above at hard analytical work.

Legal research tools with AI layers

These retrieve first and then summarize, which is a meaningfully better risk profile than a general assistant answering a legal question from memory. Better is not risk-free. The model still writes the summary, and it can still mischaracterize a case it correctly found.

Westlaw AI-Assisted Research

Thomson Reuters' AI layer over Westlaw. Natural-language questions answered against retrieved authority, with the underlying cases linked.

The habit to keep: click through. The summary is a starting point for reading, not a substitute for it.

Lexis+ AI

LexisNexis's equivalent. Conversational research, summarization, and drafting against retrieved authority.

The habit to keep: the same one. Read the case.

There is also a set of AI platforms built for law firms and legal departments, Harvey and Legora among them. They are aimed at practice workflows rather than judicial ones, and are unlikely to be relevant to chambers, but judges will hear the names from the bar.

What actually distinguishes these tools

Draft — awaiting input

Matching use cases to tools

The panel's stated goal is a matrix: for each real judicial use case, which tool actually does it best. That matrix cannot be built until the use-case list exists, so it is not here yet, and a speculative version would be worse than nothing.

It goes in this spot once the list from Hon. Maritza Dominguez Braswell arrives and the panel has vetted it.

The bar is ahead of its own policies

This tab is context rather than instruction. Nothing here tells you what to do in chambers — that is what Guardrails is for. It is here because the adoption curve outside the courthouse is already shaping what arrives inside it, and that is worth seeing plainly.

69%

of lawyers use these tools for work

8am, 2026

9%

of firms have a written, enforced AI policy

8am, 2026

30%

of lawyers, up from 11% a year earlier

ABA, 2024

Sources. 8am Legal Industry Report 2026, n > 1,300, surveyed September–October 2025; 8am is a legal software vendor, so read its levels with that in mind. ABA Legal Technology Survey Report, 2024 edition. The two disagree on how high the number is. They agree on the direction, and on the gap between use and governance.

What that gap means for a court

Filings reach you built with tools the filer’s own firm does not govern. The sanctions cases are the visible tail of that, not the shape of it. The ordinary case is a brief that was drafted with assistance nobody recorded, nobody disclosed, and nobody had a written rule about.

The schools moved first, because we had to

Legal education did not get to deliberate. Students arrived in 2023 with these tools already in hand, and a school that waited would have been setting policy for a world that had moved on. The question was never whether to allow them. It was what competence with them looks like, and how to grade it.

Penn Carey Law put AI into the 1L curriculum the year ChatGPT launched. That is not a claim to have gotten it right. It is a claim about the timeline: the schools are two full graduating classes into this, which is why the people arriving in your chambers are not where you are.

Every 1L gets a real account

Secure institutional accounts, extended to writing fellows and teaching assistants as well. Deliberately not a consumer login — the data terms and the model quality are both different, and students who learn on a free tier learn the wrong lessons about capability.

Legal-specific tools alongside general ones

Harvey sits beside the general assistants rather than replacing them. Students see where a legal-specific tool earns its keep and where it does not, which is a judgment they will need on their first day of practice.

Verification is taught first

Coursework runs from access to justice through AI in corporate practice, and Advanced Legal Research was expanded to take these tools seriously. The through-line is the same one on this portal: check the source.

What this means for your chambers

Your next law clerk has never practiced law without these tools. They will not think of using one as a decision worth mentioning, because in law school it was ordinary. That is not a discipline problem and it does not respond to a discipline answer. It responds to a stated rule in the first week, whatever the rule turns out to be — which is the point made at the end of Guardrails, arriving here from the other direction.

Sources and current guidance

Each note states what the source supports and where its limits lie.

Artificial Intelligence in Federal Courts

Supports reported use, attitudes, tasks, and chambers policies among responding federal judges.

Limitation: 112 of 502 sampled judges responded, a 22.3% response rate. Item totals vary, and only six appellate judges responded.

Read the preprint

AO Director letter, October 21, 2025

Publicly describes interim guidance distributed July 31, 2025.

Limitation: The letter summarizes the guidance; it is not the full guidance and does not establish a February 2026 update.

Read the letter

Access to Justice in the Age of AI

Documents federal pro se filing and docket trends during the generative-AI period.

Limitation: This is a working paper. It reports an association, not proof that AI caused the increase.

Read the paper

AI Hallucination Cases Database

Collects judicial matters involving alleged, implied, or found AI-related fabrication and misrepresentation.

Limitation: It changes continually, uses inclusion judgments, and does not measure incidence across all filings.

Open the tracker

Navigating AI in the Judiciary

Provides public judicial use cases and guidance on verification, confidentiality, bias, and responsibility.

Limitation: Professional guidance is not binding federal policy or product authorization.

Read the guidelines

2026 AI in Professional Services Report

Provides cross-sector adoption and use-case context.

Limitation: Toplines combine legal, tax, accounting, risk, fraud, and government respondents; they are not lawyer-only statistics.

Read the report

About this portal

Maintained by R. Polk Wagner with support from the Penn Carey Law AI Project. It does not speak for any court, the Judicial Conference, or the Administrative Office of the U.S. Courts.