AI is everywhere in law firms right now. CIOs are deploying it, managing partners are asking about it, and associates are using it – sometimes before anyone has thought carefully about what value it can actually bring.
The promise is enormous. So is the confusion.
Most of the frustration lawyers and legal professionals have with AI tools, like Harvey, Legora, or Copilot, don’t come from the technology failing. It comes from expecting technology to do something it was never designed to do. And the gap – between expectation and reality – is something IT leaders and trainers can close, if they’re willing to make the invisible visible.
Earlier this year, I walked through what really happens the instant a user submits a prompt and the instant Copilot produces an answer. I've given versions of this presentation to IT leaders and trainers across law firms because the same question keeps coming up: why isn't this working the way we expected? Understanding the sequence is foundational to setting realistic expectations, building effective training, and using AI in ways that hold up in a legal environment.
Copilot is not one thing
Copilot is not a single product. It is a family of experiences that behaves very differently depending on where you use it. Copilot in outlook is grounded in email and calendar context. Copilot in Teams differentiates between chat threads and meeting transcripts. Microsoft 365 Chat enables tenant wide reasoning across your work data.
Licensing matters, too. It doesn’t control how “intelligent” the tool is, but whether you have access to the full Copilot system at all, including its ability to ground responses in organizational data.
This means teaching users where to use Copilot is just as important as teaching them how. The same prompt entered in different applications by different people will produce different results – by design.
What really happens when you hit “Enter?”
Users see a prompt box and an answer. What they don’t see is the structured sequence of decisions that happens in between. That invisibility is the source of most confusion.
1. Context and scope
Copilot first identifies where the question comes from. Is this Word? Teams? Edge? Outlook? Where you start, shapes everything that follows.
A prompt entered in Word against an open contract will behave completely differently from the same prompt entered in Microsoft 365 Chat.
2. Orchestration
Before retrieving any data, the engine interprets the prompt and comes up with a plan to generate a response. What sources need to be consulted, what kind of task was it asked to do – summarization, drafting, extraction.
For legal workflows, this stage determines whether the system reaches only active documents, or whether it initiates a broader tenant-wide search. For ethically walled matters, this is where the system detects restrictions and stops retrieval entirely, returning a safe refusal rather than a partial, leaky answer.
3. Grounding and retrieval
The system retrieves content based on the plan, applying user identity verification and file access permissions before anything is pulled. It is strictly tenant-bound – it inherits permissions and only sees what the user can see. If the user doesn’t have access to the file, Copilot doesn’t either.
Most importantly, the retrieval process involves chunking and often prioritizes headings, defined terms, recitals, and key clauses in legal documents. This means that document structure and metadata quality directly affect what gets retrieved and how useful the output is.
4. LLM reasoning
Only after retrieval does the LLM come into play. It receives a bounded input – users prompt, permission-scoped content, system instructions – and generates a response.
It is important to note, the system operates on the limited data set it was given, not everything it “knows.” It performs transformation and synthesis, not open-ended discovery to avoid claiming specifics that are not present.
5. Output and validation
The system produces an answer, often with citations, summaries, or structured content. This is where human judgment must re-enter the workflow, especially in legal contexts.
The biggest takeaway from this explanation is that output quality is a direct function of retrieval quality. If the system retrieves incomplete or irrelevant data, the reasoning step will reflect that. Copilot isn’t wrong because the AI is unreliable – it is often limited by what it is given to work with.
What it is built for and what it is not
Part of responsible AI training in legal environments is being direct about the boundaries of the tool. Copilot excels at a specific set of tasks like summarization, extraction of structured information, content transformation, and drafting assistance grounded in an existing document.
What it is not designed to do, and where legal professionals must exercise particular caution, is anything that requires judgment. Determining whether a clause is standard, assessing risk, or drawing conclusions from incomplete facts. When a Copilot summary includes language like “this is standard,” or “this appears to be market,” that framing is a hallucination of judgment, and the model has gone beyond summarization into something it isn’t equipped to support.
These hallucinations are the most discussed concern in legal AI — and the consequences are real. Recently an Am Law 50 firm, filed an emergency letter asking a court not to sanction them over AI-generated hallucinations in a filing. This wasn't careless experimentation; it was a sophisticated practice using capable tools without the validation workflow that would have caught the error before it reached a judge. But understanding the specific conditions that cause these errors can make it easier to manage. They’re most likely to happen when the scope of the request is broad and has to stitch together content from multiple sources; source documents are poorly structured; or if the model is expected to infer facts that aren’t explicitly present.
In a legal context, these fabrications materialize in false dollar figures, incorrect case citations, or effective dates; assumed standard clauses when the source materials didn’t contain any; or mischaracterization of the legal effect that isn’t supported by the text.
To counteract these hallucinations, it is best to keep the scope narrow, have the context grounded in a single document, and make the output easy for a trained professional to validate. Traditionally, we teach users to design their workflows for narrow scope – summarizing a single contract in word – and easy validation before expanding to more complex use cases.
The IT angle: Copilot is an amplifier
For CIOs and IT Leaders at law firms, the key takeaway is that Copilot doesn’t create capability from nothing; it amplifies what already exists in your environment, for better or worse. If your access management is intentional, Copilot will reflect that. If permissions are over-broad, it will expose that. If your information architecture is well-structured, Copilot will find relevant content, reliably. If it is disorganized, or poorly named, it will retrieve noise.
Clean identity, well-structured content, properly implemented ethical walls and sensitivity labels, and clear AI policies are the foundation Copilot uses to build. And depending on how well that foundation was made determines how high the ceiling is.
The behavioral challenge
Perhaps the most important lesson learned from this kind of training is one that surprises most people: AI adoption is a behavior problem.
For legal professionals, trust is hard-earned and quickly lost. When Copilot produces an unexpected result – which it will – users need a mental model that helps them understand why. They need to know that scope, permissions, and retrieval quality explain most unexpected outputs. They need clear personal value before they’re willing to change the way they work.
This is where training strategy matters as much as technical configuration. Teaching people through the five-step sequence – context, orchestration, retrieval, reasoning, output – gives them a framework for interpreting results. Teach them what tool to use for each task; structured prompting habits rather than asking open ended questions.
CIOs and IT leaders who take on the role of translator – making the invisible process visible, explaining why AI behaves the way it does, setting appropriate expectations – are the ones who will drive real adoption.
Deploying the tool is the easy part. Building the understanding that makes people actually use it is where the work is.
This is the work I've seen Harbor's AI training and adoption programs do well: giving legal professionals the mental models, prompting habits, and workflow structure they need to use AI confidently. A dedicated training approach, role-based programs, and a cadence that doesn't end at go-live is what turns a technology investment into measurable productivity.
If your firm is working through this, Harbor has structured how it approaches AI training for law firms and legal departments.
- AI
- Training
- Tech adoption
- Show all 5



