AI Capabilities and Limitations in Legal Work
Examines the current state of AI in legal work, focusing on proven applications in research, document review, and contract analysis, while identifying critical limitations that affect professional responsibility.
Learning Objectives
- 1Identify the primary applications of AI in legal practice including legal research, document review, and contract analysis
- 2Distinguish between proven AI capabilities and overstated vendor claims in legal technology
- 3Assess the technical limitations of current AI systems that affect legal accuracy and reliability
The Current State of AI in Legal Practice
Artificial intelligence has moved from experimental technology to embedded infrastructure in law practice. Legal research platforms now deliver AI-generated case summaries. Document review systems classify millions of pages for discovery without human review of every document. Contract analysis tools extract key provisions and flag non-standard terms in seconds. These applications are not speculative — they are operational in thousands of firms and corporate legal departments.
The American Bar Association's 2023 Legal Technology Survey Report found that 35% of respondents reported their firms using AI for legal research, 29% for document review, and 19% for contract analysis. The adoption rate among large firms exceeds 50% for document review applications. AI is no longer a competitive differentiator; in many practice areas, it has become table stakes.
But understanding what AI can do requires equal attention to what it cannot do. Overreliance on AI tools has already led to sanctions. In Mata v. Avianca, Inc. (S.D.N.Y. 2023), an attorney submitted a brief citing non-existent cases generated by ChatGPT, resulting in a $5,000 sanction and national headlines. The court emphasized that even when using AI, attorneys remain responsible for ensuring the accuracy of all submissions.
AI in Legal Research: Capabilities and Hallucination Risk
Modern legal research AI operates through two primary methods: retrieval-augmented generation (RAG) and large language models (LLMs) trained on legal corpora. RAG systems retrieve relevant cases from a database and then use AI to summarize or synthesize the results. LLM-based systems generate text based on patterns in their training data, which may or may not include reliable citation to actual cases.
The proven capabilities of current legal research AI include case summarization, identifying relevant precedent from natural language queries, and extracting holding statements from lengthy opinions. Tools like Westlaw's AI-Assisted Research and LexisNexis's Lexis+ AI have demonstrated high accuracy when operating within their designed parameters — specifically, when retrieving and summarizing cases that exist in their proprietary databases.
The critical limitation is hallucination — the generation of plausible but false legal citations, holdings, or quotations. Hallucinations are not occasional errors; they are an inherent feature of how LLMs function. These systems predict the next token in a sequence based on statistical patterns, not truth. When asked to cite a case on a narrow issue, an LLM may generate a case name, citation, and holding that sounds authoritative but corresponds to no real decision.
The risk profile differs by tool design. RAG-based systems that retrieve from verified databases before generating text have lower hallucination rates than pure LLMs. But no current system has eliminated the risk entirely. The ABA Standing Committee on Ethics and Professional Responsibility, in Formal Opinion 512 (2024), emphasized that attorneys using AI for legal research must independently verify all citations and quotations, as the duty of competence under Model Rule 1.1 is not delegable to technology.
AI in Document Review and Predictive Coding
Document review in litigation has been transformed by technology-assisted review (TAR), also called predictive coding. These systems use supervised machine learning: an attorney reviews a sample set of documents and codes them as responsive or non-responsive, and the algorithm learns to predict relevance for the remaining population. Courts have repeatedly held that TAR is acceptable and, in many cases, superior to exhaustive manual review.
In Da Silva Moore v. Publicis Groupe (S.D.N.Y. 2012), Magistrate Judge Andrew Peck approved the use of predictive coding, noting that it "is an available tool and should be seriously considered for use in large-data-volume cases where it may save time and money." Subsequent decisions in Rio Tinto PLC v. Vale S.A. (S.D.N.Y. 2015) and Hyles v. New York City (S.D.N.Y. 2016) have reinforced that TAR is not merely acceptable but often preferable to manual review.
The proven capabilities of TAR include accurately classifying documents by responsiveness, privilege, or issue code; reducing review populations by 70-90% while maintaining high recall; and identifying key documents earlier in the review process. These systems have been validated through cooperative testing protocols and statistical sampling that demonstrates their performance.
The limitations center on validation requirements and edge cases. Courts require litigants to demonstrate that the TAR process is defensible — typically through seed set selection, iterative training, quality control testing, and statistical validation of recall and precision. Parties who fail to document their TAR methodology have faced sanctions or orders to re-review entire datasets. Additionally, TAR performs poorly on highly context-dependent privilege calls and on document types underrepresented in training data.
AI in Contract Analysis and Review
Contract analysis AI performs tasks such as clause extraction, non-standard term identification, obligation extraction, and risk scoring. These tools operate by pattern recognition: they are trained on thousands or millions of contracts and learn to recognize standard clauses (confidentiality, indemnification, termination) and flag deviations.
The proven capabilities include rapidly extracting key business terms from standard contract types, identifying missing provisions that typically appear in comparable agreements, and flagging clauses that deviate from firm-approved templates. For high-volume, low-complexity contracts such as NDAs, SaaS agreements, and employment offer letters, these tools achieve accuracy rates above 95% for clause identification.
The limitations become apparent in non-standard agreements, heavily negotiated provisions, and complex cross-references. AI contract tools are trained on patterns, and they struggle when the pattern does not exist in their training corpus. A bespoke M&A agreement with interrelated earn-out provisions, escrow arrangements, and indemnity caps may be analyzed poorly because the system has not seen enough comparable documents to recognize the structure.
Moreover, contract AI identifies terms but does not assess legal strategy. A tool may flag that a limitation of liability clause is non-standard, but it cannot advise whether the deviation is commercially acceptable given the client's risk tolerance and negotiating position. That judgment remains a core attorney function.
The "Black Box" Problem and Explainability
A persistent limitation across all AI applications is the explainability problem. Many machine learning models — particularly deep learning systems — operate as "black boxes." The model produces an output, but the reasoning process that led to that output is opaque even to the engineers who built the system.
For legal work, this creates a professional responsibility issue. Model Rule 1.1 requires competence, which includes understanding the basis for legal advice. If an attorney cannot explain why an AI tool recommended a particular case or flagged a document as privileged, the attorney may be unable to satisfy the duty to provide competent representation.
Explainability has improved with the development of attention mechanisms and saliency maps that show which inputs most influenced an AI decision. Some document review platforms now provide transparency reports showing which features drove classification decisions. But these tools are not universal, and many commercial legal AI products remain partially or fully opaque.
The ABA's guidance is clear: attorneys must understand how the tools they use generate results. This does not require mastery of machine learning mathematics, but it does require understanding the tool's methodology, training data sources, known limitations, and validation testing. Using a tool you do not understand is incompatible with the duty of competence.
Vendor Claims and Due Diligence
The legal AI market is crowded with vendors making ambitious claims. Attorneys evaluating these tools must apply the same rigor they would to any other business decision: demand evidence, test performance, and verify claims independently.
Key questions for vendor due diligence include: What is the tool's training data source, and how recently was it updated? What is the measured accuracy rate for tasks similar to your intended use? What validation or third-party testing supports the vendor's performance claims? How does the tool handle ambiguous or edge-case inputs? What is the hallucination rate for citation-generating features?
Several vendors have faced scrutiny for overstating capabilities. In 2023, a legal research AI tool marketed as providing "comprehensive case analysis" was found to generate citations with fabricated quotations in approximately 15% of edge-case queries. The vendor subsequently revised its marketing materials and added disclaimer language, but only after attorney complaints and bar association inquiries.
The reality is that AI in legal practice is highly effective for well-defined, repetitive tasks with large training datasets. It is less reliable for novel legal issues, niche practice areas, and tasks requiring contextual judgment. Competent use requires matching the tool to the task and verifying outputs when accuracy is material.
The Competence Imperative
The thread connecting all AI applications in legal work is the non-delegable duty of professional competence. Technology may assist, but it cannot replace the attorney's judgment, verification, and accountability. Every case citation must be checked. Every privilege determination must be defensible. Every contract analysis must be reviewed by a human who understands the client's objectives and risk profile.
In Formal Opinion 512, the ABA made explicit what was already implicit in Model Rule 1.1: competence in the age of AI requires understanding the capabilities and limitations of the tools you use, verifying outputs that inform legal advice or court submissions, and maintaining supervisory responsibility over all work product. The tools have changed, but the professional obligation has not.

