Most people are familiar with the foundational unit of electronic documents, the byte, which is a unit of digital information that consists of 8 bits. Each bit represents a binary value, either 0 or 1. We commonly see it in the gigabyte or terabyte expression used to characterize the size of document collections and it is the basis for most eDiscovery pricing models.
- 1,024 bytes =1 kilobyte (KB
- 1,024 KB= 1 megabyte (MB)
- 1,024 MB = 1 gigabyte (GB)
- 1,024 GB = 1 terabyte (TB)
But take care: the type of document is a key component of document size. By way of example, Shakespeare's complete works in plain ASCII text take up about 5.3 MB while the live version of "Cheeseburger in Paradise" from Jimmy Buffett's 2004 Fenway Park concert is an MP3 file of about 8 MB. So, in today’s ESI world of voice mail and video clips on cell phones, the number of GB is not a true indication of the number of files, which often leads to pricing confusion. To wit, why am I paying so much for so little?
AI programs exacerbate this issue even more because they use a different basic measurement for working with and pricing ESI. It is called a token and is a unit of text that the AI model processes. Tokens can be as small as a single character, like “a” or “b,” or as large as an entire word or subword, such as “hello” or “unbelievable.” For example, in the sentence “The cat sat on the mat,” a model might split each word into tokens while some models use subword tokenization, meaning a longer word like “unbelievable” could be broken into smaller parts: “un-,” “believ-,” and “able.”
So, in most AI models text is broken down into tokens, and the AI generates responses by predicting the next token in a sequence.
So much for the “intelligence” in “artificial intelligence.” In fact, since 2021, numerous computer scientists in machine learning use the term stochastic parrot for this process. They say that large language models, the most common underlying AI programs used in legal document review processes, though able to generate plausible language, do not really understand the meaning of the language they process.
If you’d like a good sleep aid, the first discussion of this position was in the paper, "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? " by Emily Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell. (Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. FAccT '21. New York, NY, USA: Association for Computing Machinery. pp. 610–623.)
Tokenization helps any AI program to process language efficiently but it is important to know how the product you are using handles tokenization because of it's impact on project pricing. Many products treat AI processing as a separate module from their standard data processing and hosting, charging accordingly, some imbed their AI processing charges in their project fees and others handle AI separately, only charging you if you want to use it.
Caveat emptor.
An increasing trend in today's legal technology is to integrate AI into a broader programs so its presence may not be readily apparent to the user. A good example is how Microsoft is increasing integration of CoPilot into its entire Office 365 platform including eDiscovery tool Purview.
A good example for lawyers is both Lexis and Westlaw have made AI part of their legal research process. This is helpful to users because it helps cut down on the well-known phenomenon of "hallucinations" or false case cites that have plagued attorneys who use a general AI tool such as ChatGPT to write briefs. Lexis and Westlaw are using AI in their own curated databases of case law that have been sorted, indexed and cross referenced with subject headings so the likelihood of a false cite is dramatically reduced.
And to further reduce that likelihood, they can also offer cite checking of cases. Lexis+ AI offers AI-powered citation checking integrated into its legal research platform as well as Shepard’s Citations to verify the validity of cited authorities.
It’s important to have this discussion with technical people, either on your firm staff or at a vendor. I recently had to explain to a 30 something techie that Shepard's Citations is a service dating back to 1873 and is used in U.S. legal research to track how cases, statutes, and other legal authorities have been cited over time. (Interesting note: the service was founded in 1873 by Frank Shepard, who noticed that attorneys kept notes on case treatment in the margins of their documents. He developed a system to organize these citations, initially publishing them as adhesive annotations that could be pasted into casebooks.)
Westlaw offers Quick Check, an AI-powered citation verification tool that helps legal professionals ensure their citations are accurate and relevant as well as KeyCite, a citation research service, which tracks case history and citing references to determine whether a legal precedent is still valid.
Be aware that many, if not all, of your legal applications may soon have AI imbedded into their systems and you need to be aware of how that is effecting the use of those applications.
The next wave of systemic integration will likely be by the use of agentic AI as many law firms and legal tech vendors are announcing a steady stream of agentic workflows and AI agents. Examples include Thomson Reuters’ addition of agentic AI capabilities to CoCounsel, Definitely’s launch of agentic AI contracting system Enhance, and Troutman Pepper Locke’s use of agentic AI workflows in its recent merger.
What is agentic AI? Remember the villainous computer system Skynet in the Terminator movies? Just kidding. Sort of.
Agentic AI is a type of generative AI that can autonomously plan and execute a multi-step project. For instance, an AI agent can write a report based on a data set, independently review and detect patterns in datasets, and open, draft and edit documents, among other functions. In essence, it limits the amount of human interaction and can even take some action on its own.
They are built on the same large language model (LLM) technology that serves as the engine for other generative AI tools; however, they combine LLMs with the tools and data access required to carry out given tasks or projects. One observer compared them to buying a bicycle where the LLM functions as the frame, and Agentic AI provides all the other components. The drawback of course is that the bicycle rider has no idea how those parts go together.
Corporations are currently using these systems for customer support inquiries. JPMorgan Chase uses agentic AI to detect fraud, Amazon integrates agentic AI to streamline supply chain optimization and IBM’s AskHR system automates over 80 common HR processes, resolving 94% of employee inquiries without human intervention. In legal, we can expect to see the development of agentic AI systems in document review.
Some of the benefits may include increased efficiency by automating time-consuming tasks, scalability for large projects and increased speed over current AI systems. Potential problems include data security and providing the human oversight that remains vital to crucial legal processes. The ethical implications and the need for this human oversight are essential for any successful implementation of these systems.
Both the Sedona Conference and the EDRM have working groups addressing AI. Sedona has released several publications on the topic including Artificial Intelligence (AI) and the Practice of Law, authored by Judge Xavier Rodriguez, Navigating AI in the Judiciary: New Guidelines for Judges and Their Chambers, and the just released Sedona Canada Primer on Artificial Intelligence and The Practice of Law.
About the Author...

Tom O'Connor
Gulf Coast Legal Technology Center
E-Discovery Committee Chair