Document Type
Article
Publication Date
9-2021
ISSN
0007-2362
Publisher
Brooklyn Law School
Language
en-US
Abstract
Legal corpus linguistics usually does something a little different. It uses datasets of language that has nothing to do with the law-articles, novels, TV shows, and so on.12 From these, it
draws conclusions about how people ought to understand language that is used in the law.13 So legal corpus linguistics takes some words used in a statute and tracks how they appear in settings that differ in genre, register, situation, and participants from that of a statute. Then, having assessed how those words are used in those nonstatutory situations, it proposes that we should understand the statutory use of those words the same way they are used in those other places. The logic of using this approach is that it answers long-standing calls for legal interpretation to be guided by ordinary language, and incorporates the recognition that it is often quite difficult to determine what constitutes ordinary language.14 Large datasets of language use promise to give empirical heft to assertions about ordinary language in legal interpretation by showing how language is used in ordinary life.
Yet, I argue below, this promise runs into problems when it confronts the fundamental questions of data-driven methodology: What counts as data for a particular inquiry, and what exactly can that data reveal? The problems stem from the fact that laws are not simply collections of words whose meaning can be copied and pasted from context to context. Rather, laws are utterances that have effects in the world-what linguists sometimes call speech acts or performative utterances.5 Indeed, this power is exactly what makes them such important objects of interpretation. And laws have those effects only because they are enacted under very particular, very unusual conditions that give them power, by speakers who themselves hold unique positions of authority to make law.16 Moreover, laws tend to be written in a genre quite different from the kinds of speech surveyed in generalist databases used by legal corpus linguistics. And academic corpus linguistics itself, as well as a century of work in related disciplines, has made it clear that genre makes a difference both to how people use language and to how they understand it.
In what follows, I consider what kind of data legal corpus linguistics offers, and what it provides data of.18 Legal corpus linguistics generally rests on one of two underlying assumptions about the nature of its data. It assumes that the datasets it uses demonstrate either (1) how ordinary people understand the terms we find in statutes, or (2) how ordinary people use the terms we find in statutes. Findings on (1) or (2) should, the reasoning goes, guide interpretations of those terms in those statutes. The following Parts respectively evaluate to what extent the data that legal corpus linguistics uses actually reveals either (1) ordinary understandings,19 or (2) ordinary uses of statutory terms.20 I conclude that this data reliably provides neither (1) evidence of ordinary understandings nor (2) evidence of ordinary usage. In the last Part, I consider another possible use for legal corpus inquiry: It can help determine whether and to what extent legal texts provide fair notice of legal standards to their audiences- that is, whether they inform people who are subject to the laws what those laws require of them.
The idea that statutes can be evaluated for the kind of notice they provide illuminates an important feature of legal language that legal corpus linguistics tends to overlook: its inherently social nature. Laws, after all, are not just collections of individual terms. They are a form of communication, a fundamentally social process that is only possible on the basis of numberless other social interactions-among those who write, analyze, enact, implement, challenge, and are constrained by the law. Both statutory texts and the social interactions that produce them, moreover, are efficacious not as a function of their word usage, but because they are embedded in particular institutions, authorities, traditions, and ideologies. I suggest that a truly data- driven approach to legal interpretation must take that social context into account. That larger context, in fact, is what makes sense of any particular kind of data about the law.22 In contrast, taking words in isolation, without reference to their role in either linguistic genres or social worlds, inhibits the use of relevant data about law. Law, in short, inheres not in individual words but in institutionally structured social interactions of many different kinds.23 That basic feature should also feature in our evaluation of what counts as data, and what it is data of.
Recommended Citation
Anya Bernstein,
What Counts as Data?
,
86
Brooklyn Law Review
435
(2021).
Available at:
https://scholarship.law.bu.edu/faculty_scholarship/4277
