As artificial intelligence transforms legal research, one enormous blind spot remains: the local laws governing everything from zoning and housing to business licensing, noise and public health. Unlike federal and state statutes, these everyday regulations remain scattered across thousands of municipal websites, scanned PDFs and proprietary platforms that were never designed for large-scale analysis.
University of Maryland computer science alumni Denis Peskoff, Ph.D. ’21, and Joe Barrow, Ph.D. ’22, are tackling that problem with LOCUS (Local Ordinance Corpus for the United States), the largest machine-readable collection of local law ever assembled in the country.
The project grew out of a striking paradox, Barrow said: Although municipal laws are in the public domain, researchers have never had a centralized dataset for studying them at scale.
“We hope it can benefit the computer science community, legal scholars, and social scientists,” Barrow said.
Working with Peskoff’s postdoctoral adviser, Diag Davenport, an assistant professor of technology policy, governance and society at the University of California, Berkeley, the researchers set out to make that vast body of local law more accessible for computational research.
Drawing from publicly available municipal and county ordinance codes, LOCUS contains 2.2 million local laws from 9,239 cities and counties nationwide. To make the data easier to navigate and compare geographically, the researchers also created a county-harmonized layer covering nearly three-quarters of U.S. counties—2,309 of 3,144 nationwide.
Although the researchers completed the project after graduating from UMD, Peskoff said the natural language processing skills required to work with data at this scale were honed during their doctoral studies in UMD’s Computational Linguistics and Information Processing (CLIP) Lab.
Peskoff was advised by Jordan Boyd-Graber, formerly a professor of computer science at UMD who recently joined Nanyang Technological University in Singapore. Barrow was co-advised by Philip Resnik, professor of linguistics, and Doug Oard, professor in the College of Information. Resnik and Oard are members of the University of Maryland Institute for Advanced Computer Studies (UMIACS), where Boyd-Graber was also a member during his time at UMD.
The large-scale language processing experience Peskoff and Barrow developed at UMD became especially relevant as they confronted one of LOCUS’s central technical challenges: organizing millions of pages of municipal records with widely varying formats.
“On a technical level, it is a challenge to organize millions of pages of vastly different documents,” Barrow said.
Using optical character recognition and custom data harmonization techniques, the team—which also included Christopher Vu, an undergraduate at UC Berkeley—converted that patchwork of records into a searchable, structured corpus designed for legal scholarship and AI applications.
According to Resnik, that emphasis on carefully structuring the data distinguishes LOCUS from datasets assembled primarily through large-scale web scraping.
“Data is the lifeblood of good research, but just collecting up data at scale isn't enough—you also need careful, thoughtful expert-informed decision-making about what to collect, how to organize it, and how to make it usable,” Resnik said. “That’s all too often absent from AI datasets and research these days, but the LOCUS project gets it right.”
Beyond assembling the corpus, the team trained specialized ModernBERT models to analyze local laws at a scale that previously wasn’t possible. The researchers evaluated characteristics including opacity—how difficult a law is for a layperson to understand—and paternalism—the extent to which a law restricts or dictates individual behavior.
Their analysis uncovered patterns across jurisdictions, including that county codes tend to be systematically more complex than municipal codes.
The LOCUS dataset and its derivative models are now publicly available on Hugging Face. The team’s foundational research paper, “Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States,” is currently under review.
“Now that we have a dataset, we are curious to use it for follow-up research projects,” Peskoff said. “We are updating and improving it and collaborating with different researchers. Recently, we’ve also had discussions about studying housing and the maintenance of sidewalks.”
The researchers are already putting LOCUS to work. They authored a follow-up paper examining local fines nationwide, which was presented at the International Conference on Machine Learning (ICML) AI4Law workshop in July.
Peskoff continues his academic research, while Barrow is a research scientist at Adobe, where he works on document processing and develops open tools for efficient vision-language model inference and document AI.
LOCUS now gives researchers a way to systematically examine a layer of American law that has long been difficult to collect, compare and analyze—down to local rules governing issues as varied as housing, fines and sidewalk maintenance.
—Story by Melissa Brachfeld, UMIACS communications group