As I’m working on these neighborhood datasets, using newly released census data (2024 American Community Survey Five-Year Estimates), I’m reminded that this process is deeper than throwing data from different census tables into unified datasets. It involves separating the wheat of significant data from the chaff of less significant and redundant data. It involves executing operations that will, in the end, allow patrons to get to the essence of the facts and save time.
It would be great to start the data curation process with source material that is high quality, just like a chef would love to rely on fresh, top-shelf ingredients. Unfortunately, some datapoints in the 2024 American Community Survey (ACS) fall short of this standard. This is especially the case with the poverty rate numbers. Poverty rates are something you have to report given how central this metric is to both housing policy and how neighborhood inequality is understood in peer-reviewed research literature. Multiple studies point out that the deficits in quality are tied to pressures related to timeliness and cost-effectiveness. Hopefully, in the near future, we will exert the pressure necessary to ensure that data so crucial for understanding economically vulnerable communities—is usable.
In terms of the steps in the curation process, data reliability checks come after the initial process of building a foundational dataset. This means combing through the raw census data and purging rows of missing data. This entails deleting those ‘9800’ census tracts and others with virtually zero useful datapoints to incorporate. The next step is purging missing data from the columned datapoints that will serve as anchors for the entire dataset. Once you purge all of the missing data, reliability becomes the focus. The goal I take seriously is balancing transparency with customer service, which translates to sharing reliability scores upfront and allowing patrons to decide if they value data reliability over data coverage, or vice versa. This is a genuine either/or choice, since a fully reliable dataset requires purging neighborhood data that do not make the cut, while emphasizing data coverage means including unreliable data. I have my preferences, but I don’t let that get in the way of client needs.
It’s serious work to translate raw census data into user-friendly neighborhood datasets that combine robustness, reliability, conciseness, and creativity. But it is worth it when it means working with clients that recognize the inherent value of high quality data directed towards equitable prosperity. Looking forward to sharing more in due time.