Data Guide
Since January 2025, PEDP PEDP has archived over 275 datasets, retrieved from federal sources either by downloading from publicly accessible websites or through FOIA requests. We store that data in external repositories such as Harvard Dataverse’s CAFE project, Zenodo, or SciOp.
We also archive federal code repositories and clone git repositories that store federal data, research, or source code, which can be accessed via our Github organization.
Notice any issues? Let us know.
How to use the PEDP data catalog
The PEDP data catalog is a browsable list of all the datasets that PEDP has archived. Its goal is to make the data more discoverable, accessible, and usable. Use the search, filters, and sort functionality to find a dataset of interest, expand the dataset record to view the dataset’s summary and archive notes, and click “Open” to view the detailed dataset information and download the dataset from the external repository.
Term definitions
Agency: The original federal agency or organization that published the dataset. Note that the dataset may have been published by a sub-agency that is more well-known than the parent agency.
Archive Notes: Any notes from PEDP archivers about the process of archiving the dataset, such as scripts used, changes made, additional documentation, missing records, etc.
Backup Location: The external storage site where the dataset archive lives. Could be Harvard Dataverse’s CAFE project, Zenodo, SciOp, Github, a file stored on Google or AWS, or other location.
Dataset Name: The name of the dataset either as documented by PEDP archivers (the “Archive Title” in the metadata) or as given by the publishing agency.
Download Date: The date that PEDP archived the dataset by storing it in an external repository.
Keywords: ISO+ tags for a given dataset, taken from a list of keywords using the ISO Category keywords, plus a few from PEDP, available here.
Metadata: Information about the dataset provided in a standardized format. Not all datasets currently have metadata–collecting and finalizing metadata is ongoing. See more information below.
Sub-Agency: The sub-agency or sub-department that published the dataset. Not all datasets were published by a sub-agency.
Time period: The dates that are covered in the data set, in MM/DD/YYYY-MM/DD/YYYY format. If time period is not available, see the dataset’s details in its backup location.
Metadata
PEDP created a new metadata standard in 2025, inspired by Harvard Dataverse’s CAFE project and with ISO metadata standards with FAIR principles at the forefront.
In doing so PEDP:
- standardized descriptions across document types and archived in different places online
- improved retrieval and sharing of information
- ensured interoperability across systems
- ensure the preservation of its information over time
- improved its archival descriptions
- demonstrated consistency across records management
PEDP’s Data Preservation Working Group are the stewards of the metadata for all archives and are ensuring that the PEDP Metadata Standard is upheld, as well as maintaining the Standard itself and updating it as necessary.
Schema
Every dataset is archived with a metadata document that can be accessed from the dataset’s record in the external repository. The schema for that metadata is documented here.