Skip to main content Accessibility help
×
Hostname: page-component-cd9895bd7-gbm5v Total loading time: 0 Render date: 2024-12-26T11:54:52.063Z Has data issue: false hasContentIssue false

Chapter 18 - Data provenance

from Interlude — Good practices for scientific computing

Published online by Cambridge University Press:  06 June 2024

James Bagrow
Affiliation:
University of Vermont
Yong‐Yeol Ahn
Affiliation:
Indiana University, Bloomington
Get access

Summary

This chapter covers data provenance or data lineage, the detailed history of how data was created and manipulated, as well as the process of ensuring the validity of such data by documenting the details of its origins and transformations. Data provenance is a central challenge when working with data. Computing helps but also hinders our ability to maintain records of our work with the data. The best science will result when we adopt strategies to carefully and consistently record and track the origin of data and any changes made along the way. For instance, we want to know where (by whom) a dataset was created and what was the process used to create it. Then, if there were any changes, such as fixing erroneous entries, we need to have a good record of such changes. With these goals in mind, we discuss best practices for tracking data provenance. While such practices generally take time and effort to implement, making them seem tedious in the short term, over time, your research will become more reliable, and you and your collaborators will be grateful.

Type
Chapter
Information
Working with Network Data
A Data Science Perspective
, pp. 289 - 292
Publisher: Cambridge University Press
Print publication year: 2024

Access options

Get access to the full version of this content by using one of the access options below. (Log in options will check for institutional or personal access. Content may require purchase if you do not have access.)

Save book to Kindle

To save this book to your Kindle, first ensure [email protected] is added to your Approved Personal Document E-mail List under your Personal Document Settings on the Manage Your Content and Devices page of your Amazon account. Then enter the ‘name’ part of your Kindle email address below. Find out more about saving to your Kindle.

Note you can select to save to either the @free.kindle.com or @kindle.com variations. ‘@free.kindle.com’ emails are free but can only be saved to your device when it is connected to wi-fi. ‘@kindle.com’ emails can be delivered even when you are not connected to wi-fi, but note that service fees apply.

Find out more about the Kindle Personal Document Service.

  • Data provenance
  • James Bagrow, University of Vermont, Yong‐Yeol Ahn, Indiana University, Bloomington
  • Book: Working with Network Data
  • Online publication: 06 June 2024
  • Chapter DOI: https://doi.org/10.1017/9781009212601.022
Available formats
×

Save book to Dropbox

To save content items to your account, please confirm that you agree to abide by our usage policies. If this is the first time you use this feature, you will be asked to authorise Cambridge Core to connect with your account. Find out more about saving content to Dropbox.

  • Data provenance
  • James Bagrow, University of Vermont, Yong‐Yeol Ahn, Indiana University, Bloomington
  • Book: Working with Network Data
  • Online publication: 06 June 2024
  • Chapter DOI: https://doi.org/10.1017/9781009212601.022
Available formats
×

Save book to Google Drive

To save content items to your account, please confirm that you agree to abide by our usage policies. If this is the first time you use this feature, you will be asked to authorise Cambridge Core to connect with your account. Find out more about saving content to Google Drive.

  • Data provenance
  • James Bagrow, University of Vermont, Yong‐Yeol Ahn, Indiana University, Bloomington
  • Book: Working with Network Data
  • Online publication: 06 June 2024
  • Chapter DOI: https://doi.org/10.1017/9781009212601.022
Available formats
×