Knowledgebase

Handling Personal Data in Pipelines Print

  • dataengineering, data, backup, guide, howto, solution, zillionkinghost, hosting
  • 0

Privacy in data engineering.

WHAT TO ESTABLISH FIRST

What personal data you hold, and why.

WHAT MINIMISATION MEANS HERE

Not ingesting fields you do not need.

WHY THAT IS THE STRONGEST CONTROL

Data never collected cannot leak, be misused, or need deleting.

WHAT MASKING PROVIDES

Replacing values with obscured versions.

WHAT TOKENISATION PROVIDES

Replacing values with references, with the mapping held separately and tightly controlled.

WHAT THAT ENABLES

Joining and analysis without holding the actual values.

WHAT TO DO AT INGESTION

Mask or tokenise sensitive fields before they land, where analysis does not require them.

WHAT AGGREGATION PROVIDES

Analysis without individual records.

WHAT TO BE CAREFUL WITH

Small groups, where an aggregate identifies an individual.

WHAT TO IMPLEMENT

Minimum group sizes for published aggregates.

WHAT DELETION REQUESTS REQUIRE

Finding every copy: raw storage, warehouse, derived tables, backups, exports.

WHY THAT IS HARD

Data is copied by every pipeline, and copies are rarely tracked.

WHAT MAKES IT POSSIBLE

Lineage, and a record of where personal data lives.

WHAT TO BUILD IN ADVANCE

The ability to locate and remove an individual's data.

WHAT OBLIGATIONS MAY APPLY

Data protection requirements, which are specific. Take advice on your position.


Was this answer helpful?
Back

Are you happy with your experience? Leave us a review on Trustpilot.


Trustpilot