People analytics runs on employee personal data, which makes GDPR a design constraint rather than a compliance afterthought. Lawful basis, DPIAs, and what to avoid.
Yes. Employee data is personal data, so every people analytics activity involving identifiable individuals falls within GDPR — including aggregate analysis built from identifiable records. The obligations are not lighter because the data subjects are employees. In several respects they are heavier, because the power imbalance between employer and employee affects which lawful bases are available.
This is not a reason to avoid people analytics. It is a reason to design it properly, and most of the design decisions that satisfy a regulator also produce better analytics — clearer purpose, less unnecessary data, and explicit thinking about who sees what.
This article is general guidance for HR and analytics teams, not legal advice. Confirm your specific position with your data protection officer or counsel.
The instinct is to ask employees to consent. For employment data this is normally the weakest option available.
GDPR requires consent to be freely given, and regulators across the EU have taken the consistent position that an employee is rarely in a position to refuse an employer freely. Consent that cannot be withheld without consequence is not valid consent. It is also withdrawable at any time, which means a retention model built on consent must be able to remove individuals retrospectively — an operational problem most teams do not want.
The bases that usually apply instead:
| Basis | When it fits people analytics | |---|---| | Contract | Data processing necessary to administer the employment contract — payroll, absence, performance records | | Legal obligation | Statutory reporting, health and safety records, pay gap reporting where mandated | | Legitimate interests | The common basis for analytics: workforce planning, retention analysis, engagement measurement — subject to a balancing test |
Legitimate interests is the workhorse, and it is not a free pass. It requires a documented balancing test showing the interest is real, the processing is necessary to achieve it, and the impact on employees does not override it. That test is exactly the discipline that stops analytics teams collecting data because it is available.
Some data used routinely in people analytics is special category data, which requires an additional condition beyond the lawful basis:
Diversity analytics runs directly into this. It is generally achievable — most jurisdictions provide a condition for equality monitoring — but it needs to be identified and documented rather than assumed. In practice this means diversity data should be separated, access-controlled more tightly than general HR data, and reported only at aggregate thresholds that prevent individual identification.
Sickness absence is the one teams miss most often. Absenteeism rate as a count of days is ordinary personal data. Absence reasons are health data, and analysing them requires the extra condition.
A Data Protection Impact Assessment is required where processing is likely to result in high risk. Several standard people analytics activities meet that bar, and the ones that do are exactly the high-value ones:
Attrition and flight risk modelling. Systematic evaluation based on automated processing, producing an assessment about individuals. A flight risk score applied to named employees is close to the textbook DPIA trigger.
Any monitoring of employee activity. Collaboration data, communication metadata, badge data, productivity telemetry. Regulators scrutinise workplace monitoring heavily.
Large-scale profiling for performance or promotion. Especially where the output influences a decision affecting the individual.
Combining datasets employees would not expect to be combined. Engagement survey responses joined to performance data joined to compensation is a common analytics ambition and a common DPIA trigger.
Do the DPIA before building, not after. Its actual value is that it forces the question "what will we do differently as a result of this model?" — and a surprising number of proposed models cannot answer it.
GDPR restricts decisions based solely on automated processing that produce legal or similarly significant effects. In an employment context that covers hiring rejection, promotion, dismissal and pay decisions.
The practical consequences for people analytics:
This is where the "explainable model" argument stops being an academic preference. A logistic regression whose drivers you can articulate is defensible. A gradient-boosted ensemble whose output nobody can explain to the affected employee is a compliance problem regardless of its accuracy.
These are treated as synonyms in most HR conversations and they are legally very different.
Pseudonymised data has identifiers replaced but can be re-linked with a key. It is still personal data and still fully in scope.
Anonymised data cannot be re-linked to an individual by any reasonably likely means. It falls outside GDPR entirely.
Most "anonymised" HR reporting is actually pseudonymised, and a lot of it is not even that. A team of six with one female engineer is not anonymised by removing the name column. This is why minimum group thresholds matter — usually five or more — and why intersectional reporting on small populations is a privacy problem as much as a statistical one.
Apply the threshold consistently and automatically. Manual judgement about whether a cell is small enough fails eventually.
GDPR requires personal data to be kept no longer than necessary. HR systems are among the worst offenders because nothing forces deletion and everything encourages keeping "in case we need it".
Set and enforce retention periods for:
The analytics-specific trap is training data. A retention model trained on five years of exit records holds personal data about every leaver, often long past the retention period applied to their HR file.
Purpose first. Write down what the analysis is for and what decision it changes before requesting the data. This single habit prevents most privacy problems and most useless analytics.
Minimise deliberately. Analytics culture favours collecting everything in case it turns out to be predictive. GDPR requires the opposite, and in practice the discipline improves models — fewer, better-justified features generalise better than everything you could join.
Report at aggregate, act at individual — with a human. Dashboards should be aggregate by default. Individual-level output should exist only where there is a named human process around it.
Separate diversity data. Different storage, tighter access, higher aggregation thresholds.
Document the balancing test. For legitimate interests, a short written assessment per processing activity. It takes an hour and it is the first thing a regulator asks for.
Tell employees plainly. A privacy notice that actually describes what analytics the organisation runs, in language employees can read, does more for trust than any amount of internal governance they never see.
Compliance sets the floor. The thing that determines whether people analytics survives inside an organisation is whether employees believe it is being done to them or for them.
A programme that measures engagement and visibly acts on it earns the next survey's response rate. A programme that scores individuals on flight risk without telling anyone loses the room the moment it leaks — and it always leaks. The governance that satisfies GDPR is largely the same governance that keeps the programme credible internally, which is the more useful reason to do it.
Can we use employee consent as the basis for people analytics? Usually not. Regulators generally consider employee consent invalid because the employment relationship makes free refusal unrealistic, and consent is withdrawable at any time. Legitimate interests, supported by a documented balancing test, is the more appropriate basis for most analytics.
Do we need a DPIA for an attrition model? Almost certainly yes. Systematic profiling of individuals to produce a risk assessment is a standard high-risk trigger. Complete the DPIA before building the model — it will also force you to answer what the organisation will do differently with the output, which is worth the exercise on its own.
What is the minimum group size for reporting survey results? Five responses is the common threshold, and many organisations use higher for sensitive topics or intersectional cuts. Below that, individual responses become inferable, which is both a privacy breach and a guarantee of poorer response honesty in future cycles.
Does GDPR stop us doing predictive people analytics? No. It constrains how. Predictive models are permitted with an appropriate lawful basis, a DPIA, explainable logic, and meaningful human involvement in any decision with significant effect on an individual. What it rules out is fully automated consequential decisions and models nobody can explain.
Analytics with governance built in. PeoplePilot Analytics applies aggregation thresholds, role-based access and explainable models by default, so the privacy design is part of the platform rather than a policy document. Talk to us or explore the HR Metrics Library.