Home Blog What Counts as "Large Scale" Processing Under GDPR?

GDPR

What Counts as "Large Scale" Processing Under GDPR?

Posted by Kevin Yun|September 9, 2026

There is no definition and no threshold number. The GDPR uses the phrase in provisions that carry real obligations and then declines to say what it means, leaving Recital 91 and regulator guidance to supply four factors. The most useful thing to know up front is that none of those factors is the size of your company. A ten-person business can process on a large scale and a large one may not.

This article covers why the term is undefined, the four factors regulators actually apply, the worked examples that anchor them, where the phrase carries consequences, and how to reach a defensible position about your own product.

The Regulation Never Defines It, Deliberately

Recital 91 is the closest the text comes, describing large-scale processing operations as those aiming to process a considerable amount of personal data at regional, national or supranational level, which could affect a large number of data subjects. It then gives one clear negative: processing by an individual physician, other health care professional or lawyer, concerning their own patients or clients, should not be considered large scale.

The Article 29 Working Party — the predecessor to the European Data Protection Board — addressed this directly in its guidelines on data protection officers, and its position is worth quoting in substance: it is not possible to give a precise number, either for the amount of data or the number of individuals, that would apply in all situations. It left open the possibility that standard practice might develop over time for common processing activities.

That refusal is not evasion. A fixed threshold would be gamed immediately and would produce absurd results across sectors with very different data intensities. The cost is that you have to reason rather than look up an answer.

The Four Factors Regulators Apply

The Working Party's guidelines recommend four factors, and supervisory authorities including Ireland's Data Protection Commission repeat them:

The number of data subjects concerned — either as a specific number, or as a proportion of the relevant population. The second half matters. Ten thousand people is a different proposition depending on whether it is ten thousand out of a hundred million or ten thousand out of forty thousand.

The volume of data and the range of different data items being processed. Breadth counts as well as depth: many fields about each person weighs more than one field about many.

The duration or permanence of the processing. A one-off collection is not the same as continuous, indefinite processing.

The geographical extent of the processing. A single town differs from several countries.

These are weighed together rather than scored. Nothing says two of four triggers anything on its own.

The Examples That Anchor The Factors

Abstract factors are hard to apply, so the worked examples do more work than the criteria. The Working Party treated the following as large scale: patient data processed by a hospital in the regular course of business; customer data processed by an insurance company or a bank; personal data processed for behavioural advertising by a search engine; and travel data on a public transport system.

Not large scale: an individual physician's patient records, or an individual lawyer's client data including criminal conviction information.

The pattern is instructive. What moves an activity into "large scale" is a combination of population coverage and data intensity, not revenue or staff. A behavioural advertising system and a solo medical practice may hold comparable sensitivity per person; they differ enormously in reach.

Why It Is Not About Your Headcount

This is the point most often got wrong, and it is worth being blunt. The Working Party framed large scale with reference to the number of data subjects rather than the size of the organisation. A company with a handful of employees can be processing on a large scale if it serves a large customer base; a company with many employees serving a small clientele is unlikely to be.

For a SaaS business that inverts the intuition entirely. A four-person analytics product with two million monitored end users is a stronger candidate than a two-hundred-person consultancy with sixty clients. The relevant population is your users and your users' users, not your payroll — which is the same reason almost nothing else in the GDPR turns on headcount either.

The other thing to hold onto is "core activities." The obligations that use "large scale" attach to core activities — the operations inextricable from what you sell — rather than to ancillary processing every company does. Payroll is not a core activity for a software company even though it is continuous and involves sensitive data.

Reaching A Defensible Position On Your Own Product

You will not find a number, so the deliverable is a reasoned, dated, written assessment. In practice that means writing down, per processing activity: how many people, what proportion of the relevant population, how many data items and of what kinds, how long you keep it, and where those people are.

Then state a conclusion and the reasoning behind it. A short document saying "we consider this large scale because of the following three factors" is worth far more than a confident verbal position, because the phrase appears in provisions where being wrong is expensive: the data protection officer obligation, the circumstances requiring a data protection impact assessment, and several places where special category data is involved. Each of those is its own subject and worth taking properly if you land near the line.

The activities that most often push a SaaS towards it are the ones running quietly in the background across the entire user base — and internal tools and dashboards are frequently where that breadth becomes visible, because they are the systems with access to everything at once. All of it belongs in your record of processing activities with the assessment attached.

Common Mistakes With "Large Scale"

Looking for a number. There isn't one, and searching for a threshold produces confident answers from sources that invented them. The absence is deliberate and the correct response is a documented assessment, not a figure.

Assuming a small company cannot qualify. The test counts data subjects, not employees. A tiny team with a large user base is precisely the profile that surprises people when the question is finally asked properly.

Counting only your direct customers. For a B2B SaaS the data subjects are usually your customers' end users, and there are far more of them. Assessing scale against your account list rather than the people in the database understates it by orders of magnitude.

Including ancillary processing in the assessment. The obligations turn on core activities. Payroll, recruitment and internal IT are real processing that belongs in your records, but they are not what makes a software company's processing large scale.

Deciding once and never revisiting. Scale is the one factor that changes on its own while nothing else does. An assessment made at ten thousand users is not evidence about a position at two million, and nobody gets a prompt when the line is crossed.

FAQ

Is there any number of users that definitely counts as large scale?

No, and any source giving one has made it up. The assessment weighs the number of people against the relevant population, alongside data volume, duration and geography. Two million users of a global consumer product and two million records covering most of a small country are treated very differently, which is exactly why a bare number cannot work.

Does processing special category data automatically make it large scale?

No — sensitivity and scale are separate questions. Special category data raises the stakes of the answer, because several obligations combine "large scale" with Article 9 data, but a small volume of health data is still a small volume. The two criteria have to be assessed independently and then read together.

We are a processor, not a controller. Does this apply to us?

Yes. The obligations that use "large scale" apply to processors as well as controllers, and a processor serving several controllers may be operating at a scale that none of them reaches individually. Assess your own aggregate processing rather than assuming your customers' assessments cover you.

Who decides whether we are right?

You do, in the first instance, and a supervisory authority does if it is ever examined. That asymmetry is the argument for writing the assessment down with a date and the factors you weighed. A documented position that a regulator disagrees with is a much better place to be than no position at all.

Closing Thought

The undefined term is doing something deliberate here, and it is worth understanding rather than resenting. A number would let organisations engineer themselves just underneath it, and it would mean the same figure applied to a national health system and a hobby forum. Leaving it open transfers the work to you and, with it, the accountability — which is the pattern the whole regulation follows. You are not told what to conclude; you are required to reason and to show your reasoning.

Which makes the deliverable a short written assessment rather than a search for the right answer. ComplyDog hosts a compliance portal on your own domain covering your DPA, subprocessor list, data subject request handling and security page, and it is the layer buyers read once you have done the underlying thinking. It does not count your data subjects or decide whether your processing is large scale — that judgement is specific to your product, and it has to be yours.