Home / Personal data, the mandatory caution
What calls for particular caution before having a document processed
Identity, health and personal-situation data are protected categories, and a secret must never pass through a conversation, even briefly.
A working file often mixes several types of information without anyone paying attention: a name, a date of birth, a health condition, a password pasted at the bottom of the page so as not to forget it. Before having this file processed by an artificial intelligence tool, some of this information calls for a caution the rest does not.
Four categories that change the rule
A person's identity, name, date of birth, address, identifies them directly. Their state of health, a diagnosis, a treatment, a situation of dependency, falls under what the CNIL calls sensitive data, regardless of how it is phrased: describing a care need in softer, more clinical language still describes the same condition, and how sensitive a fact is does not depend on the words chosen to say it. Only two things take it out of the category, and only one is within your reach: removing the information from the document, or anonymising it irreversibly, which requires that no one, anywhere, still holds the means to trace it back to the person. Leaving the category is not the only route, in any case. The regulation forbids collecting and using it, then lists five cases where it allows this anyway, including the person's express consent and the public interest in matters of health. These cases exist, and an organisation that ignores them gives up something it has the right to do: they are set out at the address at the end of this lesson, and are settled with your data protection officer, not with this course. A person's personal situation, age, household composition, resources, does not always identify them alone, but combining several of these fields in the same document often narrows the possibilities down to a single person, even when no field taken alone would be enough. A business or access secret, an API key, a password, an authentication token, is not personal data but follows the same rule of caution: a conversation window keeps a trace of what is pasted into it, so a secret that passes through it, even briefly, must be treated as already compromised.
Replacing a name is not enough
Replacing a name with a reference number gives the impression that the document becomes safe to send. That is only true if no correspondence table anywhere still links that number to the person. As long as such a table exists somewhere in the organisation, the document remains personal data, whatever it looks like on the surface. The practical rule fits in one sentence: reduce one identifying detail at a time, rather than judging the document as a whole, and remove whatever is not strictly necessary for the task entrusted to the tool.
The two wordings below are made up for the demonstration, they describe no one. The first is an example of what must not be sent.
À éviter : « Le dossier de la personne suivie, née le 4 mars 1958, domiciliée avenue de la Gare, avec un diagnostic d'insuffisance cardiaque, doit être réactualisé. »
Moins exposé : « Le dossier du bénéficiaire référence 4, tranche d'âge 65 à 75 ans, doit être réactualisé. »
Ce qui reste vrai de la seconde : tant qu'un tableau relie quelque part la référence 4 à un nom, elle porte encore une donnée personnelle.
A secret follows a different path but the same caution. It is not corrected, it is removed from the conversation entirely and kept wherever passwords are kept, in a password manager rather than in the text exchanged with the tool, and certainly not tucked into a corner of a document or an email. If a secret has already been pasted, the only action that repairs the exposure is to revoke it at the source and generate a new one, on the service that issued it. The next lesson spells out what the organisation account guarantees once this initial sorting is done, and what it does not guarantee.
Four categories of data, the example, the risk and the move
| Category | Concrete example | Risk if sent without precaution | Move before submission |
|---|---|---|---|
| Identity | Full name, date of birth, address | Direct re-identification of the person | Remove, or replace with an internal reference |
| Health | Diagnosis, treatment, situation of dependency | Protected category regardless of phrasing | Delete the fact rather than rephrase it |
| Personal situation | Age, household composition, resources | Re-identification by combining several fields | Reduce one field at a time, never the whole document |
| Access secret | API key, password, authentication token | Compromised as soon as it is pasted into a conversation | Store separately, never in the exchanged text |
An employee prepares a file to have it processed by an artificial intelligence tool. He replaces the client's name with a reference number, but keeps in the same document the department, the exact date of an appointment, and a rare job title. The correspondence table between the number and the name stays in a separate file within the organisation.
Write in one sentence what this situation establishes, and in one sentence what it does not establish.
What this establishes: It establishes that the document remains personal data, since the correspondence table still exists and the combination of the department, the date and the job title can be enough to find the person.
What this does not establish: It does not establish that the document is ready to be submitted to an external tool, nor that replacing the name with a number is enough to make it anonymous.
The three most common miscalibrations
- Too broad The document is anonymised as soon as the client's name has been replaced with a reference number.
- Too narrow This situation only concerns keeping the correspondence table, nothing else in the document is a problem.
- Beside the point It establishes that the department and the job title were necessary for the task entrusted to the tool.
- Health data keeps its sensitivity regardless of how it is phrased; removing it or anonymising it beyond recovery takes it out of the category, and five cases provided for by the regulation authorise processing it anyway.
- Several harmless-looking pieces of information brought together in the same document can be enough to identify a person, even when none of them does so alone.
- A reference number does not make a document anonymous for as long as a correspondence table still exists somewhere.
- A secret pasted into a conversation must be treated as already compromised, the only possible repair is to revoke it at the source.
- Reduction happens field by field, removing whatever is not necessary for the task, rather than judging the document as a whole.
Take a real working document, and on a copy, outside any tool, cross out everything not necessary for the task you would like to hand over. Then reread what remains, asking yourself whether the combination of remaining information would be enough to recognise the person. Do not submit anything to Claude before reading the next lesson.
These points depend on an interface or a rule that may have changed since this was written. Check them on your own screen before relying on them.
- The exact qualification of a given document, and the legal basis that authorises or forbids handing it to an external tool, fall to your organisation's data protection officer, not to this course.
- The name and location of the password manager used in your organisation should be asked of whoever administers your machines.
Every datable claim in this lesson links here to the public text behind it. A source that does not open proves nothing.
- CNIL, definition of sensitive data consultée le 2026-09-02
- CNIL, the anonymisation of personal data, and its difference from pseudonymisation consultée le 2026-09-02