Before a dataset can move through an annotation pipeline, someone has to decide what happens to the names, faces, and identifiers sitting inside it. Data anonymization strips or masks that information so the underlying patterns survive but the people behind the data don't stay identifiable. It's a compliance requirement in healthcare and finance, and a quiet expectation everywhere else. The tricky part is doing it early enough that annotators never see the raw identifiers in the first place, not scrubbing them out after the fact.