← All projects

PRIVACY·LGPD·2026

PII Detection Models (pt-BR / LGPD)

Open models for finding personal and LGPD-sensitive data in Brazilian Portuguese text, with a live demo and a benchmark of recall against false positives.

The pt-BR PII demo on Hugging Face: a Portuguese-language form for pasting text to be checked for personal data

/ Overview

Most redaction tooling is built for English and trained on the kinds of identifiers US compliance cares about. Brazilian Portuguese — and the LGPD’s specific categories of sensitive data — fall through the gaps. This project is a set of open models, a live demo, and a benchmark aimed squarely at that gap.

The point is practical privacy: catch the data that actually has to be removed before a document leaves the building, in the language and legal frame it was written in.

/ What it does

  • Tags personal and sensitive data in pt-BR text at the token level, including the categories LGPD treats as sensitive.
  • Ships as open models with a runnable demo, so the behaviour can be inspected rather than trusted on faith.
  • Publishes the predictions and metrics behind the comparison, so the trade-offs are visible, not asserted.

/ What the benchmark showed

Two model families are compared: a fine-tune of OpenAI’s Privacy Filter, and smaller GLiNER models. Neither wins everywhere. On the in-distribution Portuguese validation set they are within a point of each other (partial F1 0.897 against 0.888). On a held-out Portuguese source the GLiNER model is well ahead (0.900 against 0.704). On text that contains no personal data at all the order reverses: across 2,419 spam and phishing messages, the Privacy Filter fine-tune raised 64 false positives and the GLiNER model raised 8,666.

For a redaction tool that last number matters as much as F1: a model that flags names everywhere produces documents nobody can read, and reviewers who stop trusting the flags.

/ Approach

Evaluated in the open, with error reporting ahead of a headline number. All evaluation data so far is synthetic; a human-annotated, real-world test set and comparisons against off-the-shelf tools are the next steps, and the results here should be read with that in mind.