Independent Project: Local PII Redaction for SAP Free Text
Independent development project without a client commission. The metrics come from an evaluation on 16 held-out documents; they do not establish a production SAP integration.
Problem: Personal data can remain in free text after structured fields have been masked.
Solution: A locally runnable model adapter detects entities and returns structured output with review signals.
Business value: Technical evidence from a small test set. Value and remaining review effort need evaluation for a specific pipeline.
Frame: Independent project without a client commission. An implementation starts with a separately commissioned evaluation billed by the hour.
Independent development, with a bounded evaluation
I built this German PII-redaction model as an independent development project, without a client commission. It explores how a small, locally runnable model can identify personal data in free text that a column-based masking process may miss. It is supporting evidence of my model-development and evaluation work.
The intended setting is a SAP production-to-test data copy. Structured fields already pass through a deterministic masker. Notes, custom text fields and extracted document text need a separate check. This project explores that additional step; it does not document a deployed client pipeline.
What I built
The model is a LoRA adapter on Qwen2.5-1.5B-Instruct. It produces a structured response containing redacted_text, entities, risk_level and needs_human_review. A downstream integration can validate the response and route uncertain cases for review.
The model targets twelve German entity types, including names, addresses, phone numbers and identifiers. The adapter is available on Hugging Face. Training used PEFT/TRL on a RunPod A40. The intended inference setup is local hardware.
A deployment would need to establish which text reaches the model, how missed entities are detected, how pseudonyms remain consistent across records, and what happens when output validation fails. Those integration properties are not established by the model scores alone.
What the experiment measured
The documented v1 experiment used 75 synthetic German business documents for training and 16 held-out documents for evaluation. These are results from that small evaluation, rather than production service levels:
| Metric | Base model | Fine-tuned adapter |
|---|---|---|
| Entity F1 | 0.375 | 0.917 |
| JSON parse validity | 81% | 100% |
| Risk-level exact match | 0.50 | 0.94 |
Review-decision exact match remained at 0.69. That limitation matters: a system must reliably identify the records that still need a person to review them. The next evaluation needs more varied documents, consistent review labels and tests on representative data.
The 100% JSON figure means all responses in this evaluation parsed successfully. It does not guarantee that future outputs will parse or that every personal-data entity will be found. The evaluation does not establish regulatory compliance or safe release of a production data copy.
Commercial value: an assumption to test
A worked example makes the review effort tangible: four copies per year × 2.5 person-days per copy × €600 per day = €6,000 of annual manual-review effort. This is a model calculation, not measured client savings. Integration, compute, ongoing evaluation and remaining human review must be included before estimating a return.
Relevance to my current work
The transferable experience is local model integration, structured outputs and evaluation of failure cases. My current client focus is context layers on temporal knowledge graphs, where source evidence, access rights and authorized actions matter. The operational context-layer case study shows that work in a client setting, using fictional records in the public demo.
Frequently asked questions
Was this a client project?
No. It is an independent development project without a client commission.
What do the metrics describe?
The documented v1 experiment: 75 synthetic training documents and 16 held-out evaluation documents. These figures are not production guarantees.
Stack Stack
- Qwen2.5-1.5B-Instruct (base)
- LoRA adapter (PEFT / TRL)
- Structured JSON output (Pydantic schema)
- HuggingFace Hub (Apache 2.0)
- RunPod A40 (training)
Ähnliches Projekt auf dem Tisch? Similar project on your desk?
Am schnellsten klärt das ein Gespräch. Termin direkt hier wählen: The fastest way to scope it is a conversation. Pick a slot right here:
Your agents answer from whatever the retriever finds, and too often that is last quarter's truth. I build the context layer they answer and act from: a temporal knowledge graph that keeps every fact with its source and the time it held, reads with each person's own permissions, and writes nothing without a person's approval. On your own tenant, billed by the hour, step by step.