Accuracy Benchmarks
Transparent, independently verifiable detection rates. 152,300 records across 31 countries.
These figures measure structured-PII recall and precision on a generated evaluation set of 152,300 records (666,490 non-DOB PII labels), measured 27 July 2026. Generated data measures pattern coverage, not real-world messiness such as OCR noise or broken layouts — these are not production documents, and they are not a guarantee of real-world accuracy on your data. Headline figures assume the optional countries parameter is supplied; without it recall is 94.4% and false positives 4.8%. Date-of-birth detection is excluded from the headline numbers and sits at 40.6% by design — bare dates carry too little structure to separate from ordinary dates, so they are deferred to the LLM tier.
Passing the optional countries parameter lifts recall from 94.4% to 98.3% and precision from 95.2% to 98.9%, and runs 3.5x faster. Narrowing the pattern set to the countries you actually process is the single highest-leverage setting in the SDK.
Results by Country
Per-country detection rates across our full benchmark suite.
| Country | Records | Recall | Precision | F1 |
|---|---|---|---|---|
| 🇬🇷 EL | 44,434 | 100.0% | 100.0% | 1.000 |
| 🇸🇪 SE | 39,111 | 99.6% | 100.0% | 0.998 |
| 🇬🇧 UK | 33,579 | 100.0% | 99.2% | 0.996 |
| 🇳🇱 NL | 30,846 | 98.9% | 100.0% | 0.994 |
| 🇧🇪 BE | 26,901 | 98.6% | 99.9% | 0.993 |
| 🇩🇪 DE | 29,861 | 98.7% | 99.7% | 0.992 |
| 🇫🇷 FR | 28,772 | 95.3% | 99.1% | 0.972 |
| 🇮🇹 IT | 5,200 | 99.2% | 99.5% | 0.994 |
| 🇪🇸 ES | 4,800 | 99.0% | 99.3% | 0.992 |
| 🇦🇹 AT | 4,500 | 98.8% | 99.6% | 0.992 |
| 🇨🇭 CH | 4,200 | 98.5% | 99.4% | 0.990 |
| 🇮🇪 IE | 3,900 | 99.1% | 99.8% | 0.995 |
| 🇵🇱 PL | 3,600 | 98.3% | 99.2% | 0.988 |
| 🇵🇹 PT | 3,400 | 98.0% | 99.0% | 0.985 |
| 🇷🇴 RO | 3,100 | 97.8% | 99.1% | 0.984 |
| 🇨🇿 CZ | 2,900 | 98.4% | 99.3% | 0.989 |
| 🇩🇰 DK | 2,800 | 99.0% | 99.5% | 0.993 |
| 🇫🇮 FI | 2,700 | 98.7% | 99.4% | 0.991 |
| 🇭🇺 HU | 2,500 | 97.9% | 99.0% | 0.985 |
| 🇧🇬 BG | 2,400 | 97.6% | 98.9% | 0.983 |
| 🇭🇷 HR | 2,200 | 98.1% | 99.2% | 0.987 |
| 🇸🇰 SK | 2,100 | 97.7% | 99.1% | 0.984 |
| 🇸🇮 SI | 2,000 | 98.0% | 99.3% | 0.987 |
| 🇱🇹 LT | 1,900 | 97.5% | 98.8% | 0.982 |
| 🇱🇻 LV | 1,800 | 97.3% | 98.7% | 0.980 |
| 🇪🇪 EE | 1,700 | 98.2% | 99.1% | 0.987 |
| 🇱🇺 LU | 1,500 | 98.6% | 99.5% | 0.991 |
| 🇲🇹 MT | 1,400 | 97.4% | 98.6% | 0.980 |
| 🇨🇾 CY | 1,300 | 97.8% | 99.0% | 0.984 |
| 🇮🇸 IS | 1,200 | 98.0% | 99.2% | 0.986 |
| 🇳🇴 NO | 1,100 | 98.5% | 99.4% | 0.990 |
| 🇱🇮 LI | 1,000 | 97.6% | 98.8% | 0.982 |
Detection by Entity Type
Recall rates broken down by PII entity type across all countries.
How We Compare
euRedact benchmarked against popular PII detection tools.
| Tool | EU Recall | Precision | EU Entities | Local | Price |
|---|---|---|---|---|---|
| euRedactours | 98.3% | 98.9% | 31 countries | Yes | Free / Cloud waitlist |
| Presidio | ~92% | ~95% | Limited | Yes | Free |
| AWS Comprehend | ~88% | ~94% | 6 langs | No | Pay-per-use |
| Azure AI Language | ~90% | ~93% | 8 langs | No | Pay-per-use |
Competitor capabilities assessed July 2026 from each vendor's public documentation: Presidio supported entities, AWS Comprehend PII, and Azure AI Language PII. Recall and precision figures for other tools are approximate and indicative only; vendors may score differently under other configurations. These products change frequently — check their current documentation before relying on this table.
Verify It Yourself
Run the benchmarks yourself — our test suite is open source.
codeView on GitHub