You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
philterd/phileas#385 fixed SSN and TIN detection in the Java implementation for identifiers
written with a Unicode hyphen in place of the ASCII one, and for identifiers wrapped across a
line break. The Python port carries its own patterns in phileas/filters/ssn_filter.py, so it
still has the gap.
Verified against the current main working tree by calling SSNFilter().detect(...) directly:
Input
Result
SSN: 123-45-6789
detected
SSN: 078‑05‑1120 (non-breaking hyphens, U+2011)
not detected
SSN: 078-05-\n1120 (wrapped after the second hyphen)
not detected
totals\n123\n45\n6789
not detected (correct, keep it that way)
the tin is 11-1234567.
detected
All values above are synthetic test data.
The rule implemented in Java
SsnFilter.java builds the separator from two pieces:
A hyphen character class covering the ASCII hyphen plus the soft hyphen (U+00AD), the dashes
U+2010 through U+2015 (which include the non-breaking hyphen U+2011), the minus sign (U+2212),
and the small and fullwidth hyphen-minus forms (U+FE58, U+FE63, U+FF0D).
A line break that counts as part of the separator only when a hyphen precedes it, with
optional horizontal whitespace on either side of the break so an indented continuation line
works. A bare line break is not a separator, which is what keeps unrelated numbers on separate
lines from being joined into an identifier.
The horizontal separator was also widened from a literal space to any horizontal whitespace, so a
tab or a non-breaking space works. No normalization is applied to the input, so spans keep
indexing into the original text and the redaction covers the whole identifier, line break
included.
Intentional exclusions in the Java implementation, worth matching: a break inside a digit group
(078-05-11 then 20), non-ASCII digits (fullwidth, Arabic-Indic), and more than one whitespace
character between groups.
A separate divergence to settle first
_PATTERNS excludes 078-05-1120 and 219-09-9999 by negative lookahead, so SSN: 078-05-1120
produces no span in Python even in plain ASCII form, while Java redacts it. Both are retired
specimen numbers, but 078-05-1120 is also the value the Philter release audit used, and it is
what the new phileas regression fixtures are written against. Decide whether that exclusion stays
before writing the fixtures here, and record the decision either way.
Acceptance criteria
Decide whether the 078-05-1120 / 219-09-9999 exclusion in _PATTERNS stays, and record
the reasoning in the code or the docs.
Every hyphen substitute listed above is accepted wherever an ASCII hyphen is accepted, in
both the SSN and the TIN patterns.
A horizontal separator accepts a tab and a non-breaking space, not only a plain space.
An identifier wrapped across a line break is detected when a hyphen precedes the break,
including when the continuation line is indented and when the break is \r\n.
A line break with no preceding hyphen is not a separator: digits on separate lines produce
no span.
Span offsets index into the original text and the span covers the complete original
identifier, line break included. If normalization is introduced, offsets are mapped back.
Regression tests cover the ASCII control, each hyphen substitute, both wrap forms, and the
negative cases above, including repeated identifiers in one document and surrounding
non-ASCII text.
The final redacted text is asserted, not only the detected spans.
docs/filters.md describes the accepted separators and the intentional exclusions under the ssn section.
Description
philterd/phileas#385 fixed SSN and TIN detection in the Java implementation for identifiers
written with a Unicode hyphen in place of the ASCII one, and for identifiers wrapped across a
line break. The Python port carries its own patterns in
phileas/filters/ssn_filter.py, so itstill has the gap.
Verified against the current
mainworking tree by callingSSNFilter().detect(...)directly:SSN: 123-45-6789SSN: 078‑05‑1120(non-breaking hyphens, U+2011)SSN: 078-05-\n1120(wrapped after the second hyphen)totals\n123\n45\n6789the tin is 11-1234567.All values above are synthetic test data.
The rule implemented in Java
SsnFilter.javabuilds the separator from two pieces:U+2010 through U+2015 (which include the non-breaking hyphen U+2011), the minus sign (U+2212),
and the small and fullwidth hyphen-minus forms (U+FE58, U+FE63, U+FF0D).
optional horizontal whitespace on either side of the break so an indented continuation line
works. A bare line break is not a separator, which is what keeps unrelated numbers on separate
lines from being joined into an identifier.
The horizontal separator was also widened from a literal space to any horizontal whitespace, so a
tab or a non-breaking space works. No normalization is applied to the input, so spans keep
indexing into the original text and the redaction covers the whole identifier, line break
included.
Intentional exclusions in the Java implementation, worth matching: a break inside a digit group
(
078-05-11then20), non-ASCII digits (fullwidth, Arabic-Indic), and more than one whitespacecharacter between groups.
A separate divergence to settle first
_PATTERNSexcludes078-05-1120and219-09-9999by negative lookahead, soSSN: 078-05-1120produces no span in Python even in plain ASCII form, while Java redacts it. Both are retired
specimen numbers, but
078-05-1120is also the value the Philter release audit used, and it iswhat the new phileas regression fixtures are written against. Decide whether that exclusion stays
before writing the fixtures here, and record the decision either way.
Acceptance criteria
078-05-1120/219-09-9999exclusion in_PATTERNSstays, and recordthe reasoning in the code or the docs.
both the SSN and the TIN patterns.
including when the continuation line is indented and when the break is
\r\n.no span.
identifier, line break included. If normalization is introduced, offsets are mapped back.
negative cases above, including repeated identifiers in one document and surrounding
non-ASCII text.
docs/filters.mddescribes the accepted separators and the intentional exclusions under thessnsection.