Commit cc7c6a8
Point UNBALANCED_DELIMITER at the token it landed in
Last gap from the emitter audit. The kind carried no tokens, so the
ambiguity was locatable only by parsing an offset back out of its
detail string -- which is prose, not an API. The stray character does
survive into a token ('"Nick', 'Smith)'); nothing was pointing at it.
extract_delimited runs before tokens exist, so PendingAmbiguity gains
an `origin` character offset that tokenize resolves to the containing
token's index once the stream is built. Stages after tokenize set
indices directly and leave origin None.
Every emitted kind now carries its tokens:
'Jon "Nick Smith' unbalanced-delimiter ['"Nick']
'John Smith)' unbalanced-delimiter ['Smith)']
'Van Johnson' particle-or-given ['Van']
'John Smith MA' suffix-or-family ['MA']
'JEFFREY (JD) BRICKEN' suffix-or-nickname ['JD']
'Smith, John, Extra, Jr.' comma-structure ['Extra']
The docstring's "May carry no tokens" is narrowed to what remains true:
a character inside a masked region belongs to no token.
test_stage_field_ownership caught tokenize writing a field it did not
own before -- fixed, and classify's entry corrected in the same pass,
since its SUFFIX_OR_NICKNAME emitter had the same omission and only
passed because no case-table row triggers it yet.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>1 parent 3f0b754 commit cc7c6a8
6 files changed
Lines changed: 54 additions & 7 deletions
File tree
- nameparser
- _pipeline
- tests/v2/pipeline
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
104 | 104 | | |
105 | 105 | | |
106 | 106 | | |
| 107 | + | |
107 | 108 | | |
108 | 109 | | |
109 | 110 | | |
| |||
216 | 217 | | |
217 | 218 | | |
218 | 219 | | |
219 | | - | |
| 220 | + | |
220 | 221 | | |
221 | 222 | | |
222 | 223 | | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
47 | 47 | | |
48 | 48 | | |
49 | 49 | | |
50 | | - | |
| 50 | + | |
| 51 | + | |
| 52 | + | |
| 53 | + | |
| 54 | + | |
| 55 | + | |
| 56 | + | |
51 | 57 | | |
52 | 58 | | |
53 | 59 | | |
54 | 60 | | |
| 61 | + | |
55 | 62 | | |
56 | 63 | | |
57 | 64 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
90 | 90 | | |
91 | 91 | | |
92 | 92 | | |
| 93 | + | |
| 94 | + | |
| 95 | + | |
| 96 | + | |
| 97 | + | |
| 98 | + | |
| 99 | + | |
| 100 | + | |
| 101 | + | |
| 102 | + | |
| 103 | + | |
| 104 | + | |
| 105 | + | |
93 | 106 | | |
94 | | - | |
| 107 | + | |
| 108 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
252 | 252 | | |
253 | 253 | | |
254 | 254 | | |
255 | | - | |
256 | | - | |
| 255 | + | |
| 256 | + | |
| 257 | + | |
| 258 | + | |
257 | 259 | | |
258 | 260 | | |
259 | 261 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
239 | 239 | | |
240 | 240 | | |
241 | 241 | | |
| 242 | + | |
| 243 | + | |
| 244 | + | |
| 245 | + | |
| 246 | + | |
| 247 | + | |
| 248 | + | |
| 249 | + | |
| 250 | + | |
| 251 | + | |
| 252 | + | |
| 253 | + | |
| 254 | + | |
| 255 | + | |
| 256 | + | |
| 257 | + | |
| 258 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
46 | 46 | | |
47 | 47 | | |
48 | 48 | | |
49 | | - | |
| 49 | + | |
| 50 | + | |
| 51 | + | |
| 52 | + | |
50 | 53 | | |
51 | | - | |
| 54 | + | |
| 55 | + | |
| 56 | + | |
| 57 | + | |
52 | 58 | | |
53 | 59 | | |
54 | 60 | | |
| |||
0 commit comments