Skip to content

Latest commit

 

History

History
329 lines (279 loc) · 15.3 KB

File metadata and controls

329 lines (279 loc) · 15.3 KB

Locale packs

A locale pack is an opt-in bundle of policy — and, when a naming tradition needs it, vocabulary — for one specific pattern: East Slavic patronymics, Turkic patronymic markers, and so on. Packs apply only when you ask for one by name, and every pack makes the same promise: it never changes a name outside the shapes it declares. Read that precisely — a shape is a pattern, not a language, so a pack can act on a name from a tradition it was not written for. The warning under Using a pack shows what that costs.

What works without a pack

Most international names need no pack at all. The default vocabulary covers five scripts — Latin, Cyrillic, Greek, Arabic and Hebrew, plus Devanagari titles — so honorifics, conjunctions and name particles in those scripts are recognized out of the box:

>>> from nameparser import parse
>>> name = parse("الشيخ محمد بن سلمان")
>>> name.title, name.given, name.family
('الشيخ', 'محمد', 'بن سلمان')
>>> parse("عبد الرحمن محمد").given
'عبد الرحمن'

Native-script entries are safe to enable by default precisely because they cannot collide with Latin-script names. That rules out the reverse: transliterations like sri/shri are not in the default vocabulary, because they are also ordinary given names in Latin script. The same conservatism holds back a handful of native entries that collide within their own script — bare سيد and شيخ (common given names), Hebrew רב (an ordinary word), the Arabic د. abbreviation (it would swallow initials).

If your data is homogeneous enough that a collision can't occur, the reasoning behind the default doesn't apply to you — add the entry:

>>> from nameparser import Lexicon, Parser
>>> lex = Lexicon.default().add(
...     titles={"سيد"}, given_name_titles={"سيد"})
>>> name = Parser(lexicon=lex).parse("سيد محمد")
>>> name.title, name.given
('سيد', 'محمد')

Both fields, because given_name_titles is a marker over titles rather than a separate vocabulary: titles makes the word a title at all, and listing it in given_name_titles says the honorific precedes the given name — as Arabic ones do — so the word after it isn't read as a family name. Listing it in given_name_titles alone raises ValueError rather than quietly doing nothing.

Two East Asian behaviors are on by default for the same reason, except that what selects them is the script rather than the word. The background (covered fully under :ref:`east-asian-names` in :doc:`usage`): Chinese, Japanese, and Korean names put the family name first in native script and are usually written with no space between the parts. Both defaults follow from facts the script alone establishes. A name written wholly in Han or Hangul — or one mixing kanji with kana, a combination only Japanese produces — is assigned family-first, because every language written in those scripts orders names that way; no language guess is involved. An unspaced hangul name is additionally split into surname and given name, because hangul is written by nothing but Korean and Korean surnames are a closed census set that ships as default vocabulary. Splitting an unspaced Han name is the one behavior the script cannot license — the same characters could be a Chinese or a Japanese name, and a Chinese surname list splits Japanese names in the wrong place — so it is not a default: the zh and ja packs below turn it on for data whose language you can declare.

A pack is for something different: a structural rule, like reordering a patronymic, that vocabulary alone can't express.

Using a pack

:func:`~nameparser.parser_for` folds one or more packs onto a base :class:`~nameparser.Parser` (the module default, unless you pass base=). Here the Russian pack reads "Сидоров Иван Петрович" (Sidorov Ivan Petrovich — surname/given/patronymic order) the way a formal Russian document intends:

>>> from nameparser import locales, parser_for
>>> ru = parser_for(locales.RU)
>>> ru.parse("Сидоров Иван Петрович").given
'Иван'

Packs stack: pass more than one pack and their policies fold together in order.

>>> both = parser_for(locales.RU, locales.TR_AZ)
>>> sorted(rule.name for rule in both.policy.patronymic_rules)
['EAST_SLAVIC', 'TURKIC']

Find what's shipped with :func:`~nameparser.locales.available`, and look one up dynamically by its lowercase code with :func:`~nameparser.locales.get` — the same code the --locale flag takes:

>>> locales.available()
('ja', 'ru', 'tr_az', 'zh')
>>> locales.get("ru") is locales.RU
True

The command line accepts the same codes: python -m nameparser --locale ru --json "Сидоров Иван Петрович" applies the pack before parsing, equivalent to parser_for(locales.get("ru")).

Shipped packs
Code Turns on
ja Japanese segmentation — activates division for unspaced Japanese names, which needs a segmenter to act: install nameparser[ja] and pass parser_for(locales.JA, segmenter=locales.ja_segmenter()) (山田太郎 → family 山田, given 太郎). The pack ships no surname list, because no list divides a kanji name, and no order: a Japanese name already reads family-first by default.
ru East Slavic patronymic order — detects a formal given/patronymic/family shape (Cyrillic and transliterated -ovich/-ovna-style endings) and assigns it accordingly.
tr_az Turkic patronymic markers — detects a standalone marker token (oglu, qizi, uulu, and their Latin- and Cyrillic-script variants) and reads the name around it as given/middle/family.
zh Chinese surname segmentation — splits an unspaced Han name into surname and given name (毛泽东 → family , given 泽东) against the surname list the pack ships. It sets no name order: native-script Han already reads family-first without a pack.

ja, ru and tr_az are policy-only — they carry no vocabulary of their own. zh is both halves at once: a surname list, plus the one policy field that turns segmentation on for the script it covers. See :doc:`concepts` for how that split (language vocabulary vs. behavior) is drawn, and Contributing a pack to nameparser for which half a new naming rule belongs in.

Warning

A pack declares a name shape, not a language, and it cannot tell whose name it is looking at. Any surname that happens to end in a patronymic suffix matches the East Slavic rule, including names that are not East Slavic at all:

>>> ru = parser_for(locales.RU)
>>> name = ru.parse("David Michael Abramovich")
>>> name.given, name.family
('Michael', 'David')

The default parser reads that as given David, family Abramovich. This is the trade the pack asks you to make, and it is why packs are opt-in rather than automatic: enable one only for data that is predominantly in the order it detects. If your input mixes traditions, parse the subsets separately with different parsers rather than enabling a pack over all of it.

Segmenters

ja is policy-only in a second sense: it turns division on for Japanese text without supplying anything to divide with. That job goes to a segmenter, which is passed to :func:`~nameparser.parser_for` rather than carried by the pack — a :class:`~nameparser.Locale` is pure data, and a third-party callable is neither pure nor data. :func:`~nameparser.locales.ja_segmenter` is the shipped one; writing your own is worth it for any script whose divisions you know better than a surname list does.

A :data:`~nameparser.Segmenter` is any callable taking a token's text and returning a :class:`~nameparser.Segmentation` — the interior offsets to cut at, plus how confident you are — or None to decline, leaving the token whole. Declining is the load-bearing half of the contract, because segment_scripts unions across packs: your segmenter is offered every token of every activated script, not only the ones its own pack turned on. Recognize the text you can actually read and return None for the rest, rather than answering for a script you never meant to handle. Exceptions are the one thing that does not stay inside the parse — a segmenter is your code, so its errors propagate out of parse() instead of being absorbed as content errors.

The pack half of that arrangement carries no words at all, which makes it the shortest kind of pack there is:

>>> from nameparser import Lexicon, Locale, PolicyPatch, Script
>>> mine = Locale(code="myscript", lexicon=Lexicon.empty(),
...               policy=PolicyPatch(
...                   segment_scripts=frozenset({Script.HAN})))

nameparser/locales/ja.py is the shipped example of exactly that shape: activation is the pack's entire contribution, and a pack applied without a segmenter simply divides nothing.

Creating your own Locale

You don't need to touch nameparser's registry to use your own pack — :class:`~nameparser.Locale` is a plain, constructible value: Locale(code=..., lexicon=..., policy=PolicyPatch(...)). A :class:`~nameparser.PolicyPatch` is a :class:`~nameparser.Policy`-shaped patch: every field defaults to :data:`~nameparser.UNSET` (leave it alone) instead of to a concrete value, so a pack only ever states what it changes. A patch can also be applied directly, without a pack — see :meth:`Policy.patched() <nameparser.Policy.patched>`.

The policy half works that way, but the lexicon half does not. A pack's :class:`~nameparser.Lexicon` is a complete value in its own right and is validated on its own, before it is unioned onto the base — so a fragment that marks a word must also carry the word it marks. To make an existing base title precede the given name, restate the title in the fragment rather than listing it in given_name_titles alone. zh is the shipped worked example: its Lexicon(surnames=...) has to satisfy every Lexicon rule standing alone, before anything unions it onto the base.

>>> from nameparser import Lexicon, Locale, PolicyPatch, parser_for
>>> lex = Lexicon.empty().add(titles={"kapitan"})
>>> mine = Locale(code="mycorp", lexicon=lex,
...                policy=PolicyPatch(middle_as_family=True))
>>> name = parser_for(mine).parse("Kapitan Anna Maria Schmidt")
>>> name.title, name.given, name.family
('Kapitan', 'Anna', 'Maria Schmidt')

That pack does two things at once: the :class:`~nameparser.Lexicon` fragment teaches the parser that kapitan is a title, and the :class:`~nameparser.PolicyPatch` turns on middle_as_family so any remaining given-position words after the first fold into family instead of middle — compare this to the default parser's reading of the same string, which has no title and splits given='Kapitan', middle='Anna Maria', family='Schmidt'.

When parser_for folds one or more packs onto a base, lexicons union (a pack's words are added to the base's, never removed); policy fields declared as set-valued in :class:`~nameparser.PolicyPatch` (patronymic_rules and the delimiter fields) union the same way; and every other, scalar field is later-wins — if two packs (or a pack and an explicit conflicting value) set the same scalar field, the last one applied wins and a UserWarning is raised so the conflict isn't silent.

Contributing a pack to nameparser

Shipping a pack in nameparser itself (rather than keeping it local to your own code) means meeting the in-repo contract, checked mechanically by tests/v2/test_locales.py:

  1. Add a registry entry in nameparser/locales/__init__.py — a "CODE": ("module.path", "ATTR") row in _REGISTRY, so the pack loads lazily on first access (importing nameparser.locales never imports pack modules).
  2. Declare a module-level DEVIATES(name) predicate: given a name string, return whether this pack alone might parse it differently from the default parser. Over-declaring is safe; under-declaring is not — when in doubt, DEVIATES should say yes. A pack whose scope is a script builds the predicate with _script_matcher from nameparser/_policy.py, the way zh and ja do — never by compiling its own character ranges: the factory's docstring explains how the pack-contract test enforces this.
  3. Add a rotator list to tests/v2/test_locales.py. Every pack needs one, but what it has to contain follows from how the pack declares its scope. A pack declaring by marker regex (ru, tr_az) needs at least one name exercising every alternation branch of every regex it defines — test_rotators_cover_every_marker_branch fails until each branch is hit. A pack declaring by codepoint range (zh, ja) has no branches to sweep and drops out of that test, so its rotators have to carry the same weight by hand: the unspaced names the pack must split, one per shape of the vocabulary it ships — single surname, compound surname, and any spelling variant it means to cover. A pack that ships no vocabulary lists the shapes its segmenter must divide instead, and marks the rotator tests to skip when the optional dependency is absent, so the contract tests still run everywhere.
  4. Keep the non-interference gate green over the shared corpus plus your rotators: every name the packed parser parses differently from the default must be one your DEVIATES predicate flags — no silent, undeclared side effects on names outside the pack's stated scope.
  5. Decide which layer the vocabulary belongs in, if the pack carries any. Vocabulary that is self-selecting — able to match only text of the tradition it came from, the way a hangul surname can only ever match hangul — is default-safe, and belongs in the default lexicon (nameparser/config/) rather than in a pack: a pack nobody knows to ask for is vocabulary nobody gets. Vocabulary that declares a language its script does not — a Chinese surname list, which silently mangles the Japanese names written in the same characters — belongs in the pack, where asking for it is the declaration. ja, ru and tr_az need no vocabulary at all and ship an empty :class:`~nameparser.Lexicon`; nameparser/locales/zh.py is the template for one that does.
  6. Curate vocabulary conservatively, the same rule as :doc:`customize`: when you're unsure whether a word or a marker belongs, leave it out.

nameparser/locales/ru.py is the reference implementation for a policy-only pack, nameparser/locales/zh.py for one that carries vocabulary, and nameparser/locales/ja.py for one whose whole contribution is turning a stage on. Packs still in progress are tracked in issue #146 (Vietnamese).