parse() turns a name string into a
:class:`~nameparser.ParsedName`. This page explains the model behind
that call: how a string becomes tokens and tokens become fields, how
those tokens get their roles, where configuration lives and why it is
split the way it is, why parsers are plain values, and what happens
when a name is genuinely ambiguous. The task pages all build on these
five ideas.
Every parse follows the same path: the input string is split into
:class:`tokens <nameparser.Token>`, each token is assigned one of the seven roles — title,
given, middle, family, suffix, nickname,
maiden — and every string you read off the result is computed
from those tokens at read time.
Parsing "Dr. Juan Q. Xavier de la Vega III" produces eight
tokens. The first is Dr. with the title role; de, la,
and Vega each carry the family role, which is why
name.family returns "de la Vega" — the field is a view that
joins the family-role tokens in order, not a stored string.
Each token also records where it came from. A :class:`~nameparser.Span` is a pair of
character positions bounding the token in the original string:
Dr. has span (0, 3), and name.original[0:3] is exactly
"Dr.". Internally, spans let the pipeline refer to a token by
position instead of by text. v1 re-found name pieces by searching for
matching text, so a name with a repeated word could make the parser
rewrite the wrong occurrence (issue #100 and its
relatives); a position cannot be confused with a look-alike.
:class:`~nameparser.ParsedName` is frozen: there is no attribute
assignment, ever. If a parse is almost right and you want to fix one
field, you call .replace(), which returns a new
:class:`~nameparser.ParsedName` with that field changed and everything
else — tokens, spans, the rest of the roles — carried over unchanged
(replace() tokens carry no vocabulary tags; :meth:`Parser.revise
<nameparser.Parser.revise>` is the tag-preserving form).
str() renders the default view; nothing about calling it mutates
the value you called it on.
Roles are not assigned in one pass. A vocabulary layer runs first and
claims words for what they are, wherever they sit: titles, particles,
conjunctions, recognized suffixes, and anything set off by nickname or
maiden delimiters. Titles chain, so "Asst. Vice Chancellor" is one
title; particles join forward, so de la attaches to Vega.
Whatever the vocabulary layer has not claimed is left to a positional
layer, which assigns purely by where a word sits: the first unclaimed
word is the given name, the last is the family name, and anything
between them is the middle name. name_order, an explicit comma,
and — for a name written wholly in one East Asian script —
script_orders change what "first" and "last" mean here; the
Chinese interpunct · dividing such a name walks the last of those
back, marking a transcription that keeps its source order; nothing
else does.
This is the whole parser in two sentences, and it explains its
character. A word nameparser has never seen still gets a sensible role,
because the positional layer does not need to recognize anything. The
same word can play different parts in different places — Dr. is a
title before a name and a suffix after it, which is why the field names
title and suffix are really "pre-nominal" and "post-nominal".
And nothing is statistical: there is no model and no training data, so
the same input always parses the same way, and a parse that is wrong is
wrong reproducibly, which is what makes it fixable by configuration.
The split also tells you which container a setting belongs in, before you look anything up: if you are teaching the parser a word, it goes in the :class:`~nameparser.Lexicon`; if you are changing how unclaimed words are arranged, it goes in the :class:`~nameparser.Policy`.
Every piece of nameparser configuration falls into exactly one of three places, and which one is decided by a single question: what does this setting vary with?
- :class:`~nameparser.Lexicon` holds everything that varies by language: the vocabulary — titles, particles, suffixes, conjunctions, and the rest of the word lists the parser matches against.
- :class:`~nameparser.Policy` holds everything that varies by data source or application: the behavior switches — name order, patronymic rules, delimiters, strip flags — anything that changes how the pipeline runs, not what words it recognizes.
- :ref:`Rendering arguments <rendering-arguments>` cover everything
that varies by output destination: the
specyou pass torender(spec), or a keyword toinitials()/capitalized().
"Dean" is a common given name, so it is not in the default titles
vocabulary. But in some data it is more common as a title. The right
reading is a fact about the domain the names come from: that makes it a
:class:`~nameparser.Lexicon` entry. A CRM that wraps nicknames in
square brackets — John [Johnny] Smith — instead of quotes is a
fact about that one data source's export format, not about the
language of the names in it — that's a
:class:`~nameparser.Policy` (nickname_delimiters={('[', ']')}).
One particular report
wanting names formatted as "Family, Given" while every other consumer
of the same parsed data wants "Given Family" is a fact about where the
string is going next, decided at the moment you render it — that's a
rendering argument, not something baked into how the name was parsed.
This replaces v1's single Constants object, which mixed all three
concerns — vocabulary, behavior, and output formatting — into one
mutable bag plus a string_format template string. Sorting a
setting into the right container is largely mechanical once you ask
the "varies by what?" question above; see :doc:`migrate` for the
attribute-by-attribute mapping from the old Constants to the new
containers.
parse() is a convenience function over a module-level default
:class:`~nameparser.Parser`. You only need to build your own
:class:`~nameparser.Parser` (directly, or via
:func:`~nameparser.parser_for`) when you want non-default vocabulary
or behavior — and when you do, build it once and reuse it. Constructing
a :class:`~nameparser.Parser` validates its configuration up front, so
it's cheap but not free; parsing individual names is the hot path, and
a :class:`~nameparser.Parser` is designed to be called many times
without reconstruction.
:class:`~nameparser.Lexicon`, :class:`~nameparser.Policy`, :class:`~nameparser.Parser`, and :class:`~nameparser.ParsedName` are all frozen and hashable. That means they're safe to share across threads without locking, safe to use as dict keys or cache keys, and equality means exactly what it says — two :class:`~nameparser.Parser` instances built from equal configuration are equal values, not merely two objects that happen to behave the same. Every piece of configuration in the 2.0 API is a frozen value — including the module-level default parser itself.
Parsing never raises. Pass in a string that doesn't look like a name
at all, and you get back a :class:`~nameparser.ParsedName` with empty
fields, not an exception. The parser's job is to make a reasonable
call on real-world text, not to reject it. The single exception is
code you supplied yourself: a :data:`~nameparser.Segmenter` passed to
Parser(segmenter=...) runs inside the parse, and its own
exceptions propagate rather than being swallowed — a failure there is
a bug in your callable, not a fact about the name.
Some calls are irreducibly ambiguous — both readings are legitimate,
and no amount of rule-tuning resolves them without breaking some other
name. Those surface as entries on ParsedName.ambiguities instead
of being silently guessed away. The canonical example: a leading "Van"
reads as a given name — the right call for the actor Van Johnson, the
wrong one for a bare "Van Buren", and nothing in the two-word shape
distinguishes them — so the parse records a particle-or-given
ambiguity alongside its answer. You can inspect ambiguities to decide, case
by case, whether your data needs a second look.
An ambiguity records a decision, not a word. The same token in a
different position may present no fork at all: do is in the
ambiguous post-nominal vocabulary, but in "Joao da Silva do Amaral
de Souza" it sits mid-name, where nothing has to choose between
readings — so nothing is recorded. A comma can settle the question
before it arises, too: "Ma, Jack" fixes the family name, so the
credential reading never comes up, while "John Smith MA" has to
call it and says so.
An empty ambiguities is therefore not a certificate of certainty.
Reporting is deliberately partial: :class:`~nameparser.AmbiguityKind`
lists the forks worth flagging, and even those are not reported
everywhere they occur — the comma paths stay quiet on purpose, since a
comma usually settles the structure before the question arises.
Coverage grows over releases. Treat a non-empty ambiguities as a
signal to act on; do not read an empty one as a guarantee.
:class:`Tokens <nameparser.Token>` also carry tags — a second, independent label alongside their
role, recording how a token was classified rather than what part of
the name it belongs to — but only a handful of them are part of the
stable API, collected in :data:`~nameparser.STABLE_TAGS`: particle,
conjunction, initial, and joined.
Any tag written with a namespace prefix, like vocab:..., is
provenance information for debugging how a token got classified — it
can change shape between releases and isn't something to match against
in your own code. If you need to branch on how a token was
classified, branch on role or on one of the four stable tags above,
not on a namespaced one.