Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions docs/customize.rst
Original file line number Diff line number Diff line change
Expand Up @@ -197,6 +197,17 @@ lost and those words fall back to the all-capitals acronym repair:
``john smith phd`` would give ``John Smith PHD`` rather than
``John Smith PhD``.

A mask also outranks the lowercase that case repair gives a surname
particle (``de la Vega``). The shipped map uses that for the Irish
particles, which are written capitalized:

.. doctest::

>>> str(parse("SEÁN Ó MURCHÚ").capitalized())
'Seán Ó Murchú'
>>> str(parse("JUAN DE LA VEGA").capitalized())
'Juan de la Vega'

How a key matches a word
^^^^^^^^^^^^^^^^^^^^^^^^

Expand Down
16 changes: 15 additions & 1 deletion docs/design/decisions.md

Large diffs are not rendered by default.

28 changes: 27 additions & 1 deletion docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -954,6 +954,27 @@ P6. Rationale: a particle ending the name has nothing to link
should expect that test, not this file, to say so first.
history: decisions.md#P6 · interacts: A1, C1, P1, S2, P5, M2 · implemented: nameparser/_pipeline/_post_rules.py

P7. Rationale: a one-letter particle is spelled with the same letter
as an initial, and the period is what tells them apart. An
initial is an abbreviation and is written with one; a particle is
a whole word and is not. The Irish Ó is never written with a
period, while Ó. is the initial of Óscar, Ólafur or Ólöf.
A one-letter word written with its period is an initial and not a
particle, whatever particle the same letter spells; written bare,
it is the particle. It is an initial in every position, so it
neither opens a surname, nor joins a chain, nor attaches to the
family behind a comma.
"Juan Ó. Pérez" → middle="Ó."
"J. Ó. Pérez" → middle="Ó."
"Ó. Pérez" → given="Ó."
"Pérez, Juan Ó." → middle="Ó."
"Juan Ó Pérez" → family="Ó Pérez" · boundary
Accepted: an accented initial written WITHOUT its period reads as
the particle. Such initials normally carry the period, and the
bare letter is exactly how the particle is written.
"Juan Ó Pérez" → middle=""
history: decisions.md#P7 · interacts: P1, P2, P3, P6, R3, R4 · implemented: nameparser/_lexicon.py, nameparser/_pipeline/_classify.py

## Suffixes: generational & credentials (S)

Background: what follows a name is one of two different things — generational suffixes (Jr., III), which attach to the name itself, and credentials (PhD, MD, MBA), which are earned attachments. The suffix sets match one written word at a time: a multi-word entry there can never match anything and is warned about at configuration. Only two sets are exempt from that rule -- given_name_titles and, since #434, maiden_markers -- and neither is a suffix set, so within the suffix vocabulary the one-word rule is absolute. That limit is on STORAGE, not on the shape a name may have: adjacent suffix tokens are reassembled after matching (`_vocab.is_wholly_suffix`), so a multi-word credential is reachable as its component words -- `John Smith, MD PhD` has read suffix `MD PhD` since 1.4.0 -- and a caller reaches an unshipped one by adding the words it is made of rather than the phrase (#433). The eight multi-word entries that shipped dead for years span the suffix sets and the titles alike and are the Excluded story in decisions.md. CLDR personNames keeps them as separate fields (`generation`, `credentials`) and formats them differently; this library currently reports both in one `suffix` field, a merge #326 examines. The vocabulary is largely split already: a generational word list and a credential acronym list, plus a short list of acronyms that are also ordinary names (MA, BA) and so are AMBIGUOUS as bare words.
Expand Down Expand Up @@ -2586,7 +2607,10 @@ R4. Rationale: case repair is a display concern, applied only on
Case repair returns a repaired copy and never mutates the parse.
Where it acts at all — R5 decides where — the copy honors the
casing a vocabulary entry records (PhD, BSc) and the Mac/Mc
convention (McDonald), not only ordinary word-by-word casing, and
convention (McDonald), not only ordinary word-by-word casing; a
particle is written in lowercase, but a particle the vocabulary
records a casing for takes that casing, the entry being how the
word is written wherever it stands; and
a part whose every word is particle vocabulary is repaired as
ordinary name words, since none of them is doing a particle's
work there (R2). A CONNECTIVE the parse placed among the name
Expand Down Expand Up @@ -2662,6 +2686,8 @@ R4. Rationale: case repair is a display concern, applied only on
"Juan McDonald" → capitalized_forced="Juan McDonald"
"ANH DO" → capitalized="Anh Do"
"anh van do" → capitalized="Anh Van Do"
"SEÁN Ó MURCHÚ" → capitalized="Seán Ó Murchú"
"ina binti navalamar" → capitalized="Ina binti Navalamar" · boundary
"john smith phd" → capitalized="John Smith PhD"
"john smith ph.d." → capitalized="John Smith Ph.D."
"JOHN SMITH PH.D." → capitalized="John Smith Ph.D."
Expand Down
2 changes: 2 additions & 0 deletions docs/release_log.rst
Original file line number Diff line number Diff line change
Expand Up @@ -74,6 +74,8 @@ Release Log

- **Fix halfwidth corner brackets not being read as a nickname.** ``HumanName("山田 「タロー」 タロウ")`` gives nickname ``タロー``, last ``山田``, first ``タロウ``, where every release gave middle ``「タロー」`` and first ``山田`` (the order moves with the halfwidth katakana change above; the bracket would otherwise have blocked it). The halfwidth ``「」`` are the corner brackets of legacy JIS X 0201 data, the same punctuation as ``「」``, and are now a default nickname pair in both APIs: ``DEFAULT_NICKNAME_DELIMITERS`` gains ``("「", "」")`` and the 1.x ``nickname_delimiters`` gains the key ``halfwidth_corner_brackets``. They are not limited to Japanese text: ``John 「Jack」 Smith`` gives nickname ``Jack`` where it gave middle ``「Jack」``. A ``Constants`` restored from a pickle keeps the keys it was saved with, as it did when 2.0 added the other typographic pairs. See the ``N1`` entry of ``docs/design/decisions.md`` (closes #597)

- **Fix the Irish particles Ó, Ní and Ua and the Malay binti being read as a middle name.** ``HumanName("Liam Ó Murchú")`` gives first ``Liam``, last ``Ó Murchú``, where 1.4.0 through 2.3.0 gave middle ``Ó``, last ``Murchú``; ``Sinéad Ní Mhurchú``, ``Seán Ua Buachalla`` and ``Ina binti Navalamar`` (and the Singapore spelling ``binte``) move the same way. ``Ó`` and ``Ní`` are never given names, so ``Ó Murchú`` alone is all last name, where every release gave first ``Ó``. ``Ua`` and ``binti`` can be, so a leading one stays the first name and ``parse()`` reports ``particle-or-given``: ``Ua Buachalla`` gives first ``Ua``, last ``Buachalla``, as before. ``Ó.`` written with a period is still an initial: ``Juan Ó. Pérez`` keeps middle ``Ó.``, while ``Juan Ó Pérez`` gives last ``Ó Pérez``. Case repair writes the Irish particles capitalized, ``SEÁN Ó MURCHÚ`` repairing to ``Seán Ó Murchú``, and ``binti`` in lowercase: ``INA BINTI NAVALAMAR`` repairs to ``Ina binti Navalamar``, where 2.3.0 gave ``Ina Binti Navalamar``. The Irish casing comes from new ``capitalization_exceptions`` entries, which now outrank the lowercase case repair gives a particle, and that holds for your own entries too: with ``constants.capitalization_exceptions['van'] = 'Van'``, ``ludwig van beethoven`` repairs to ``Ludwig Van Beethoven``, where 1.4.0 through 2.3.0 kept ``van``. See ``P7`` and the #604 entries under ``vocabulary-collisions`` and ``R4`` in ``docs/design/decisions.md`` (closes #604)

**Additions**

- **Add Lexicon.conjunctions_ambiguous, the one-letter connectives that read as initials.** A subset of ``conjunctions`` holding ``e`` and ``i`` by default; it is the knob for the change above rather than a switch. Portuguese data, where ``e`` links surnames the way ``y`` does in Spanish, takes it out: ``Lexicon.default().remove(conjunctions_ambiguous={"e"})`` restores the joining reading. Dutch data, where a bare single letter is an initial and never a connective, adds the other one: ``Lexicon.default().add(conjunctions_ambiguous={"y"})``. A v1 ``Constants`` has no manager of its own for it -- deleting the word from ``conjunctions`` is what turns the marking off, the same rule the glued-honorific tails follow. See ``docs/customize.rst`` (#383, #479)
Expand Down
6 changes: 4 additions & 2 deletions nameparser/_facade.py
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@

import nameparser._render as _render
from nameparser._config_shim import CONSTANTS, Constants, _cached_parser
from nameparser._lexicon import _normalize
from nameparser._lexicon import _normalize, _spells_an_initial
from nameparser._parser import Parser
from nameparser._types import (FOLDED_TAG, UNCLASSIFIED_TAG,
UNJOINED_CONJUNCTION_TAG, ParsedName,
Expand Down Expand Up @@ -485,7 +485,9 @@ def given_names(self) -> str:

def _is_particle(self, text: str) -> bool:
self._resolve()
return _normalize(text) in self._lexicon.particles
n = _normalize(text)
return (n in self._lexicon.particles
and not _spells_an_initial(n, text))

def _token_is_conjunction(self, tok: Token) -> bool:
# #528: the PARSE's answer, not the vocabulary's. A token the
Expand Down
25 changes: 25 additions & 0 deletions nameparser/_lexicon.py
Original file line number Diff line number Diff line change
Expand Up @@ -175,6 +175,31 @@ def _normalize(word: str) -> str:
word = stripped


# rules.md#P7: "A one-letter word written with its period is an
# initial and not a particle"
def _spells_an_initial(n: str, text: str) -> bool:
"""Whether `text` is ONE letter written with its period ('Ó.',
'j.'), which makes it an initial whatever particle the letter also
spells (#604). `n` is `_normalize(text)`, folded once by the caller.

Every particle test asks this AFTER its membership hit, so a word
no particle set holds never pays the frame. It lives here rather
than in _pipeline._vocab because the two views that re-read the
particle vocabulary over a parsed name -- _render's case repair and
the facade's initials and last-name split -- must ask it too, and
neither may import the pipeline.

The ASCII period only, as `_vocab.is_initial` takes it: a CJK full
stop marks no initial ('Ó。' is the particle). The other half of
is_initial, the scripts that write no initials (#320), is NOT asked
-- its script table is _policy's, which this module may not import
-- so a caller's one-character particle in such a script, written
with a period, is vetoed where is_initial would call it no initial
(decisions.md#P7). No shipped particle is one character of those
scripts."""
return len(n) == 1 and text[-1] == "."


def _fold_words(words: Iterable[str]) -> list[str]:
"""The words of a title run, folded for storage and lookup.

Expand Down
11 changes: 7 additions & 4 deletions nameparser/_pipeline/_classify.py
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@
from __future__ import annotations


from nameparser._lexicon import _normalize
from nameparser._lexicon import _normalize, _spells_an_initial
from nameparser._policy import CapsSuffixes
from nameparser._pipeline._state import (
AMBIGUOUS_ACRONYM_TAG, SHAPE_ACRONYM_TAG, ParseState, PendingAmbiguity,
Expand Down Expand Up @@ -117,10 +117,13 @@ def _tags_for(token: WorkToken, n: str, state: ParseState,
tags.add("vocab:suffix-word")
if n in lex.suffix_acronyms_ambiguous:
tags.add(AMBIGUOUS_ACRONYM_TAG)
if n in lex.particles:
# rules.md#P7: "A one-letter word written with its period is an
# initial and not a particle" -- both particle tags, so 'Ó.' is
# neither half of the vocabulary
if n in lex.particles and not _spells_an_initial(n, token.text):
tags.add("particle")
if n in lex.particles_ambiguous:
tags.add("vocab:particle-ambiguous")
if n in lex.particles_ambiguous:
tags.add("vocab:particle-ambiguous")
# rules.md#P3: "a single-letter connective reads as an initial
# where the writing says so: written as a bare Latin capital in a
# name that is not written wholly in one case, or — in a name
Expand Down
4 changes: 2 additions & 2 deletions nameparser/_pipeline/_vocab.py
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@
from typing import Literal

from nameparser._lexicon import (
FULL_STOPS, Lexicon, _VOCAB_FIELDS, _normalize,
FULL_STOPS, Lexicon, _VOCAB_FIELDS, _normalize, _spells_an_initial,
)
from nameparser._policy import (CapsSuffixes, Policy, Script, _JA_SCRIPTS, _NO_INITIALS,
_SCRIPT_RANGES, _script_matcher)
Expand Down Expand Up @@ -998,7 +998,7 @@ def surname_unit_tags(text: str, lexicon: Lexicon,
"." in text and n not in lexicon.titles
and period_joined_vocab(text, lexicon) == "suffix")
particle = n in lexicon.particles and not (
leading and n in lexicon.titles)
(leading and n in lexicon.titles) or _spells_an_initial(n, text))
if suffix:
return SURNAME_UNIT_TAGS if particle else _SUFFIX_ONLY
return _PARTICLE_ONLY if particle else _NO_TAGS
Expand Down
10 changes: 8 additions & 2 deletions nameparser/_render.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,8 @@
import unicodedata
from collections.abc import Callable

from nameparser._lexicon import FULL_STOPS, Lexicon, _normalize
from nameparser._lexicon import (FULL_STOPS, Lexicon, _normalize,
_spells_an_initial)
from nameparser._types import (FOLDED_TAG, SHAPE_ACRONYM_TAG,
UNCLASSIFIED_TAG, UNJOINED_CONJUNCTION_TAG,
UNJOINED_TAG, Ambiguity, ParsedName, Role,
Expand Down Expand Up @@ -441,8 +442,13 @@ def _cap_word(word: str, role: Role, tags: frozenset[str],
generation = role is Role.SUFFIX and "vocab:suffix" in tags
if not generation and (
(normalized in lex.particles
and not _spells_an_initial(normalized, word)
and role in (Role.MIDDLE, Role.FAMILY)
and UNJOINED_TAG not in tags)
and UNJOINED_TAG not in tags
# rules.md#R4: "a particle the vocabulary records a
# casing for takes that casing" -- the mask below
# outranks the particle's lowercase (#604)
and normalized not in lex.capitalization_exceptions_map)
or "conjunction" in tags
or (UNCLASSIFIED_TAG in tags
and _reads_as_conjunction(word, lex))):
Expand Down
10 changes: 10 additions & 0 deletions nameparser/config/capitalization.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,16 @@
'rph': 'RPh',
'thd': 'ThD',
'thm': 'ThM',
# Irish patronymic particles (#604). Case repair writes a particle
# in lowercase ('de la Vega'), but Irish writes these capitalized
# ('Liam Ó Murchú', 'Sinéad Ní Mhurchú'), and a mask outranks the
# particle lowercase (rules.md#R4). Each mask is the word's own
# title case, so a bare word in any other role repairs as it would
# anyway; the dotted 'u.a.' and 'n.í.' now repair to 'U.A.' and
# 'N.Í.' (were 'U.a.' and 'N.í.').
'ní': 'Ní',
'ua': 'Ua',
'ó': 'Ó',
}
"""
Words whose case ``str.capitalize()`` gets wrong, each mapped to a
Expand Down
25 changes: 25 additions & 0 deletions nameparser/config/particles.py
Original file line number Diff line number Diff line change
Expand Up @@ -90,7 +90,19 @@
# Donald" gave given='Mc'. Also in SUFFIX_ACRONYMS as
# Master of Ceremonies, but that is trailing position and
# unaffected -- "John Smith MC" still reads suffix (#360).
'ní', # Irish "daughter of" ("Sinéad Ní Mhurchú"), the female
# patronymic beside 'ó'; never a given name (#604). Its
# sibling 'nic' is EXCLUDED: an abbreviation of Nicolas and
# Nicole and a given name of its own (Nic Cage), so it
# belongs in an opt-in Irish pack, not the defaults.
'op', # Dutch "at/on" ("op den Berg")
'ó', # Irish "descendant of" ("Liam Ó Murchú"); never a given
# name (#604). The one ONE-LETTER particle, so the letter
# written with its period is still the initial it
# spells -- 'Juan Ó. Pérez' is Óscar, not a surname
# (rules.md#P7). Unaccented 'o' is EXCLUDED: bare 'O' is
# a common periodless initial, and the anglicized form
# is glued ("O'Neil") anyway.
'ste', # Contraction of 'Sainte'; has a vowel, which is why the
# criterion is "abbreviation of a word that is never
# itself a name" rather than the vowel shape #360 first
Expand Down Expand Up @@ -196,6 +208,12 @@
'bin', # Arabic "son of", but kept ambiguous deliberately: #269
# judged the Latin transliteration separately from the
# native-script بن, which IS never-given
'binte', # Singapore spelling of 'binti', kept beside it (#604)
'binti', # Malay "daughter of" ("Ina binti Navalamar"), the
# female partner of 'bin'. Binti is also a given name
# (Swahili, "daughter"), so ambiguous like 'bin' (#604).
# The abbreviations 'bt'/'bte' are left out: two- and
# three-letter words nobody has reviewed
'bon',
'da',
'dal',
Expand Down Expand Up @@ -239,6 +257,13 @@
# is why attestation rather than etymology decides
'tho',
'thoe',
'ua', # Irish "descendant of", the older spelling of 'ó'
# ("Seán Ua Buachalla"). No given-name use found, but
# searched less thoroughly than 'ó' and 'ní', and C-i
# defaults to AMBIGUOUS under uncertainty -- so the
# leading "Ua Buachalla" keeps the given reading and
# reports it, while mid-name the chain is the same
# either way (#604)
'van', # Vietnamese Văn, and Van Johnson -- the canonical
# particle-or-given ambiguity this whole flag exists for
'vande',
Expand Down
29 changes: 29 additions & 0 deletions tests/v2/cases.py
Original file line number Diff line number Diff line change
Expand Up @@ -1469,6 +1469,35 @@ def _check_cjk_shape_purity(self) -> None:
{"given": "John", "middle": "Q", "family": "Smith",
"suffix": "MA"},
ambiguities=("suffix-or-name",)),
# #604: the one-letter particle against the initial it spells
# (rules.md#P7). The period decides, in every position a particle
# reading could otherwise claim the letter.
Case("one_letter_particle_written_bare", "Juan Ó Pérez",
{"given": "Juan", "family": "Ó Pérez"},
classification="fix(#604)",
notes="the bare letter is the Irish particle and opens the "
"surname; 1.4.0 read it as a middle initial. The "
"contrast row for the two below"),
Case("one_letter_particle_with_period_is_an_initial",
"Juan Ó. Pérez",
{"given": "Juan", "middle": "Ó.", "family": "Pérez"},
notes="Ó. is Óscar's initial: 'ó' in the particle vocabulary "
"alone gave family 'Ó. Pérez', the particle match "
"folding the period away (#604)"),
Case("one_letter_particle_with_period_after_family_comma",
"Pérez, Juan Ó.",
{"given": "Juan", "middle": "Ó.", "family": "Pérez"},
notes="P6's site: without P7 the trailing 'particle' attached "
"to the family behind the comma, family 'Ó. Pérez' "
"(#604)"),
Case("ambiguous_particle_ua_leading", "Ua Buachalla",
{"given": "Ua", "family": "Buachalla"},
ambiguities=("particle-or-given",),
notes="ua joined the AMBIGUOUS half, not the never-given one "
"(#604, decisions.md#vocabulary-collisions): leading, "
"it keeps the given reading and reports the fork, "
"where never-given would have made the string "
"surname-only"),
Case("titled_ambiguous_particle_does_not_chain", "Dr. Van Johnson",
{"title": "Dr.", "given": "Van", "family": "Johnson"},
classification="fix(#367)",
Expand Down
Loading
Loading