Methodology

How place data is collected, how etymological interpretations are produced, and how the site communicates certainty and uncertainty.

Geographical data

Place locations, names and administrative boundaries are sourced from published authoritative datasets. The primary source for coordinates and place types is OS Open Names (Ordnance Survey), published under the Open Government Licence v3. Additional names and variant forms come from the Gazetteer of British Place Names (GBPN), published under CC BY 4.0.

Etymological sources

Etymological information comes from several tiers:

  1. Attested records — interpretations drawn directly from published place-name dictionaries or county surveys. These are marked as Source-backed or Researched.
  2. AI-assisted interpretations — where no attested source is available, a language model (Google Gemini 2.5) is used to suggest an interpretation based on linguistic patterns. These records are marked Automated or Provisional.
  3. Lexicon matching — a structured lexicon of known place-name elements is used to identify components. Matches are ranked by strength.

Automated processing

The site processes over 110,000 British place names. Manual verification at this scale is not possible in full, so an automated pipeline generates first-pass interpretations. The pipeline:

  • normalises and tokenises place names
  • matches tokens against a lexicon of Old English, Gaelic, Old Norse, Brittonic and other language elements
  • calls a language model to produce a candidate summary and language identification
  • assigns a numeric confidence score
  • records sources so that provenance is traceable

Automated output is not treated as authoritative. It is a starting point for further review, not a final answer.

Source-backed vs inferred interpretations

Source-backed

At least one published source directly supports the interpretation.

Inferred / Automated

Generated by software. May be linguistically plausible but has not been verified by a specialist.

An automated result with a high confidence score is still an automated result.

Confidence levels

  • Documented — Supported by published sources with historical attestation.
  • Probable — Consistent with known linguistic patterns; plausible but not fully attested.
  • Suggested — A plausible candidate; limited direct evidence.
  • Uncertain — Origin unknown; no reliable interpretation established.

Page quality and indexing

Each place page is assessed before being included in search engine sitemaps. A page must have either a named etymological source or lexicon evidence for every displayed name element; geographic data and automated output alone are not enough. Pages without either evidence signal carry a noindex directive. Advertising also requires a higher content-quality score.

How corrections are reviewed

Corrections submitted through the site are reviewed by the editor. Supporting evidence — a published source, a historical spelling, or a specific reference — is required for changes to attested records. Provisional records may be updated on the basis of well-reasoned corrections.

Not every submitted correction will result in a change. See the Corrections Policy for more information.

Current limitations

  • Most records are AI-generated and have not been individually checked by a specialist.
  • Coverage is uneven: areas with richer source documentation are better represented.
  • Historical spelling records are available for only a subset of places.
  • The lexicon of place-name elements is growing but is not exhaustive.