Blogging

Governance-by-Naming: When Your Only Metadata Layer Is a Filename Convention

You know the document exists. You watched the program manager upload it in March. You type the project code into SharePoint search and get nothing — then browse to the library, scroll, and find it sitting there as Final_v3_PROJ-7720.docx. The search didn’t fail. The naming convention did, and it failed quietly, one file at a time, for three years.

We call this condition governance-by-naming: the structural substitution of a filename convention for an actual metadata layer. It is one of the most common failure modes we see in enterprise and public-sector knowledge systems, and one of the least diagnosed, because nothing looks broken until the day someone needs to find, retain, or disclose a specific set of documents. This piece names the condition, walks through a SharePoint evidence case, and ends with an export you can run on Monday.

The condition, defined

Governance-by-naming is what happens when an organization encodes its taxonomy — its structured vocabulary for describing content — inside filenames instead of in managed metadata fields. The convention looks like governance: it has rules, it has a documented pattern, it may even have a two-page PDF explaining it. But it lacks the two things that make a real taxonomy durable: an enforcement layer (something that rejects non-conforming entries) and a machine-readable validation mechanism (something that can check conformance at scale).

Compare this to a SKOS taxonomy — the W3C’s Simple Knowledge Organization System, a standard for publishing controlled vocabularies as linked data. A SKOS concept scheme is machine-readable: every concept has a URI, broader/narrower relationships are explicit, and a validator can tell you when a label is missing or a relationship is orphaned. A filename convention has none of that. It is a taxonomy expressed in a medium — the filesystem name — that was never designed to carry structure, validated only by the memory of whoever last read the PDF.

The evidence case: a SharePoint library with one metadata layer

The case we’ll walk through is a composite example, not a firsthand account: a public-sector agency, roughly 2,000 employees, with a program document library in SharePoint Online. The library had exactly three columns: Name, Modified, and Created By. Everything else lived in the filename, governed by a convention documented in 2019:

PROJ-DEPT-YEAR-Title.ext
e.g., PROJ-7720-FIN-2023-Quarterly-Budget-Review.docx

By 2024, a full export of the library showed the convention had drifted into at least six recognizable dialects:

Figure 1 — Drift in one naming convention, one library (n ≈ 4,100 files)

PROJ-DEPT-YEAR-Title      38%  ─ the 2019 pattern, intact
PROJ-YEAR-Title           21%  ─ DEPT dropped after a reorg
DEPT-PROJ-Title           14%  ─ segment order swapped
PROJ-DEPT-Title           11%  ─ YEAR dropped; retention breaks here
Title-PROJ-DEPT            9%  ─ title-first, search breaks here
free text / Final_v3       7%  ─ no pattern at all

Each dialect is internally consistent. Each was invented by a reasonable person under deadline pressure. None of them can talk to the others, and none of them can be queried without first being parsed — which is the whole point of metadata, and the thing the convention was substituting for.

Why the drift is structural, not careless

It’s tempting to read Figure 1 as a discipline problem. It isn’t. The drift follows a predictable structural logic:

First, the convention had no enforcement layer. SharePoint document libraries accept any filename. Unlike a managed metadata column backed by a term set — a controlled vocabulary, ideally aligned to something like ISO 25964, the international standard for thesauri and interoperability — a filename field will never reject Final_v3_PROJ-7720.docx. The system’s permissiveness guarantees eventual divergence.

Second, the convention encoded unstable values in fixed positions. Department codes change in reorganizations; years roll over; project codes get recycled. A convention that hard-codes volatile values into a positional grammar inherits every organizational change as filename debt. When the reorg happened, staff faced a choice: keep writing a department code that no longer existed, or drop it. They dropped it, correctly, and the convention fractured.

Third, there was no feedback signal. Nobody ever saw a report saying “21% of new files this quarter did not match the pattern.” Drift is invisible at the moment of upload and only visible at the moment of retrieval — which is to say, at the worst possible time: a FOIA request, a retention purge, a migration.

How the drift breaks search and retention

Search breaks first. SharePoint’s full-text index does tokenize filenames, but a user searching for “project 7720 budget” is relying on the code having survived into the filename in the expected position and format. The 9% of files with title-first names and the 7% with no pattern effectively vanish from code-based retrieval. Multiply a 16% blind spot across a 4,100-file library and you have roughly 650 documents that exist, are permissioned correctly, and are functionally unfindable by the convention’s own logic.

Retention breaks second, and worse. The agency’s records schedule — mapped to NARA General Records Schedules, the U.S. National Archives’ baseline categories for federal record disposition — keyed retention on project and year. Files that dropped the year segment (11% of the library) could not be matched to a retention period by any automated rule. Under a governance-by-naming regime, the retention policy is only as good as the filename parse, and the parse is only as good as the drift allows. A Microsoft Purview retention label applied by rules that expect YEAR in position three will simply never fire on PROJ-7720-FIN-Quarterly-Budget-Review.docx.

What makes any naming system durable

Stepping back from SharePoint: the question of what makes names durable is not specific to document management. Any naming system — a taxonomy, a records classification, even a fictional character-naming scheme — survives on the same three properties: embedded rules (the grammar is explicit, not remembered), constrained vocabularies (each segment draws from a closed list, not free text), and consistency checks (something periodically verifies conformance and reports exceptions).

Pattern-based name generators make the grammar-versus-meaning distinction concrete. A character name generator, for example, produces names from an explicit rule system — which is exactly the property a filename convention needs and usually lacks: the rules are encoded in the generator, not in a PDF nobody reads. The names a generator emits follow a grammar; whether they carry meaning is a separate question, and that separation is the diagnostic point. A filename convention that only follows a grammar (correct segments, correct order) can still fail to carry meaning if the vocabulary inside each segment is unconstrained. Durability requires both.

The trade-off to name explicitly: constrained vocabularies cost flexibility. A closed list of department codes means a new department requires a vocabulary change, not just a new filename. That friction is the point — it forces the change through the governance layer where it belongs, instead of letting each uploader improvise.

The lightweight, reversible fix

You do not fix governance-by-naming by renaming 4,100 files. That’s a migration-scale project with its own semantic drift risks. The reversible fix has three steps, none of which destroys anything:

Step 1 — Export and measure. Export the library (SharePoint’s “Export to Excel” or a PowerShell Get-PnPFile inventory is sufficient) and run a pattern-coverage report: for each file, test the filename against each known dialect and record which one matched, if any. Figure 1 above is that report. The artifact is a CSV with columns for filename, matched pattern, and unmatched segments.

Step 2 — Add managed metadata columns alongside the names. Create three site columns — Project, Department, Date — backed by a term set, and populate them by parsing the filenames you can parse (the 84% that match a known dialect). Leave the filenames untouched. You now have a real metadata layer that search and Purview retention labels can act on, without touching a single document.

Step 3 — Make the convention advisory, not load-bearing. Keep the naming convention for human readability, but stop depending on it for retrieval or retention. Publish the coverage report quarterly. When drift appears in a new dialect, that’s a signal someone needs a vocabulary value that doesn’t exist yet — which is a term-set change, not a filename problem.

This fix is reversible because it adds a layer without modifying the old one. If the term sets prove wrong, you delete the columns and lose nothing but the parsing effort.

What to inspect on Monday

Run these against your largest convention-governed library. Each is tied to an exportable artifact:

  • Export the library inventory (filename, created date, author) and compute pattern coverage. If more than one dialect exceeds 10% of files, you have active drift, not historical drift.
  • Check for missing segments. Count files where the year or department segment is absent. That count is your retention blind spot — the files no automated disposition rule can match.
  • Pull a query log excerpt (SharePoint search usage reports, or the Microsoft 365 search analytics export) and check how many queries contain a project code or department code. Those queries are the ones your drift is silently failing.
  • Review your Purview label coverage report for the library. If auto-apply rules depend on filename patterns, compare the label coverage rate to your pattern coverage rate — the gap between them is the set of documents your retention policy believes in but cannot reach.
  • Find the convention document. If it’s a PDF last modified more than two reorgs ago, the convention is already historical fiction; treat it accordingly.

Governance-by-naming fails not because people are careless but because filenames are a medium with no enforcement, no validation, and no feedback. Give the structure a layer that has all three, and the names can go back to being what they were meant to be: something a human reads, not something a retention policy depends on.