Blogging

Why Enterprise Knowledge Management Systems Fail Repeatedly: Six Conditions and the Exports That Expose Them

An enterprise knowledge management system is any platform an organization uses to store, describe, and retrieve its working knowledge: SharePoint sites, Teams channels, OneDrive folders, Confluence spaces, OpenText repositories, an Elasticsearch or Solr index sitting in front of all of them, or a public-sector data portal built on CKAN or Socrata. Failure in these systems rarely looks like an outage. It looks like a person who knows a document exists, has every right to read it, and still cannot find it in the time they are willing to spend. I review these systems for a living, and I have stopped being surprised by the pattern: the same handful of structural conditions appears in every organization, on every platform, and each one leaves evidence in a file somebody can export this afternoon.

That is the argument of this article, and it is narrower than it sounds. I am not going to tell you that knowledge management fails because of culture, or because of one vendor, or because people resist change. Those explanations are comfortable because they demand no evidence. What follows instead are six conditions I can name, the platform where each is easiest to observe, the standard that applies where one does, the trade-off each fix asks you to accept, and the export that settles the question. The failures repeat not because the technology is weak but because the causes are portable. They travel with the content, the permissions, and the habits, and they survive every migration intact.

A small team reviewing shared documents around a conference table with laptops

The six repeating conditions

1. Taxonomy drift

A taxonomy is a controlled vocabulary: a fixed list of terms with defined relationships — broader, narrower, preferred term, synonym. Taxonomy drift is what happens when nobody owns that list. Two teams file the same project under “Client A” and “A Client”; a merged department keeps both legacy labels; near-synonyms accumulate without ever being recorded as equivalents. The standard that applies is ISO 25964, which governs thesaurus structure, and SKOS, the W3C vocabulary for expressing term relationships in machine-readable form.

Drift is invisible from inside the term store and obvious from two exports. The first is the term set itself — the SharePoint Term Store exports to CSV, and a quarterly diff of that file shows every added, renamed, or orphaned term. The second is the query log, the record of what users actually searched for: when the same concept appears in queries under three different spellings and none of them matches a term, drift has a findability cost you can count in zero-result searches.

The trade-off is ownership. A taxonomy costs someone real, recurring time, and most organizations decline to name that person. The fix is deliberately lightweight — a quarterly review of the term-set export, merging duplicates and recording synonyms — and fully reversible, since every merge can be undone from the same file.

2. Metadata decay

Metadata is the structured description attached to content: the columns, tags, and labels that make a library filterable. Metadata decay is the slow emptying of those fields. Required columns sit blank because the dialog allowed “skip”; OneDrive folders, which enforce no column discipline, absorb the documents that would have carried the descriptions; labels applied in a burst of enthusiasm cover a shrinking share of what was created since. Dublin Core — the fifteen-element baseline for descriptive metadata — is the standard most often cited here, and its value is mainly as a floor: title, creator, subject, date, the minimum a document needs to be findable by anything other than luck.

The evidence is a label coverage report: in Microsoft Purview, the share of items carrying each retention label; in SharePoint, a list view export showing populated versus blank required columns. I have yet to review one in which coverage of any meaningful field exceeded sixty percent, and the true number is usually lower.

The trade-off is friction. Every mandatory field slows the person uploading, and slowed uploaders drift to OneDrive, where no fields exist at all. The fix is to cut required fields down to the two or three that carry real filter weight, then backfill them on the hundred documents people actually search for — a reversible edit, proven out by re-exporting the same coverage report next quarter.

3. Permission-induced findability gaps

Search results are security-trimmed: the index removes anything the seeker lacks permission to open. This is correct behavior, and it is also a quiet failure mode. A document exists, is current, and would answer the question — but it sits in a site the seeker cannot read, so the search page returns nothing, and the seeker reasonably concludes the document does not exist. In SharePoint and Teams this concentrates around private channels, sites with broken permission inheritance, and OneDrive shares that outlived the projects that needed them.

No widely adopted standard covers this condition, which is itself worth noticing; the artifact is the export. Site permissions export cleanly from the SharePoint admin center, and the query log supplies the demand side: the queries that returned nothing, or results nobody clicked. Lay the two files side by side and the gaps name themselves — the terms people searched, the sites they could not see. I have written a separate guide to reading permission exports against query logs elsewhere on this site.

The trade-off is political rather than technical: closing a gap means either granting access or admitting that restricted content exists. The lightweight fix is a known-hidden inventory — a published list of restricted sites and what they broadly contain, so a seeker can request access instead of assuming absence. It costs a paragraph to reverse and answers the question the empty result page never did.

4. Records and retention friction

Retention is the schedule by which records are kept and then disposed of; a hold is a legal instruction suspending that disposal. Friction is what happens when both are half-applied: retention labels cover a fraction of the estate, holds outlive the matters they were created for, and because nobody trusts the labels, nobody deletes anything — so the pile grows until search quality decays under its own weight. In the Microsoft estate this lives in Purview; in US public-sector organizations the schedule itself comes from the NARA General Records Schedules, and ISO 15489 is the general standard for records management practice.

Stacks of labeled document folders and archive boxes in an office storage room

The evidence is a hold inventory — Purview exports every hold with its date, scope, and the matter it serves — read alongside the label coverage report from the previous section. In the inventories I have reviewed, holds attached to matters closed years earlier are the single most common finding, and nobody involved knew they were still active.

The trade-off deserves honesty: deletion is the one fix in this article that is genuinely irreversible, which is why organizations avoid the whole subject and the pile keeps growing. The move I recommend is not deletion at all. Close the holds whose matters have ended — which merely re-enables the schedule that was already published — export the inventory first, and record what you closed and why.

5. Migration semantic drift

Semantic drift is the change in meaning that happens when content moves between systems. A Confluence label becomes a SharePoint column nobody maps; a multi-value field is flattened to its first entry; an OpenText category is dropped because the target has no equivalent; an Elasticsearch or Solr index is rebuilt with a different analyzer, so the same words tokenize differently and yesterday’s first result becomes today’s twentieth. Where a standard applies, it is the crosswalk — a documented mapping between the source vocabulary and the target, ideally expressed in SKOS or Dublin Core terms.

The evidence is a crawl log and a query regression set. Freeze a few hundred real queries from the old system’s query log before the migration, run them against the new index afterward, and diff the top results. The crawl log — the record of what the indexer reached and recorded — explains the differences: pages the crawler could not access, fields it did not index, encodings it misread.

The trade-off is scope. Perfect fidelity is not available at any budget, and migrations that pretend otherwise are the ones that stall. The realistic goal is to know what changed, which the frozen query set delivers — and the concern is reversible in practice, because the legacy system can be kept read-only for a quarter while the differences are worked through.

6. Governance by convention

Governance by convention is the state in which the rules that make content findable live in habits rather than in systems: “final versions go in that channel,” “the fiscal-year prefix goes first.” It works while the founding team is present and decays the moment they are not, because nothing enforces it and nothing records it. Public-sector portals built on CKAN and Socrata show the condition in its purest form: departments publish datasets freely, the catalog grows, and the DCAT export — the machine-readable description of every dataset, per the W3C’s Data Catalog Vocabulary — shows required fields like publisher, license, or update frequency blank or inconsistent across departments that were certain they had agreed.

The evidence is the catalog export itself, which both platforms produce. Count the required fields that are populated; that count is the condition, stated as a number.

The trade-off is participation: enforcement deters publication, and an empty open-data portal fails more publicly than an inconsistent one. The fix is to shrink the required set to what the catalog cannot function without, publish it as a one-page standard, and measure compliance monthly from the export. Every element of that can be revised the following month.

Why the failures repeat

The cycle repeats because the causes are organizational, and organizations treat the symptom as a purchasing decision. The sequence is always a version of this: search disappoints; a committee blames adoption; a replacement platform is bought; the content is migrated — and the permissions, the half-empty metadata, the orphaned terms, and the unwritten conventions are migrated with it; eighteen months later the new system disappoints for the same reasons, now with a layer of migration drift on top. I have watched this cycle run twice inside one organization. The second failure was faster.

The repeating part is not the technology. The conditions above are properties of the content and the habits around it, and no procurement replaces those. The organizations that break the cycle do something less dramatic: they pull the exports, name the conditions, and assign one owner to each artifact. That is unglamorous work. It is also the only version of this work I have seen succeed.

The diagnostic sequence: five exports in order

If you suspect your system is failing, do not start with a survey or a vendor briefing. Start with five files, in this order:

  1. Query logs — the demand signal. What people searched for, what returned nothing, what was offered and never clicked. If you have never opened one, I have written a companion piece on how to read a query log.
  2. Permission exports — restricted supply. What exists that seekers cannot see.
  3. Label coverage reports — how much of the estate carries the retention and metadata structure you believe it does.
  4. Hold inventories — what is frozen, and whether it should still be.
  5. Crawl logs — what the index actually reached, which explains whatever remains unexplained.

Each file takes an afternoon to produce, and none requires a project. Re-run the comparison in ninety days; the diff between the two runs is your progress report, and it is evidence rather than opinion.

An analyst comparing printed reports and spreadsheets across two screens at a desk

The trade-offs, stated plainly

Every fix above asks you to accept a cost, and it is more honest to say so than to pretend otherwise. Taxonomy ownership takes recurring staff time. Fewer required fields mean less descriptive richness. Closing permission gaps means either granting access or admitting hidden content exists, and both carry political weight. Closing stale holds means signing your name to a records decision. A documented crosswalk slows a migration by the days it takes to write. A published data standard tells departments their current practice is not good enough.

None of these costs is large. All of them are visible, bounded, and reversible in a way that a failed platform replacement is not. That is the whole argument of this article compressed into one sentence: the small, inspectable fixes are the ones you can undo.

Frequently asked questions

Why do knowledge management systems fail even after user training?

Training acts at the point of contribution, but the six conditions are structural — they live in the taxonomy, the permissions, the labels, and the crawl, none of which a training session touches. The proof is in the query log: zero-result searches continue after training, which tells you the users were never the problem.

Which export should I pull first?

The query log, always. It is the demand signal, and it tells you which of the other four exports matter in your case. An organization whose queries succeed but whose labels are empty has a different problem from one whose queries fail on content that exists but is locked away.

Is this a tooling problem or a people problem?

Neither, cleanly. It is a structure problem that both the tools and the people inherit. The same conditions appear across SharePoint, Confluence, OpenText, and CKAN with different interfaces, which is the strongest evidence I know that the platform is not the cause.

How do I know if my taxonomy has drifted?

Export the term set, diff it against last year’s export, and count the terms added without an owner, a synonym, or a merge. More than a handful is drift, and the query log will tell you what it is costing in failed searches.

A closing note

I have described these conditions as though they were exotic, and in practice they are the opposite — they are the ordinary weather of organizational knowledge. The weariness in this line of work is not about the problems, which are simple and countable, but about how rarely anyone counts them. The exports are sitting there, produced by systems you already run. Pull the query log this week, name the largest gap, fix the smallest thing that closes it, and re-export in ninety days. Repeat that four times and you will have done more for findability than any replacement project I have reviewed — and unlike the replacement project, you will be able to prove it, from a file, to anyone who asks.