Blogging

The Crawl Log’s Error Column: What Indexing Failures Explain About ‘Search Is Broken’

“Search is broken” is almost never a search problem. It is an indexing problem that search is faithfully reporting. The crawl log — the record of what the indexer attempted to read, parse, and admit into the index — has an error column, and that column is the closest thing an enterprise knowledge system has to a truthful account of why users cannot find things. The error column is not a list of search bugs. It is a list of structural information failures that were present before anyone typed a query.

This article is about reading that column diagnostically: what the errors mean, which platform produces which kind, which standard is being violated, and what lightweight, reversible fix is anchored to an artifact you can export and inspect. The recurring theme is that the fix is rarely a re-index. It is usually a small correction to a schema, a label, a permission, or a retention rule — and the crawl log is the evidence that tells you which.

What a crawl log actually is

A crawl log is the per-item record produced when a search indexer walks a content source. Each row typically carries the item URL or identifier, the time of the crawl, the outcome (success, warning, error), and — in the error column — a code or message describing why the item was not indexed or was indexed incompletely. In SharePoint and Microsoft 365, the crawl log is exposed through the search administration interface and can be filtered by error type. In Elasticsearch and Solr, the equivalent is the indexing response and the cluster log, where rejected documents and mapping conflicts surface. In Confluence, OpenText, and CKAN, the pattern is the same: a per-item ingestion record with an outcome and a reason.

The error column matters because it is item-level. A dashboard that says “index health: 94%” tells you nothing actionable. A crawl log that says 4,000 items failed with the same error code tells you exactly which structural condition to fix. The diagnostic move is always the same: export the error column, group by error code, and read the top three groups.

Error category one: schema and mapping conflicts

The most common indexing failure in modern search platforms is a mapping conflict — the indexer encountered a field value whose type does not match the type already declared for that field. In Elasticsearch, a field mapped as date that receives a string, or a field mapped as keyword that receives an object, produces a rejected document and a log entry naming the field and the conflicting type. In Solr, the analogous failure is a schema field type mismatch during document add. In Microsoft Search, the equivalent is a crawled property that cannot be mapped to a search attribute because its type is incompatible.

Mapping conflicts are diagnostic of schema drift: the content source changed its metadata shape, and the index schema did not. A source system that once emitted createdDate as ISO-8601 and now emits it as a Unix epoch integer will produce a steady stream of mapping conflicts until the schema is reconciled. The error column names the field. That is the whole diagnosis.

The standard that applies here is the one governing the schema itself. In Microsoft Search, the search schema defines crawled properties, search attributes, semantic labels, and aliases, and the documentation is explicit that “Incorrectly mapping labels degrades the search experience” and that “Adding a search capability requires a full crawl” (Microsoft, Manage search schema). The trade-off is stated plainly: marking too many properties as retrievable “increases search latency,” so the fix is not to make everything searchable but to map the right properties correctly.

Reversible fix: export the crawl log, group by field name in the error column, and reconcile the schema for the top offending field. In Microsoft Search, this means adding or correcting the crawled property mapping and publishing schema changes, then running a full crawl. In Elasticsearch, it means either correcting the source data or introducing an explicit mapping that accepts the new type. Both are reversible: schema changes can be reverted, and a full crawl can be re-run. The artifact is the crawl log export itself, which becomes the before-and-after evidence.

Error category two: permission-induced findability gaps

A second class of crawl log error is the access denied or insufficient permissions entry. This is the most misread error in the column, because it looks like a security success. It is not. It is a findability failure with a security cause.

When the indexer’s service account cannot read an item, the item is not indexed. When the item is not indexed, no user can find it through search — including users who do have permission to read it. The crawl log records the denial; the search index records the absence. The user experiences this as “search is broken” because a document they know exists does not appear. The permission export — the list of items and their effective permissions — is the artifact that explains the gap.

This is a permission-induced findability gap: a structural mismatch between who can read an item and who the indexer is allowed to crawl. It is common after a migration, after a site collection is moved, or after a service account is changed. The crawl log’s error column will show a cluster of access-denied entries concentrated in a site, library, or content type. That concentration is the diagnosis.

Reversible fix: export the permission report for the affected scope, compare it against the crawl log’s denied items, and grant the indexer read access to the specific scope. This is reversible because permissions can be reverted, and the crawl log export provides the evidence that the gap existed and was closed. The trade-off is explicit: broadening indexer access increases the surface the index can see, which may surface content that was previously invisible to search. That is a governance decision, not a technical one, and it should be made deliberately.

Error category three: metadata decay and label coverage

A third class of error is the missing or malformed metadata entry. The item is crawled successfully, but a field the index expects — a title, an author, a classification label, a retention code — is absent or unparseable. The item may still be indexed, but it is indexed poorly: it cannot be refined, cannot be sorted, cannot be filtered, and may not surface in the result cluster at all.

This is metadata decay: the gradual erosion of metadata quality as content is created, migrated, and edited without enforcement. It is visible in the crawl log as a rising count of items with missing or defaulted fields. The standard that applies depends on the metadata scheme in use. Dublin Core defines a minimal set of descriptive elements — title, creator, date, subject — and an item missing all of them is effectively unsearchable by any of those facets. SKOS governs concept labels and their relationships; a taxonomy term with no prefLabel is a label that cannot be matched. ISO 25964 governs thesauri and their interoperability; a thesaurus whose terms have drifted from the source vocabulary produces exactly the kind of label mismatch that shows up as a metadata error.

The diagnostic artifact here is a label coverage report: a count of items by whether each required field is populated. The crawl log’s error column gives you the numerator — items with missing fields — and the total item count gives you the denominator. The ratio is the coverage. A coverage report that shows 60% of items missing a classification label explains why faceted search returns nothing useful, without any need to investigate the search engine itself.

Reversible fix: generate the label coverage report from the crawl log export, identify the top missing field, and either backfill it from a source of truth or relax the schema requirement. Both are reversible. The trade-off is that backfilling metadata is labor-intensive, while relaxing the requirement reduces the precision of refinement. The coverage report makes the trade-off visible and lets the decision be made on evidence rather than assumption.

Error category four: retention and hold friction

A fourth class of error is less visible but structurally important: the item that is under a retention hold or subject to a disposition rule that prevents the indexer from processing it in the expected way. In Microsoft 365, items under a Purview retention label or an eDiscovery hold may be excluded from certain processing paths. In records management systems, items in a closed file may be locked against modification, which can prevent the indexer from updating their metadata.

This is records and retention friction: the tension between the obligation to preserve records and the operational need to index them for findability. The standard that applies is the retention schedule itself. In the United States federal context, the NARA General Records Schedules define how long categories of records must be kept and when they may be destroyed. A crawl log entry showing an item excluded from indexing because of a hold is not an error in the conventional sense — it is the retention rule working as designed. But it is a findability gap, and it should be recorded as such.

The diagnostic artifact is a hold inventory: a list of items under hold, their retention category, and their indexing status. Cross-referencing the hold inventory against the crawl log’s error column shows which holds are producing findability gaps and whether those gaps are acceptable. Some are; a closed investigation file should not surface in general search. Others are not; a policy document under a long retention period should still be findable by the people who need it.

Reversible fix: export the hold inventory, identify holds that are producing unintended findability gaps, and adjust the indexing scope for those specific holds. This is reversible because holds can be re-scoped, and the hold inventory provides the audit trail. The trade-off is between preservation integrity and findability, and it should be resolved by policy, not by default.

Error category five: migration semantic drift

A fifth class of error appears after a migration: items that were indexed correctly in the source system but fail in the target because a field name, a value format, or a taxonomy term changed during the move. This is migration semantic drift: the meaning of a metadata field shifted between systems, and the index schema did not follow.

The crawl log’s error column will show this as a cluster of mapping conflicts or missing-field errors concentrated in migrated content. The pattern is distinctive: the errors begin at the migration cutover date and affect a specific content set. The diagnostic artifact is the migration mapping document — the record of how source fields were mapped to target fields — compared against the crawl log’s error column. Where the mapping document says Author maps to dc:creator but the crawl log shows dc:creator missing, the mapping was not applied or was applied incorrectly.

The standard that applies is the metadata crosswalk used in the migration. Dublin Core and DCAT both define element sets that are commonly used as crosswalk targets. A migration that maps to Dublin Core without validating that the target fields are populated will produce exactly this error pattern.

Reversible fix: compare the migration mapping document against the crawl log error column, identify the fields that did not survive the move, and re-run the mapping for the affected content set. This is reversible because the mapping can be corrected and the crawl re-run. The trade-off is that re-crawling migrated content takes time and may surface additional errors, but the crawl log export provides the scope of the problem before the re-crawl begins.

What the error column does not tell you

The crawl log is a record of indexing outcomes, not of user intent. It will not tell you what users searched for and did not find. That is the query log’s job. The two artifacts are complementary: the crawl log explains why an item is not in the index, and the query log explains what users expected to find. A diagnosis that uses only one of them is incomplete.

The practical procedure is to export both, align them by content scope, and look for the intersection: queries that return zero results for terms that appear in items the crawl log shows as failed. That intersection is the strongest evidence that “search is broken” is actually an indexing failure with a specific, fixable cause.

Why the fix is usually small

The reason “search is broken” complaints persist is not that the fixes are hard. It is that the evidence is not read. The crawl log’s error column is exportable, groupable, and specific. It names the field, the item, and the reason. The fix is almost always one of: correct a schema mapping, grant indexer access to a scope, backfill a missing label, re-scope a hold, or re-run a migration mapping. Each of these is reversible, each is anchored to an exportable artifact, and each can be verified by re-exporting the crawl log after the fix.

The discipline is to treat the error column as the primary diagnostic record, not as a background log. Group by error code. Read the top three groups. Export the artifact that explains each group. Apply the smallest reversible fix. Re-export and compare. That procedure resolves more “search is broken” complaints than any search configuration change, because it addresses the structural information failures that search is faithfully reporting.

FAQ

How often should the crawl log be reviewed?

Review the error column after any migration, after any permission change, and on a regular cadence tied to content publishing. The cadence should be frequent enough that error clusters are caught before they become user-visible findability gaps. The export is cheap; the diagnosis is the work.

What is the difference between a crawl log error and a search relevance problem?

A crawl log error means the item is not in the index or is indexed incompletely. A relevance problem means the item is in the index but ranks poorly. The crawl log cannot diagnose relevance; it diagnoses presence. If the item is absent, fix the crawl. If the item is present but ranked low, that is a separate analysis using query logs and ranking configuration.

Can the crawl log be used as a retention record?

The crawl log is an operational record, not a records schedule. Its retention should be governed by the same schedule that applies to the content it describes, or by an operational retention rule if it is treated as a system log. The NARA General Records Schedules provide the framework for federal records; other jurisdictions have equivalent schedules. The hold inventory, not the crawl log, is the artifact that maps items to retention categories.

What if the error column shows no errors but users still cannot find things?

Then the problem is likely in the query layer, not the index layer. Export the query log and look for zero-result queries. If the terms users search for do not appear in any indexed item, the issue is vocabulary mismatch — the users’ language and the content’s language have drifted apart. That is a taxonomy problem, and the fix is a label or synonym mapping, not a re-index.

Which artifact should be exported first?

The crawl log error column, grouped by error code. It is the fastest way to distinguish between the five failure categories described above. Once the category is known, the second artifact — permission export, label coverage report, hold inventory, or migration mapping document — follows directly.