IN THIS ARTICLE

The Document Chaos Problem — SharePoint, DWG Files, ERP Exports, and the Excel Log from 2012

If you want to understand why most industrial AI projects struggle and quietly are not delivered, don't start by looking at the AI.

Start with the documents.

Because inside most engineering-driven organizations, knowledge isn't missing. It's fragmented, duplicated, and structurally unusable.

And deep down, everyone knows it.

The Reality Check: "We Have It Somewhere…"

Ask any engineer or project manager: "Have we done something similar before?"

The answer is rarely "no." It's usually:

"Yeah… I think so… let me check."

That "check" turns into:

  • Digging through SharePoint folders
  • Searching file servers with vague keywords
  • Opening random PDFs
  • Scanning email attachments
  • Asking senior engineers directly
  • Searching ERP

This is not inefficiency. This is systemic knowledge fragmentation.

5.1 What the Document Landscape Actually Looks Like

Let's move beyond theory and describe what truly exists inside industrial environments.

1. Engineering Drawings (PDF + DWG/CAD)

These are often the most valuable intellectual assets a company owns. But also the least accessible.

What's happening:

  • CAD files (.DWG, .STEP) are not text-searchable
  • They require specialized tools (AutoCAD, SolidWorks) to open
  • Key insights (dimensions, tolerances, annotations) are embedded visually
  • Multiple revisions exist across disconnected folders

Operational consequence:

  • Engineers rely on memory instead of systems
  • Similar designs are recreated unnecessarily
  • Historical design intelligence is effectively locked away

2. OEM Manuals & Manufacturer Spec Sheets

These documents are essential for system design, integration, compliance, and troubleshooting.

But structurally:

  • Many are scanned PDFs (no searchable text layer)
  • Formats vary widely across vendors
  • Key specifications are buried in tables, diagrams, and footnotes

Impact:

  • Engineers manually scan 100+ page manuals
  • Time is wasted locating basic specifications
  • Risk of using outdated or incorrect data increases

3. ERP Outputs (SAP, Oracle, etc.)

ERP systems generate critical business intelligence:

  • Pricing history
  • Procurement data
  • Vendor performance
  • BOMs (Bill of Materials)

But in practice:

  • Data is exported as static PDFs or spreadsheets
  • File names lack meaning or context
  • No linkage to real-world projects or applications
  • No metadata enrichment

Result:

  • Data exists but lacks usability
  • Historical insights are not available
  • Decision-making becomes reactive instead of informed

4. Project Knowledge (The Most Undervalued Asset)

This includes proposals, bid documents, costing sheets, project reports, and internal discussions and analysis.

Where it lives:

  • SharePoint
  • Personal desktops
  • Local drives
  • Email attachments

Core issues:

  • No standardized structure
  • No tagging by industry or application
  • No connection between similar past projects
  • Not able to store discussions or analysis

What this causes:

  • Teams unknowingly solve the same problem repeatedly
  • Pricing strategies are rebuilt from scratch
  • Proposal quality varies widely
  • Always scrambling to find a person who knows it all

5. The Excel Log from 2012 (The "Ghost System")

Almost every organization has one — a spreadsheet that once tracked past projects, solutions implemented, customer requirements, and key learnings.

Then it stopped being updated. Its version became obsolete.

Yet paradoxically: it still contains the most structured knowledge in the company.

The dilemma:

  • It's outdated
  • It's incomplete
  • It's not trusted

But there is no better alternative.

5.2 The Bigger Problem: Document Quality Is Inconsistent

Even within PDFs — which often make up 80–90% of enterprise document repositories in industrial firms — quality varies significantly. (https://www.businessinsider.com/sc/how-unstructured-enterprise-data-is-limiting-ai-performance)

  • Some documents are fully searchable
  • Some are scanned images requiring OCR
  • Some contain broken tables when extracted
  • Many are duplicated across systems

Why this matters:

AI systems depend on clean text and metadata extraction, consistent structure, and contextual knowledge. If documents are inconsistent, AI outputs will be too.

Version Chaos: The Trust Breakdown

This is not just a productivity issue — it's a risk issue.

In many companies:

  • Multiple versions of the same document exist
  • File names don't indicate authority
  • No clear "source of truth"

Employees spend an average of 1.8 hours per day searching for information. (https://www.mckinsey.com/industries/technology-media-and-telecommunications/our-insights/the-social-economy)

That's nearly 20% of the workweek lost. (https://www.forbes.com/councils/forbestechcouncil/2021/10/14/flying-blind-how-bad-data-undermines-business/)

But the deeper issue is trust:

  • "Is this the latest version?"
  • "Did this spec change?"
  • "Which drawing was approved?"
  • "Is it up to date with the latest regulations?"

When answers depend on who you ask, not what the system confirms, you have a knowledge reliability problem.

This gap between stored knowledge and uncaptured knowledge is what creates the Knowledge Cliff.

Case Insight: Siemens — Preventing Engineering Knowledge from Retiring with the Workforce

Global industrial giant Siemens has publicly highlighted the growing risk of engineering knowledge loss caused by retiring experts and fragmented documentation systems.

In a Siemens industry whitepaper focused on aerospace and defense engineering, the company warned:

"Your E/E engineers are retiring. What's your plan to prevent their decades of knowledge and expertise retiring with them?"

(https://resources.sw.siemens.com/en-GB/white-paper-aerospace-defense-knowledge-loss-electrical-electronic-design)

Siemens identified several major operational risks:

  • Retirement of senior engineering talent
  • Difficulty transferring institutional knowledge
  • Fragmented technical documentation
  • Disconnected engineering systems
  • Slow onboarding of new engineers

The company emphasized that critical engineering knowledge was often spread across:

  • CAD systems
  • Product lifecycle management systems
  • PDFs and technical documentation
  • Legacy databases
  • Disconnected departmental repositories

As engineering complexity increased, Siemens noted that organizations struggled to:

  • Retrieve historical project knowledge
  • Reuse prior engineering decisions
  • Maintain consistency across revisions
  • Onboard engineers efficiently

The business impact included slower engineering workflows, duplicated effort, reduced productivity, and higher operational risk.

To address this, Siemens invested heavily in centralized knowledge management and PLM (Product Lifecycle Management) systems designed to preserve engineering intelligence and improve discoverability across the enterprise.

GrayCyan Case Study: From Document Chaos to Manufacturing Intelligence

A mid-sized North American manufacturing company approached GrayCyan with a growing operational problem that many manufacturers quietly face but rarely discuss openly.

Over nearly two decades, the company had accumulated engineering drawings, OEM manuals, ERP exports, maintenance logs, bid documents, project reports, and CAD files across SharePoint libraries, local servers, email archives, and disconnected engineering systems. The information existed, but finding the right information at the right time had become increasingly difficult.

Engineers often spent hours searching for previous solutions to problems the company had already solved years earlier. In several cases, teams unknowingly recreated designs, troubleshooting processes, and proposal calculations simply because the engineer who originally handled the project was no longer with the company and the historical context could not be located quickly enough.

As Nishkam Batta, CEO of GrayCyan, explained in his Forbes Business Council article Unlocking Insight Through Custom Retrieval-Augmented Generation:

"Manufacturers generate decades of engineering insight, but it can get buried inside maintenance notes, PDFs, change orders and disconnected systems. When that knowledge cannot be surfaced in context, it might as well not exist."

When GrayCyan began its audit, the scale of the problem became clear. The organization had nearly 1.8 million documents spread across multiple environments. Approximately 65% of the files were PDFs, many of them scanned and non-searchable. Another 20% consisted of CAD and DWG engineering drawings that could not be searched without manually opening each file.

The remaining records included spreadsheets, data in ERP, maintenance reports, and historical project documentation stored with inconsistent naming conventions and duplicate versions.

The issue was not a lack of information. It was the absence of structure, context, and retrieval logic.

A maintenance report referencing a PLC shutdown existed in one system, while the corrective engineering action existed in another. ERP contained useful PLC operational data but lacked metadata or context that linked them to real-world engineering events.

Rather than immediately deploying AI, GrayCyan focused first on understanding the company's knowledge ecosystem. Documents were mapped, classified, and prioritized based on operational importance. Engineering drawings, maintenance logs, OEM manuals, and recurring equipment failure records were identified as high-value knowledge sources requiring structured ingestion and contextual retrieval.

The project also revealed a deeper operational reality that many manufacturers now face.

First, manufacturers need to actively document the conversations, decisions, and troubleshooting discussions that often remain uncaptured. Much of the most valuable operational knowledge exists in emails, shift handovers, engineering discussions, maintenance calls, and the experience of senior employees. If that institutional knowledge is never documented, AI systems cannot retrieve or reason over it effectively.

Second, GrayCyan helped the manufacturer define its true sources of truth. Instead of treating every file equally, the team identified which engineering systems, maintenance records, OEM manuals, ERP data, and project documents should serve as authoritative references for operational decision-making. This reduced confusion caused by duplicate files, outdated revisions, and disconnected repositories.

This reflected another key point from Nishkam Batta's Forbes article:

"RAG surfaces what exists. If maintenance records are inconsistently formatted or institutional knowledge was never documented, the system will reflect those gaps."

Instead of creating a generic chatbot layered over disconnected documents, GrayCyan developed a manufacturing-specific retrieval framework designed around operational context. The system connected engineering records, maintenance history, project documentation, and asset-level information so engineers could retrieve grounded answers instead of isolated files.

Within the initial deployment phase, engineering search time dropped from hours to seconds. Teams began reusing previous solutions instead of rebuilding them from scratch, proposal turnaround improved significantly, and onboarding new engineers became faster because expertise was no longer dependent entirely on senior employees.

But the most important transformation was cultural.

Engineers stopped relying exclusively on tribal knowledge and memory. Instead of asking "Who knows this?" — they got to work.

For GrayCyan, the project reinforced a larger reality now emerging across manufacturing: the competitive advantage is no longer simply collecting data. It is transforming decades of disconnected operational knowledge into contextual intelligence that engineers can actually use when decisions matter most. (https://www.forbes.com/councils/forbesbusinesscouncil/2026/04/07/unlocking-insight-through-custom-retrieval-augmented-generation/)

This Is Not a Technology Step

Before any AI delivers value, companies must answer:

  • What documents do we actually have?
  • Which ones matter most?
  • What condition are they in?
  • What should be prioritized first?

These are not IT questions. They are strategic decisions about knowledge infrastructure.

For manufacturers, this challenge goes even deeper. Some of the most valuable operational knowledge never exists inside formal systems at all. Critical troubleshooting discussions, engineering decisions, maintenance conversations, shift handovers, and lessons learned often remain undocumented and trapped inside emails, meetings, or employee memory.

Manufacturers must actively find ways to capture and structure these uncaptured conversations and discussions if they want AI systems to deliver meaningful operational intelligence.

GrayCyan worked with the organization to define its true sources of truth across engineering systems, maintenance records, OEM manuals, ERP exports, project documentation, and operational reports. This created clarity around which data sources could be trusted for different operational and engineering decisions.

From there, GrayCyan built a system capable of reasoning across connected knowledge sources instead of simply searching isolated files. The platform connected engineering records, maintenance history, asset-level information, CAD documentation, ERP data, and historical troubleshooting records to provide contextual, grounded answers engineers could actually use in real-world situations.

The result was not simply faster search. It was the transformation of disconnected operational data into usable manufacturing intelligence.

What Smart Industrial Companies Are Doing

Instead of rushing into AI tools, leading firms:

  1. Audit their document ecosystem
  2. Define classification frameworks
  3. Fix quality gaps (OCR, duplication, structure)
  4. Prioritize and capture high-impact knowledge
  5. Establish governance for future data

Only then do they deploy systems like RAG (Retrieval-Augmented Generation). (https://www.forbes.com/councils/forbesbusinesscouncil/2026/04/07/unlocking-insight-through-custom-retrieval-augmented-generation/)

Because in manufacturing, the real value of AI does not come from simply adding another chatbot or automation layer. It comes from building a knowledge foundation that allows decades of engineering expertise, operational history, and institutional intelligence to become searchable, connected, and usable in context.

The manufacturers that will lead the next decade are not necessarily the ones collecting the most data. They are the ones creating systems that can understand, connect, and reason across the knowledge they already possess.

Looking for AI advice at your company? Talk to our Editor-in-Chief

Nishkam Batta

Nishkam Batta

Editor-in-Chief – HonestAI Magazine (400,000+ Readers)
HonestAI magazine’s Editor-in-Chief is Nishkam Batta. HonestAI focuses on practical, credibility-first AI adoption, with clear standards for human-in-the-loop systems, no black box AI (explainable AI), measurable outcomes, and governance built for manufacturing and enterprise environments. The magazine covers applied topics such as agentic ERP systems, auditability, integration into existing operations, and the distinction between helpful automation and risky hype, emphasizing what decision makers can verify, measure, and implement.

Unlock the Future of AI -
Free Download Inside.

Get instant access to HonestAI Magazine, packed with real-world insights, expert breakdowns, and actionable strategies to help you stay ahead in the AI revolution.

Download Edition 16 & Level Up Your AI Knowledge