Gutenbrain: An Architecture for Equipment Technical Attributes Extraction from Piping & Instrumentation Diagrams

Marco Vicente, João Guarda, Fernando Batista

2022

Abstract

Piping and Instrumentation Diagrams (P&ID) are detailed representations of engineering schematics with piping, instrumentation and other related equipment and their physical process flow. They are critical in engineering projects to convey the physical sequence of systems, allowing engineers to understand the process flow, safety and regulatory requirements, and operational details. P&IDs may be provided in several formats, including scanned paper, CAD files, PDF, images, but these documents are frequently searched manually to identify all the equipment and their inter-connectivity. Furthermore, engineers must search the related technical specifications in separate technical documents, as P&ID usually don’t include technical specifications. This paper presents Gutenbrain, an architecture to extract equipment technical attributes from piping & instrumentation diagrams and technical documentation, which relies in textual information only. It first extracts equipment from P&IDs, using meta-data to understand the equipment type, and text coordinates to detect the equipment even when it is represented in multiple lines of text. After detecting the equipment and storing it in a database, it allows retrieving and inferring technical attributes from the related technical documentation using two question answering models based on BERT-like contextual embeddings, depending on the equipment type meta-data. One question answering model works with free questions of continuous text, while the other uses tabular data. This ensemble approach allows us to extract technical attributes from documents where information is unstructured and scattered. The performance results for the equipment extraction stage achieve about 97,2% precision and 71,2% recall. The stored information can be later accessed using Elasticsearch, allowing engineers to save thousands of hours in maintenance engineering tasks.

Download


Paper Citation


in Harvard Style

Vicente M., Guarda J. and Batista F. (2022). Gutenbrain: An Architecture for Equipment Technical Attributes Extraction from Piping & Instrumentation Diagrams. In Proceedings of the 14th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management (IC3K 2022) - Volume 1: KDIR; ISBN 978-989-758-614-9, SciTePress, pages 204-211. DOI: 10.5220/0011528500003335


in Bibtex Style

@conference{kdir22,
author={Marco Vicente and João Guarda and Fernando Batista},
title={Gutenbrain: An Architecture for Equipment Technical Attributes Extraction from Piping & Instrumentation Diagrams},
booktitle={Proceedings of the 14th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management (IC3K 2022) - Volume 1: KDIR},
year={2022},
pages={204-211},
publisher={SciTePress},
organization={INSTICC},
doi={10.5220/0011528500003335},
isbn={978-989-758-614-9},
}


in EndNote Style

TY - CONF

JO - Proceedings of the 14th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management (IC3K 2022) - Volume 1: KDIR
TI - Gutenbrain: An Architecture for Equipment Technical Attributes Extraction from Piping & Instrumentation Diagrams
SN - 978-989-758-614-9
AU - Vicente M.
AU - Guarda J.
AU - Batista F.
PY - 2022
SP - 204
EP - 211
DO - 10.5220/0011528500003335
PB - SciTePress