Overview
The growing interest in Big Data in recent years has led to a host of questions about how this trend influences or is impacted by applications in manufacturing operations. Often, the first area of attention is the role played by how data historians are evolving with the repositories for the large volume of process-related information generated in a typical industrial plant.
It is worthwhile to review the situation and some of the positions and opinions offered. Some have predicted that the traditional role of the historian in enabling integration of information and processes between enterprise and operations systems will change as that scope expands to include so-called “community systems” that exist at a logical level above the individual enterprise.
Meanwhile, others have predicted that the traditional historian will ultimately disappear or transform to an extent as to make it unrecognizable. This is an unlikely development. To paraphrase Mark Twain, “Reports of the death of the data historian have been greatly exaggerated.”
Certainly, the emergence of Big Data and IIoT have blurred the role of the historian, but with the corresponding growth in the volume of process data, this component is likely to continue to play a major role.
Historians: An Essential Component
The data historian has been an essential component of process automation systems for decades. As early as the 1970s, companies began to introduce general-purpose computers into their facilities to augment dedicated electronic controllers. This new technology allowed controller inputs and outputs and other process data to be displayed and trended for analysis in real time. This soon became the task of the data historian. Early historians may have only collected a few hundred process signals every five to ten seconds.
Since this initial introduction, this component and its underlying technology has continued to evolve and improve. Methods such as data compression enabled ever larger amounts of information to be collected. Supported data types have expanded to include complex data types. Finally, the continuous advancement in computing performance has vastly increased both the volume and resolution of data collected. In larger facilities, it is common to collect tens of thousands of data elements, with scan frequencies as high as a thousand times per second. Today’s data historians can collect, contextualize, and store data to provide workers with timely insights on easy-to-understand digital dashboards.
Advancements in technology and performance have brought us ever closer to the vision of the comprehensive “flight recorder” for the plant or enterprise, sometimes described as all data, all the time, forever. Current systems may even exceed some aspects of this vision, in that they can collect derived and calculated information, along with data from other sources.
So Much Data, So Many Possibilities
The sheer quantity of data collected in operating facilities can be daunting. One consulting company stated that manufacturing stores more data than any other sector. It estimated that approximately two exabytes of new manufacturing data was collected and stored in 2010 alone.
With every advance in the quantity, quality, and range of data that can be collected, new potential applications come to mind. These include applications that combine operational data with geodata, weather data, financial data, and so on. While process data may have traditionally been only of interest to operations and maintenance personnel, these new applications have attracted the attention of users from across the enterprise.
Operations staff have long used historical data to identify and analyze the root causes of equipment failure and other process-related events. People in functions ranging from research and development to supply chain have come to realize that the same tools, methods and data can also be applied to a much broader range of opportunities.
Regulated or quality-sensitive industries have also found historical data to be useful – even essential – in tracking and tracing raw materials and products, as well as recording products for later reporting.
Various engineering disciplines also make extensive use of the data contained in historians for applications such as advanced control and optimization or energy optimization.
Finally, historians are often the information source of record for regulatory compliance (and other) reporting purposes.
Access According to Purpose
The variety of applications and potential users of historical data presents a challenge to the deployment model that includes only a single data historian. Requirements and constraints related to performance, capacity, and security have led to the development of models that include multiple historians, each with a specific purpose.
The general purpose data historian still remains an important source of information for operations personnel, who make extensive use of various trend displays on the operator interface of the control system. In addition, a high-speed, high-fidelity historian may be used to collect smaller volumes of data for applications such as sequence of events analysis; while a lower resolution, long-term historian may summarize process data for long-term analysis of production performance. This information may be aggregated from more than one data historian.
This diversity of uses can lead to a requirement for a hierarchy of historians, positioned at various levels of the physical architecture. Access and security considerations may require that individual historians be separated by a firewall or similar network segmentation device. In such cases, synchronizing data and management of change becomes critically important.
How Big Is Big?
Considering all of the above, it is not unreasonable to consider the collection and use of historical process data to be an example of Big Data that existed long before this term came into common use. One case cited in a supplier whitepaper describes a company where manufacturing a personal care product generates 5,000 data samples every 33 milliseconds, resulting in:
- 152,000 samples per second, or
- 9 million samples per minute, or
- 545 million samples per hour, or
- 4 billion samples per shift, or
- 13 billion samples per day, or
- 4 trillion samples per year, or
An average refinery collects 100,000 points per second. This works out to nine billion points per day and three trillion points per year. This is for a single product, in a single facility. These numbers virtually explode when considering all of the signals and data in a larger or extended operation. As new sensor and communications technology becomes available, the amount and variety of data available for collection and analysis expands as well.
Historically, there were serious challenges associated with analyzing and applying such large volumes of information. Often, there has been a temptation to collect the data simply because it is available, without giving serious consideration to how it might be used. Without proper tools, the traditional practice was to simply collect and log large amounts of data, moving it to offline storage (e.g., tapes) when online storage was exhausted. Although, technically, the data was available, it was not commonly used.
Continuing Evolution
Faced with the above developments, most major suppliers continue to evolve their data historian products. Most have positioned their products as “enterprise historians,” emphasizing a variety of new and improved features and capabilities. Just as with other elements of process automation, security remains a concern. Current historians must be able to ensure data confidentiality, while providing both high availability and high levels of system integrity.
Industrial Internet of Things (IIoT) gives rise to additional opportunities for historians, which can now be located in the plant, in the cloud, on the edge, or in hybrid configurations. ARC expects that most process plants in the hazardous industries will utilize a combination of all of these locations for historian data. Both the amount and diversity of data will increase as more devices are employed, with broader capabilities for data generation and transmission. More data means more and more diverse opportunities for value-added applications, as well as new stakeholders.
Applications such as predictive analytics and decision support can take considerable advantage of the information collected in larger and more powerful historians. Process historians or data platforms will continue to evolve and adapt to new initiatives such as Industrie 4.0, Smart Manufacturing, and Open Process Automation.
From Component to Ecosystem
One major supplier has described this evolution as the “Historian to Infrastructure Journey.” This reflects the elevation of the historian from a system component to the heart of a much larger and complex ecosystem. Simple data collection and storage is no longer adequate, nor is the practice of “holding data hostage” within single devices or components. There must be sophisticated tools and capabilities for data organization, management, and advanced access from a variety of sources.
This infrastructure or ecosystem must make optimum use of available technologies. For example, storage methods must extend to use cloud services where appropriate. With data distributed across multiple locations, there is also a requirement for a single logical view for the purpose of analysis.
Observations and Conclusions
Ultimately, it’s up to the asset owners to determine the extent to which these and other trends evolve into solid trends and developments in technology and usage. Rather than purchasing and deploying capabilities simply because they are available, the first step must be to define a data collection and usage strategy that considers what is possible, what is available, and how and by whom the information is to be used. There seems to be little doubt that the quantity of information available in manufacturing operations can be characterized as “big,” but data collection leads to data ownership and management. These incur a cost for operation that must be offset by solid benefits and business return.
Recommendations
Based on ARC research and analysis, we recommend the following actions for asset owners and other technology users:
- Define the data strategy – This includes identifying what data is available and whether it should be collected, stored, and managed. A subset of this may have immediate, value-added value. Define this value in terms of not only the ability to collect, but also the cost of ownership and the benefits associated with its use.
- Review historian positioning – If the current configuration includes one or more historians, conduct a detailed review of their functions and positioning within the system and technical architecture. Pay particular attention to the types and quantity of data collected in each repository, as well as how, and by whom these data are employed.
- Consider non-traditional sources and data types – It is possible that other data and/or data types may be useful for analysis. Review the configuration to identify additional opportunities, or to remove data already collected in the historian that may not be needed.
- Confer with data scientists, knowledge workers, engineers and analytics experts – Discuss additional applications – beyond normal operations – that may have been identified by those who have a need to analyze process data. This includes people in business and support functions.
If you would like to buy this report or obtain information about how to become a client, please Contact Us
Keywords: Analytics, Big Data, Data Management, Historian, IIoT, ARC Advisory Group.