Distributed Analytics Are Beneficial for Industry

Author photo: Michael Guilfoyle

 

Table of Contents

  • Edge Data Takes Center Stage
  • Rethinking Cloud-centricity
  • Distributed Intelligence Evolves
  • Distributed Analytics Methods
  • A Few Use Cases
  • Recommendations

Edge Data Takes Center Stage

Distributed analytics extends data processing and computing close to or at the data source.  In many instances, data for Distributed Analyticsdistributed analytics comes from IIoT devices located at the edge of the operational network.  These devices can be located near or embedded in a wide variety of edge machines and equipment, such as robots, fleet vehicles, and distributed microgrids.

These analytics can reside within one or more intelligent, two-way communication devices, such as sensors, controllers, and gateways.  While the terms “edge” and “fog” are often used interchangeably to describe these types of distributed analytics, as we’ll see, there are some distinctions between the two.

Analytics at the edge of the operational network is common in some industries, such as retail and finance.  However, until recently, analytics at the edge wasn’t possible in industrial settings due to a mix of cost, complexity, security, and technology barriers.  Instead, edge data-to-cloud analytics was the default starting point.

Integrating analytics, and distributed analytics into industrial setting offers significant and far reaching benefits.  These include:

  • Revenue generation from new methods of serving existing customers and ways of reaching new ones
  • Asset optimization through improved, proactive, and highly-automated management of infrastructure, resources, and capital
  • Higher satisfaction and retention by engaging customers with highly-valued products and services where and when they need them
  • Improved operational flexibility and responsiveness through better and faster data-driven decisions

When considering and developing an analytics strategy to drive these benefits, companies often begin with centralized solutions in mind.  These typically leverage cloud solutions due to their ability to provide powerful analytics to complex and large data sets.  However, ARC has observed a recent trend toward edge data and analytics, with most solution providers, including those with industrial platforms, heavily marketing edge capabilities.  That makes sense.  The growth of IoT devices, supporting systems, and related data has skyrocketed and will continue.  Digitization will occur in brownfield environments by adding on devices or greenfield ways by building intelligent infrastructure.

For companies engaged in digital transformation, analytics will be key to turning that data into value both for their own operations as well as the customers they serve.  These companies need to think very carefully about where and how that data is and will be used.

In some instances, data needs to be processed centrally, such as in a cloud, to drive strategic decisions.  In other situations, decisions will need to be instantaneous, meaning that centralized solutions cannot be used.  From the edge to the cloud, how should companies think through the different ways to employ analytics?

This report will provide useful ways of thinking about distributed analytics today.  It will outline the different methods for deploying analytics, including in the cloud, at the edge, and via fog methods.  To help ARC clients better understand how to consider the right analytics solutions for their respective businesses, we’ll provide multiple use case examples where edge data are leveraged

Rethinking Cloud-centricity

As business leaders wrestle with “data explosion,” cloud computing has been seen as the solution for any associated volume, speed, and complexity issues.  While that is largely true, the edge environment also presents arguments for distributed data intelligence.

Certainly, cloud solutions can lessen or eliminate many storage, access, and sharing constraints.  Also, cloud solutions can handle a host of other issues that a traditional system solution architecture cannot, even those constructed modularly.  Most importantly, cloud solutions can bring massive computational power to problem solving, such as using analytics for asset performance or process quality optimization.  

One of the most accepted benefits of the cloud is its viability for complex analytics.  The related thinking is that edge data should be captured and fed to the cloud.  However, even though cloud solutions can manage complex, high-volume data, many businesses are reconsidering the role of edge data, particularly for analytics.  Use cases are being developed that increasingly demonstrate the value of keeping edge data localized.

Some major justifications for doing so include:

  • Bandwidth and immediacy:  These are the most obvious arguments for keeping analytics as close to the edge device as possible.  In many instances, operational situations and/or the remote location of the network edge constrain bandwidth.  Also, instantaneous decision making is often critical, so latency must be minimized.  Military field use and autonomous mining vehicles are common examples.  In these instances, the data and analytics are deployed tactically to ensure immediate benefits such as ensuring safety or identifying hazards.
  • Cost of noise:  As IIoT devices proliferate, the amount of data they produce could overwhelm even the most robust cloud-based application.  A major reason is the sheer noise and intermittency of the data generated by edge devices.  Often, cloud analytics solutions filter this noise from the computational process as they are irrelevant to supporting the use case.  In this instance, a company using a cloud-based approach still must shoulder the cost and process complexity of transporting, ingesting, storing, and filtering that noisy data, which provides no real business value.
  • Privacy:  In some cases, regulations, laws, or customer choice prohibits transferring data away from an edge source, but analytics are still needed to run the product or service correctly.  An example might be safe operation of an autonomous taxi.  Here, the operating agreement with customers might limit the amount of individual passenger information (such as locations visited) passed along to the enterprise data layer.
  • Security:  From a cyber threat perspective, limiting physical and cyber access to data-generating assets is a viable approach for minimizing vulnerabilities.  This includes developing processes that wall off data, including at the network edge, to prevent mis-operation or unauthorized use.  An example is North American Reliability Council’s Critical Infrastructure Protection standard (NERC CIP) to limit cybersecurity vulnerabilities of the bulk electric system (BES).  Also, many utilities and solution providers add improved data security capabilities to components of distributed energy resources (DER), such as microgrids, in distribution grids.  In both instances, access points from the edge to the enterprise can be reduced to minimize threat scalability.  At the same time, localized analytics can continue to support local network and asset performance optimization.   

Distributed Intelligence Evolves

Until recently, advanced analytics for industrial purposes couldn’t be done in a distributed way due largely to cost, data management, and technology limitations.  Instead, advanced analytics using edge data relied on centralized Distributed Analyticsarchitecture.  This required data to be transported to a centralized resource for processing, often far away from the source.  This included instances using cloud-based analytics engines.

That dynamic is quickly changing.  The main drivers of change are improvements in edge technology compute power and communication (at an increasingly lower cost), combined with a better understanding of edge data.  Analytics can be done right at the edge, close to or in the device or machine.  Also, solutions are being offered that can distribute analytics across the layers of a business.

Because of this shift, more use cases are being built on the principle of “data stays or goes where it is of most value.” This flexible use of data for analytics is the idea upon which fog computing is built.

Distributed Analytics Methods

Cloud Only:  Edge Data to Cloud Analytics

As mentioned, the default starting point for advanced analytics in industrial settings has typically been edge data-to-cloud analytics.  With this approach, edge data is sent to a cloud for storage and processing.

Models can be created and data run through an analytics engine.  Results are shared via back-office system integration or through reports to help humans make informed decisions.  In some instances, results can be sent back to the edge devices or, more commonly, to edge control systems to modify device performance. 

Distributed Analytics

Additional data from the enterprise or third parties may also be stored, processed, and analyzed in the same cloud environment.  If a use case supports it, that additional data could be used for edge analysis.  This centralized model provides the most power for processing, storage, and analytics.

Local Edge:  Analytics Computed Locally

Using a local cloud or server, this deployment leverages the power of a traditional server architecture or centralized cloud.  The edge gateway communicates with the edge device, ingests data, and transports that information to the local server or cloud.  The local server or cloud manages data readiness, storage, analytics, and visualization.

Distributed Analytics

Security, privacy, data-related cost, and regulatory constraints are often the reasons cited for keeping the analytics local.

Edge Only:  Embedded and Edge Network Analytics

This instance is the technical definition of edge analytics as it excludes the cloud or any other centralized analysis, even if local.  The analytics are embedded within an edge machine or device, delivered via an add-on edge technology, or supported by a nearby gateway.  Data storage and analytics are done on the device or within a network of edge devices.  This is beneficial when bandwidth requirements are small and the analytics less complex than what is typically processed at a platform level.

Distributed Analytics

Speed is typically cited as the chief reason for using edge analytics where results data is streaming or results must be delivered in real-time or near-real time.  Proponents of this approach also mention that a lot of edge data is not useful for the rest of the enterprise.  Edge analytics can then be used to eliminate the cost of transporting, storing, and processing of this “noise.”

Fog:  Analytics Distributed Throughout the Network

The most recent innovation related to edge data envisions intelligence distributed at any data node.  This approach is referred to as fog analytics or fog computing.

Distributed Analytics

Fog analytics operates on the principle that processing should be done where it provides value.  In this instance, flexibility is the overriding requirement, as the analysis might occur at any layer of the operational footprint, extending to any point both horizontally or vertically.  A use case may require any mix of device-only, local, network layer, or enterprise analytics. 

Fog analytics provides agility: data is measured and processed based on its ability to provide value, and not just for the processes they might support.  The data can be considered for analytics across multiple time scales, from streaming to investment planning, for example.  If the use case requires the data to stay in a contained environment, the organizations benefits by limiting data cost and process complexity.  Or, it can be shared and broadly used across the enterprise, should a use case develop where it provides business value to do so. 

A Few Use Cases

As companies create use cases that include edge data, they will increasingly embrace the value of distributed analytics.  In some cases, there will be quantifiable tactical benefits to contain data locally and use it for instantaneous edge analytics.  In other instances, some or all data will be shared to support “big picture” analysis.  These analytics, processed in the cloud, will deliver business model, process optimization, competitive, customer satisfaction, and other strategic-level benefits.

A likely high-growth area for distributed analytics is with solutions for asset performance management (APM).  Given how many high-value, high-risk assets operate at the industrial edge, it’s a logical jumping off point.  Here are a few examples of use cases that process edge data both locally and in a centralized cloud:

  • Wind turbines: Devices on the individual turbine or within the edge network layer can use edge analytics to optimize wind output, enabling real-time automated management of factors such as speed and pitch.  Information can then also be sent to a cloud for more complex processes that will typically include other data sources.  Analysis can uncover patterns that could lead to cascading failure.  Also, this layer of analytics can support a wide range of continual process improvement, including capital planning, predictive maintenance, customer programs, and reliability planning, to name just a few.
  • Vehicle fleets (autonomous or piloted): On-board sensors contain edge analytics that monitor engine performance, such as shift patterns, and make real-time adjustments to improve performance, such as reduced emissions or improved fuel consumption.  Data from the engine and vehicle are then sent to the cloud for more complex processing.  Centralized analytics could be used to improve routing and supply chains, identify manufacturing or service issues across fleets, or deploy more targeted maintenance support.
  • Smart buildings: Intelligent infrastructure and devices in buildings have provided a plethora of benefits using a combination of edge and cloud analytics.  Security, HVAC, lights, elevators/escalators, thermostats, fire alarms, windows, and more leverage analytics to reduce daily operating cost and safety incidents.  Using connected devices and local analytics (embedded within equipment or on a supporting edge IoT network device like a gateway or router), an entire building ecosystem can dynamically adjust and react to human presence, intrusions, and operational activity.  Using a fog analytics approach, edge data can be processed at a network of buildings or campuses level to improve and automate decisions related to issues such as worker time-shifts, time zones, security needs, cost of energy, etc.  In tandem, data can be sent to a cloud.  Centralized analytics can inform resource and budget allocations, uncover enterprise supply chain issues, optimize productivity across sites, and drive sustainability efforts.

Recommendations

Integrating analytics into a business is challenging.  It requires businesses to rethink industrial and corporate settings in which processes are rigidly embedded.  However, the benefits far outweigh the risk of change.  As analytics become increasingly integral to digital economies, businesses that aren’t effectively using data to drive decisions won’t be able to compete.

Based on ARC research and analysis, we recommend the following for industrial organizations:

  • Embrace analytics as a core competency.  Despite the challenge presented by having to sort through all the different approaches to analytics, it should be integrated within the business as a cornerstone capability.  Analytics can help reduce risk, avert catastrophic operational disruption, increase customer satisfaction and loyalty, and improve profitability.
  • Design an analytics strategy with the edge in mind.  Every company experiencing digital transformation will reach a point where some aspects of operational or service processes need to be instantaneous.  This requirement means some analytics must be done at the edge.  Solutions that don’t account for edge data will have limited value.
  • Let the use case drive the data and analytics approach.  Find situations where operations, assets, or processes are in continual crisis.  Then use the business objective to determine the data and analytics used.  For example, if looking to automate a response based in an asset based on performance conditions, you can use engineered analytics at the edge.  In contrast, if you are trying to predictively identify failure issues of high-value assets, you will want to take a much broader view of the data involved beyond performance.  Consider using deep learning analytics that can include structured data (e.g., historian) as well as unstructured formats such as audio logs, video, and work order notes.  By starting with the use case, you correctly align solution cost and complexity with the correct data and analytics solution.

 

If you would like to buy this report or obtain information about how to become a client, please  Contact Us

Engage with ARC Advisory Group

Representative End User Clients
Representative Automation Clients
Representative Software Clients