Beyond the Single "North Star": Are Native Apps or Enterprise Alliances the True Core of the Industrial Data Fabric?

Author photo: Colin Masson
ByColin Masson
Category:
Industry Best Practice

Executive Takeaway

Over the past two years, as industrial enterprises scrambled to escape "pilot purgatory" and bridge the chasm between messy plant-floor operational technology (OT) and enterprise IT, horizontal cloud data platforms—most visibly the Databricks Data Intelligence Platform, but also Snowflake, Microsoft Fabric, AWS, and Google Cloud—have been heralded as the singular Industrial Data Fabric "North Star."

However, as the market matures, a critical architectural reality has set in: horizontal data platforms are not vertical industrial software suites. To deliver real operational value, they rely on an ecosystem of specialized partners.

As I’ve been exploring in my concurrent series, "The Industrial Copilot R(E)volution: Beyond the Chat Widget," you cannot drop a shiny generative AI assistant or an autonomous reasoning agent onto the factory floor and expect miracles if the underlying data architecture is a mess. When an AI agent reasons over plant data, it doesn't care how elegant your cloud marketing slides look—it cares whether the tag is frozen, whether the engineering units match, and whether the physical context is intact. That grounding comes entirely from an Industrial-Grade Data Fabric (IDF).

Today, industrial technology leaders are not faced with a simplistic binary choice between "building everything natively in the lakehouse" or "buying monolithic vendor suites." In practice, forward-thinking organizations are assembling an Industrial Data Fabric across three complementary architectural layers:

  1. Enterprise Domain Alliances (e.g., SAP, AVEVA CONNECT, Kinaxis, Siemens): Ideal for organizations with established, distributed centers of gravity seeking governed, federated data exchange across core business and operational systems of record.

  2. Industrial Context & Edge DataOps Platforms (e.g., Cognite Data Fusion, HighByte, Litmus, Cybus, inmation): The indispensable "first-mile" bridge that extracts, sanitizes, contextualizes, and models chaotic brownfield OT signals before exposing them upstream.

  3. Lakehouse-Native Applications & Accelerators (e.g., Nexalis Cloud, Covasant Auraa, and SI accelerators such as Bizmetric): Purpose-built software modules running directly inside the customer’s lakehouse tenant under unified governance (such as Databricks Unity Catalog), executing zero-copy analytics and low-latency agentic workflows.

Crucially, none of these models should mean lifting and shifting petabytes of raw, uncontextualized sensor telemetry across the factory wall. True industrial data architectures respect Data Gravity, contextualizing signals at the edge and querying data in place to avoid cloud egress bill shocks, eliminate latency penalties, and prevent the creation of an expensive "Centralized Data Swamp."

Retrospective: Modulating the Databricks "North Star" Perspective

Reflecting on my coverage following recent Databricks Data and AI Summits, my perspective on Databricks’ role in the industrial ecosystem has undergone a necessary and healthy calibration. Two years ago, when the Lakehouse vision promised to collapse data warehouses and data lakes into a single governed estate, it was tempting to view Databricks as the ultimate one-stop shop for Industrial AI.

However, subsequent real-world deployments have highlighted what I often describe to ARC clients as the "Some Assembly Required" reality of industrial software. (And as anyone who has tried to assemble flat-pack furniture knows, the phrase "some assembly required" usually hides a fair amount of sweat, a missing Allen wrench, and a few bruised knuckles.)

While Databricks provides an exceptional horizontal computational engine, open Delta Lake storage, and robust Unity Catalog governance, it relies heavily on ecosystem partners to build the deep vertical domain ontologies, time-series contextualization tools, and asset-specific physical intelligence required on the plant floor.

This partner-led reality has created a richer, multi-layered ecosystem where industrial leaders—spanning IT, OT, engineering (ET), and data science—must decide where data lives, where context is applied, and who controls the operational logic.

Plain English for the OT Exec: Demystifying "Zero-Copy" & "Delta Sharing"

To a plant manager, VP of operations, or chief automation officer, cloud terms like "Zero-Copy Architecture" and "Delta Sharing" often sound like software vendor sleight-of-hand—or worse, an IT magic trick that threatens plant uptime. Let’s strip away the buzzwords and explain what is actually happening in plain engineering terms.

The Old Way: The "Photocopy & Truck" Model (Traditional ETL)

For the last 20 years, whenever corporate IT, a central data science team, or a third-party software vendor wanted to run analytics on factory data, they built an ETL (Extract, Transform, Load) pipeline.

Think of traditional ETL like this: Every time someone at corporate headquarters wants to check a compressor’s vibration trend or analyze batch quality, the system photocopies a 500-page paper logbook at the plant, puts it in a box, trucks it to a corporate data warehouse, and refiles it in a brand-new filing cabinet.

The operational fallout is painful:

  • You pay double or triple for storage: You pay to store the original data in the plant historian, pay the cloud provider for network transport and egress "trucking" fees, and pay a SaaS vendor a 300 percent markup to store duplicate files in their proprietary cloud silo.

  • Your data is instantly out of date: By the time the batch photocopy lands in the central database hours or days later, the machine's physical state has moved on.

  • Operational context is lost in transit: The photocopy shows raw numbers (e.g., "45.2"), but strips away the ISA-95 asset hierarchy, operator shift logs, calibration states, and batch IDs that give the number physical meaning. On the plant floor, a number without context isn't just useless—it's dangerous.

Demystifying "Zero-Copy": The Governed Digital Library Card

Zero-Copy is not a magic trick. It simply means you stop photocopying and trucking the logbook.

Instead of duplicating petabytes of high-frequency factory telemetry into multiple vendor databases, Zero-Copy architectures leave the authoritative data in one governed storage location (typically an open Delta Lake or Apache Iceberg table inside your company's cloud storage tenant or edge repository). When an application—whether it’s SAP Business Data Cloud, AVEVA CONNECT, or a native tool like Nexalis—needs to analyze that data, it is issued a secure, digital "library card" (a direct data pointer).

The application reads the exact records it needs directly off the shelf, executes its analytics, and closes the book.

(Note: In technical terms, "zero-copy" means avoiding separate, persistent, vendor-managed data replicas. Queries still consume network bandwidth, compute, and temporary memory caching during execution.)

How Zero-Copy Eliminates Data Duplication

Because open table formats (Delta Lake with UniForm or Apache Iceberg) act as a universal file format, multiple compute engines (Spark, SQL, Python, Presto) can query the exact same physical files simultaneously. You no longer pay third-party SaaS vendors to re-host your own operational telemetry.

How Zero-Copy Minimizes Access Latency

Latency in industrial analytics rarely stems from raw network speed; it is introduced by intermediate pipeline hops. In legacy systems, data lands in a staging area, waits for a nightly batch transformation, gets converted into a proprietary SQL schema, and lands in an analytical dashboard days later. Zero-Copy eliminates intermediate landing zones, enabling near-real-time query readiness for autonomous AI agents and process engineers alike.

Demystifying "Delta Sharing": The Universal Delivery Pipe

If Zero-Copy is the "library card," Delta Sharing is the open, secure protocol that makes the library card work across different software companies.

Historically, if Vendor A wanted to read data from Vendor B, they forced you to pay hundreds of thousands of dollars for custom API integration middleware—or you ended up relying on the world's most ubiquitous, duct-taped industrial system: the manual Microsoft Excel spreadsheet export.

Delta Sharing is an open-source standard (originated by Databricks and widely supported across the industry) that allows your enterprise to share live data tables securely with external vendors, partners, or AI agents—across different clouds, different regions, and different applications—without copying the data or writing brittle glue code.

(Hyperscaler alternatives like Snowflake Data Sharing, AWS Zero-ETL, and Google Cloud Analytics Hub provide analogous zero-copy sharing mechanisms within their respective platform ecosystems.)

De-Mythologizing the "Three Industrial Data Risks"

A pervasive fear among industrial CIOs, CISOs, and operations executives is that adopting a cloud-native lakehouse requires streaming unthrottled, raw 1,000 Hz sensor waveforms straight into the public cloud.

At ARC Advisory Group, we actively advise clients against this "brute force" approach. Streaming raw, uncontextualized telemetry directly to cloud storage creates three acute failure modes:

Simplified Three-Layer Industrial Data Fabric Architecture

Rather than forcing the market into a rigid binary choice between pure SaaS suites and raw DIY cloud builds, modern Industrial Data Fabrics assemble across three complementary building blocks:

Industrial Data Fabric Architectures: Detailed Comparison

Deep Dive: The Three Building Blocks of an Industrial Data Fabric

1. Enterprise Domain Alliances (Federated Centers of Gravity)

Major incumbent software providers have established strategic alliances with horizontal data platforms to bridge operational technology, supply chain, and enterprise resource planning without displacing established systems of record:

  • AVEVA CONNECT & Databricks: Integrates AVEVA CONNECT and the AVEVA PI System with Databricks using Delta Sharing, allowing high-frequency time-series data to be accessed alongside enterprise datasets without breaking OT operational guardrails.

  • SAP & Databricks (SAP Business Data Cloud): Combines SAP Databricks with SAP Business Data Cloud Connect for Databricks, enabling governed, bidirectional sharing of SAP data products (ERP, financial, and supply chain context) and Databricks data without ETL tax.

  • Kinaxis & Databricks: Connects Kinaxis Maestro supply chain orchestration with Databricks to strengthen Maestro's supply chain data fabric, unifying internal and external telemetry to support scalable AI across planning and scenario orchestration.

  • Siemens, FFT DataBridge & Databricks: Delivers an edge-to-cloud discrete manufacturing pattern where contextualized production telemetry from Siemens Industrial Edge is streamed directly to Databricks for AI development and advanced analytics.

Strategic Callout — AVEVA, Cognite, and TwinThread: Schneider Electric has announced a definitive agreement to acquire Cognite. Once completed, adding Cognite Data Fusion, its Industrial Knowledge Graph (IKG), and Atlas AI to the broader Schneider Electric and AVEVA software portfolio will significantly expand CONNECT's contextualization footprint. Separately, AVEVA maintains a strategic technology partnership with TwinThread, leveraging TwinThread's predictive twin technology within AVEVA Advanced Analytics.

2. Industrial Context & Edge DataOps Platforms (The First-Mile Bridge)

Industrial data rarely arrives cloud-ready. This vital middle layer connects, filters, validates, structures, and semantically models chaotic plant telemetry before exposing it to enterprise analytics (get a much wider vendor perspective in our Assembling Your Industrial-Grade Data Fabric blog series and report):

  • Cognite (Cognite Data Fusion): An industrial DataOps and contextualization platform that auto-connects heterogeneous OT sources, constructs dynamic Industrial Knowledge Graphs (IKG), and exposes contextualized data products to Databricks, Snowflake, and Microsoft Fabric.

  • HighByte (Intelligence Hub): An Industrial DataOps platform focused on modeling, contextualizing, and delivering industrial data at the edge. HighByte standardizes PLC tag streams into unified namespace (UNS) payloads and acts as an Industrial MCP Server for AI agents.

  • Litmus, Cybus, inmation, Velotic: Representative edge and industrial data-hub platforms that normalize diverse PLC protocols, enforce local industrial namespaces, and publish clean event streams to enterprise platforms.

3. Lakehouse-Native Applications & Accelerators (In-Tenant Execution)

A growing class of software applications, platforms, and accelerators is designed to deploy directly within or execute closely alongside a customer’s lakehouse environment:

  • Nexalis Cloud (nexalis.io): A cloud historian and data-enablement platform built on Databricks for utility and renewable energy operators. It automates high-frequency time-series data ingestion, organization, and direct lakehouse access via Delta Sharing without external data duplication.

  • Covasant (Auraa): A Databricks-native, agent-driven data engineering platform that automates source discovery, ingestion, pipeline construction, data quality enforcement, and Unity Catalog registration—accelerating the creation of AI-ready lakehouse foundations.

  • Bizmetric (SI & Solution Accelerator Partner): A certified Databricks consulting, systems integration, and solutions partner. Bizmetric offers prebuilt industry lakehouse accelerators, IoT pipelines, and GenAI frameworks for energy, utilities, and manufacturing.

Follow the Money: Who Holds the Budget for Industrial AI?

Architectural selection is driven as much by organizational budget ownership and governance mandates as by technical software benchmarks:

Most large industrial organizations do not select a single isolated path. Instead, they fund a hybrid architecture that balances central governance, plant safety, and commercial product differentiation.

This budgetary dynamic serves as the direct bridge to our companion analysis ("Beyond 'Making Stuff': How Servitization, Product-in-Use Telemetry, and Industrial Data Fabrics are Turning Manufacturers into Custom AI Builders").

Matching Architectural Strategy to Organizational Governance

To help executive teams determine where their architectural center of gravity should sit, consider how organizational maturity and system footprint intersect:

1. Favor Enterprise Domain Alliances when:

  • Your enterprise operates strong, decoupled centers of gravity (such as plant operations in AVEVA, enterprise planners in SAP, supply chain teams in Kinaxis).

  • Software vendors are expected to maintain and certify complex vertical domain models and regulatory compliance standards.

  • The primary goal is cross-domain business insights without disturbing frontline plant workflows.

2. Favor Industrial Context & Edge DataOps Platforms when:

  • OT data is fragmented across brownfield sites, disparate historians, PLCs, and fieldbus protocols.

  • You require reusable, contextualized data products that serve multiple cloud platforms (Databricks, Snowflake, AWS, Azure), analytics tools, and AI agents simultaneously.

  • Edge filtering, deterministic safety boundaries, and vendor-neutral data access are non-negotiable.

3. Favor Lakehouse-Native Applications & Accelerators when:

  • A centralized Data & AI CoE has standardized development and governance on a selected platform (such as Databricks Unity Catalog).

  • You mandate strict in-tenant execution to minimize egress fees, storage duplication, and external SaaS silos.

  • Your team is building custom agentic workflows and proprietary product-in-use telemetry models for servitized business models.

Diagnostic Questions for C-Suite Leaders

Before committing capital to an Industrial Data Fabric architecture, ask prospective software vendors and internal architecture teams these six diagnostic questions:

  1. Data Gravity & Location: Where is authoritative data stored? Does this design create persistent replicas, use governed sharing, or query data in place? What network transfer, caching, and derivative storage still occur?

  2. First-Mile Edge DataOps: How are raw OT signals filtered, validated, timestamped, buffered, and mapped to ISA-95/i3X asset hierarchies at the plant boundary before enterprise transmission?

  3. Contextualization Effort: How much manual engineering work is required to create and maintain industrial semantics? Which tag mappings are automated by AI, and how are automated mappings reviewed and governed?

  4. Cost Governance & Tokenomics: What explicit controls exist for ingestion volume, retention, egress, compute, model selection, token consumption, and autonomous agentic loop rate-limiting?

  5. Servitization & IP Protection: As we deploy product-in-use telemetry to support outcome-based SLAs, who owns the contextualized schemas, feature stores, model weights, and derived intellectual property?

  6. Portability & Lock-In: If a vendor relationship ends, which data, schemas, knowledge graphs, lineage, models, and application artifacts remain fully accessible in open formats (Delta Lake/Apache Iceberg)?

Centers of Gravity, Not Winner-Take-All

The Industrial Data Fabric is not a winner-take-all contest between enterprise alliances and native applications. It is an architectural composition.

Enterprise domain platforms preserve operational and business workflows. Industrial context and edge DataOps suppliers make distributed OT information trustworthy and reusable. Lakehouse-native applications and accelerators bring development, analytics, and AI closer to governed enterprise data.

The most resilient strategy is hybrid. Industrial technology leaders should select components according to workload, data gravity, latency requirements, domain context, organizational maturity, and long-term data portability. Success depends less on naming a single "North Star" than on governing how multiple centers of gravity exchange trusted data to support operational autonomy.

Engage with ARC Advisory Group

The Industrial AI (R)Evolution is moving faster than ever. To dive deeper into the frameworks and data shaping the future of the industrial sector, explore my latest research:

Where do you Stand in the Industrial AI (R)Evolution?

Take our Industrial AI Assessment to benchmark your organization's maturity, identify critical gaps in your IT/OT/ET convergence, and get actionable recommendations to accelerate your path to becoming an Industrial AI Pacesetter (and download the 2026 Report). If you think you’re already a Pacesetter, nominate your team for the ARC Industrial Pacesetters Awards!

Don't guess what your global operations or prospective customers need. Use empirical data to align your stakeholders and de-hype the market with ARC Advisory Group's Voice of Market Service.

For tailored recommendations on governing and guiding major people, process, and technology decisions across the enterprise, cloud, industrial edge, and AI, please contact Colin Masson at [email protected].

Or, set up a meeting with my fellow Analysts and I at ARC Advisory Group to find out more about our Executive Insights Service for Industrial organizations and our Industrial AI Insights Service for Vendors.

Engage with ARC Advisory Group

Representative End User Clients
Representative Automation Clients
Representative Software Clients