
ARC Advisory Group's Industrial Data Fabric
As Director of Research for Industrial AI at ARC Advisory Group, I focus on the dynamic intersection of operational technology (OT), information technology (IT), and artificial intelligence. A key enabler of this convergence is the Industrial Data Fabric—more than just a buzzword, it serves as the essential architecture for integrating and leveraging data from sensors, machines, enterprise systems, and supply chains to power the next wave of Industrial AI innovations.
My ongoing research into the Industrial Data Fabric market—examining its scale, growth trends, and the evolving architectural models adopted by industrial clients—has uncovered a vibrant ecosystem of technologies and vendors. We’re seeing clients craft purpose-built, industrial-grade Data Fabrics, recognizing them as foundational to addressing persistent skills shortages and driving the next phase of productivity through Industrial AI.
Amid this evolving landscape, one company continues to stand out: Databricks. Long seen as an enterprise IT player, Databricks is now making notable advances and forming strategic partnerships that position it at the heart of the Industrial AI movement.
This blog, part of my ongoing series on the Industrial Data Fabric market, explores why Databricks is increasingly gaining relevance for industrial organizations as they advance on their AI journeys.
From Spark to Intelligence: Databricks' Open Source DNA
To understand Databricks, it is important to trace its roots. The company was born out of the AMPLab at UC Berkeley and founded by the original creators of Apache Spark—the open-source distributed computing framework that transformed big data processing. This academic and open-source foundation isn’t just background—it continues to shape Databricks’ philosophy and product strategy today.
Commitment to Openness: From the beginning, Databricks has advocated for open platforms to foster innovation, build community, and avoid vendor lock-in—a principle that remains central to its approach.
Core Ecosystem Pillars: Databricks has anchored its platform around several influential open-source projects it either originated or actively supports:
Delta Lake: Adds ACID transactions, reliability, and time travel to data lakes using open formats like Parquet, forming the backbone of the Lakehouse architecture.
MLflow: An open-source standard for managing the complete machine learning lifecycle, from experimentation to deployment and monitoring.
Delta Sharing: Enables secure, real-time data sharing across organizations and platforms without the overhead of replication.
Unity Catalog: Offers unified governance for data and AI assets, aligning with Databricks’ open and integrated philosophy.
DBRX LLM: The release of this powerful open-source foundational language model reflects Databricks’ ambitions to lead at the AI model layer as well.
This open-source strategy encourages widespread adoption and cultivates a thriving ecosystem, while Databricks generates revenue through its managed cloud platform, which delivers enterprise-grade features, robust security, and user-friendly experiences.
The Lakehouse Evolves: The Data Intelligence Platform
Databricks introduced the "Lakehouse" architecture to unify the scalability of data lakes with the reliability and governance of data warehouses. This approach is especially well-suited for industrial environments, which deal with a wide variety of data types—from structured ERP data and semi-structured logs to unstructured OT data such as sensor outputs, images, and time-series information.
Databricks recently rebranded its platform as the Data Intelligence Platform—a move that goes beyond marketing and signals a strategic evolution. By integrating generative AI capabilities, particularly following its acquisition of MosaicML, Databricks is building directly on its Lakehouse architecture to power the next generation of AI-driven solutions.
Unified Foundation: The platform brings together data engineering, SQL analytics, business intelligence, data science, and machine learning in one seamless environment.
AI-Centric Approach: It’s designed to help organizations harness and govern their proprietary data—including IT/OT convergence—to develop and deploy advanced AI applications, such as generative AI agents.
Enterprise-Wide Access: A core focus is on democratizing both data and AI, making powerful capabilities available across the organization.
Although Databricks doesn’t frequently use the term “Decision Intelligence,” its platform embodies its essence—moving beyond traditional analytics to deliver diagnostic, predictive, and prescriptive insights powered by AI and machine learning.
Governance for Industrial AI: The Critical Role of Unity Catalog
In Industrial AI applications—where decisions directly affect physical operations, safety, and high-value assets—governance is absolutely critical. Trust, transparency, and regulatory compliance aren’t optional; they’re essential. This is where Databricks' Unity Catalog stands out as a key differentiator, offering centralized governance to manage data and AI assets with the rigor industrial environments demand.
Unified Governance: Unity Catalog offers a centralized control plane to manage all data and AI assets—tables, files, models, features, notebooks, and dashboards—across multiple clouds and workspaces.
Fine-Grained Access Control: It supports granular permission settings at every level—catalog, schema, table, row, and column—using familiar SQL syntax, with both Role-Based Access Control (RBAC) and Attribute-Based Access Control (ABAC).
Automated Lineage: Especially valuable in industrial settings, Unity Catalog automatically captures end-to-end data lineage across queries and programming languages, including column-level lineage for machine learning models. This ensures traceability for audits, debugging, and understanding model origins.
Auditing: Comprehensive audit logs record all user interactions with governed assets, supporting strict compliance and security requirements.
Data Discovery: A built-in, searchable catalog makes it easy for users to locate trusted, governed data and AI resources across the organization.
This comprehensive governance framework is essential for establishing reliable DataOps and MLOps practices. By unifying the management of data and AI assets, Unity Catalog streamlines the complexity of developing trustworthy AI solutions—making it particularly well-suited to meet the rigorous demands of the industrial sector.
Strategic Alliances: Building Bridges in the Industrial Ecosystem
Databricks clearly recognizes that succeeding in the industrial market demands more than a robust platform—it calls for deep domain expertise and seamless integration with existing industrial and enterprise systems. To achieve this, the company leans heavily on strategic partnerships with established industry leaders.
AVEVA: This partnership directly addresses the challenge of IT/OT convergence. By integrating AVEVA CONNECT—AVEVA’s industrial intelligence platform—with Databricks, the collaboration aims to bring together AVEVA-managed operational data (from historians, MES, SCADA) and enterprise data stored in Databricks. Using Delta Sharing, it establishes a unified foundation for advanced analytics and AI applications such as predictive maintenance, process optimization, and sustainability reporting—all governed through Unity Catalog. This represents a major step in linking the operational floor with cloud-based AI capabilities. For deeper insights, see my coverage from AVEVA World 2025: Unifying Legacies and Forging Ecosystems for the Industrial AI Era.
Kinaxis: For industrial companies, supply chain visibility and resilience are mission-critical. The partnership between Databricks and Kinaxis addresses this by integrating Kinaxis’ Maestro™—an AI-native supply chain orchestration platform—with Databricks’ scalable data infrastructure and robust governance. Delta Sharing enables the smooth exchange of core supply chain data and external signals within Maestro, supporting faster, AI-driven decision-making. The result: enhanced agility and stronger, more resilient operations. For more, check out How Kinaxis AI Agents are Making Supply Chain Heroes using Databricks.
SAP: Arguably the most transformative alliance for enterprise integration, the partnership between Databricks and SAP embeds Databricks directly within SAP’s Business Data Cloud. Branded as “SAP Databricks,” the integration leverages Delta Sharing to enable zero-copy access to SAP data—from S/4HANA, Ariba, SuccessFactors, and more—alongside non-SAP data, all under the governance of Unity Catalog. This addresses a long-standing challenge for enterprises by creating a unified, governed foundation for building AI agents—such as SAP Joule or custom-built models—that can tap into rich business context and Databricks’ AI capabilities. For a deeper dive, see ARC Advisory Group’s Logistics Viewpoints blog on why The SAP Databricks Alliance is Truly Significant.
These partnerships go beyond technical integrations—they represent a strategic ecosystem approach. By positioning itself as the core data and AI engine, Databricks connects seamlessly with vital data sources through trusted partners, creating a more powerful and cohesive solution that resonates strongly with industrial clients.
Databricks' Place in the Industrial Data Fabric: A Pivotal Hub
So, returning to the central question: Is Databricks the leading Industrial Data Fabric? According to ARC’s definition—which differentiates between Enterprise Data Fabrics (focused on transactional and application data) and Industrial Data Fabrics (focused on complex OT data)—the answer is nuanced.
Databricks as the Enterprise Hub: Databricks stands out as the enterprise analytics core within a broader Industrial Data Fabric architecture. Its strengths lie in unifying diverse data once ingested, delivering advanced AI/ML capabilities, and enforcing robust, centralized governance through Unity Catalog. Strategic partnerships with SAP and Kinaxis further reinforce its enterprise-centric role.
The Need for OT Specialists: The “first mile” of OT data—capturing data from PLCs and SCADA systems, performing protocol translation, edge processing, and deeply contextualizing raw sensor data—is typically managed by OT-native platforms. This is where partners like AVEVA (with AVEVA CONNECT), Industrial Analytics players like Seeq, or Industrial DataOps players like HighByte, HiveMQ, TwinThread, and XMPro, play an essential role. They prepare and structure OT data for downstream consumption by platforms like Databricks.
In conclusion, while Databricks doesn’t represent the full end-to-end Industrial Data Fabric—especially in terms of native OT edge integration—it plays a pivotal role at the convergence point of IT and OT. For many ARC Advisory Group clients, Databricks increasingly serves as the high-performance engine where industrial data meets AI, backed by strong governance and a growing ecosystem of strategic integrations.
Looking Ahead: Databricks and the Future of Industrial AI
Databricks’ emphasis on open standards, unified governance, advanced AI capabilities, and strategic partnerships positions it as a powerful platform for organizations laying the groundwork for Industrial AI. Its role as a centralized, governed hub for analyzing converged IT/OT data is a key differentiator.
ARC Advisory Group clients can expect Databricks—and its partnerships with AVEVA, Kinaxis, and SAP—to feature prominently in our upcoming Industrial Data Fabric Market Landscape and Archetypes reports. Together, they represent a significant force in shaping how industrial organizations harness data and AI to drive the next wave of innovation.
However, it's important to recognize that the Industrial Data Fabric landscape is both complex and continuously evolving. No single vendor today can meet the full range of requirements across the diverse spectrum of Industrial AI use cases. Achieving success demands a well-considered architectural approach—one that integrates multiple best-of-breed solutions working together seamlessly.
Moreover, technology alone won’t close the gap. Overcoming persistent skills shortages and realizing the next level of industrial productivity will require more than robust platforms. It calls for investment in people, refinement of processes, and the cultivation of a culture rooted in data-driven decision-making. The journey is only just beginning.
Engage with ARC Advisory Group
For ARC Advisory Group recommendations for Navigating the AI Wars, Closing the Digital Divide by Embracing Industrial AI, assembling your Industrial-grade Data Fabric, and governing and guiding major decisions about enterprise, cloud, industrial edge, and AI software, please contact Colin Masson at [email protected] or set up a meeting with me, or my fellow Analysts at ARC Advisory Group.