Welcome back to our ongoing coverage of the 2026 ARC Industry Leadership Forum. On Tuesday, we explored the gritty "trench warfare" of scaling Industrial AI and escaping pilot purgatory. By Wednesday morning, the conversation shifted to something more fundamental: the heavy machinery required to actually win that war.

I hosted the Track 1 breakout session on Wednesday morning: "Assembling the Industrial Data Fabric Foundations for Industrial AI." Let me be blunt: you cannot execute advanced agentic AI or predictive operations if your underlying data infrastructure is a chaotic web of point-to-point integrations. As I told the packed room to open the session:
"We've seen a massive rediscovery of the data quality problem. Garbage in, garbage out is going to get exponentially worse with the power of AI layered on top. Ensuring data quality remains the absolute number one barrier to scaling out AI today." — Colin Masson, ARC Advisory Group
To prove this isn’t just theoretical architecture, we brought in leaders from International Paper, Plains, and Baker Hughes, alongside their ecosystem partners Radix, Streamline Control, and HighByte. While I hesitate to label anyone definitively outside formal assessments, the velocity at which International Paper and Baker Hughes are moving strongly suggests Pacesetter-level maturity—particularly through platforms like Cognite Data Fusion and SAP Databricks.
Here is how these leaders are systematically dismantling legacy data silos.
International Paper: Breaking the "Exponential Wall"
Stephen Krassick, who leads the Data Fabric digital transformation program at International Paper, described a challenge that resonated across both OT and IT: the "Multi-Threaded Exponential Wall."

"We are moving towards a multi-factor, multi-threaded, exponential wall. The more data sources you have and the more applications you build, the less sustainable it becomes. You must engineer a way to make that sustainable." — Stephen Krassick, International Paper
International Paper recognized that routing data directly from source systems to individual applications was a dead end. Instead, it partnered with Radix to implement Cognite Data Fusion, building a central knowledge graph that contextualizes everything from historian data to unstructured engineering artifacts like P&ID drawings.
The result is scale. Krassick highlighted their "Golden Run" application, which optimizes over 200 constantly changing variables on a paper machine to determine the optimal production path—much like a Formula 1 team tuning a race car. With a data fabric foundation in place, International Paper deployed 18 production applications in just six months—execution speed that places it firmly among top-tier industrial operators.
Plains: Untangling the Point-to-Point Chaos

Juan Yactayo (Plains) and Matthew Vana (Streamline Control) on simplifying legacy integration architectures
Juan Yactayo and Matthew Vana (from integration partner Streamline Control) walked through how Plains is addressing similar integration challenges in pipeline and storage operations. Juan highlighted the compounding effect of legacy acquisitions:
"We had numerous point-to-point integrations and a disparate number of solutions across our assets. We never truly simplified our internal architecture, and over time, it became incredibly challenging to support and troubleshoot." — Juan Yactayo, Plains
To bring order to its architecture, Plains deployed the HighByte Intelligence Hub to establish a true Industrial DataOps foundation. Acting as an OT broker, it standardized payload structures and eliminated brittle, custom integrations. The success of this deployment also prompted me to bring HighByte’s Chief Product Officer, John Harrington, into the panel later in the session. The gains in agility were immediate. Matthew Vana highlighted the speed of the new architecture:
"From our initial proof of concept, we quickly scaled and configured the data pipeline for approximately 1,400 EFM meters in a matter of days. We finished well ahead of our project timeline, which is a testament to the intuitive configuration." — Matthew Vana, Streamline Control
Baker Hughes: Asset Strategy Over Asset Health
Carlos Gomez from Baker Hughes presented a compelling case for building a next-generation Asset Performance Management system. SAP had specifically pointed to Baker Hughes as a strong example of leveraging the SAP Databricks alliance—a timely topic as organizations struggle to unlock siloed SAP data for Industrial AI.
Beyond the architecture, Gomez emphasized a critical distinction between detection and resolution:

"Just because you can detect an asset health problem doesn't mean you've corrected it. You can catch a failure and fix it, but it will happen again in three to six months because you didn't change your overarching asset strategy to solve the root cause." — Carlos Gomez, Baker Hughes
By bringing Databricks natively into the SAP environment, Baker Hughes reduced complex entity mapping timelines from 12 months to 30 days. But Gomez also issued a forward-looking warning:
"Just because we have connected information does not mean we have intelligent information."
The data fabric is the starting line. Reasoning and context are the race.
The Panel: Swamps, Missing Data, and the "Big Bang" Debate

Industrial Data Fabric panel discussion at ARC Industry Leadership Forum. Panel (l-R): Colin Masson, John Harrington, Simon Sierra, Stephen Krassick, Carlos Gomez, Matthew Vana, and Juan Yactayo
When John Harrington (Chief Product Officer, HighByte) and Simon Sierra (VP of Sales, Radix) joined the panel, the discussion quickly shifted to deployment strategy: build everything at once (the "Big Bang"), or scale use case by use case.
Simon Sierra delivered one of the most pointed observations of the session:
"One of the biggest mistakes the industry made over the past few years was moving everything into the cloud just for the sake of doing it—turning data lakes into data swamps. The true complexity within a plant is on a scale that is nowhere near appreciated by IT." — Simon Sierra, Radix
John Harrington expanded on why this cloud-first mindset often fails. Many IT teams assume plant historians (such as PI) provide a complete data foundation. Harrington challenged that assumption directly:
"IT assumes historians have all the answers. But most of the decisions on what data went into PI were based around managing and documenting a specific process. There is a massive amount of untapped data on the factory floor, the processing plant, or the oil rig. Layering an API on top of that and making it available when needed is critical." — John Harrington, HighByte
This reinforced the pragmatic approach taken by end users. Stephen Krassick and Juan Yactayo both emphasized that specific business problems—such as Golden Run optimization or EFM meter scaling—must drive data strategy, rather than attempting to centralize everything at once.
As Harrington summarized: "You must start deriving tangible value from your systems immediately... but you [also] have to establish that underlying data strategy from day one."
As moderator, I pushed the panel on a persistent reality: legacy environments often lack sufficient sensors and clean data.
"If you lack the sensors, or your data quality is poor, how do you handle those use cases? Do you infer the data, or can you actively clean it up?"
Carlos Gomez brought the conversation back to the real objective: automated reasoning.
"Beyond building a connected fabric and establishing a well-contextualized capability set, we need the reasoning layer—understanding why something is happening to make an autonomous decision. Just because we have connected information does not mean we have intelligent information." — Carlos Gomez, Baker Hughes
As we closed, I pointed to what lies ahead:
"Next year, we will undoubtedly be diving deeper into AI agents, MCP services, and AI for DataOps to address these missing data types. However, I predict we will still be having this conversation, because not everyone will have solved their foundational Data Fabric problem." — Colin Masson, ARC Advisory Group
The Verdict: Build the Foundation or Get Left Behind
Wednesday’s session reinforced a core tenet of ARC’s Cyber-Physical Industrial Architecture framework: AI is only as good as the context feeding it.
Whether you are using Cognite, HighByte, SAP, or a hybrid ecosystem, assembling a unified, contextualized Industrial Data Fabric is now the non-negotiable cost of entry.
Stop dumping unstructured data into cloud swamps. Start mapping your assets, securing your namespaces, and engineering your context.
The Data is Clear
If you need further proof that the Data Fabric is the primary battleground of 2026, look no further than our recently published Industrial AI Pacesetters 2026 Report and insights from our Q4 2025 Industrial AI, Energy, and Robotics Survey. Across the board, our data shows an overwhelming consensus: without a decoupled, contextualized Industrial Data Fabric, organizations will remain trapped in pilot purgatory.
Engage with ARC Advisory Group
The Industrial AI (R)Evolution is moving faster than ever. To dive deeper into the frameworks and data shaping the future of the industrial sector, explore my latest research:
Navigating the AI Wars and the escalating Industrial Robot Wars
Closing the Digital Divide by Embracing Industrial AI
Assembling your Industrial-Grade Data Fabric
Charting the new frontier of Physical Intelligence and transitioning to a Cyber-Physical Industrial Architecture (CPIA)
Mapping your maturity and strategy with ARC's 3-Axis Industrial AI Models Taxonomy
Where do YOU stand in the Industrial AI (R)Evolution? Take our Industrial AI Assessment to benchmark your organization's maturity, identify critical gaps in your IT/OT/ET convergence, and get actionable recommendations to accelerate your path to becoming an Industrial AI Pacesetter.
Don't guess what your global operations or prospective customers need. Use empirical data to align your stakeholders and de-hype the market with ARC Advisory Group's Voice of Market Service.
For tailored recommendations on governing and guiding major people, process, and technology decisions across the enterprise, cloud, industrial edge, and AI, please contact Colin Masson at [email protected].
Or, set up a meeting with my fellow Analysts and I at ARC Advisory Group to find out more about our Executive Insights Service for Industrial organizations and our Industrial AI Insights Service for Vendors.