ADR-26-04: Non-Validated xRED Snowflake
Status |
Accepted |
Impact |
Medium |
Demand Owner |
Imran Khan |
ADR Contributors |
Anish Kejariwal |
Reviewers |
Architecture Council |
Informed |
Architecture Council, Heads of CSCoE-DDC |
Publication Date |
June 9th, 2026 |
š Background
The Computational Sciences Center of Excellence (CS-CoE) was established to unify and modernize the computation and data ecosystems across Rocheās Research and Early Development (RED) organizations, including pRED and gRED. In late February 2026, Accelerator C was kicked off to define a strategy for a future data backbone to support the REDs. Following May 29, 2026 feedback from the Heads of the DDC regarding the Accelerator C work, the decision was to move forward, but have a phased approach that delivers immediate, tangible business progress in iterative stages. The immediate priorities identified by leadership are:
-
Supporting the Lab Automation teamās efforts on Instrument Data Acquisition (IDA), by ensuring instrument data can be stored, accessed, and made AI data ready
-
Standing up a shared data store capable of handling structured and unstructured data to initially enable PCI (unstructured) and gRED assay plate data (structured) and corresponding catalog
To effectively support these priorities, a solution for managing structured data for analytical purposes is required. In early May, a 4-week āsprintā was initiated to answer the following questions:
-
Can we leverage a commercial off-the-shelf platform for accelerating how we deliver value for the organization? (rather than building everything ourselves on AWS)
-
What aspects of the end-to-end data problem do these solutions address?
-
As it pertains to our priorities around structured data (e.g., assay data, structured scientific results, and metadata), which solution will best address our needs?
Note: This evaluation was conducted in partnership with the Lab Automation team during their concurrent review of Instrument Data Acquisition (IDA) platforms, as some IDA vendors also offer data platform capabilities.
ā¹ļø Context
These are the following business drivers:
-
gRED assay data is extremely fragmented across numerous different data warehouses and data lakes, including a variety of technical solutions such as Postgres, Redshift, Oracle, Athena, DataOne/Snowflake.
-
Some gRED assay data is in systems (e.g., SMDI) that will have to be retired in the coming months
-
Project teams continue to build their own data lakes, warehouses, and lakehouses to solve their specific needs āĀ this leads to a duplication of efforts, without an aligned data strategy
-
AI/ML/Computational scientists have difficulty accessing assay data. One example is how AI4DD has had to build their own data lake to get from upstream assay systems to enable Lab-in-the-Loop
-
Speed to implementation: A strict requirement was given to deliver fast, incremental value to downstream consumers rather than waiting on long-term platform development cycles
-
Decision must account for fiscal and operational efficiencies
-
FDFM (Foundational Data for Foundational Models) and other gRED programs leverage Mnemos for metadata, and Mnemosā backend of Postgres is beginning to have scaling challenges (in ~1 year, 17M assets have been cataloged, with the peak being 2.3M assets in a single day)
-
To accommodate the vast volumes of unstructured data in the REDs, Accelerator C recommended a lakehouse approach to unify structured and unstructured data, metadata, and lineage under a single framework.
š Consequences
The strategic pivot to a phased data layer approach, combined with the urgent business timelines for the two priority workstreams, presents significant risk if a standardized structured data platform is not established quickly. We risk:
-
Timelines: Failing to deliver on the two priority leadership workstreams within a critical 3ā4 month window
-
AI/ML Bottlenecks: Continued challenges for AI/ML scientists trying to find, access, and analyze data
-
Metadata scalability: Continued scalability challenges for storing metadata
-
Continued data fragmentation in the organization, instead of beginning to establish a foundational data layer.
š¤ Outcome / Decision
Recommendation and scope
The recommendation is to leverage a nonāGxP-validated instance of Snowflake provided by the RDT to address analytical structured data needs for the two priority areas identified by the DDC leadership, but with an eye towards being broadly used across the REDs as part of an eventual xRED data layer. We will be positioning Snowflake in the CS-CoE as a specialized tool in our overall toolkit rather than an all-encompassing platform, allowing us to accelerate analytics while keeping our core data storage within AWS S3. We will begin by selecting projects/use-cases from the two data layer priorities identified by leadership.
This xRED Snowflake provided by RDT would:
-
Be located in us-west-2 initially, and eu-central-1 soon after
-
Leverage RDT enterprise discounts for Snowflake
-
Leverage RDTās current strong partnership with Snowflake
-
Leverage RDTās already built integrations (ex: with Pingfed, networking, alerting and monitoring, etc.)
-
Leverage L3 (infrastructure) support from RDT (Defining L1-L3 support between RDT and CoE, and leveraging TSO, is a work-in-progress)
Solutions Considered
The following were considered
-
AWS - Leverage iceberg, S3 Tables, and Redshift for structured data
-
Databricks - Roche does not currently have a license for Databricks
-
DataOne - GxP validated version of Snowflake provided by RDT
-
Snowflake - vanilla version provided by RDT for these evaluations
-
TetraScience - platform for instrument data acquisition, but also has a data platform built on Databricks. This ADR focuses only on the data platform aspect of TetraScience.
The only option not evaluated was DataOne (the GxP-validated version of Snowflake provided by RDT). While DataOne is already being used in production today by the CoE, it wasnāt evaluated for the following reasons:
-
RDT and DataOne platform leadership explicitly align that a GxP-validated environment introduces unnecessary rigidity for exploratory, research-driven workflows in the REDs
-
Every minor modification or permission update requires navigating slow, centralized change management processes.
-
Is optimized for more static use cases, rather than experimental workflows and short-lived use cases
-
Historical precedents in pRED show that platform updates are often pushed with minimal notice, breaking existing downstream data products.
-
Prior prototyping efforts were hindered by complex governance and slow onboarding
-
Teams are tightly bound to the initial schema and database structures they originally requested, but the inability to modify these leads to redundancies in the data model
-
Strict validation and PII guardrails mandate highly complex integration patterns, frequently driving users to copy data out and into S3
-
New functionality from Snowflake is gated in DataOne due to validation processes. These hinder innovation and have resulted in Snowflake being used as a traditional data warehouse
There is precedent for not using DataOne, as Chugai has their own instance of Snowflake provided by RDT.
Evaluation Process
To evaluate the different platforms, it was determined that we would identify 5 business use cases and execute PoCs on each of these across the four platforms. The goal was not to do end-to-end PoCs nor to deliver things to production. Rather, the goal was to do enough to compare the different platforms side-by-side.
To identify the business use cases for the evals, the following process was leveraged:
-
Accelerator C identified potential use cases for the xRED data layer. This was used as a starting point.
-
As it was understood there was a gap in structured storage for gRED assay data, these use cases were reviewed with the gRED Domain Heads for prioritization
-
Meetings were conducted with key stakeholders to identify data readiness, feasibility for a rapid PoC, stakeholder availability, and relevance to evaluating the different platforms
Based on this, five use cases were selected:
-
pTox Cell Viability Assay (gRED LW&D)
-
BCP Functional Assays Large Molecule Data Warehouse (gRED TD)
-
ROCKS (pRED Scientific Knowledge for Unstructured Data)
-
Preclinical Imaging metadata for OPS data (xRED PCI)
-
Single Cell Atlas (gRED TIE)
A team was constructed for each use case involving two engineers and at least one stakeholder. There were daily standups and weekly show & tell demo meetings. Lastly, to ensure robust, objective evaluations of the different platforms:
-
Each vendor was asked to fill in the evaluation criteria
-
Each vendor had at least one office hours per week with the evaluation team
-
A Slack channel was set up per vendor to ensure the evaluation team could ask direct questions of the vendor and get support as needed
To conclude the evaluation period:
-
Each technical team member was asked to fill out the evaluation spreadsheet for each of the platforms. Team members were not required to rate all evaluation criteria, just the criteria they evaluated.
-
Each use case team presented their learnings across the platforms
-
Each member presented their learnings and their recommendation
Rationale for Decision
As described earlier, each of the technical evaluators was asked to fill out the evaluation spreadsheet for each of the platforms. Below is the summary of the averaged scores across all the evaluators for all the evaluation criteria:
| TetraScience | Databricks | Snowflake | AWS (Redshift) | |
|---|---|---|---|---|
AVG SCORE |
1.81 |
2.39 |
2.45 |
2.31 |
These average scores are provided only to give a high-level representation of the scoring results. These scores were not used to make the final decision.
During the evaluation process, AWS and TetraScience were quickly ruled out:
-
AWS S3 Tables, Iceberg, Redshift, Athena:
-
AWS Redshift was deemed not viable due to the significant engineering required to stitch services together
-
Redshift was significantly slower than Snowflake and Databricks on JSON/VARIANT queries, and had difficulty handling large JSON data structures
-
-
TetraScience (Data Platform only; not the Instrument Data Acquisition components)
-
TetraScienceās data platform had too narrow a scope, as its data platform is only for data acquired from instruments and is not as robust at supporting externally generated data (e.g., CROs, licensed data, public data)
-
Out-of-the-box data transformation pipelines created complex, difficult-to-understand data structures
-
Note: these notes are not reflective of TetraScienceās instrument data acquisition capabilities.
-
The evaluation team was evenly split on Databricks vs. Snowflake, as most considered both to be excellent options capable of delivering strong business benefits while supporting a foundational, long-term technical strategy (specifically a Lakehouse strategy leveraging Apache Iceberg tables on AWS/S3).
Ultimately, Snowflake is recommended for the following reasons:
-
While Databricks offers an expansive platform with robust ML and computational tools (e.g., MLFlow, Spark), these are not the largest gaps in the organization today, and these were not part of the evaluation criteria; Snowflake, in contrast, is a more narrowly focused product, while still addressing the high-priority data gaps we have in our organization today
-
Snowflake provided the best support during the evaluations
-
Snowflake had the most credible global, multi-region architecture proposal
-
CoE could leverage the existing relationship and experience that RDT has with Snowflake
-
Roche already has an enterprise license with Snowflake āĀ this enables the CoE to hit the ground running without having extensive procurement discussions
-
RDT has already negotiated significant enterprise discounts (note: we very roughly estimate Databricks to be at least 2 times more expensive)
-
RDT already has a strong partnership with Snowflake. This will be leveraged to a direct partnership between the CoE and Snowflake, and secure a Snowflake solutions architect to help design a modern implementation with best practices.
-
RDT already has strong expertise with Snowflake
-
CoE already has pockets of strong expertise with Snowflake, and already has workloads running on Snowflake (via DataOne)
-
While both Snowflake and Databricks do not currently solve PBAC (policy-based access control), PBAM (project-based access management), and ABAC (attribute-based access control), RDT has already begun discussions with Snowflake on PBAC
-
CoE can leverage RDT for infrastructure support (ex: L3)
-
Rather than the CoE having to determine how to implement base infrastructure, RDT can spin up Snowflake for the REDs with many mandatory integrations already provided (integration with PingFed, networking, alerting and monitoring, etc.)
-
Lastly, to limit vendor lock-in, we will anchor the architecture on a Lakehouse strategy leveraging open-source Apache Iceberg tables on AWS/S3. Furthermore, Snowflake mostly utilizes ANSI SQL, ensuring that core queries and pipelines can be portable across alternative data platforms. This strategy ensures long-term flexibility, allowing us to point other engines at our S3 buckets in the future.
š Resources
-
Use Case PoC Readouts
-
Bonus Use Case: SADR