Data Architecture

Information Models

Purpose

Information models give an overview of the information and data within an area of interest using the common terms and language of that area. The intent is not to show all details but ensure that key concepts and their relations are clearly documented and concisely communicated.

Usage

Information models are to be used when designing business processes or applications to ensure a consistent language across the organisation. Relationships between concepts are important to ensure the right context is included when information is captured or used and to identify the information which is central to joining data together.

Exploration

The xRED information models are a growing collection which will have regular updates. They can be explored by clicking on the following image. Deeper dives into the model can be made by clicking on the key concepts of Molecule, Therapeutic Target and Gene

xRED_Information_Model

Information models follow standard UML format.

Each box represents a data object

Relationships have the following meanings

 — ◆ Aggregation: Indicates that a concept groups a number of other concepts.

 —  Association: An association models an unspecified relationship.

 — ◇ Composition: The composition relationship indicates that a concept consists of one or more other cpncepts.

 — 🢖 Generalization: Indicates that a concept is a particular type of another concept.

Data Flows

Purpose

Data flows represent the movement of information between a source application and a target application.

Usage

Data flows can be used for several use cases including the following common examples

Impact assessments: When changing the data model of an application information flows will indiciate which applications may be impacted.

Data lineages: To undersand from where an application or report obtains its data the data flows can be used to trace back to the sources

Trusted sources of data: The true sources of data can be indetified as those applications which do not have an incoming data flow for that data

Complexity: The relative complexity of data flows can be calculated and compared by considering the volume, velocity, variety and technical aspects of the data flow.

Integrity: Tracing the flow of data through applications from the original source to the final target can identify any risks to loss of integrity.

Exploration

Data flow report

All data flows can be viewed in AMR via the Data Flow Report. The data flows can be explored by clicking on the following image. (must be on Roche network)

xRED_Data_Flow_Report

Data flow complexity heatmap

The xRED AMR Data Flow Complexity Heatmap shows the complexity of data flows per data object. A colour scale indicates the relative complexity and the size of each box indicates the number of applications per data object. The heatmap can be explored by clicking on the following image.

xRED_AMR_Data_Flow_Complexity_Heatmap

Simplification Opportunities

As part of the xRED Reference architecture goal opportunities for simplification across a variety of use cases are being identified, Below are the observations and recommendations.

Safety Data Integration

The Safety Data Integration offers a comprehensive solution for the ingestion, integration, storage, and sharing of study safety data. An extensive data model and data quality rules ensure good data can be captured as required. Integration with the study system of record ensures a transparent catalogue. Reuse of terminology and reference data services ensures alignment with internal and external data standards.

Opportunities for simplification include

  • Terminology: The scope of the solution is to store safety data, which is broader than data generated by animal toxicology studies. This leads to the data structures of the system using terms which are more generic than the actual usage, or exposing APIs with no data, e.g., visits

  • Data Quality: Rules that have been added to correct previously detected errors on future loads, so-called ‘edit rules’, are inconsistently applied. Sometimes they fire, and sometimes they don’t.

  • Data Integration: After a data load has been validated, it can take many hours (10+) for a data package to be mapped into the universal data model.

  • Data Integration: Synchronizing with FISH is an optional process performed after data loading by the end users

  • Data Sharing: The reuse of data by other solutions is not transparent due to insufficient tracking of consumers.

  • Data Quality: The solution offers many data quality rules, which may or may not be required

Biological Data system

Work in progress

Electronic Lab Notebook

Work in progress

Project data ecosystem

Work in progress

Research Data Management Capabilites

A collection of data capabities have been defined under the business capability Research Data Management. These can be used to map applications which have a primary function related to data. The data capabilities can be explore via the Business Capability Taxonomy (must be on Roche network)

  • Research Data Management: The overall capability to govern, organise, store, maintain and use data assets throughout their lifecycle to ensure they are accurate, available and fit for purpose.

    • Research Data Acquisition: The capability to identify, access and obtain data from external or internal sources for use in research and clinical activities, incl. negotiating access, establishing connections, and receiving data feeds.

    • Research Data Collection: The capability to gather data at the point of generation: capturing data from experiments, clinical procedures, patient interactions, and instrument outputs into governed systems.

    • External Research Data Exchange: The capability to securely transfer data between the organization and external partners such as CROs, central laboratories, regulatory agencies, health authorities and data providers using standardised formats with appropriate governance and audit trails.

    • Research Data Ingestion: The capability to receive, validate and load data from external or instrument sources into the research data platforms which includes format validation, schema mapping and initial quality checks before data enters governed stores.

    • Research Data Integration: The capability to combine data from disparate sources that have different systems, formats and domains into unified, coherent and consistent datasets which can be used for analysis, reporting or decision making.

    • Research Data Storage: The capability to provide governed, structured storage environments where research data is held, organized, indexed and made accessible for retrieval and analysis which includes databases, data lakes and domain-specific scientific repositories.

    • Research Data Cataloguing: The capability to create and maintain an inventory of available data assets by documenting what data exists, where it is held, what it contains, who owns it and how it can be accessed thereby making data discoverable across the organization.

    • Research Data Curation: The capability to actively maintain and improve data quality over time by selecting, organising, annotating, enriching and preserving data assets to ensure their ongoing fitness for scientific and regulatory use.

    • Research Data Governance: The capability to govern the availability, usability, integrity and security of the data in enterprise capabilities, based on internal data standards and policies that also control data usage. Effective data governance ensures that data is consistent and trustworthy and doesn’t get misused.

    • Research Master Data Management: The capability to define, maintain, and govern the authoritative master records and reference datasets that provide consistent, standardised values used across all research and clinical systems which includes compound identifiers, patient identifiers, site codes and standard reference lists.

    • Research Reference Data Management: The capability to govern the standardised vocabularies, coding systems and controlled terminologies used across research and clinical data which includes MedDRA, WHO Drug Dictionary, CDISC controlled terminology, SNOMED and ICD codes thereby ensuring consistent application and version management.

    • Research Data Security & Access Control: Govern data classification, role-based access control, encryption and audit trail management for research data.