UPMC Enterprises’ Data Platform Speeds Research with Widely Used Common Data Model for Health Care

For researchers, pharmaceutical companies, and AI developers, gaining access to large datasets to generate valuable insights is only part of the challenge. Equally important is the ability to work with that data efficiently and consistently. Ahavi, the real-world data platform from UPMC Enterprises, addresses this challenge by using a widely used common data model across its structured datasets, reducing the time spent preparing data and increasing the time spent generating insights.

The approach enables faster research, smoother collaboration, and greater confidence that analyses can be reproduced and scaled across projects, organizations, and use cases. In a landscape where speed and interoperability are critical, the benefits extend far beyond standardization itself.

By applying a common data model to UPMC’s vast stores of data, “we’re able to simplify, standardize, and stratify our data,” said Michele Mueser, senior director of data engineering at UPMC Enterprises. “This allows us to gather clinically meaningful information from the patient record that can support responsible research, analytics, and model development.”

Ahavi, the de-identified, secure, and compliant data platform with wraparound expert services, generates insights from millions of patient care journeys across UPMC for pharmaceutical and biotechnology companies and also offers a robust training and validation environment for health care AI models. UPMC is among the nation’s largest non-profit provider-payer systems with more than 40 hospitals and 800 outpatient sites.

Structured data is standardized following the Observational Medical Outcomes Partnership (OMOP) common data model, designed to enable large scale observational analytics across multiple data sources. This model creates opportunities for faster delivery times, richer data sets, greater collaboration, and new capabilities for partners utilizing UPMC’s real-world data capabilities.

“With these advancements, there are significant opportunities for us to go faster and integrate more broadly in the health care ecosystem with Ahavi,” Mueser said. “We’re very excited to see where OMOP takes us in the future.”

Mueser led the adoption of OMOP, along with many others across the UPMC Enterprises’ technology team, including individuals in software engineering, data engineering, data analytics and informatics, solutions architecture, cloud infrastructure services, and operations representing both the complexity of the problem space as well as the expertise available to Ahavi analytic and solution building customers.

“The projects were a melting pot of every one of the skills that our technology and data services have been developing and delivering to support our partners for many years,” Mueser said.

What is OMOP and Why Adopt It?

OMOP is a standardized data model and analytics framework maintained by the Observational Health Data Sciences and Informatics (ODHSI) community that allows researchers, health care organizations, and regulators to perform large-scale, reproducible observational studies across multiple institutions without needing to re-format data for each analysis.

ODHSI community includes more than 4,700 researchers across 88 countries and collaborates on a single open data standard and the development of open-source standardized analytics tools.

UPMC Enterprises adopted the OMOP standard for multiple reasons. It is battle-tested with wide adoption across health care, existing tools reduce time to deliver customer value, and it addresses market needs in both model development and analytics.

For a health system of UPMC’s scale, normalization is not a technical preference — it is a prerequisite for responsible, scalable use of real-world data. Source systems such as Epic, Cerner, and other clinical platforms capture similar clinical concepts in different structures, vocabularies, and workflows. By mapping those data sources into OMOP, Ahavi creates a consistent analytic layer that makes cohort definition, feature generation, reproducible analytics, and model validation more efficient and comparable across a depth and breadth of clinical data.

“OMOP is already broadly established — we didn’t want to build our own standard from the ground up,” Mueser said. “Instead, we’re leveraging the technologies, applications, and wisdom that’s already been created.”

How the OMOP Common Data Model Benefits Ahavi

Ahavi provides de-identified health care data to enable researchers, scientists, and developers to create curated datasets for accelerating research, clinical trial design, and AI development. By leveraging UPMC’s extensive datasets, Ahavi ensures that AI models, predictive analytics, and clinical decision-making tools are built on trusted, real-world data, helping solve some of the biggest challenges in health care.

The platform is a repository of structured and unstructured data based on the care journeys of more than 5 million patients. It includes structured data going back to 2019 and unstructured to 2012. On top of those impressive numbers, Ahavi boasts an industry-leading linkage rate between structured and unstructured data of at least 80%.

Christopher Horvat, MD, vice chair of clinical informatics and digital health for the Department of Critical Care Medicine, University of Pittsburgh and the UPMC ICU Service Center, sees significant advantages for Ahavi by adopting the OMOP standard.

Dr. Horvat explained that traditionally 90% of the work of analyzing data was spent cleansing and curating it. That meant only about 10% of the effort would actually be spent generating insights. “A key advantage of OMOP is it inverts that proportion,” Dr. Horvat said.

With OMOP, he said, “it’s a common schema that the world is familiar with. So, 90% of your time and data scientists’ talent and effort can go into generating key insights.”

What is OMOP and Why Adopt It?

OMOP is a standardized data model and analytics framework maintained by the Observational Health Data Sciences and Informatics (ODHSI) community that allows researchers, health care organizations, and regulators to perform large-scale, reproducible observational studies across multiple institutions without needing to re-format data for each analysis.

ODHSI community includes more than 4,700 researchers across 88 countries and collaborates on a single open data standard and the development of open-source standardized analytics tools.

UPMC Enterprises adopted the OMOP standard for multiple reasons. It is battle-tested with wide adoption across health care, existing tools reduce time to deliver customer value, and it addresses market needs in both model development and analytics.

For a health system of UPMC’s scale, normalization is not a technical preference — it is a prerequisite for responsible, scalable use of real-world data. Source systems such as Epic, Cerner, and other clinical platforms capture similar clinical concepts in different structures, vocabularies, and workflows. By mapping those data into OMOP, Ahavi creates a consistent analytic layer that makes cohort definition, feature generation, reproducible analytics, and model validation more efficient and more comparable across use cases.

“OMOP is already broadly established — we didn’t want to build our own standard from the ground up,” Mueser said. “Instead, we’re leveraging the technologies, applications, and wisdom that’s already been created.”

Ahavi Allows AI Developers to Go Faster

While EHRs are the source for the most robust data sets, it is not practical or necessary for scientists and developers to use everything to answer research questions or train models.

“OMOP takes the best of the best elements that the ODHSI community has agreed is important — it takes what would be hundreds of thousands of data tables and fields from the EHR and distills them down to essentially 40 tables,” Mueser said. “The result is you get tremendous insights at the patient level without having the complexity of an EHR.”

Because OMOP is broadly used in health care research and real-world evidence generation, many life sciences and AI teams already understand its structure. That familiarity reduces onboarding friction and allows partners to focus less on data wrangling and more on cohort design, feature engineering, model training, validation, and interpretation.

“We’ve standardized the data in Ahavi and that allows companies to accelerate their work because they know what to expect,” Mueser said.

Interaction with Ahavi’s Unstructured Data

While Ahavi’s unstructured data is not currently incorporated into the OMOP model, OMOP benefits the retrieval and matching of unstructured data from Ahavi. Cohorts developed from the structured data in OMOP help identify patient populations for whom relevant documents can be retrieved from the unstructured data repository within Ahavi. Integrating unstructured insights into the OMOP model space as a native capability is firmly on the Ahavi roadmap.

“Eventually, we hope to leverage natural language processing and large language models to extract metadata from unstructured documents, which we’ll then insert back into OMOP where’s there’s a space for notes and note metadata,” Mueser said, adding that extracting metadata from unstructured notes would enrich the OMOP data for direct analytics as well as reduce costs and time for customers.

“There are significant opportunities for us to leverage the vast amount of unstructured data that we have by using models to pull out key data,” she said.

Integration of Claims, Imaging, Device, and Omics Data

As OMOP and related OHDSI extensions continue to evolve, the Ahavi team can evaluate high-value data domains, including imaging, device, and multi-omic data for incorporation into standardized or OMOP-aligned structures where appropriate.

Looking ahead, the same normalization strategy can support Ahavi’s expansion beyond traditional clinical data into molecular and multi-omic datasets, including genomics, transcriptomics, proteomics, and related assay-derived features. While OMOP’s support for high-throughput omics is still evolving, emerging OHDSI community work is demonstrating how clinical and molecular data can be harmonized into OMOP-compatible structures.

For Ahavi, this creates a path toward combining longitudinal clinical context with molecular signals, an important foundation for precision medicine analytics, biomarker discovery, cohort enrichment, and AI model development.

“OMOP gives us a standardized foundation that extends beyond traditional clinical data,” Mueser said. “As standards and community-supported approaches continue to evolve, it creates opportunities to connect longitudinal clinical information with emerging data types, including molecular and other multi-modal datasets. That combination has the potential to enable more advanced analytics, precision medicine research, and next-generation AI applications.”

Ahavi is Accelerating Innovation

With these upgrades and capabilities, UPMC Enterprises and the Ahavi data platform are poised to deliver significant advantages to its partners leveraging real world data for clinical trials and AI development.

Already a tremendous resource for real-world health care data, Ahavi is creating even more opportunities for pharmaceutical and biotechnology companies and AI developers to accelerate their innovations.

Next Steps

You Might Also Like…

Read More