Scale matters for clinical research and model development, and OMNY Health just moved from large to industrial scale. On July 31, 2025, the company said its de-identified real-world data network covers more than 100 million patients, roughly 30 percent of the U.S. population, and reaches every one of the 50 states. The release says the dataset includes eight years of historical records, 6.5 billion clinical notes, billions of clinical encounters from about 650,000 providers, and contributions from more than 46 healthcare organisations. Named contributors include St. Luke's University Health Network, Bon Secours Mercy Health and Baptist Health System KY & IN. That scale, OMNY argues, lets life sciences companies, health systems and AI developers query standardised longitudinal patient journeys at scale in a HIPAA-compliant, de-identified form.

The read here is simple. Scale matters for clinical research and model development, and OMNY Health just moved from large to industrial scale. Reaching 100 million de-identified patient records across every U.S. state shortens the path from asking a question about rare outcomes to actually finding cohorts big enough to answer it. For companies training AI on clinical text, the 6.5 billion clinical notes OMNY reports are the headline metric that changes the calculus on feasibility and generalisability.

OMNY Health positioned the milestone as an expansion of accessible, HIPAA-compliant de-identified real-world data for research and AI development. The July 31, 2025 press release states the network aggregates electronic medical records, unstructured clinical notes and medical claims spanning inpatient and outpatient care. The platform, the company says, now spans more than eight years of historical records and coverage from more than 650,000 providers.

The release also lists named provider contributors that anchor the dataset: St. Luke's University Health Network, Bon Secours Mercy Health and Baptist Health System KY & IN. The announcement emphasises therapeutic-area breadth, noting most areas are represented by data from more than 10 million patients each. OMNY framed scale and depth as the platform's primary advantages, arguing that longitudinal data across institutions helps reduce fragmentation that commonly limits single-system research.

That multi-institutional aggregation matters for a few concrete reasons. First, cohort construction is faster when systems standardise records from many sites into a single queryable model. Second, deep note corpora make it easier to train natural language models to recognise clinical concepts that billing codes don't capture. Third, broad provider coverage improves geographic diversity, which matters for external validity when algorithms are applied across different hospitals and populations. Matthew Fenty, Managing Director for Innovation & Strategic Partnerships at St. Luke's University Health Network, said OMNY provides a secure route for health systems to contribute and collectively build data foundations for AI development.

Mark Townsend, MD, MHCM, Chief Clinical Digital Ventures Officer at Bon Secours Mercy Health and Accrete Health Partners, said shared de-identified data helps address information gaps exposed by the COVID-19 pandemic.

The technical inventory behind those claims is familiar to anyone who has worked with clinical data. Electronic medical records and extracts remain central, but unstructured clinician notes, diagnostic and procedure codes, laboratory results, medication codings, provider metadata, device outputs and raw biomedical signals and images all play roles. Common code sets include ICD-10-CM/PCS for diagnoses and procedures, CPT for billing and services, LOINC for lab tests, and RxNorm and NDC for medication encoding. Administrative claims, surveillance systems, population surveys and primary data collection such as cohort studies and clinical trials are cited as complementary sources rather than competitors.

OMNY's expansion is chiefly a U.S.-centred commercial story, but the practical considerations map directly onto Canadian practice. The Government of Canada Open Data portal hosts datasets relevant to Canadians and offers search and reuse mechanisms. British Columbia maintains a BC Data Catalogue and BC Data portals, and PopDataBC supplies longitudinal, de-identified, person-level datasets covering provincial populations, with some holdings dating back to 1985. PopDataBC also operates a Data Scout query service that returns aggregated cohort-level feasibility metrics to researchers planning studies, while full person-level access requires formal data access requests.

For technologists and researchers choosing where to get data, the choices are concrete and predictable. Large commercial networks such as OMNY concentrate multi-institutional, longitudinal clinical data and deep unstructured note corpora, which simplify cohort construction and model training when access is granted. Public-sector datasets in Canada supply population-scale administrative and registry data suited to epidemiology and policy analysis, often with pre-established governance and data access pathways.

There are trade-offs. Primary data collection yields bespoke variables and control over measurement, but it's costly and slow. By contrast, administrative and EMR-derived sources offer scale and temporal depth at the expense of variable completeness and potential coding bias. That matters for model builders who must choose between cleaner, bespoke variables and larger, messier datasets that better reflect real-world practice. For policy analysts the question is different: administrative claims and registries often cover whole populations and long time spans, which is exactly what population-level policy questions require.

The practical implication for Canadian teams is to be explicit about the question they want to answer. If the goal is to train a clinical language model that recognises rare presentations, a multi-institutional, note-rich commercial network reduces the time to a usable dataset. If the aim is to measure incidence trends or evaluate health policy, provincial registries and federal open data holdings provide the canonical public-sector route. PopDataBC's Data Scout, for example, is a built-in feasibility tool that helps investigators estimate cohort sizes before they invest in the formal access process.

Access, governance and privacy remain core constraints in both sectors. OMNY emphasises its HIPAA-compliant, de-identified approach to enable sharing at scale. In Canada, established data access pathways, ethical review and formal request processes are the mechanisms that balance research utility with privacy obligations.

Related Articles

OMNY Health's July 31, 2025 announcement that its de-identified network exceeded 100 million patients and spans every U.S. state is a clear commercial milestone for U.S. clinical research and AI development. The next question is whether OMNY will extend partnerships beyond the U.S., and if so how it will reconcile international privacy and governance differences. For now, Canadian researchers continue to rely on the Government of Canada Open Data portal, the BC Data Catalogue and PopDataBC for governed, population-level datasets.

This article was created with AI assistance.