---
title: "Synthetic Data: Solving the Clinical Trial Privacy Paradox"
description: "Solving the Privacy Paradox: Generating Anonymous Clinical Trial Data with AI"
image: https://blog.jmir.org/hubfs/Pierre-Antoine%20Gourraud%20%20blog.png
---

[![JMIR Publications logo](https://blog.jmir.org/hs-fs/hubfs/JMIR-Publications_logo_2020_original.png?width=550&height=109&name=JMIR-Publications_logo_2020_original.png "JMIR Publications logo")](http://www.jmirpublications.com)

Open main menu Close main menu

[Research News](https://blog.jmir.org/tag/research-news)

# Synthetic Data: Solving the Clinical Trial Privacy Paradox

![Reviewed by Kayleigh-Ann Clegg, PhD](https://blog.jmir.org/hs-fs/hubfs/Headshot.jpg?width=45&name=Headshot.jpg) 

[Reviewed by Kayleigh-Ann Clegg, PhD](https://blog.jmir.org/author/reviewed-by-kayleigh-ann-clegg-phd)

Published Date: 17 December 2025 5:13 p.m.

[Share this blog post on Twitter](https://twitter.com/intent/tweet?text=I+found+this+interesting+blog+post&url=https://blog.jmir.org/synthetic-data-solving-the-clinical-trial-privacy-paradox) [Share this blog post on Facebook](http://www.facebook.com/share.php?u=https://blog.jmir.org/synthetic-data-solving-the-clinical-trial-privacy-paradox) [Share this blog post on LinkedIn](http://www.linkedin.com/shareArticle?mini=true&url=https://blog.jmir.org/synthetic-data-solving-the-clinical-trial-privacy-paradox)

![Solving the Privacy Paradox: Generating Anonymous Clinical Trial Data with AI](https://blog.jmir.org/hubfs/Pierre-Antoine%20Gourraud%20%20blog.png)

The foundation of modern, data-driven medicine rests on high-quality empirical evidence, primarily derived from **Randomized Clinical Trials (RCTs)**. However, sharing the individual patient data (IPD) from these trials is heavily restricted by regulatory frameworks. For example, the EU’s General Data Protection Regulation (GDPR) which aims to protect patient privacy. This "privacy paradox" can hinder innovation, preventing researchers from reusing valuable data to develop predictive models or external control arms.

 

A novel solution is emerging from the realm of generative artificial intelligence: **synthetic data**. Instead of sharing sensitive data from real patients, researchers can generate shareable *virtual* patient populations as proxies. But can these synthetic datasets accurately replicate complex clinical outcomes while guaranteeing anonymity?

A recent study tackled this head-on. In their paper, "[Privacy-by-Design Approach to Generate Two Virtual Clinical Trials for Multiple Sclerosis and Release Them as Open Datasets: Evaluation Study](https://www.jmir.org/2025/1/e71297/)," published in the [*Journal of Medical Internet Research*](https://www.jmir.org/), Pierre-Antoine Gourraud and a multi-site research team from France demonstrated a successful method for achieving both high utility and satisfactory privacy, specifically for Multiple Sclerosis (MS) trials.

### **Avatars: A Privacy-by-Design Approach**

The study utilized a privacy-by-design technique called the **"avatars"** technique, which generates synthetic data points using a multidimensional reduction and nearest neighbors algorithm. Unlike typical AI generators, the avatars technique is designed specifically as an anonymization method, enabling the team to perform an explicit privacy assessment.

The researchers tested their method against data from two phase 3 MS RCTs: CLARITY (Merck) and ADVANCE (Biogen), which involved a total of over **2,300 patients**. The goal was to select a configuration that could successfully replicate all reported main and secondary results  across all patient subgroups, while satisfying demanding privacy metrics like the **Hidden Rate (HR)**, a measure of how well individual patients are protected against re-identification attacks.

### **Replicating Results, Guaranteeing Anonymity**

The results were a **game changer** for clinical research:

- **Satisfactory Privacy:** The selected datasets achieved Hidden Rates (HR) of **85.0%** and **93.2%**, meaning an attacker would likely fail if they tried to confirm a patient's membership in the trial data. This explicit privacy assessment allows the synthetic datasets to be legally qualified as non-personal data, effectively meeting GDPR restrictions for data sharing.
- **High Utility:** The optimization process successfully yielded synthetic datasets that **replicated all efficacy endpoints** (both primary and secondary results) for the placebo and approved treatment arms of the trials. This included complex post-hoc subgroup analyses and safety outcomes, demonstrating that the synthetic data acts as an accurate proxy for the original information.

This study proved that while a trade-off exists between privacy and utility, optimization allows researchers to select datasets that meet both ethical and analytical requirements. Generating synthetic data is not just about reusing data; it's about secondary uses of data contributing to the data value chain for innovation, a hallmark of 21st century healthcare.

To demonstrate the full potential of this method to unlock health data sharing for the global community, the researchers took it a step further: they released the **placebo arms of both synthetic datasets as open-access resources**. This action allows any researcher, without complex credentialing or restrictive analysis plans, to use high-quality clinical trial data for feasibility studies, sample size estimation, or predictive model development.

Don't just take their word for it: [**read the paper**](https://www.jmir.org/2025/1/e71297/)to discover how synthetic data can safely accelerate personalized medicine. [**Watch the video**](https://www.youtube.com/watch?v=Ng54PAYi5Wk) to hear Dr. Pierre-Antoine Gourraud discuss how these cutting-edge methods are shaping the future of health data sharing.

 

 

Subscribe Now

![Reviewed by Kayleigh-Ann Clegg, PhD](https://blog.jmir.org/hs-fs/hubfs/Headshot.jpg?width=150&name=Headshot.jpg)

#### Reviewed by Kayleigh-Ann Clegg, PhD

[Follow me on my website](https://www.researchgate.net/profile/Kayleigh-Ann-Clegg) [Follow me on LinkedIn](https://www.linkedin.com/in/kayleighclegg/)

Scientific News Editor @ JMIR Publications with a diverse background in clinical psychology, academic research, and digital health program development. I'm passionate about knowledge translation and accessibility, as well as data- and values-driven innovation. My goal is to make vital scientific knowledge clear and accessible to everyone so that we can better our lives and the world around us.

### Leave a Comment

## Related Articles

[![Beyond Shiny Toys: Building Mature and Equitable AI Triage in Primary Care](https://blog.jmir.org/hubfs/StoryTap%20Blog%20-%20Siaw-Teng%20Liaw.png)](https://blog.jmir.org/beyond-shiny-toys-building-mature-and-equitable-ai-triage-in-primary-care)

[Research News](https://blog.jmir.org/tag/research-news) [Artificial Intelligence](https://blog.jmir.org/tag/artificial-intelligence)

### [Beyond Shiny Toys: Building Mature and Equitable AI Triage in Primary Care](https://blog.jmir.org/beyond-shiny-toys-building-mature-and-equitable-ai-triage-in-primary-care)

Artificial intelligence is fast becoming the digital front door of modern medicine. In the face of overwhelming patient demand, healthcare systems are rapidly deploying...

![Reviewed by Kayleigh-Ann Clegg, PhD](https://blog.jmir.org/hs-fs/hubfs/Headshot.jpg?width=45&name=Headshot.jpg) 

[Reviewed by Kayleigh-Ann Clegg, PhD](https://blog.jmir.org/author/reviewed-by-kayleigh-ann-clegg-phd) 

[Read More](https://blog.jmir.org/beyond-shiny-toys-building-mature-and-equitable-ai-triage-in-primary-care)

[![Decoding Informed Consent: Evaluating AI Disclosures in Clinical Research](https://blog.jmir.org/hubfs/StoryTap%20Blog%20-%20Hankun%20Su.png)](https://blog.jmir.org/decoding-informed-consent-evaluating-ai-disclosures-in-clinical-research)

[Research News](https://blog.jmir.org/tag/research-news) [Artificial Intelligence](https://blog.jmir.org/tag/artificial-intelligence)

### [Decoding Informed Consent: Evaluating AI Disclosures in Clinical Research](https://blog.jmir.org/decoding-informed-consent-evaluating-ai-disclosures-in-clinical-research)

Artificial intelligence is rapidly shifting from experimental code to point-of-care medical decision-making. As machine learning models gain autonomy in diagnosing...

![Liana Ramos, Marketing Associate](https://blog.jmir.org/hs-fs/hubfs/Liana%20Ramos.jpeg?width=45&name=Liana%20Ramos.jpeg) 

[Liana Ramos, Marketing Associate](https://blog.jmir.org/author/liana-ramos) 

[Read More](https://blog.jmir.org/decoding-informed-consent-evaluating-ai-disclosures-in-clinical-research)

#### About

- Menu Item One
- Menu Item Two
- Menu Item Three

#### Services

- Menu Item One
- Menu Item Two
- Menu Item Three

#### News

- Menu Item One
- Menu Item Two
- Menu Item Three

Follow us on Facebook Follow us on LinkedIn Follow us on Twitter Follow us on Instagram

```json
{
  "@context" : "https://schema.org",
  "@type" : "BlogPosting",
  "author" : {
    "@type" : "Person",
    "name" : "Reviewed by Kayleigh-Ann Clegg, PhD",
    "url" : "https://blog.jmir.org/author/reviewed-by-kayleigh-ann-clegg-phd"
  },
  "dateModified" : "2026-09-17T18:23:15.274Z",
  "datePublished" : "2025-12-17T22:13:56.000Z",
  "headline" : "Synthetic Data: Solving the Clinical Trial Privacy Paradox",
  "image" : [ "https://blog.jmir.org/hubfs/Pierre-Antoine%20Gourraud%20%20blog.png" ],
  "mainEntityOfPage" : {
    "@id" : "https://blog.jmir.org/synthetic-data-solving-the-clinical-trial-privacy-paradox",
    "@type" : "WebPage"
  },
  "publisher" : {
    "@type" : "Organization",
    "logo" : {
      "@type" : "ImageObject",
      "url" : "https://blog.jmir.org/hubfs/JMIR-Publications_logo_2020_original.png"
    },
    "name" : "JMIR Publications"
  }
}
```