The pharmaceutical ecosystem is seeing rapid growth of open-source R tools and AI-enabled applications across clinical development, analysis, and regulatory submissions.
However, objective benchmarking and evaluation of these tools is currently constrained by the lack of high-quality, publicly available clinical trial datasets.
Existing public datasets, such as CDISC Pilot 1, have important limitations:
- They are outdated and not well aligned with current CDISC standards or current industry practice
- They are limited in scale and complexity
- They are insufficient for evaluating modern workflows, including AI-assisted analysis, automation, and end-to-end submissions
As a result, there is a clear need for modern, realistic, and reusable synthetic clinical trial data that can support:
- Tool demonstration
- Method evaluation
- Community development and education
The initial scope of this project was to build benchmark datasets and evaluation test cases for open-source tools relevant to clinical data science.
The goal is to create a shared and practical foundation for assessing how well open-source tools perform on realistic pharma use cases, especially in settings where modern workflows require more representative and reusable benchmark data.
The scope has now been expanded to include a second workstream: benchmark test cases for pharma open-source skills.
These skills are more focused, task-oriented, and reusable capabilities that support specific workflows in clinical data science. The first skill in this effort is group sequential design (GSD).
This expansion is being developed in collaboration with BBSW, which is supporting the work by sponsoring shared tokens for the automated evaluation pipeline.
The project now includes two complementary parts:
- Part A: Benchmark datasets and test cases for evaluating open-source tools
- Part B: Benchmark test cases for evaluating pharma open-source skills, starting with GSD
Together, these efforts aim to build a stronger public foundation for rigorous, scalable, and transparent evaluation of open-source capabilities in clinical data science.
Pilot 7 holds weekly standups three times a month on Fridays from 8-9 AM PST. We also host monthly Submissions Working Group meetings with FDA staff, bringing together participants across different pilot subgroups.
Everyone is welcome to join. To access our calendar and join the Slack workspace, please see here
To learn more about the R consortium Submissions Working Group, visit here
All data under .archieve/CWMM-LAA1/ and kn564/ was generated using the clinical-trial-ipd-sim skill, based on the original protocol.
Each study lives in its own top-level folder, organized by pipeline stage. The data flow is ODM -> SDTM -> ADaM -> TLF: an ODM export of the collected CRF data is mapped to SDTM tabulation datasets, those are derived into ADaM analysis datasets, and the ADaM feed the tables, listings, and figures.
<study>/ a study name
+-- data/
| +-- odm/ source ODM v2.0 export
| +-- sdtm/ SDTM tabulation datasets
| +-- adam/ ADaM analysis datasets
+-- program/
| +-- odm/ ODM XML -> per-form CRF data
| +-- sdtm/ CRF data -> SDTM
| +-- adam/ SDTM -> ADaM
+-- spec/ dataset specifications and standards metadata
| +-- define/ define.xml files, when present
| +-- yamaa/ yamaa derivation specifications, when present
+-- tlf/ tables, listings, and figures
+-- log/ execution logs, one per program
Stage folders that are not yet populated are tracked with an empty .gitkeep,
without which git would not carry the folder at all.
This layout is enforced in CI by the Structure Check workflow. The template
it checks against lives in .github/repo-structure.json, so adding a stage
folder to every study is a one-line edit there, and adding a new study needs
no change at all -- any top-level folder is discovered and checked
automatically. Run it locally with:
python3 .github/scripts/check_structure.py| Folder | Contents |
|---|---|
.github/ |
CI workflows (CDISC CORE validation, structure check) and the folder-structure manifest |
.automation/ |
Scheduled Claude Code routine prompts and bootstrap scripts |
.archieve/ |
Earlier studies, retained for reference |
- Original sdtm: json version
- Original xpt versions: see pilot 5 repo
- CSR: https://github.com/cdisc-org/sdtm-adam-pilot-project/blob/master/updated-pilot-submission-package/900172/m5/53-clin-stud-rep/535-rep-effic-safety-stud/5351-stud-rep-contr/cdiscpilot01/cdiscpilot01.pdf