In collaboration with Leiden University Medical Center

The data backbone for one of the world's largest longitudinal Parkinson's studies.

The ProPark Data Platform is secure, consent-governed infrastructure that turns the multimodal data of the ProPark study — clinical examinations, questionnaires, medication records, and continuous wearable sensor recordings — into analysis-ready material for Parkinson's disease research.

Get in touch

By the numbers: 1 trillion+ individual data points held on the platform · 2,000+ participant-weeks of continuous sensor data · 100 Hz wrist and lower-back motion recording, at home, for week-long periods · All primary storage and processing resides in the European Union

A study that observes Parkinson's where it happens — daily life

ProPark (Profiling Parkinson's Disease) is a multi-year observational study led by Leiden University Medical Center with participating sites across the Netherlands, following approximately 900 people with Parkinson's disease alongside healthy controls. Beyond clinic-based motor examinations and questionnaires, participants wear motion sensors on the wrist and lower back during repeated week-long home assessments — capturing how the disease actually presents across medication states and ordinary daily activity, rather than only as it appears in a brief clinic visit.

Data of this scale and variety creates a corresponding infrastructure challenge: trillions of sensor measurements, heterogeneous clinical records, and the practical messiness that is normal in real clinical research. Since 2023, OccamzRazor has served as the study's data engineering partner — designing, building, and operating the platform that receives this data, organizes and cleans it, protects it, and gives researchers a controlled way to analyze it.

What the platform does

Data moves through five governed stages, from raw study files to research-ready analysis.

  1. Ingestion of raw study data. Sensor files, clinical exports, and medication forms are received into dedicated secure storage configured so that uploaded files cannot be altered or deleted after submission — preserving an unmodified record of what was received.
  2. Processing in a controlled compute environment. Incoming data is validated, cleaned, and standardized. Wrist and lower-back recordings are time-synchronized to a common clock, and validated algorithms extract clinically meaningful measures — tremor, bradykinesia, gait, and overall activity — with thousands of processing jobs running in parallel.
  3. Organization into a queryable data store. Processed data is organized into linked sensor, feature, and clinical tables, so that measures collected for the same participant at the same point in the study can be brought together into a single analysis dataset.
  4. Controlled access for researchers. Researchers query the data on demand in standard SQL — retrieving exactly the participants and time periods relevant to an analysis, without operating any infrastructure of their own. Analytical queries that once took days or weeks of dataset preparation now run in seconds.
  5. Human review inside the governed environment. Where analysis requires expert judgment — such as labeling motor tasks within a recording — data segments are routed through a structured, secure annotation workflow, and the resulting labels are returned to the database without ever leaving the governed environment.

Governance is enforced by the system, not by policy documents

Research infrastructure for patient data is only as trustworthy as its weakest control. The platform was built from the ground up so that its central commitments do not depend on any individual applying the rules correctly.

Access follows participant consent. The platform maintains a record of each participant's data-sharing permissions. Researchers are assigned to defined categories and query controlled views that return only the participants whose consent permits that category of use — enforced automatically by the platform itself.

European data residency. All primary storage and processing take place within a European Union cloud region, consistent with the GDPR and with the privacy commitments made to study participants.

Encryption throughout. Data is encrypted at rest and in transit. Ingestion pathways are versioned and deliberately constrained, and administrative access runs through controlled, recordable channels.

Reproducibility and audit. Every transformation applied to a clinical value is recorded in a transformation log, and access to data is logged. Any result can be traced back to the originally recorded value, and any access can be reviewed.

From bottleneck to instrument

Before the platform, the path from collected data to analysis-ready dataset was the limiting step of the research. Treating data management as a formal part of the scientific process changed the pace of what the study can ask.

  • 12 weeks → under 30 minutes. Cohort-wide tremor feature extraction, previously an ad-hoc workflow across local machines, now runs in the platform's parallel compute environment in a fraction of the time.
  • Days → seconds. Creating a study dataset once took days to weeks of manual assembly. Researchers now pose analytical queries directly against the governed data store and receive answers interactively.
  • AI-ready by design. Because the data is clean, documented, and governed, AI-assisted analysis can operate directly on it. In a live demonstration for Dutch research funders, a published peer-reviewed analysis from the study was reproduced end-to-end — figures included — through plain-language interaction with the data.

Evidence, not promises

The platform's purpose is published science. Its data engineering underpins peer-reviewed research from the ProPark study, and its outputs are built to be traceable from figure back to raw signal.

npj Parkinson's Disease · 2025 — Added value of a wrist-worn device for assessing tremor in Parkinson's disease: reliability and validity of tremor evaluation at home. In 219 ProPark participants, wearable-derived tremor measures showed excellent test–retest reliability and agreed with clinical assessment — and detected tremor that patients reported but brief clinic exams missed. The platform ingested and processed the raw sensor data behind the study and integrated it with clinical data to enable the analyses.

Wearable sensors don't need to be the final biomarker. They are the instrument that finds the patient subgroups — so that better molecular signatures, better trial designs, and better treatment decisions can follow.

Why we do this

OccamzRazor is a digital biotech company focused on complex diseases of brain aging, starting with Parkinson's disease. Parkinson's is a spectrum of conditions, not a single disease — and progress toward better treatment depends on measuring it precisely enough to tell its subtypes apart.

Our work with LUMC is the working model of that conviction: govern the data properly, engineer it to research grade, and keep it continuously ready — so that as analytical methods advance, the science is never waiting on the infrastructure. The capability we build alongside our academic partners is theirs to use, on data that never has to leave its governed home.


OccamzRazor · Johnson & Johnson Innovation — JLABS · 101 6th Avenue, New York, NY 10013 · info@occamzrazor.com · www.occamzrazor.com

The ProPark Study is led by Leiden University Medical Center · www.proparkinson.nl

This page describes research infrastructure developed and operated by OccamzRazor under its collaboration with Leiden University Medical Center for the ProPark study. The platform itself is not publicly accessible; access to study data is governed by participant consent and the study's data-sharing agreements. © 2026 OccamzRazor. All rights reserved.