The data, and the rules that govern them
A measurement without provenance and without uncertainty is not a measurement. This is the proposed technical setup.
What is recorded with every value
Every number kept carries: the instrument and its calibration state, the version of the algorithm that produced it, the estimated uncertainty, and the chain of provenance back to the raw signal.
Raw data are never overwritten. Missing data are flagged; they are not filled in by interpolation at the raw level. Reconstructions live only in derived layers, marked as such.
Standards instead of inventions
No proprietary formats where shared ones exist. The proposal is to build on what the community already uses, so the data stay readable without us: file structures already adopted in research, open formats for signals, a time-series store for high frequencies, and a common data model for the clinical part.
For collection there are open, maintained platforms built precisely for multi-sensor studies. Reusing them makes more sense than rebuilding them.
The acquisition layer — nodes, synchronisation, message format — is dealt with separately on the Instruments page.
Before the terabytes, the catalogue
Millions of measurements do not automatically contain the synchronised stimuli needed to compute a response and a recovery. A minute-level average heart rate, for instance, does not allow inter-beat intervals to be reconstructed.
So the first archive to build is not an archive of signals but a catalogue of available evidence: for each study, which signals, at what resolution, with which events recorded, which outcomes, which calibrations, and under what access conditions. From there a dataset is chosen to suit a question, instead of accumulating in hope.
Access to existing collections
The large public cohorts with wearable data cannot be downloaded at will. They require a titular institution, formal authorisation and, in some cases, analysis inside a controlled environment. One of them currently has applications suspended.
This constraint has to be faced at the start, because it determines who can do what: it is not an administrative detail.
People
Health data, therefore: an impact assessment before starting, pseudonymisation with separated keys, storage on controlled infrastructure, and a consent process that explains in plain language what is measured and what it is for.
Anonymisation alone is not realistic for continuous physiological recordings: they must be treated as identifiable.