Data Standards, Ontologies, and Compliance
Two things determine whether a field value will validate: whether it conforms to a published data standard ImmPort templates were built from, and whether it's checked at parse time against a controlled vocabulary (lk_* lookup table). This page covers both, plus the HIPC guidelines and Subject De-Identification policy — it doesn't restate every value list.
Published standards behind the templates
| Standard | Covers | Source |
|---|---|---|
| HIPC Bioinformatics Standards Working Group guidelines | Flow Cytometry / Mass Cytometry (CyTOF), Metabolomics & Proteomics Mass Spectrometry, Transcriptional Profiling, Subject Immune Exposure | See HIPC guideline-to-template mapping below |
| NIH-endorsed Common Data Elements (CDEs) | Many non-assay templates (study/subject/clinical fields) | Common Data Elements — template-by-template CDE mapping. Mapping is done by ImmPort at ingest; no extra action is needed on submission, it's informational only |
When a template and a published standard appear to disagree
The template's own field constraints (required/conditional, controlled vocabulary, length) always win for validation purposes — the published standard explains why a field exists, not what will pass the validator.
HIPC Bioinformatics Standards Working Group Guidelines
ImmPort works with the HIPC Bioinformatics Standards Working Group to develop guidelines for metadata and results.
Controlled vocabulary mechanism (how validation actually checks values)
Any template column marked CONTROLLED VOCABULARY in its requirements (see How to Read the Template Documentation) is checked against an lk_* lookup table at parse time:
- The check is case-insensitive —
"Male","male", and"MALE"all pass ifMaleis a validlk_human_sexentry. - On success, the value is rewritten to the lookup table's canonical case spelling — don't assume the byte-for-byte value you submitted is what ends up stored/returned.
- On failure, validation fails with an error referencing the field and its lookup table.
- Some columns additionally have a preferred vocabulary — a secondary, non-blocking normalization layer (e.g. population marker names) separate from the required-match controlled vocabulary check.
Where to browse the actual value lists
- Ontologies overview — a searchable table of every
lk_*table, the ontology it's sourced from, and which ImmPort table/column references it. - Lookup Tables under Template Documentation — per-table browsable term lists (the same tables the controlled-vocabulary check above validates against).
Don't hardcode a vocabulary snapshot
Lookup tables are maintained independently of template schema versions and can gain new terms between releases. If you cache a vocabulary list locally, re-fetch it periodically (or re-check via Ontologies) rather than assuming a list captured once is exhaustive.
Subject De-Identification
The ImmPort team follows the HHS Safe Harbor Method in data capture templates and in data submission by:
- Rounding any “ages” submitted > 89 to 90 and add a comment indicating that this was done
- Not recording any dates within ImmPort templates; all time points and schedules are referenced to study day 0 (or a designated day in the study as determined by the data provider)
- Not recording any geographic location subdivisions information below the state level in ImmPort templates
- Not capturing any subject identifying information in ImmPort templates itemized in the Safe Harbor Method (Name, accounts, phone or fax numbers, email addresses, etc)
The ImmPort team removes the connection to subject identifiers in the source data management system by:
- If possible, discuss with the data providers to determine that local subject identifiers that were used in the submission are deidentified
- Not displaying subject source identifiers, sample source identifiers, or experiment sample source identifiers in the ImmPort application to cut linkage with the source system
- Not exporting the subject source identifiers, sample source identifiers, or experiment sample source identifiers in any ImmPort generated data packages delivered online, at sFTP sites, or in the future, via Amazon AWS S3 storage
- Not providing original, unstructured result files submitted by sites if they (1) haven’t stated concurrence with the disconnect with source identifiers; and/or (2) if the data can be parsed into standard ImmPort templates and source identifiers hidden
The ImmPort team does not normally accept or share GWAS level SNP human genotyping data nor Next Gen human genome sequencing data, and recommends depositing such data in dbGAP or SRA.
For data provided in more “native” format (i.e., not in ImmPort template format), such as result files from closed clinical trials (assessments, lab tests, adverse events, etc):
- If the data source is ITN TrialShare, the dates are already obfuscated and are not identifying as are the subject identifiers; no special treatment is required
- For other data sources with closed clinical trial data (or the equivalent), data files are reviewed by the ImmPort Curation team who will 1) Remove identifying information such as names, dates, locations, etc...2) Normalize dates to study day 0; 3) Comment and Description fields are removed unless determined to be completely devoid of identifying information
Concrete Constraints When Filling Out Templates
- No calendar dates. All timepoints/schedules are relative to study day 0 (or a study-defined reference day) — never submit an absolute date in a template field.
- No subject-identifying fields. Don't populate name, phone, fax, email, account numbers, or any other Safe Harbor identifier — these have no corresponding template columns, but free-text
description/commentfields must not be used to smuggle them in. - No geographic detail below state level.
subject_location-type fields should not contain city/county/zip. - Ages > 89 must be rounded down to 90. This is enforced by the Batch Uploader itself, not just policy — see below.
- Local/source subject identifiers must not be exposed.
subject_id(the user-defined ID) must not be a value traceable back to the source system's own subject identifier for human subjects.
The Age-Rounding Rule, as Actually Implemented
Transcribed from the subjectHumans/subjectAnimals template requirements (PreProcesses, set-to-value):
If
age_unit=years(case-insensitive) andmin_subject_age(ormax_subject_age) > 89, the Batch Uploader sets that age to90and appends;age rounded down to 90 by ImmPort Upload toolto the subject'sdescriptionfield.
Practical guidance
Don't pre-round ages yourself and don't rely on client-side logic to detect ages > 89 — submit the true value in years and let the uploader's post-process do the rounding and annotate description. If you submit an age already capped at 90 you'll lose the distinction between "actually 90" and "was >89, rounded" that the appended comment preserves.