Skip to content

Validation Rule & Process Type Glossary

Every ImmPort template's validation behavior is driven by a small set of reusable rule types (pass/fail checks) and process types (data transformations) defined once in XmlUtil.java and implemented in Rules.java / Processes.java (in the immport-data-upload-service repository). These names appear throughout the generated requirements documents and the template XML, but are not explained in plain English anywhere else — this page exists to close that gap.

Confidence level

Each entry below was verified by reading the actual Java implementation this session. Where noted, a rule's effective behavior is more specific/subtle than its name implies — read the note, not just the name.

Rule types (pass/fail checks, run as PreRules/PostRules)

Rule type What it actually checks
all-empty-or-nonempty Across a set of columns, either all must be empty or all must be non-empty — a "fill in this whole group together or not at all" constraint.
at-least-one-nonempty At least one column in the set must be non-empty (e.g. amount or duration or temperature for a treatment).
empty-value Fails if any listed column has a value — i.e. these columns must be blank (often paired with a condition, e.g. "blank unless X").
not-empty-value Fails if any listed column is blank — a straightforward required-field check.
not-equal-to-value The column's value must not equal one of the given values.
values-equal The two (or more) columns must contain equal values. Used for things like "this computed status must equal this constant" — e.g. subjectHumans enforces subject_required == "yes", which is the actual mechanism forcing every Subject ID in that template to be new rather than a reference to an existing subject.
values-not-equal Inverse of values-equal.
less-than Numeric column must be less than a given value/column.
check-date-format Validates DD-MMM-YYYY / DD-MMM-YY date format (rejects ISO/numeric dates).
check-not-contains-case-insensitive Two columns (or a column and a list) must not match case-insensitively — typically a duplicate/self-reference guard.
check-value-in-entity The column's value must be found in a referenced set — either loaded from the database via a named Validation Query, or accumulated in-memory from earlier rows in the same upload package. This is the general "must already exist" / "must be consistent with" check (e.g. a subject referenced in a result row must actually belong to the stated study).
check-value-not-in-entity Inverse of the above — fails if the value is found. Used e.g. in study_design_edit's arm_2_subject section to block submitting the identical subject+study pair twice within the same package (it is not a general cross-study duplicate blocker — see Subject Sharing Across Studies for the full story and why this doesn't prevent legitimate cross-study subject reuse).
check-result-schema-for-result-file / check-result-schema-for-accession Cross-checks that a result file's declared schema/accession is consistent with what's already registered for that result set.
check-population-markers Assay-specific (flow cytometry/CyTOF-style) check that reported cell-population marker expressions are well-formed/consistent.
check-immport-template Validates that an uploaded file actually matches the expected ImmPort template structure.

Process types (data transformations, run as PreProcessing/PostProcessing)

Process type What it does
determine-combined-status Used by combined templates (e.g. subjectHumans) to decide, per row, whether an ID column refers to a new record or an existing one — checks the current upload package first, then the database. Its output (..._required = yes/no) is what other rules key off of.
set-and-check-study-accession Pins the single study_accession for an entire upload batch the first time it's seen; fails if a later row in the same batch tries to use a different study. (A batch upload can only ever target one study.)
add-values-to-entity Appends a value to an in-memory, package-scoped lookup table (EntityToValuesManager) — e.g. recording "this subject has now been assigned to this study" for later duplicate checks within the same package only. Not a database write.
determine-user-defined-id-to-parent Resolves a child record's user-defined ID scoped under its parent entity (used for things like study_categorization, which is only unique within a given study).
set-preferred-value Implements the "preferred vocabulary" pattern: if a reported free-text value matches an entry in a *_pref_map table, fills in the companion preferred-value column automatically.
initialize-columns Sets a group of columns to a default (usually empty) state under given conditions — e.g. clearing demographic fields when a row is just referencing an already-defined subject.
set-to-value Unconditionally (or conditionally) sets a column to a fixed value — e.g. capping reported ages over 89 down to 90 for de-identification, with an explanatory note appended to the description field.

How to use this page

When a template's own reference page says something like "Conditionally required... when defining a new X rather than referencing an existing one", that phrasing is almost always backed by a determine-combined-status process plus a values-equal (or absence thereof) rule. If the behavior described doesn't match what you're seeing, check that template's XML for its actual PostRules/ConstantColumns rather than trusting the prose alone — see Subject Sharing Across Studies for a worked example of exactly this trap.