# Application and future research data

Worldview Sorter is an application that gives source-linked, authored interpretations of answers to philosophical questions. The frozen pilot is a content candidate, not a validated psychological scale. The application does not currently run a research study. Participation in future research reuse is optional and does not change quiz or result access.

## Data boundary

| Layer | Current behavior | Access and use |
| --- | --- | --- |
| Operational application | The public quiz saves its raw session and pinned release references in the user's browser. Results are rebuilt from those answers and versioned rules; a user may separately download a summary or complete backup. No account or automatic answer submission exists. | Browser owner; browser storage may be cleared. The optional legacy development collector is off by default and must not be used as a production research store. |
| Product analytics | No per-user telemetry or analytics collector is currently enabled. Server health is operational, not a research measure. | Future aggregate instrumentation needs its own purpose, retention, and review. |
| Research contribution | After results, an explicit checkbox and button can send one pilot attempt to a separate private store when the server operator enables this feature. | Pseudonymized raw answer rows may later enter a reviewed research export. Consent is recorded per attempt. |

The contribution record contains exact assigned item IDs and revisions, raw responses, response state, item order, routing state, answer change counts, completion state, contribution date, and bank, instrument, form, model, result, derived-rule, affinity, and available localization versions. It does not contain the application session ID, respondent key, email, demographics, IP address, precise response times, or per-answer timestamps. The server keeps a one-way session fingerprint only for duplicate prevention; it is never exported. The submission limiter temporarily holds connection IPs in process memory for a ten-minute abuse-control window, without writing them to the research store or operational event log. Ordinary hosting or proxy request logs are outside this research store and require separate operator governance.

## Consent, withdrawal, and linkage

The contribution API is disabled unless `WORLDVIEW_ENABLE_RESEARCH_CONTRIBUTIONS=true`. The user must affirm `research-consent-1.0.0` for each attempt. The exact terms are served from a [versioned, hash-pinned artifact](../data/research/consent-v1.json) and included in the export. The endpoint rejects extra identity fields and non-pilot packets. The public client creates and saves a private receipt before submission; if the response is lost, the user can retry with the same receipt without creating another contribution. If browser storage cannot retain the receipt, the client does not submit. The token is shown in a downloadable receipt, and anyone holding it can withdraw the contribution. Withdrawal replaces the server record with a minimal tombstone, removing the response rows from later exports. A completed export or copy already shared externally cannot be recalled automatically. No external sharing occurs through this application code; any release needs separate privacy, governance, and researcher access review.

The current private store has **no automatic expiry**. Active contributions remain until withdrawal or an operator-run retention action; minimal withdrawal tombstones and export audit entries also remain until operator retention action. Before enabling real collection, the operator must set a documented retention schedule, backup deletion procedure, and a way to honor withdrawal in held packages. The application has no built-in account recovery for a lost withdrawal receipt.

Longitudinal linkage is off by default. A separate checkbox permits a browser-held random link secret to produce the same research pseudonym across voluntary attempts. The link secret is not transmitted to the exported dataset; clearing browser storage ends future linkage. Export includes the UTC contribution date, not session timing, so same-day order and precise intervals are unavailable. A stable pseudonym is still potentially linkable through answers or external information. No demographic questions are collected: no current analysis purpose justifies them, and they are not used for philosophical classification.

## Instrument and inference provenance

The active candidate is `pilot-candidate-1.1.0`: the same 238 fixed, revision-pinned questions in `worldview-pilot-1.0.0`, from bank `0.9.0`, with the SO09 interpretation retired in model release 1.2.0. The frozen route has branching; one response row is exported for **every assigned route position**, including unshown and unreached items. The original session in the browser also retains presentation timestamps and response timings, but these are deliberately excluded from the research contribution. `not_reached` with `completionStatus=in_progress` identifies partial administration; `branch_not_shown` identifies routing absence. Invalid submissions are rejected, not turned into an `invalid` response category.

Questions are original authored items built from the source-linked concept inventory and content review, with IDs and revisions preserved rather than rewritten in place. The construct registry is an authored ontology. The model contains exact proposition and evidence rules, including direct, derived, research-only, and unmeasured distinctions. Its conclusions are **not** ground-truth labels for validating those rules. The catalog compares interpreted propositions to source-backed traditions; affinity is comparison, not identity. Historical results require the matching versioned bank, route, model, semantics, and catalog. Reinterpretation under a later model must be a separate analysis of the unchanged raw answers. [Pilot limitations](PILOT_V1.md), [construct audit](UNMAPPED_AUDIT.md), [affinity contract](AFFINITY_CATALOG_V1.md), and the package snapshots explain these boundaries.

## Reproducible private package

On the server, an authorized operator runs:

```bash
node scripts/export-research-package.mjs --store /private/contributions --out /private/new-package
```

The contribution store must already exist and be private; a missing path fails instead of being created as an apparently empty dataset. The output path must be new and outside the repository and contribution store. Optional paired `--from YYYY-MM-DD --through YYYY-MM-DD` arguments select an inclusive UTC consented-contribution window; the manifest records it. Without them, the command checks every active contribution. It checks pinned pilot/catalog/consent bytes and each **selected** active contribution's version tuple, excludes tombstones, writes a file-hash manifest, and records an export audit in the private store. Invalid consent timestamps fail closed even when a window was requested. It produces `respondents.ndjson`, `administrations.ndjson`, `responses.ndjson`, `items.ndjson`, `versions.json`, a [field dictionary](RESEARCH_DATA_DICTIONARY.md), a JSON schema, a loading example, and snapshots of the consent terms, full item bank, registry, interpretation model, source ledger, affinity catalog, instrument, pilot manifest, form policy, response scales, and localization catalogs/bundles. The directory and files are created with restricted permissions where supported. This is an internal handoff mechanism, **not** permission to publish or share data. An empty package can be generated from an existing empty store or a genuinely empty selected window.

Researchers can join response rows to administrations by `researchAdministrationId`, to respondents by `researchRespondentId`, and to canonical wording by `(position, itemId, itemRevision)`. Newer sessions also pin `textVersion`, `variantId`, and a localization bundle; pre-localization sessions explicitly report `historical_canonical_unpinned`. Presentation locale is not nationality or ethnicity, and no cross-language measurement equivalence is implied. Newer administrations also pin `modelReleaseVersion`, identifying the complete immutable model configuration included in the package. Older contributions retain their explicit component-version tuple with a null model-release pointer. `response-scales.json` defines numeric categories; `items.ndjson` defines item-specific choices. The [dictionary](RESEARCH_DATA_DICTIONARY.md) specifies every tabular and metadata field. A future package schema/version change requires a new exporter and retained old package. Source snapshots identify academic basis but do not establish content or psychometric validity.

## Sampling and potential analyses

Opt-in application users are self-selected and cannot be treated as representative of any population. Route design, UI wording, voluntary consent, item comprehension, and repeated-use selection can affect the data. No response collection, live deployment, external data release, or empirical validation is part of this repository change. A later approved analysis could examine response-category use, missingness, branch effects, repeated administrations with linkage, item pairs, revision effects, and disagreement between authored propositions and raw-answer patterns. These are research questions, not automatic item edits or scientific findings.

## Versioned independent-research handoff

An authorized operator can inspect an existing private contribution store **without printing raw answers**:

```bash
node scripts/research-inventory.mjs --store /private/contributions --out /private/research-inventory.json
```

Use optional paired `--from YYYY-MM-DD --through YYYY-MM-DD` arguments to inspect the same inclusive UTC contribution-date window planned for a snapshot. The default console output contains only aggregate counts. The optional detailed report includes per-item category cells and requires a private output directory. It reports consented active administrations in the selected window, whole-store withdrawal tombstones, completion, version/locale and response-format distributions, missingness, and linked repeats; it does not supply an unobserved nonconsent denominator, incident history, or a power analysis. A missing or inaccessible store is **unknown production data**, not zero administrations.

After quiescing writes and reviewing known test, automation, duplicate, and incident records, create a new restricted snapshot:

```bash
node scripts/create-research-snapshot.mjs --store /private/contributions \
  --out /private/releases/wvs-research-2026-09-29-v1 \
  --id wvs-research-2026-09-29-v1 --from 2026-01-01 --through 2026-09-29 \
  --quiesced --exclusions /private/reviewed-exclusions.json
```

The dates above are **examples**, not a declared production window. The optional exclusion file is private mode-0600 JSON, e.g. `[{"researchAdministrationId":"a-…","reason":"known_test"}]`; reasons are `known_test`, `known_automation`, `known_incident`, or `known_duplicate`. It must list actual in-window records. Unusual philosophical answers are not an exclusion reason. Snapshot output must have a new directory whose basename equals its ID, outside the repository and contribution store. The exporter checks source files before and after extraction, then rekeys respondent and administration pseudonyms for this release and verifies all file hashes and authored replay. Concurrent writes still require the operator to quiesce the store; `--quiesced` is an attestation, not a software lock.

For a correction, create a new `-vN` ID and pass `--supersedes OLD_ID --correction-note "reason"`; retain the previous package and publish the relationship in the release register. Publication also writes a restricted `release-linkage-<snapshot ID>.json` to the private contribution store. This operator-only file joins withdrawal receipt IDs to that snapshot's rekeyed administration IDs and pins the manifest hash. It is backed up with the store and never enters the researcher package. After a withdrawal, run `npm run research:withdrawal-impact -- --store /private/contributions --contribution UUID` on the private host, notify recipients of affected releases, and issue a corrected snapshot; previously distributed copies cannot be retracted automatically. Protect and retain this linkage for as long as release copies may need withdrawal notices. Older releases without linkage still require manual release-register review.

The snapshot separates `responses.ndjson` (observed answers) from `derived.ndjson` (reproducible authored interpretations) and `quality-flags.ndjson` (mechanical metadata flags, never a sincerity judgment). It includes item and authored-claim codebooks, aggregate diagnostics, structural readiness classifications, frozen source/model/localization artifacts, exact engine source for replay, Python and R loaders, an integrity checker, researcher documentation, and a hash manifest. The authored-claim index exposes exact answer mappings, route inclusion, source-link status, and affinity uses without presenting legacy scope or topic citations as reviewed propositions. Corrections require a new ID; do not edit a cited release. A snapshot is restricted by default even after rekeying. The [researcher README](../research/RESEARCHER_README.md), [analysis possibilities](../research/ANALYSIS_POSSIBILITIES.md), [warnings](../research/ANALYSIS_WARNINGS.md), and [snapshot dictionary](../research/SNAPSHOT_DICTIONARY.md) are copied into each release. [Handoff status and operator gates](RESEARCH_HANDOFF.md) distinguish what this repository can establish from facts requiring access to the production store.
