DATACORE GUIDE
From a research table to evidence.
Use your own eligible nonconfidential research tables, or start with the clearly synthetic examples below. Personal or confidential information is not permitted.
- Request access and sign in with your individual pilot account. There is no public registration.
- Create a project with a clear research question. Select Synthetic demonstration for artificial data, or Nonconfidential research for your own eligible table and record your permission or public-data license reference.
- Download the approved CSV cohort or approved XLSX cohort. Alternatively upload your own CSV/XLSX/TSV/TXT/Parquet or flat JSON/JSONL file within the limits. Nested JSON is rejected; no automatic flattening occurs.
- Open Datasets, choose the file, declare interpretation and column types/roles, then Validate & import. TXT requires confirmed delimiter and encoding. Inspect missingness, duplicates, categories, outlier suggestions and identity roles before proceeding.
- Analyze the version. The default grouped logistic baseline uses outcome, age_years, marker and subject_key. Settings and partitions are preserved.
- Choose a guided analysis/model and review its settings. To clean data, append explicit preparation steps, preview their changes, then confirm an immutable derived version. Originals remain preserved. Learned model preparation fits training partitions only. Inspect recorded assumptions, exclusions, diagnostics and warnings; compare compatible frozen evaluations or export HTML/DOCX/JSON and an evidence ZIP.
- Register a model, record research review and acknowledge its limitations before project-private publication. For the default two-feature model, use the approved prediction JSON in the input form. Published models require role authorization and schema-valid inputs. Model explanations are associations, not causes or diagnoses.
- Retire a published endpoint when finished. Sign out on shared devices.
Access and limits
Owners manage projects and memberships. Analysts import and run analyses. Reviewers review and export permitted evidence. Viewers can read granted projects. Account provisioning and recovery are operator-assisted; project invitations do not send email.
Hosted limits: eligible nonconfidential files, 1 MiB upload, 10000 rows, 100 source columns, 100000 source cells, 16 MiB decoded/workbook expansion, 64 MiB estimated model matrix, 64 MiB original-input quota per project, one worker job per host, 15-minute job deadline and two-hour sessions. General large-dataset performance is unqualified.
Dataset deletion is an audited tombstone, blocked while active jobs or published endpoints use it. Immediate physical purge and automatic retention are not implemented. Backups preserve earlier project data.
Research-only methods need scientific judgement. The pilot does not provide diagnoses, treatment advice, institutional SSO or regulatory certification.
Synthetic format and method examples
All examples below are artificial. They contain no real people.
- TSV cohort example
- TXT cohort example
- PARQUET cohort example
- XLSX cohort example
- JSON cohort example
- JSONL cohort example
- Synthetic statistical/model workflow CSV
The TXT example uses pipe delimiter and cp1252 encoding. The workflow contains x/z measurements, binary outcome, multiclass multi, count response, positive exposure, pair keys, grouping and dates for guided examples.
SPSS/Stata/SAS, mixed/repeated inference, survival, panel, forecasting, association rules, nested evaluation and multiple imputation remain unavailable. Chronological holdout is available; chronological hyperparameter tuning is unavailable. These limits must not be bypassed by substituting an independent-row model.