An identified, patient-led dataset
I share these data under my own name for research and reanalysis. Sequencing, processed calls and native images are freely downloadable; some files remain pending. Genomic data can identify me and reveal inherited information.
Evidence has a source and a status
Stable identifiers link datasets to sources and specimens. I keep originals privately and publish reviewed copies with provider attribution and preparation notes. Findings are labelled clinical, research-derived, patient-reported or unresolved. Missing data do not mean a negative result.
Older tissue may not represent a current node. Check specimens, collection/report dates and reference genomes before combining assays. Do not combine GRCh37/GRCh38 coordinates without documented conversion.
Reuse and attribution
My curated tables and explanatory text use CC BY 4.0; website and release-tool code use MIT. Third-party reports, assay outputs and vaccine designs retain their own rights. Sharing vendor-processed VCFs, expression and alignments grants no license to proprietary methods. Research candidates do not establish a final manufactured product.
Cite: Alexander Young. Open Cancer Data, release 1.15, 8 October 2026. Include the dataset filename and SHA-256 from the manifest in analyses. A DOI has not been assigned.
Reproducibility
The manifest and SHA-256 checksums document the data archive. The source ZIP contains the website, structured data, HTML report assets and release tools; PDFs download separately. Serve the extracted Invoke report through a local static server for offline use.
Large files are outside both ZIPs. Find them in the manifest’s external_files and the sequencing, imaging and pathology catalogues, with checksums. Downloads require no account.
For large files, use curl -L -C - -O URL to resume downloads, then verify the published SHA-256 checksum.
Known gaps
Inspiration
This independent project was inspired by osteosarc.com.