Transferring genomic data to Northwestern's HPC cluster.
Public genomic repositories such as the Gene Expression Omnibus (GEO), Sequence Read Archive (SRA), and the European Nucleotide Archive (ENA) are invaluable resources but can be challenging to use due to their diverse structures, metadata formats, and download protocols. This workshop will introduce key tools and workflows for accessing raw sequencing and processed genomic data from major public databases to work with on Quest. Participants will learn how to map between GEO and SRA accessions and retrieve metadata and sequence files using command-line tools. We will also discuss best practices for data management, metadata parsing, and troubleshooting common issues in retrieval workflows.
This workshop requires a laptop with a terminal application installed and expects familiarity with working from the command line or attendance at the ‘Working from the Command Line’ workshop.
Audience
- Faculty/Staff
- Student
- Post Docs/Docs
- Graduate Students
Contact
Leticia Vega
Email
Interest
- Academic (general)
- Data Science & AI