Mots-Clés
cancer
genomics
precision-medicine
pipelines
snakemake
translational-medicine
multiomic
Description
Junior Bioinformatics Engineer
Gustave Roussy, Europe’s leading cancer centre is seeking a junior bioinformatics engineer.
This is a 5 year contract position on Gustave Roussy terms and conditions of employment.
Summary
The Clinical Discovery Bioinformatics group within the L’Institut Hospitalo-Universitaire (IHU) de médecine de précision PRISM at Gustave Roussy Institute seeks to recruit a collaborative and technically minded Bioinformatics Data Engineer with experience of data analysis workflows and data management in a research environment. IHU PRISM is a multidisciplinary, cross-tumor program aimed at reducing cancer mortality by better understanding the biology of each patient’s cancer. In this key role, and as part of our dynamic team, you will apply your bioinformatics, programming and data science skills to help generate results from large multiomic patient datasets produced by the Institute’s clinical programs. You will play a key role in helping to answer the latest questions in cancer translational research and help facilitate collaborations with research teams to develop advanced analyses and AI tools for translational medicine.
Project summary
In this role, you will be responsible for running and developing the data processing component of our integrated data analysis platform. You will apply cutting-edge curated data processing pipelines to generate robust results from the latest spatial-transcriptomic platforms, single-cell technologies, bulk RNA-seq, WGS, exome and imaging modalities. In this key role, you will work closely with the clinical discovery teams to integrate new patient-sample data into the platform and deliver results used to drive translational research. You will be responsible for developing and maintaining these pipelines as well as the integrated platform itself, giving you the opportunity to contribute high-quality reproducibly results to our translational programmes. There is a rich bioinformatics and data science community at GR within which you will play an active role. You will work closely with this community to develop new pipelines, contribute to metadata management, interact with state-of-the-art HPC and data storage systems and help research groups use these data to drive forward their research. You will also be using the latest technologies and methodologies to ensure data processing reproducibly and transparency across the platform.
Key responsibilities
• Process sample data generated by the Institute’s clinical programs.
• Deliver results and provide expertise on QC and results interpretation to clinical teams.
• Manage a suite of curated multiomic data processing pipelines.
• Develop automation strategies for the efficient processing of samples within the data analysis platform.
• Contribute to the development and running of a data processing and analysis tracking system.
• Develop and implement innovative ways of delivering results to help the clinical teams with their research.
• Manage the quality of results across the platform.
• Apply the latest technologies and methodologies to ensure full data processing reproducibility across the data analysis platform.
• Validate and implement new analysis and data processing steps in collaboration with other bioinformatics groups.
• Work closely with our data management team to develop results storage and archiving strategies.
• Contribute to the development of metadata management infrastructure.
• Continue professional development through maintaining awareness of developments in the wider bioinformatics and research communities.
• Participate in Gustave Roussy meetings, workshops and seminars.
• Contribute to publications when required.
Key experience and competencies
The post holder should be open, dynamic and collegial, in addition to:
Essential Qualifications, experience and competencies:
• A degree in a relevant subject with an extensive analytical component e.g. bioinformatics, statistics, molecular biology or mathematics.
• Excellent coding skills in a relevant language e.g. Python, Java, R.
• Excellent Linux and HPC skills.
• Experience of running and developing bioinformatics pipelines using a workflow management system such as Nextflow or Snakemake.
• A good understanding of high-throughput sequencing data processing software and bioinformatics data processing tools.
• An understanding of multi-omics data and data analysis approaches.
• Knowledge of reporting software e.g. quarto, markdown.
• Knowledge of technology ensuring data processing reproducibility e.g. containerisation, environments, change control software.
• Understanding of software development processes and management.
• The ability to organise and prioritise workload
• Excellent scientific communication skills within a research environment.
Desirable Qualifications, experience and competencies:
• Familiarity with cancer biology.
• Familiarity with the application of statistical techniques to biological data.
• Familiarity with omics data generation protocols
• Familiarity with visualising omics data analysis results.
• Knowledge of relational database technology.