Mots-Clés
Pangenome
Haplotype
Phasing
Imputation
Bioinformatics
Simulation
Description
Internship aim:
Refine the parameters that drive haplotype phasing and genomic imputation quality when working with a pangenome graph reference.
Context
This internship is part of the PANHCHO ANR project: PANgenomic-enabled High-Capacity Haplotyping Objective.
PANHCHO aims to build a new pangenome reference by sequencing 20 individuals of French ancestry using ultra-long reads sequencing.
This reference will be represented as a graph-structured genome capable of capturing complex genetic rearrangements, with the goal of enabling more accurate representation of human variation than linear references.
As pangenome-based genomics is emerging in healthcare, PANHCHO will develop statistical methods that leverage the new pangenome to expand existing resources through imputation of complex variants.
Using a pangenome can indeed give a 100% increase in structural-variant detection as well as decreasing overall error rates for variant discovery by 34% and an overall increase of 9% of total genomic regions covered including an additional 2000 genes.
The new reference data and new methods will therefore be of use in both clinical and research applications.
Internship goal
You will contribute to improving the quality of imputation by studying how modeling choices and pangenome graph properties affect phasing and imputation performance.
Main tasks
1) Benchmark on controlled data:
- Build synthetic chromosome X datasets from two male donors.
- Use these data to benchmark phasing and imputation accuracy.
2) Identify graph-driven factors
- Quantify how pangenome graph characteristics (e.g., structure/complexity, local graph topology, and coverage of alternative paths) influence imputation results.
- Turn observations into imputation model optimisation.
Key Skills
Here are the skills that will be the most valuable for this internship:
- Comfortable with UNIX/Bash to run and orchestrate bioinformatics workflows.
- Able to use R for analysis and visualization.
- Solid understanding of key genomics concepts, including notably haplotype phasing and sequencing alignment.
- Able to use Git/GitHub for version control and collaboration.
Skills appreciated
The following skills would be a plus:
- Previous experience in manipulation of genomic data and manage tool inputs/outputs (e.g. plink, bcftools, samtools).
- Some mathematical background, a key concept in the internship will be Hidden Markov Models (HMMs).
- Comfortable applying FAIR principles to data organization and reproducible analysis.
- Ability to write unit tests and use CI/CD practices where relevant.
- Strong collaborative skills, including working effectively through pull requests.
Candidature
Procédure : Please sent a CV and a cover letter to anthony.herzig@inserm.fr and louis.le-nezet@inserm.fr
The internship starting date can be adjusted to suit the candidate’s needs
Date limite : 30 septembre 2026
Contacts
Louis Le Nézet
loNOSPAMuis.le-nezet@inserm.fr
Anthony Herzig
anNOSPAMthony.herzig@inserm.fr