Stage de Master2 en développement d'outil de visualisation

 Stage · Stage M2  · 6 mois    Bac+5 / Master   DIADE - Diversité, adaptation, développement des plantes · Montpellier cedex 5 (France)

 Date de prise de poste : 4 janvier 2027

Mots-Clés

visualization epigenomics software development

Description

Visualising transcription factor binding on Pangenomes
Decreasing costs of high-throughput and long-read DNA sequencing over the last decade led to the rise of Pangenomics, where for example genomes of different varieties of one species are compared instead of just sequencing one representative “model” variety. Therefore, we now often compare of fully assembled genomes in contrast to the previous standard method where variants were determined by mapping short read sequencing data to a reference genome. As a consequence, Pangenomics has to handle as many genome coordinate systems as there are varieties analysed in the population of interest. Using whole-genome alignment, the multitude of insertions, deletions and inversions can be represented in a Pangenome variation graph and studied for functional association. While this greatly improves the accuracy of the analysis, it also creates challenges for the visualisation of results because genomes can no longer be displayed as a linear strand, which is the basis for all standard genome browsers. To answer this challenge, the PANEEC team at the DIADE research unit has developed visualisation tools for pangenome variation graphs such as SaVanache (Mohamed et al. 2026). 
Another challenge is the visualisation of functional features on multiple genomes, such as transcription factor binding. In the REPROD team, we use transcription factor (TF) foot- printing by MOA-seq to study gene regulatory processes during plants exposure to stress and during developmental processes. Within a collaboration we analysed TF binding in 26 diverse Maize varieties to generate a Pan-cistrome, comprising all TF binding sites (Engelhorn et al., 2025). We could successfully link SNPs and small InDels to changes in TF occupancy, and identify candidate regions for crop improvement from the data. However, visualisation of such data remains challenging (Figure 1). Most genome browsers allow only a single genome to be displayed, while some newer versions (e.g. Jbrowse2, (Diesh et al., 2023), or SyRI (Goel et al., 2019)) allow for two or more genomes. Synteny is displayed as a ribbon, and insertions can be collapsed. However, loading several genomes into the browser is resource intensive and can create challenges on a local computer and restrict usability. 
This Master project will explore new possibilities for the visualisation of pan-cistrome data by
• Generation of a Pangenome graph for the 26 maize lines of our first pan-cistrome
• Implementation of gene models and TF binding track data into SaVanache, including file size and format considerations
• Evaluation of visualisation tools on selected loci in the maize genome where we discovered sequence variants associated with binding differences (bQTL)
Once successfully applied to one plant, the solution can be implemented for various other species of interest for which we are currently preparing Pan-cistromes, e.g. Pearl millet, Barley, African rice and European Maize. 
We are looking for a candidate with a strong interest in programming of visualisation tools and understanding of gene regulation. Prior experience with whole genome alignment and NGS data is a plus but not strictly necessary. Basic understanding of Unix is highly desirable.

References: 
Diesh, C., Stevens, G. J., Xie, P., De Jesus Martinez, T., Hershberg, E. A., Leung, A., Guo, E., Dider, S., Zhang, J., Bridge, C., et al. (2023). JBrowse 2: a modular genome browser with views of synteny and structural variation. Genome Biol. 24, 74. 
Engelhorn, J., Snodgrass, S. J., Kok, A., Seetharam, A. S., Schneider, M., Kiwit, T., Singh, A., Banf, M., Doan, D. T. H., Khaipho-Burch, M., et al. (2025). Genetic variation at transcription factor binding sites largely explains phenotypic heritability in maize. Nat. Genet. 57, 2313–2322. 
Galli, M., Chen, Z., Ghandour, T., Chaudhry, A., Gregory, J., Feng, F., Li, M., Schleif, N., Zhang, X., Dong, Y., et al. (2025). Transcription factor binding divergence drives transcriptional and phenotypic variation in maize. Nat. Plants 11, 1205–1219. 
Goel, M., Sun, H., Jiao, W.-B. and Schneeberger, K. (2019). SyRI: finding genomic rearrangements and local sequence differences from whole-genome assemblies. Genome Biol. 20, 277. 
Mohamed, M., Durant, É., Rouard, M., Muller, C., Conte, M. and Sabot, F. SaVanache: indexing and visualizing pangenome variation graphs. 
Contacts:

François Sabot (UMR DIADE, PANEEC Team, IRD Montpellier), francois.sabot@ird.fr

Julia Engelhorn (UMR DIADE, REPROD Team, IRD Montpellier/University Montpellier), julia.engelhorn@umontpellier.fr

Candidature
Contacts

 Francois Sabot
 frNOSPAMancois.sabot@ird.fr

 Julia Engelhorn
 juNOSPAMlia.engelhorn@umontpellier.fr

Offre publiée le 25 août 2026, affichage jusqu'au 27 novembre 2026