Mots-Clés
statistical learning
clustering
genetics
neurological disorders
autism
Description
Probing Genotype-Phenotype Relationships in UK Biobank Data using guided clustering
Understanding the complex relationships between genetic variation and phenotypic traits remains one of the central challenges in modern genetics. Traditionally, the association between traits and variants is calculated for each trait through Genome-Wide Association Study (GWAS). More recently, Litman et al., proposed to identify patterns of core, associated and co-occurring traits using a generative mixture modeling approach [1]. In such case, data are grouped by similarity within a single feature space (e.g., phenotypic traits), but these groupings lack explanatory power. They do not reveal why elements behave similarly, ignoring the underlying causal features that drive observed patterns; enrichment analyses for variants must be conducted a posteriori [1]. To address this critical gap, we developed STOIC (Statistical learning TO Inform Clustering), a framework which introduces guided clustering, a novel paradigm that jointly performs unsupervised clustering and supervised feature learning [2]. STOIC represents a fundamental shift in clustering methodology by (i) simultaneously clustering data points in one feature space (e.g., phenotypic measurements) while learning predictive models in another (e.g., genetic variants), and (ii) leveraging supervised signals to guide the clustering process, ensuring clusters are both statistically coherent and biologically interpretable. The UK Biobank represents an unparalleled resource for genotype-phenotype studies, containing in-depth genetic and health information from half a million UK participants. With comprehensive genomic data (including array-based genotyping and exome sequencing) linked to extensive phenotypic measurements (clinical assessments, biomarker data, lifestyle factors, and health outcomes), UK Biobank provides an ideal platform to apply STOIC’s guided clustering approach to unveil novel genotype-phenotype relationships. The primary objective of this internship is to adapt and apply the STOIC framework to UK Biobank data to identify and characterize genotype-phenotype relationships that traditional clustering methods cannot reveal [1], with special emphasis on variants linked to neurological disorders and autism.
Specifically, the candidate will:
. Compare the STOIC methodology with existing methods (first semester - bibliography)
. Select genetic features (e.g. eQTLs targeting genes involved in brain development) and phenotypic features for guided clustering
. Modify STOIC to handle the scale and structure of UK Biobank data
. Apply guided clustering to identify phenotypically coherent groups with shared genetic determinants
. Interpret and validate discovered clusters through functional enrichment and orthogonal validation
This internship is a joint collaboration between CNRS (https://www.igmm.cnrs.fr/en/service/joint-igmm-lirmm-imag-computational-regulatory-genomics-team/) and Maynooth University (https://www.maynoothuniversity.ie/faculty-science-engineering/our-people/lorna-lopez), with the possibility of a short research visit to Ireland during the internship period.
The candidate will work in a multidisciplinary team (combining biology, computer sciences and statistics) and in a very active international environment (this internship will be realized in collaboration with Maynooth University, but the team is also member of the FANTOM consortium based at RIKEN Yokohama, Japan). She/he will have a good knowledge of programming and statistical learning, with an interest in genetics and genomics. We seek highly motivated students, craving to learn and discover, ready to take the doctoral school exams and to apply to other PhD fundings. Individual qualities such as adaptability, perseverance, creativity and teamwork are expected.
References
[1] Litman, A., Sauerwald, N., Green Snyder, L., Foss-Feig, J., Park, C. Y., Hao, Y., … & Troyanskaya, O. G. (2025). Decomposition of phenotypic heterogeneity in autism reveals underlying genetic programs. Nature Genetics, 57(7), 1611-1619.
[2] Cassan, O., Raynal, J., Vroland, C., Yasuzawa, K., Kouno, T., Chang, J. C., … & Lecellier, C. H. (2025). A genome-wide, machine learning-guided exploration of the cis-regulatory code involved in neuronal differentiation. bioRxiv, 2025-05.