Stage Master 2 - Explaining protein language model hallucinations

 Stage · Stage M2  · 6 mois    Bac+5 / Master   Inria d'Université Côte d'Azur · Sophia Antipolis (France)  Gratification selon barème

 Date de prise de poste : 4 janvier 2027

Mots-Clés

Protein language models Hallucinations Interpretability Protein structure prediction

Description

Description

Transformer-based models for protein structure prediction (TBMs) hallucinate a non-negligible fraction of their targets – they confidently produce physically implausible outputs whose failure modes are poorly understood and almost entirely opaque [1].

We want to develop explainability methods that identify which input features, attention patterns, and latent representations drive hallucination events. To do so, we need to first compile a dataset of structure prediction hallucinations: we recently have released abstraqt, a scoring function based on local packing able to catch some very common hallucinated structure – extremely long and bent alpha helices, exotic beta secondary structure – but other and more subtle hallucination types are still undetected.

This project aims to build FLoPS, a curated dataset of structure prediction hallucinations that will be used as a base to develop interpretability methods for structure predictors such as the AlphaFold and ESMFold families. The main tasks will be:
• Conceive HPC scripts to run state-of-the-art protein structure predictors on entire genomes;
• Analyze the predictions with abstraqt and extend the performance of the scoring function;
• Collect internal data (activations, latent representations, etc.), organize and analyze them to find common traits between different failure modes.

The student will familiarize with structure quality assessment methods, protein language models and explainable AI (XAI) concept and practices.

[1] E. Sarti and F. Cazals. “Fold or flop: quality assessment of AlphaFold predictions on whole proteomes”, Bioinformatics Advances (2026), vbag190, doi:10.1093/bioadv/vbag190.

Requirements

Essential

• Following the 2nd year of a Master program in bioinformatics, computational biology, bio- physics, machine learning or a closely related field.
• Solid Python programming skills.
• Experience with Linux operating systems.
• Familiarity with computational structural biology data or with deep learning architectures.

Nice to have

• Background in computational structural biology or deep learning.
• Prior experience with explainability or interpretability methods for large neural networks.
• Familiarity with protein language models (ESM family, ProtTrans, or similar).
• Background in statistical mechanics, molecular dynamics, or enhanced sampling, even at a foundational level.
• Experience with HPC environments and GPU-accelerated training.

Opportunities and contact

Depending on the candidate’s motivation, possibilities to prolong the stage with a Ph.D. fellowship will be explored.
Please write an email to edoardo.sarti@inria.fr attaching your resume and M1 grades.

Candidature

Procédure : Please edoardo.sarti@inria.fr

Contacts

 Edoardo Sarti
 edNOSPAMoardo.sarti@inria.fr

 https://team.inria.fr/abs/files/2026/09/M2_XAI-sarti-2027.pdf

Offre publiée le 2 octobre 2026, affichage jusqu'au 30 novembre 2026