Contractor for molecular OTU clustering and target-gene classification services

 CDD · Autres  · 12 mois    Bac+3 / Licence   Global Biodiversity Information Facility · Copenhagen (Danemark)  total sur la durée du contrat: 100333 euros

 Date de prise de poste : 4 janvier 2027

Mots-Clés

sequence search/clustering tools sequence analysis distributed management platforms eDNA metabarcoding GIT-based workflows

Description

As part of SPLICE, the multi-year initiative funded by the Novo Nordisk Foundation aimed at integrating environmental DNA (eDNA) data into its global biodiversity data infrastructure, GBIF is seeking a contractor to design and build a workable, sustainable approach to two problems: Firstly, a large and growing share of eDNA sequences cannot be resolved to a formal scientific name. Without a stable unit by which to group and reference them, this “dark taxon” data is hard to name, compare, track or build on across studies and over time. Secondly, no globally agreed method reliably determines which gene region or sub-region (e.g. ITS1, ITS2, 16S-v4, 18S-v9) a given sequence represents—a determination that downstream search, reannotation and clustering all depend on.

The challenge

This is deliberately not a fully specified engineering brief: the SPLICE project team is not convinced an off-the-shelf tool or narrow technical fix will hold up at scale, across the diversity of taxa and marker genes involved. GBIF seeks a contractor who can bring their own judgment to developing a solution that is scalable, flexible across gene regions and low-maintenance in production.

GBIF expects meaningful early effort spent exploring and comparing existing approaches and tools rather than treating tool selection as already decided. The project team is open to being persuaded on architecture; what matters is a solution that stays scalable and sustainable for years to come and integrates well with GBIF’s current processing pipelines.

Indicative scope of work

The two tasks may be treated as linked, as reliable gene-region classification may facilitate better sequence clustering. In broad terms:

  1. Target-gene / gene-region classification
    Develop a service or tool that can reliably classify an incoming DNA sequence to its gene region and, where relevant, sub-region (e.g. distinguishing ITS1 from ITS2, or different hypervariable regions of 16S/18S), to support downstream clustering, reannotation and filtering options.

  2. Molecular OTU (sequence) clustering
    Develop a pipeline that can cluster sequences into referenceable and indexable molecular operational taxonomic units (mOTUs) across multiple marker gene regions, at a scale suited to GBIF’s growing sequence archive—including consideration of how units can be indexed, versioned and remain linked to the GBIF taxonomic backbone, reference libraries and other sequence-related services.

Conceptually, the proposed development sits alongside existing curated clustering systems such as UNITE Species Hypotheses and BOLD BINs. GBIF intends to collaborate with, not replace, these systems, and will defer to them wherever they apply. The gap to fill is the remaining space of target-gene/taxonomic-group combinations that no such externally curated system currently covers. Whatever is built should be designed to sit comfortably alongside SHs and BINs—combinable with them, not a competing or incompatible standard.

Within this, GBIF anticipates the engagement will typically involve:

Reviewing the problem and relevant existing work and tools, and proposing an architecture and approach
Building and iterating on a working prototype
Testing against representative, real-world GBIF/eDNA datasets
Documenting the design, its limitations and how it should be operated and maintained
Supporting GBIF on integration into its existing data infrastructure and software stack
The exact division of effort between design/exploration and implementation, and between the two problem areas, is expected to be proposed by the contractor as part of the engagement and refined together with GBIF early in the contract.

Deliverables

An architecture/design proposal, agreed with GBIF, covering both the classification and clustering components
A working prototype implementation of each service, tested against representative datasets
Documentation covering the design, known limitations, and operational/maintenance requirements
Recommendations for integration into GBIF’s production infrastructure
Deliverables and payment instalments will be finalized with the selected contractor at contract signature.

Skills and experience

Ideal candidates should be able to demonstrate most of the following:

  • Strong background in bioinformatics, computational biology or a closely related field, with hands-on experience in sequence analysis
  • Practical experience with sequence search/clustering tools (e.g. VSEARCH, USEARCH, SWARM, MMseqs2, DNAclust) and/or sequence classification approaches (e.g. Kraken2, HMMER, or similar)
  • Practical experience with DAG-based orchestration systems, such as Airflow or Nextflow
  • Familiarity with Kubernetes, Cloud or distributed management platforms consistent with GBIF’s infrastructure.
  • Comfort working with GIT-based workflows
  • Experience with Spark or distributed pipeline design, critical for the solution to hold up against our growing eDNA archive.
  • Comfort with HDFS/object storage, or any distributed file systems as the underlying data layer
  • Comfort working with ambiguous, evolving requirements and proposing their own technical direction
  • Experience with metabarcoding/eDNA data specifically is an advantage but not a strict requirement
  • Willingness to produce a documented handoff/runbook and technical documentation as an explicit deliverable, not just code, since our team owns this after the 12-month engagement ends.

Application process

Candidates should submit, in English:

  • A CV
  • A short (max 2-page) outline of how they would approach the problem, including any initial thoughts on architecture or tooling
  • Examples of relevant past work (code, publications or systems built) from the last five years

Applications and questions should be sent to dna-contractor-2026@gbif.org by 30 October 2026, 17:00 CET (UTC+1). Shortlisted candidates may be invited to an interview.

Application deadline: 30 October 2026, 17:00 CET (UTC+1)
Expected start: 4 January 2026 or soon thereafter
Duration: up to 12 months
Location: Remote (from contractor’s home base), working closely with the GBIF Secretariat
Remuneration:
Up to €100,333, paid in instalments against agreed milestones

https://www.gbif.org/news/6DchNRkO4YWxR4pMVF776E/gbif-seeks-contractor-for-molecular-otu-clustering-and-target-gene-classification-services

Candidature

Procédure : Candidates should submit, in English: A CV A short (max 2-page) outline of how they would approach the problem, including any initial thoughts on architecture or tooling Examples of relevant past work (code, publications or systems built) from the last five years Applications and questions should be sent to dna-contractor-2026@gbif.org by 30 October 2026, 17:00 CET (UTC+1). Shortlisted candidates may be invited to an interview. Application deadline: 30 October 2026, 17:00 CET (UTC+1) Expected start: 4 January 2026 or soon thereafter Duration: up to 12 months Location: Remote (from contractor's home base), working closely with the GBIF Secretariat Remuneration: Up to €100,333, paid in instalments against agreed milestones

Date limite : 30 octobre 2026

Contacts

 GBIF Secretariat
 dnNOSPAMa-contractor-2026@gbif.org

 https://www.gbif.org/news/6DchNRkO4YWxR4pMVF776E/gbif-seeks-contractor-for-molecular-otu-clustering-and-target-gene-classification-services

Offre publiée le 6 octobre 2026, affichage jusqu'au 30 octobre 2026