Join WALDO
Community Initiative

WALDO

Hidden in Plain Sight

community-Wide Annotation of γδ Lymphocytes in single-cell Data, Openly benchmarked

An open community of expert γδ T cell and cancer immunology laboratories building, benchmarking, and sharing computational pipelines that help the broader immunology research community reliably identify and properly annotate γδ T cells in single-cell RNA sequencing data. In the long run, WALDO will provide community-wide consensus γδ T cell subset annotation guidelines across health and disease.

10Expert Labs
6Countries
3Pipelines
6Roadmap Phases

The problem

γδ T cells carry no unique marker gene and are routinely mislabeled, absorbed into neighboring T or NK clusters, or discarded as unassignable in single-cell RNA-seq data.

The approach

Community-built identification pipelines with deliberately different design philosophies, benchmarked side by side on shared, community-curated datasets.

The outcome

Open, versioned pipelines, practical usage guidance, and consensus γδ T cell annotation guidelines across tissues, conditions, and species.

Background & Motivation

γδ T cells are a conserved lineage of unconventional lymphocytes that combine innate-like mechanisms of activation and response with adaptive antigen recognition through a dedicated γδ T cell receptor (TCR)[1]. They typically constitute only 1–10% of peripheral blood CD3+ T cells[1,2], but are markedly enriched at epithelial barriers and in tissues, accounting, for example, for roughly 20–30% of intestinal intraepithelial lymphocytes (IELs) in humans and up to 50–60% in mice[3], where they respond rapidly to infection, stress, and malignant transformation[2,4].

γδ T cells are therefore increasingly recognized as promising effectors in cancer immunotherapy[4]. Despite their biological importance, they remain surprisingly under-represented and misannotated in publicly available single-cell RNA sequencing (scRNA-seq) datasets, where they typically constitute a small fraction of profiled cells and are frequently mislabeled, absorbed into neighboring T or NK clusters, or discarded as unassignable[5]. Laboratories in this community have met the problem directly: γδT-omics profiling across blood, healthy colon, primary tumor and liver metastases in colorectal cancer paired γδTCR sequencing with transcriptomics to validate and refine γδ T cells reliably[6], and γδ T cell identification was a prerequisite for the analysis of a longitudinal single-cell CAR T cell atlas in multiple myeloma[7].

This difficulty is structural rather than incidental. γδ T cells carry no unique marker gene, and their few defining T cell receptor transcripts (TRDC, TRGV, TRDV) are prone to the well-known dropout problem of scRNA-seq data. Moreover, their expression profile overlaps extensively with cytotoxic αβ T cells, natural killer (NK) cells, and MAIT cells, so that score- and threshold-based classifiers, and reference-mapping tools trained on atlases in which γδ T cells are themselves under-annotated, fail with a lot of false negatives.

No single laboratory can fix this alone: the problem spans wet-lab biology, computational method design, and the need for benchmark datasets across tissues, species, and disease states. WALDO, a world-wide community of γδ T cell and cancer immunology laboratories, is an initiative to tackle this problem collectively, developing identification pipelines together, testing them on shared data, and giving the wider research community clear guidance on what to use, when, and why.

Objectives & Scope

  • Charting γδ T cell identification pipelines. Gather, document, and evolve a portfolio of γδ T cell identification methods with deliberately different design philosophies.
  • Benchmarking. Evaluate every pipeline side by side on shared, community-curated datasets (both 3′ and 5′ technologies) spanning blood and tissue, healthy and diseased, human and mouse samples.
  • Interpretability and reproducibility. Every pipeline exposes its decision logic and is versioned, so results stay auditable and comparable over time.
  • Practical guidance. Document features, strengths, and ideal use cases so that researchers can choose the right approach for their sample type, size, and scale.
  • Consensus nomenclature. Provide community-wide consensus γδ T cell subset annotation guidelines across health and disease.

The WALDO Consortium

WALDO is open to any laboratory or researcher working on γδ T cells, single-cell genomics, or computational method development. Members contribute pipelines, datasets, expert annotations, benchmark results, or simply use cases and feedback. The effort is developed by the community, for the community.

WALDO currently comprises researchers from 10 research laboratories across 6 countries, spanning Europe (Netherlands, Belgium, Germany, United Kingdom), Asia (China), and North America (the United States).

0
Expert Laboratories
0
Countries
0
Charted Pipelines
Member laboratories · worldwide
Hong Kong Cleveland Princeton Europe ×7
Hover a pin for laboratory details
Europe · detail
Seven laboratories across Europe
University Medical Center Utrecht The Netherlands
Université Libre de Bruxelles Belgium
The Chinese University of Hong Kong China
University of Glasgow United Kingdom
University Medical Center Hamburg-Eppendorf Germany
Hannover Medical School Germany
University of Warwick United Kingdom
Princeton University USA
Leiden University Medical Center The Netherlands
Case Western Reserve University USA

Join WALDO

WALDO is actively looking for new contributing laboratories. Whether your strength is γδ T cell biology, single-cell genomics, or computational method development, there is a way to take part, and contributions of any size shape the community benchmark and the consensus guidelines.

Contribute a pipeline

Share your γδ T cell identification method, and we chart, document, and benchmark it side by side with the existing portfolio.

Share data

Contribute curated or expert-annotated datasets across tissues, species, technologies, and disease states.

Expert annotation

Help define ground truth, harmonized subset definitions, and the consensus nomenclature guidelines.

Use cases & feedback

Run the pipelines on your own data, report edge cases, and help shape the practical usage guidance.

Get in touch to join

or write directly to Dr. Farid Keramati (University Medical Center Utrecht), f.keramati-2@umcutrecht.nl

WALDO Roadmap

Click a phase to expand its details.

Phase 1 Pipeline Archive Ongoing
Gather γδ T cell identification methods from community members, chart their design philosophies, and collect them in a shared repository with fully documented, runnable scripts. Three pipelines (WALDO-ML, WALDO-Diff, WALDO-Net) have been charted and documented so far; further community submissions are welcome.
Phase 2 Benchmark Ongoing
Side-by-side evaluation of all charted pipelines on the community's curated datasets, both 3′ and 5′ scRNA-seq technologies; blood- and tissue-derived; healthy, cancer, autoimmune, aging, and viral-infection cohorts; human and mouse, using harmonized ground truth (expert annotation or paired TCR sequencing) and pre-registered metrics.
Phase 3 Usage Guideline Ongoing
Curate practical guidance on each pipeline's features, strengths, and ideal use cases, including which approach to choose for a given sample type, size, and scale.
Phase 4 Packaging Q1 2027
Hardened, cleaned and ready-to-dispatch releases of all charted pipelines (R and Python packages, containers) with continuous testing.
Phase 5 Consensus Annotation Q1-2 2027
Community-wide consensus γδ T cell subset annotation guidelines, harmonized subset definitions, marker panels, and nomenclature across health and disease, informed by the community experts' review.
Phase 6 Publication Q3 2027
Submit for peer-reviewed publication reporting the community benchmark and the consensus annotation guidelines, superseding the versioned Zenodo preprints, a first step toward a community-wide γδ T cell reference across tissues, conditions, and species.

Current Pipeline Portfolio

WALDO is currently reaching out to research laboratories with expertise in γδ T cell biology or computational scRNA-seq annotation, inviting contributions of insights, pipelines, and benchmarking datasets. The goal is to chart each method and curate fully documented, runnable scripts in a shared repository. Three main pipelines (described briefly below) have already been charted and documented, and researchers from the contributing member laboratories are actively refining them. Detailed mathematical descriptions of all three pipelines will be provided in the accompanying publication.

WALDO-ML plug and play · extremely fast · atlas scale

A fast, pre-trained machine-learning classifier shipped as a single frozen model artefact (model + gene panel + threshold). It streams a standard .h5ad file, no preprocessing, embedding, or reference atlas needed, and returns per-cell scores and calls, with a "death penalty" that attenuates scores of likely compromised cells and a recorded model fingerprint (SHA-256) for full reproducibility.

WALDO-Diff graph diffusion · continuous scores

A framework that treats markers as high-confidence seeds, not final labels. It gates γδ, αβ T, and NK seed families, weights seeds by signature strength (UCell) and graph-derived quality, then diffuses each family's evidence separately over a shared cell-cell graph via random walks with restart. Per cell, the strongest competitor walk is combined with the γδ walk into two tunable, fully auditable axes, support and contrast, with labels calibrated from the run itself.

WALDO-Net interpretable · high precision

An interpretable hybrid pipeline combining marker gating with network topology. Five inspectable stages, conservative marker nomination, island/branch pruning on the nearest-neighbor graph, Leiden sub-clustering, adaptive NK-contaminant removal, and staged network rescue of dropout-affected cells, produce high-precision call sets in which every inclusion or exclusion can be traced to a named step. The two-pass nature of the protocol re-embeds nominated cells in a TCR-focused feature space for deeply confounded samples to reach high precision.

Current Benchmarking Dataset Portfolio

In parallel, ground truth datasets for the formal side-by-side benchmark are being nominated and collected from contributing laboratories, encompassing both publicly available and privately held resources. These span healthy and diseased states, blood- and tissue-derived samples, and human and mouse cohorts.

Paired scRNA-seq + scTCR-seq 5′ chemistry · same-sample αβTCR and γδTCR

Both public datasets (e.g., GSE144469) and privately generated datasets from contributing laboratories (Jurgen Kuball, Seth Coffelt, Noel Miranda) are being collected, spanning human and mouse, healthy and diseased, PBMC- and CRC tissue-derived samples. Paired TCR sequencing on the same cells provides a direct molecular ground truth for αβ/γδ lineage assignment.

FACS-sorted αβT and γδT cells clean ground truth labels

αβT and γδT cells are sorted from the same biological sample and sequenced together, ensuring clean, sort-based ground truth labels. Privately generated datasets from contributing laboratories (David Vermijlen, Noel Miranda) are being collected for this dataset class.

Expert-consensus public datasets both 3′ and 5′ chemistries

Publicly available datasets are being curated for γδT cell identity by expert consensus: multiple community members independently identify γδ T cells in agreed-upon datasets (e.g., GSE271896), with labels reconciled across annotators.

CITE-seq RNA + surface protein · definitive αβ/γδ assignment

Datasets containing αβ, γδ T, and NK cells with surface protein markers, where combined RNA and protein evidence enables definitive αβ/γδ assignment. Privately generated datasets from contributing laboratories (Immo Prinz) are being collected for this dataset class.

Bonus datasets activation-state stress tests

Longitudinal samples (e.g., pre/post stimulation with phosphoantigen or IL-2/IL-15) and TCR-stimulation samples, to benchmark γδT identification across different activation states.

Research laboratories able to contribute to any of the above ground truth dataset categories, whether through published resources or privately held datasets shared under confidentiality, are invited to participate.

References

  1. Vantourout P, Hayday A. Six-of-the-best: unique contributions of γδ T cells to immunology. Nat Rev Immunol. 2013;13(2):88–100. doi:10.1038/nri3384
  2. Ribot JC, Lopes N, Silva-Santos B. γδ T cells in tissue physiology and surveillance. Nat Rev Immunol. 2021;21(4):221–232. doi:10.1038/s41577-020-00452-4
  3. Rampoldi F, Prinz I. Three layers of intestinal γδ T cells talk different languages with the microbiota. Front Immunol. 2022;13:849954. doi:10.3389/fimmu.2022.849954
  4. Silva-Santos B, Mensurado S, Coffelt SB. γδ T cells: pleiotropic immune effectors with therapeutic potential in cancer. Nat Rev Cancer. 2019;19(7):392–404. doi:10.1038/s41568-019-0153-5
  5. Mullan KA, de Vrij N, Valkiers S, Meysman P. Current annotation strategies for T cell phenotyping of single-cell RNA-seq data. Front Immunol. 2023;14:1306169. doi:10.3389/fimmu.2023.1306169
  6. Gatti LCDE, Nicolasen MJT, Keramati F, Brazda P, Meringa AD, De Bont DA, Aarts-Riemens T, Brandwijk WJC, van der Wijst MIM, Stuut AHG, Gasull Celades L, Zawal D, Vazaios K, Daudeij A, Huismans MA, Parigiani MA, Straetemans T, Beringer DX, Kranenburg O, Berlin C, Minguet S, Kesselring R, Stunnenberg HG, Roodhart JML, Sebestyen Z, Kuball J. Functional γδT-omics pipeline reveals compartmentalization of Vδ1+ T cell migration, tumor-reactivity, and clonality in human colorectal cancer. bioRxiv. 2025. doi:10.1101/2025.08.19.671055
  7. Rade M, Fandrei D, Kreuz M, Seiffert S, Grahnert A, Friedrich M, Wiemers T, Born P, Fischer L, Weidner H, Hofbauer LC, Baber R, Wang SY, Bach E, Hoffmann S, Scolnick J, Friedrich M, Keramati F, Brazda P, Sebestyen Z, Kuball J, Alb M, Scheller L, Hudecek M, Einsele H, Metzeler KH, Herling M, Herling CD, Jentzsch M, Franke GN, Boldt A, Köhl U, Platzbecker U, Vucinic V, Reiche K, Merz M. A longitudinal single-cell atlas to predict outcome and toxicity after BCMA-directed CAR T cell therapy in multiple myeloma. Cancer Cell. 2026;44(3):586–603.e9. doi:10.1016/j.ccell.2025.10.014

How to Cite

Keramati, F., Rezwani, M., Tan, L., Hamilton, D., Song, Z., Yang, T., Wang, C., Moreno Vicencio, D., de Miranda, N. F. C. C., Brubaker, D. K., Sebestyen, Z., Davey, M., Ravens, S., Coffelt, S., Prinz, I., Vermijlen, D., & Kuball, J. (2026). WALDO: community-Wide Annotation of γδ Lymphocytes in single-cell Data, Openly benchmarked. Zenodo. https://doi.org/10.5281/zenodo.21888765