Publication

P588: AI and drug discovery with 100 million cells of genome-wide Perturb-seq

Jan 1, 2026 · 29 authors · 3 topics

Abstract

Understanding gene function at scale requires systematic perturbation and single-cell profiling across diverse cellular contexts. Here, we describe the generation of the largest genome-wide Perturb-seq atlas to date, comprising over 100 million single cells across dozens of human cell types. Methods: The dataset was generated using Illumina's emulsion-based single-cell kit with direct guide capture (PIP-seq), enabling profiling of over 1 million single cells per reaction. We integrated genome-wide CRISPR interference (CRISPRi), and CRISPR activation (CRISPRa) across multiple cell types. To test the reproducibility of the CRISPR-based perturb-seq results, we also performed orthogonal perturbation experiments using a genome-wide shRNA library. Results: CRISPRi Perturb-seq achieved 75% guide assignment and a median knockdown efficiency of 80% in genome-wide screens. CRISPRa screens across diverse cell types enabled robust activation of thousands of lowly expressed and lineage-specific genes, expanding functional annotation of regulatory networks. Genome-wide CRISPRi Perturb-seq in iPSCs revealed numerous biologically relevant cell state transitions and novel transcriptional regulators of pluripotency. Despite mechanistic differences between CRISPR-based and shRNA perturbations, we observed strong concordance in downstream gene expression programs, supporting the robustness of the approach. Conclusion: This Perturb-seq atlas provides a foundational resource for functional genomics and AI-driven modeling, enabling prediction of gene function, inference of pathway activity, and prioritization of therapeutic targets. By integrating CRISPRi, CRISPRa, and shRNA perturbations across diverse cell types, this dataset establishes a benchmark for reproducibility and accelerates discovery in human biology. Introduction: Impaired speech and language development may be the first sign of autism spectrum disorders (ASD) and neurodevelopmental delays (NDD). More than 1/3 of the individuals diagnosed with autism may remain nonverbal (NV -no consistent oral spoken expressive language at 18 months of age) or minimally verbal (MV-less than 50 words of oral expressive language at 30 months of age or older). Clinical genome sequencing in 117 NV to MV individuals (70 males, 47 females) yielded a molecular diagnosis for 20% of individuals. Here, we sought to reanalyze the clinical genomes to determine the molecular architecture and genomic profile to identify underlying genetic etiology in NV and MV ASDs and NDDs. Methods: GS was performed using Illumina short-read technology (2×150 bp, ≥30× depth). Variants in coding and splicing regions were identified using Illumina DRAGEN, annotated with ANNOVAR, and custom R scripts, and prioritized based on population frequency, conservation, probability of being loss-of-function intolerant (pLI), and ASD/NDD relevance. Gene ontology (GO) enrichment analysis was conducted using topGO R package. Results: We identified 422-1701 low frequency variants per proband (<5% in gnomAD v4.1), and 7-37 variants per proband when restricted to ASD/NDDassociated genes. GO enrichment of genes impacted by low-frequency variants revealed terms related to transcriptional regulation and neural development. Filtering for very rare variants (<0.001% in gnomAD v4.1) highlighted terms associated with synaptic function and glutamatergic signaling. We identified 52 protein-truncating variants in 37 likely haploinsufficient genes, including 25 variants across 16 ASD/NDD genes. High-confidence, pathogenic variants were identified in ANKRD11, SETBP1, STAG1, FOXP1, CACNA1A, CNOT1, SRRM2, PDHA1 and WDR45. Variants in genes belonging to transcription/chromatin and RNA processing (ANKRD11, CNOT1, SRRM2, TAF4, ZMYND8, MAX, YBX1), ion channel/signaling (CACNA1A, KCNMA1), and mitochondrial metabolism (PDHA1) networks were identified. Additional ultra-rare predicted loss-of-function variants in several genes including NR2F2, ZSWIM6, CAND1, YBX1, HIVEP1, HSP90B1, ATXN7L1, EPB41L4B, CNOT7, PRDM16, CDK16, FNDC3A, USP8, ATXN1, PIAS2, ADCY2, RPS6KA6, SETDB1, TENM2, TAF4, ZMYND8, and SRRM2 were identified. Conclusion: Nearly, a third of autism and profound neurodevelopmental delays are caused by rare and ultrarare monogenic pathogenic variants. However, we and others have observed additional burden of rare deleterious variants across key neurodevelopmental networks and transcriptional/chromatin remodeling genes suggesting the potential role of multiple gene variants belonging to specific pathways. Additional confirmatory testing and analyses are ongoing to decipher the complex molecular architecture of autism and neurodevelopmental delays expanding our understanding of pathophysiology of autism and NDDs and providing potential treatment options. https://doi.org/10.1016/j.gimo.2026.104080

Showing the abstract — retrieve the full paper via the Exa API.

Authors

8 of 29
Kwontae YouJiang ZhuDulguun AmgalanEyal Ben DavidAlejandro Mendez MancillaEmily LaubscherJonatan PerezLenka Dohnalova

Topics

Single-cell and spatial transcriptomicsCancer Genomics and DiagnosticsCell Image Analysis Techniques

About

PublishedJan 1, 2026
Citations0

Powered by the Exa API