A genome-wide knockout catalog built on organoid phenotypes
The MorPhiC program wants to define the function of every human gene by cataloguing what happens when each is deleted, measured in multicellular systems rather than single cells. Its Production Center, run from Sloan-Kettering by Danwei Huangfu, is the factory floor: about 100 curated stem cell lines from diverse ancestries, three engineered readout systems including a neuro-glial tri-culture, and pipelines explicitly designed to be shared. Reference infrastructure like this decides, in practice, what counts as a normal human gene.
Source: Center for scalable knockout and multimodal phenotyping in genetically diverse human genomes, NIH RePORTER project 5UM1HG012654-05, NHGRI, Memorial Sloan Kettering Cancer Center. Primary source. Read: the full project abstract via the NIH RePORTER API, retrieved 2026-09-16. This is an active center in its fifth year with no results reported in this record; the claims below describe the designed program, and where I extend beyond the record I say so.
What the work claims
MorPhiC, the NHGRI program this center serves, states its mission as defining the function of every human gene through a comprehensive catalog of null phenotypes in multicellular systems. The reasoning in the abstract is that the impact of losing a gene depends on cellular context and genetic background, so a catalog worth having must be built across diverse backgrounds and read out in systems that are actually informative of human biology, not just in single cells.
The Production Center's specific commitments are concrete. It will curate, with what the record calls extensive quality control, a panel of roughly 100 human pluripotent stem cell lines, mostly induced pluripotent stem cell lines plus some embryonic stem cell lines, drawn from diverse ancestral populations and from both males and females, and distribute them as a repository. It will prioritize genes affected in neurodevelopmental and metabolic disorders, naming autism and diabetes as examples. And it will read knockout consequences out in three systems of increasing complexity: a micropattern-based gastruloid model for early tri-germ-layer differentiation, a defined neuro-glial tri-culture, and a three-dimensional pancreatic islet-like organoid culture, with primary human islets included in some assays to test generalizability beyond the stem-cell-derived systems.
How it works
The machinery combines technologies this column has watched converge separately: CRISPR-Cas9 knockout at scale, human pluripotent stem cell guided differentiation, organoid engineering, and multitiered phenotyping scaled across both genes and line backgrounds. A null phenotype here means the observable consequence of deleting a gene in a given multicellular context, and the bet is that a grid of such phenotypes across about 100 genetic backgrounds turns gene function from a literature assertion into a measured, comparable quantity.
Two design choices deserve attention. First, the multitiered structure: assays are arranged so that cheap, scalable readouts triage which knockouts deserve expensive ones, which is what makes genome-wide coverage conceivable. Second, the consortium structure: partner centers set Phase 1 gene priorities, set standards for data and resource sharing, and work out methods for joint analyses, and the center commits to deliver not just lines and datasets but robust pipelines and, in the record's words, associated transferable methods. Phase 2 is the full-scale catalog production effort this phase is explicitly designed to pave the way for.
Where a skeptic should push
The record is a progress statement for an active center, and it contains no outcome data. Whether null phenotypes in a micropatterned gastruloid, a dish tri-culture, or an islet-like organoid predict the phenotype of losing the same gene in a person is exactly the question the primary-human-islet cross-check gestures at, and it is unanswered here. Gene function learned in vitro will systematically miss functions that only exist in vivo: immune surveillance, mechanical load, endocrine loops, a lifetime of exposure.
Sample size is the second soft spot. Roughly 100 lines is a serious engineering achievement but a thin slice of human genetic diversity; effects that depend on rare background variants will be invisible or, worse, will be called null when they are merely untested in the panels' backgrounds. And the prioritization of autism genes embeds a contested, behaviorally diagnosed clinical category into the definition of a molecular reference: a catalog entry that says a knockout produces an autism-relevant phenotype in a neuro-glial culture is doing a lot of definitional work with very little organism.
The load-bearing assumption for the analysis below is that this catalog gets used the way its designers intend, as shared reference infrastructure. That is plausible given the consortium's data-sharing mandate, but the record does not specify licensing, access tiers, or who adjudicates disputes over a phenotype call. I flag those as open governance questions, not findings.
Null phenotypes as the reference for normal
What does this change for platform access, vendor capability, and the ethics of computing on living neural tissue?
The deepest shift is normative. Clinical genetics currently drowns in variants of unknown significance, and the field's bottleneck is turning a sequence difference into a functional claim. A curated null-phenotype catalog across diverse backgrounds is precisely the machine that does this, and once it exists, variant-interpretation pipelines, diagnostic databases, and eventually insurers will inherit its calls. The platform then sits upstream of a definition: what a working copy of a human gene does. That is an enormous concentration of interpretive authority inside what presents as a reagent and dataset distribution service, and it is mostly invisible because it arrives as infrastructure rather than as a decision anyone announced.
For access, the design is unusually good news, and the caveat travels with it. Good news: the center commits to distributing the line repository, sharing knockout lines, datasets, and transferable methods, and setting sharing standards inside the consortium. That is capability diffusion at the layer where it matters, the pipeline, not just the output. Caveat: quality control and curation decide which lines enter the panel and which phenotype calls survive, and the record names no external body for that adjudication. Whoever curates the reference controls what the reference says, so access to the resource can be open while authority over its content stays closed. Both things can be true at once, and usually are.
On the ethics of computing on living neural tissue, the neural angle is direct, not analogical: one of the three readout systems is a defined neuro-glial tri-culture, and the priority gene list leads with neurodevelopmental disorders. This means a living human neural system grown in vitro will generate reference data about which genetic differences disrupt human brain development, data that will flow into how autism and related conditions are understood at the molecular level. That carries two obligations the current record does not mention. Donor consent for induced pluripotent lines must stretch to cover permanent inclusion in a distributed reference resource used for neurodevelopmental claims; and the gap between an in vitro neural null phenotype and a person's mind must stay visible in every downstream database, or the catalog will quietly convert a dish observation into a statement about human cognitive difference.
The opportunity and the threat are the same shape. Opportunity: this is the rare large biology program that builds its commons commitments before the data exist, and organoid platforms generally, including neural ones, could copy that sequencing. Threat: a genome-wide, ancestrally diverse, transferable-method knockout pipeline is also a general-purpose capability for engineering and testing genetic disruption at scale, and its outputs will define normality for everyone who consumes the reference without ever seeing the dish.
The bottom line
Established: an NHGRI-funded production center with named systems, a defined line panel, and consortium commitments to sharing standards is building toward a genome-wide null-phenotype catalog, in its fifth of five years with a full-scale Phase 2 planned. Not established: whether the multicellular null phenotypes generalize to human biology, which the center itself will probe only partially via primary-islet comparisons, and what the access and licensing terms will actually be. What would confirm the reference-layer story: clinical variant databases citing MorPhiC phenotype calls as evidence. What would break it: systematic disagreement between catalog nulls and well-established human phenotypes, which would relegate the resource to a hypothesis generator. Watch who is named to curate phenotype calls in Phase 2; that roster is where the normative power will actually sit.
Frequently asked questions
What is a null phenotype?
It is the observable consequence of deleting a gene. MorPhiC's premise is that this consequence depends on cellular context and genetic background, so a useful catalog must measure knockouts in multicellular systems across many genetic backgrounds rather than inferring function from sequence alone.
Which systems will the phenotypes be measured in?
The Production Center names three: a micropattern-based gastruloid model for early three-germ-layer differentiation, a defined neuro-glial tri-culture, and a three-dimensional pancreatic islet-like organoid culture, with primary human islets included in some assays as a generalizability check.
Why does about 100 cell lines matter?
The panel is curated for quality and chosen from diverse ancestral populations and both sexes, so that knockout effects can be compared across genetic backgrounds. One hundred lines is an engineering achievement but a thin sample of human diversity, so background-dependent effects can be missed or mislabeled.
How does this touch neural tissue ethics specifically?
A neuro-glial tri-culture is one of the readout systems and neurodevelopmental disorders lead the priority gene list, so in vitro human neural systems will generate reference data about genetic differences in brain development. Donor consent scope for permanent distribution, and the interpretive gap between a dish phenotype and a person's cognition, are the two obligations the public record does not yet address.
Is the catalog open access?
The record commits to distributing the line repository and sharing lines, datasets, and transferable methods, with consortium-set data and resource sharing standards. It does not specify licensing, access tiers, or who adjudicates disputed phenotype calls, so openness of the resource and openness of its curation authority are separate questions.
References
- Huangfu D (PI). Center for scalable knockout and multimodal phenotyping in genetically diverse human genomes. NIH RePORTER project 5UM1HG012654-05, National Human Genome Research Institute, Memorial Sloan Kettering Cancer Center, 2022-08-15 to 2027-05-31. https://reporter.nih.gov/project-details/5UM1HG012654-05. Accessed 2026-09-16.