Reward-modulated plasticity as an organoid training protocol
A Johns Hopkins team has NSF funding to train brain organoids through reward-modulated spike-time-dependent plasticity inside shell micro-electro-fluidic arrays, with the explicit goal of teaching them Atari-like games and robot control. The engineering plan is specific, but the most consequential finding already published is from the project's embedded neuroethics team: the more consciousness Americans attribute to a biocomputer, the more they support researching it.
Source: EFRI BEGIN OI: Improving Learning of Embodied Organoid Intelligence Through Reward-Based Training, NSF award 2515214, awarded 15 August 2025; and Boyd, J. L. et al., "Ethical concerns about embodied brain organoids shaped by foundational distinctions and perceptions of consciousness", Scientific Reports 16, 2026. Primary source. Read: the full NSF award record via the NSF API and the full text of the Scientific Reports article via its DOI.
What the work claims
The grant is a research plan, not a result. Thomas Hartung at Johns Hopkins, with co-investigators David Gracias, Lena Smirnova, Erik Johnson and Jonathan Boyd, holds 1,999,508 dollars in fiscal 2025 plus a 2026 increment, running from August 2025 to July 2029 under NSF's Emerging Frontiers in Research and Innovation office.1 The proposal is to build a complete training stack for embodied organoid intelligence: three-tier brain assembloids that couple cortical, dopaminergic and striatal regions inside self-folded shell micro-electro-fluidic arrays, or SMEFAs; closed-loop software that maps observations to stimulation codes and decodes action signals from high-density recordings; and a curriculum of Atari-like tasks and embodied robot control to benchmark the system against deep-reinforcement-learning baselines.
The central hypothesis is that reward-modulated spike-time-dependent plasticity, R-STDP, will enable data-efficient, continual reinforcement learning in living neural networks.1 R-STDP is the biological rule that strengthens or weakens a synapse depending on whether a spike shortly before or after a reward signal preceded a postsynaptic spike. In conventional artificial reinforcement learning, the reward is computed externally and back-propagated through a static graph; here the reward would be delivered as a patterned electrical or chemical signal to the tissue itself, and the network would reshape its own connectivity.
The companion ethics paper, published in Scientific Reports in 2026, reports an experimental-bioethics survey of American attitudes toward biocomputers.2 The headline claim is that public support for biocomputer research increases as respondents attribute more consciousness to the systems, a pattern that runs against the standard philosophical assumption that consciousness should constrain permissible research.
How it works
The hardware is the SMEFA. It is a self-folded shell micro-electro-fluidic array: a curved electrode and microfluidic shell that wraps around a brain organoid or assembloid rather than sitting beneath it on a flat plate. The proposal says this geometry can deliver millisecond-precision electrical stimulation and micromolar-resolution neuromodulator gradients while recording three-dimensional neural activity.1 Compared with a conventional planar microelectrode array, the shell should improve contact with curved tissue and allow stimulation and chemical delivery from multiple directions, though at the cost of more complex fabrication and less optical access.
The biological substrate is a three-tier assembloid. Cortical, dopaminergic and striatal organoids are fused so that the reward pathway circuits thought to support reinforcement learning are present from the start. The dopaminergic tier provides the neuromodulatory reward signal; the striatal tier handles action selection; and the cortical tier supplies the sensory-motor integration. Whether these three regions actually wire together functionally in a dish, and whether the resulting circuit resembles the mammalian basal-ganglia loop closely enough to support R-STDP, is the biological load-bearing assumption.
The training loop is closed through software. Observations from a virtual environment or robot sensors are converted to stimulation patterns; the organoid's recorded activity is decoded as an action; the action changes the environment; and a reward signal is fed back to the tissue. The benchmark suite includes Atari-like tasks and small-robot navigation, with comparison to deep-reinforcement-learning agents on training efficiency, power consumption and task performance.1
The ethics component is threaded through the project rather than bolted on. Thread 3 proposes an experimental neuroethics program that defines measurable capacities, sentience, agency and evaluative cognition, and implements tiered safeguards if any of those capacities is detected.1 The published survey is part of that program.
Where a skeptic should push
Start with the central hypothesis. R-STDP is a well-documented synaptic learning rule in biological preparations, but moving from "synapses can change in response to reward timing" to "a brain organoid can learn to play a video game" requires an enormous leap in scale, stability and read-write bandwidth. No published result from this team, or from the field, has shown sustained reward-driven task learning in a 3D human brain organoid. The NSF record lists two papers associated with the award: a methods paper on 3D neuromodulation with shell MEAs, and the ethics survey.1 Neither reports task learning. A reader should treat the learning claim as a hypothesis until a preprint or paper demonstrates it.
The survey result is real but narrow. The validation sample is N=1,238 respondents, nationally representative of the United States, collected in November 2025 with demographic weights based on the 2023 American Community Survey five-year estimates.2 The effect is clear: consciousness attribution correlates positively with support for biocomputer research, and the correlation survives controls for application area and described cognitive attribute. But a cross-sectional survey is not a policy mandate. It tells us what Americans think after reading a short vignette, not what protections the systems will need, and not what the same respondents would think if asked about specific harms or specific uses.
The most load-bearing assumption in the whole plan is long-term stability. The SMEFA and the assembloid must remain healthy and electrically addressable for weeks to months while being repeatedly stimulated and dosed with neuromodulators. Brain organoids are notorious for necrotic cores, batch-to-batch variability and drift in firing properties. Thread 1 of the proposal is devoted to standardized culture protocols and SMEFA stability, which is an honest acknowledgement that the problem is unsolved.
What reward-modulated learning means for organoid platforms
For platform access, this project is an attempt to define a reference stack. A complete embodied-organoid-intelligence platform needs four things: a living neural substrate, a bidirectional neural interface, closed-loop training software, and a benchmark curriculum. The grant proposes all four and promises open-source hardware and software plus citizen-science courses.1 If the team delivers, the stack becomes reusable across laboratories, lowering the barrier to entry in exactly the way that open-source machine-learning frameworks lowered the barrier to artificial neural networks. The risk is the opposite: if the stack is tied to a specific assembloid recipe or a custom-fabricated SMEFA that few labs can replicate, the platform advantage evaporates and the work becomes a boutique demonstration.
For vendor capability, the SMEFA is the component most likely to become a product category. Planar MEAs are already commodified; a 3D shell electrode with integrated microfluidics is not. The proposal describes a device that records, stimulates and delivers chemicals around a curved tissue volume, which is a plausible next-generation neural interface for organoids and, eventually, for ex vivo tissue slices. Vendors should watch whether the team can publish performance specifications: channel count, impedance, stimulation voltage range, microfluidic flow rates and long-term viability metrics. Without those numbers, the SMEFA remains a concept, not a catalog item.
The governance implication is the most important and the most unsettling. Conventional bioethics has assumed that evidence of consciousness in a brain organoid would trigger stronger protections and, by extension, public opposition to certain kinds of research. The survey suggests the public does not follow that script. Respondents who attributed higher consciousness to biocomputers also reported higher support for the research, higher perceived benefits relative to risks, and greater willingness to see the technology commercialized.2 The authors offer two interpretations: the public may not share the ethicists' intuition that consciousness constrains research, or the public may share the intuition but weigh medical and computational benefits more heavily.
Either interpretation is bad news for governance-by-polling. If consciousness does not reduce public support, then public attitudes cannot be relied on to apply brakes as organoid capabilities advance. If support rises because benefits are salient, then the same public may be unprepared for the harder cases, such as organoids that show signs of suffering, or organoids whose training involves punishment-like negative feedback, or organoids derived from donors who did not anticipate their tissue being turned into a game-playing controller. The project's own tiered-safeguard framework is a more robust response than polling, because it is triggered by measurable capacities rather than by public mood.
The genuine threat is regulatory capture by optimism. A field in which consciousness attribution increases support is a field in which hype about benefits can outrun evidence of safety. The genuine opportunity is that a transparent, capacity-based safeguard system, published as open-source protocols alongside open-source hardware, could become a de facto governance standard before any regulator writes a rule.
The bottom line
Established: the Hartung team has a detailed, four-year plan to train brain assembloids through R-STDP inside SMEFAs, and its ethics team has published a nationally representative US survey showing that public support for biocomputer research rises, not falls, with attributed consciousness.
Hypothesis, not established: that R-STDP in a three-tier assembloid can support data-efficient reinforcement learning on Atari-like or robot tasks, or that the SMEFA can remain stable long enough to make such training meaningful. No task-learning results have been published from this award.
What would confirm the technical claim: a preprint or paper showing that an organoid trained through the SMEFA improves its task performance with reward history in a way that cannot be explained by simple adaptation or drift, with replication across organoid batches. What would break it: failure to maintain viable, electrically addressable assembloids for the weeks required by the curriculum, or equivalence of trained and untrained controls on the benchmarks. On the governance side, what would overturn the reading is evidence from other countries or more detailed questions that the public does, in fact, draw a hard line once specific harms are described.
Frequently asked questions
What is R-STDP?
Reward-modulated spike-time-dependent plasticity is a biological learning rule in which the strength of a synapse is adjusted based on the timing of pre- and postsynaptic spikes and an external reward or neuromodulatory signal. It is one candidate mechanism for how biological networks learn from feedback.
What is a SMEFA?
A shell micro-electro-fluidic array is a curved, self-folded device that wraps around a brain organoid to record neural activity, deliver electrical stimulation, and apply chemical or neuromodulator gradients from multiple directions.
Has the team actually taught an organoid to play a game?
Not in any published work associated with this award as of August 2026. The game-playing and robot-control claims are part of the funded research plan, not demonstrated results.
What did the public-attitudes survey find?
In a nationally representative US sample of 1,238 respondents, people who attributed more consciousness to biocomputers were more supportive of biocomputer research, more likely to say benefits outweigh risks, and more open to commercialization. This ran counter to the standard bioethical assumption that consciousness would reduce support.
Why does this matter for governance?
If public support rises with perceived consciousness, then public opinion cannot be counted on to slow research as organoids become more capable. Governance needs capacity-based safeguards and transparent protocols rather than reliance on polling.
What would make this a reusable platform?
Open-source release of the SMEFA design, the closed-loop training software, the assembloid protocols and the benchmark tasks, together with reproducible performance specifications, would turn the project into a platform others can adopt.
References
- Hartung, T. et al. EFRI BEGIN OI: Improving Learning of Embodied Organoid Intelligence Through Reward-Based Training. NSF award 2515214. 2025. https://www.nsf.gov/awardsearch/showAward?AWD_ID=2515214. Accessed 2026-08-30 via the NSF award API.
- Boyd, J. L., Jensen, E. A., Jensen, A. M. and Lipshitz, N. Ethical concerns about embodied brain organoids shaped by foundational distinctions and perceptions of consciousness. Scientific Reports. 16, 10687623 (2026). https://doi.org/10.1038/s41598-026-43243-y. Accessed 2026-08-30.