The primate benchmark organoid computing does not have
Systems neuroscience just got a 50-fold larger dataset for modeling the primate dorsal visual stream: 2,244 neurons recorded while two rhesus macaques watched thousands of short natural videos. The numbers are solid, the limitations are candidly listed, and the dataset's real significance for organoid intelligence is what its existence reveals about who governs claims made on living neural tissue.
Source: STSBench: A Large-Scale Dataset for Modeling Neuronal Activity in the Dorsal Stream of Primate Visual Cortex, Trepka et al., arXiv:2607.15631 [q-bio.NC], submitted 17 July 2026. Primary source. Read: the full arXiv HTML version, retrieved 2026-09-14.
What the work claims
This is a dataset-and-benchmarks paper, a methods contribution rather than a new biological result, and it should be read as one. The claim has two parts. First, scale: the authors release single-neuron recordings from 2,244 neurons in the superior temporal sulcus (STS), in areas MT and MST, collected across eight sessions from two male rhesus macaques while the animals viewed roughly 4,500 unique natural videos. They state this is nearly a 50-fold increase over the largest existing dorsal-stream dataset, which contained 45 neurons from area MT1. Second, utility: the dataset is shown to work as a benchmark, supporting encoding models that predict single-neuron firing rates from video, and a reconstruction model that generates the viewed frame from population activity.
The dorsal-versus-ventral framing matters. The ventral stream, which supports object recognition, already has multiple large benchmark datasets and a mature culture of convolutional encoding models. The dorsal stream, which supports motion and spatial processing, had been benchmark-poor; the paper's argument is that the field's theory of dorsal processing is underdeveloped because its data is, and that releasing the data fixes the bottleneck.
How it works
Recordings used a Neuropixels 1.0 NHP probe, a high-density extracellular electrode with 384 recording contacts, targeted to the STS with anatomical MRI guidance and confirmed by the functional properties of the neurons. All procedures were approved by the Stanford IACUC. In the main task, monkeys fixated a central spot and were shown sequences of full-screen videos, each 200 milliseconds long, five frames at 24 frames per second, drawn from a pool of 4,533 clips (4,493 used for model training, 40 held out for testing), receiving juice for maintaining fixation1.
On the encoding side, the authors benchmark feature extractors that are hand-tuned, pretrained, or trained end-to-end, feeding a readout that predicts each neuron's z-scored firing rate. End-to-end 3D convolutional networks beat both a 3D ResNet pretrained on a self-motion estimation task and a hand-tuned 3D Gabor filter pyramid; performance improved with depth and plateaued after five layers, and was better for regular-spiking (putatively excitatory) than fast-spiking (putatively inhibitory) neurons, a difference the authors report with a two-sample t-test at p below 0.00001. Inspection of the first-layer filters found 10 of 32 resembling drifting Gabor patches, consistent with textbook MT models, plus chromatic structure the textbook model does not predict1.
On the reconstruction side, a neural-conditional latent diffusion model generates the first frame of a viewed video from population spiking activity. Reconstructions from dorsal-stream activity capture low-spatial-frequency luminance contours, the edges of desks and shelves, while high-frequency object identity is absent, matching the dorsal stream's expected where-not-what coding. Quantitatively the diffusion model scores 14.16 PSNR and 0.668 LPIPS against mean-image and shuffled nulls of 11.34/0.873 and 9.85/0.751 on the dorsal stream; a ventral-stream V4 reference from a companion dataset shows the reverse profile, with better perceptual fidelity but worse pixel fidelity1.
Where a skeptic should push
The most load-bearing assumption is the neuron count itself. The 2,244 figure comes from Kilosort 4.0, an automated spike-sorting pipeline that the authors themselves note can split one neuron's spikes into two (inflating counts) or merge two neurons into one. No manual curation audit is reported, and no cross-session stability check that would bound the error. Given that the paper's headline is a 50-fold scale increase, an unknown fraction of that scale could be sorting artifact; the benchmark remains useful either way, but the count should be read as an upper estimate1.
Second, breadth: two animals, eight sessions, one brain region, one highly constrained behavior. The authors are direct about anatomical contamination (some superficial neurons in early sessions may be area 7a or white matter) and about cell-type classification being crude. Third, the stimuli are naturalistic in content but unnatural in structure: 200 millisecond clips with no inter-stimulus interval, passively viewed under fixation. That is a thin slice of vision, and encoding models fit to it inherit its constraints. Fourth, the reconstruction results, while above nulls, are modest in absolute terms; the ventral comparison is explicitly caveated by the authors because neuron counts, preprocessing (spike-sorted versus threshold-crossing), and stimuli all differ, so no clean dorsal-versus-ventral decoding claim survives. The demonstrated core is a well-documented dataset with working baselines; the asserted parts are what it will eventually explain about dorsal stream theory.
Benchmarks as governance for neural compute
The non-obvious implication for platform access and vendor capability is that this paper is really about who owns the yardstick. Before STSBench, the largest dorsal-stream dataset was 45 neurons; after it, the reference standard for a whole subfield is whatever one Stanford lab chose to record, preprocess, host, and metricize. The paper's ventral-stream examples (BrainScore, MacaqueITBench, TVSD) show the pattern generalizes: once a benchmark exists, capability claims in that space are made in its vocabulary or not made at all. Organoid intelligence and neural organoid platforms currently have no equivalent shared benchmark for the only question that matters commercially and scientifically: what is this living tissue actually computing, and how would we know? In that vacuum, vendor capability claims are self-graded, and the first party to ship a credible organoid benchmark will set the terms under which every other platform is evaluated. The opportunity is enormous; so is the authority, and nothing in the current landscape decides who is allowed to hold it.
The access model deserves equal attention. A federally funded primate recording dataset, produced under NIH support, is distributed through Kaggle, a commercial machine-learning platform, with code on GitHub. The evaluative commons now runs on vendor terms of service; access to the reference standard is only as durable as a commercial platform's business model. Anyone building governance for living-tissue computing should notice that the infrastructure layer of open evaluation has quietly become platform-dependent.
On ethics, the paper is the incumbent against which organoid substitution claims will be judged, and the comparison is less flattering than the substitution narrative assumes. The ground truth of primate systems neuroscience is two macaques with implanted probes, the moral cost externalized onto animal subjects and normalized as the price of a leaderboard. Organoid computing promises to replace exactly this, but the authors' own caveats show why substitution is hard: their dorsal and ventral datasets cannot be directly compared even within the same paper, because preprocessing, stimuli, and counts differ. If organoid benchmarks inherit primate eval culture without solving comparability, the moral-status advantage becomes marketing rather than measurement. And the reconstruction result points at a sharper dual-use concern: the field is actively improving its ability to read out what neural tissue was exposed to. Optimizing donor-derived neural tissue for decodability, in a benchmark culture that rewards exactly that, is a consent question nobody's benchmark currently asks.
The bottom line
Established: a documented, accessible dataset of 2,244 (likely somewhat fewer, pending sorting-error audits) primate STS neurons with working encoding and reconstruction baselines, roughly 50 times prior dorsal-stream data, honestly caveated. Hypothesis: that this unblocks dorsal-stream theory; plausible but unproven. What would confirm it: independent labs re-using the data and converging on models; a sorting-error audit firming the neuron count; replication in more animals. What would break it: finding that a large fraction of units are sorting artifacts, or that the encoding gains do not transfer beyond these two animals. For organoid platforms the takeaway is structural, not biological: evaluation infrastructure is governance infrastructure, the primate field has it and the organoid field does not, and the window in which the first credible organoid benchmark could be community-owned rather than vendor-owned is still open, but not for long.
Frequently asked questions
What is STSBench?
A dataset of single-neuron recordings from 2,244 neurons in the superior temporal sulcus of two rhesus macaques, collected with Neuropixels probes while the animals viewed about 4,500 unique short natural videos, released by Stanford researchers in July 2026.
Why is a 50-fold increase significant?
The largest prior dorsal-stream dataset contained 45 neurons from area MT. The new dataset is nearly 50 times larger, and its authors argue this scale difference, not a lack of ideas, is why dorsal-stream models lag ventral-stream models.
Can the recordings reconstruct what the monkey saw?
Partially. A diffusion model conditioned on population activity reconstructs low-spatial-frequency contours such as the edges of furniture, but not object identity, consistent with dorsal stream coding of motion and spatial layout. Absolute fidelity is modest, well above null models but far from veridical.
How reliable is the 2,244 neuron count?
Uncertain at the margins. The count comes from the automated Kilosort 4.0 pipeline, which the authors note can split or merge neurons, and no manual audit is reported. The dataset is useful regardless, but the count should be treated as an upper estimate.
What does this have to do with organoid platforms?
Benchmarks define the vocabulary in which capability claims are made. Primate neuroscience has shared benchmarks; organoid computing has none, so vendor claims are currently self-graded. The first credible shared organoid benchmark will effectively govern the field's evaluation standards.
Where is the dataset hosted?
On Kaggle, a commercial machine-learning platform, with accompanying code on GitHub. Access to this publicly funded reference dataset therefore depends on commercial platform terms of service.
References
- Trepka EB, Xia R, Zhu S, Saleki S, Abreu Lopes D, NiƱo Cital SJ, Willeke KF, Kim M, Moore T. STSBench: A Large-Scale Dataset for Modeling Neuronal Activity in the Dorsal Stream of Primate Visual Cortex. arXiv:2607.15631 [q-bio.NC]. 2026. https://arxiv.org/abs/2607.15631. Accessed 2026-09-14.