One pipeline from raw recording to neuromorphic chip
Brain-computer interface research has long been slowed by fragmented data formats, incompatible decoders, and hardware-specific deployment tools. BCIJelly, a new open-source ecosystem from a Chinese Academy of Sciences-led team, bolts the whole chain together: standardized datasets, a library of reusable modules, large-language-model-guided decoder design, and a single-command compile step that ports trained models onto neuromorphic silicon at a fraction of the power.
Source: BCIJelly: An integrated ecosystem for brain-computer interface research, bioRxiv, 2026. Primary source. Read: the full bioRxiv text, abstract, results, limitations, and code availability statements.
What the work claims
Han, Yang, Zheng, Yang, and colleagues, with senior authors including Mu-ming Poo, Bo Xu, and Tielin Zhang from the Institute of Neuroscience, the Institute of Automation, Lingang Laboratory, and affiliated University of Chinese Academy of Sciences schools, present BCIJelly as a unified computational ecosystem for brain-computer interface research.1 The components are concrete: 18 curated BCI datasets standardized into AI-ready form, 15 complete benchmark decoders spanning the major methodological families, a library of 80 reusable modules (fully connected, convolutional, backbone, and attention families), an automated architecture search that assembles task-specific decoders without manual design, and a hardware-aware pipeline that compiles trained decoders for neuromorphic chips. An interactive visualization platform provides code-free exploration of recordings and decoding outputs.
Three headline results anchor the validation. First, across five paradigms (motor, visual, speech, emotion, and auditory) in humans, macaques, and mice, no single decoder dominates, which the authors frame as the justification for search rather than fixed models. Second, the LLM-driven closed-loop extension of the architecture search, tested with three commercial model backends, produces decoders competitive with state-of-the-art pretrained models such as NDT1, NDT2, and POYO, and exceeds them in selected cross-species joint-training settings, while the three LLM backends do not differ significantly from one another. Third, deployment on neuromorphic hardware cuts power consumption roughly 30-fold on the commercial Lynxi chip and 50-fold on the in-house TaiBai design compared with a GPU, while preserving decoding performance.
How it works
The architecture search samples and assembles combinations from the 80-module library and evaluates them per task. On a representative classification dataset under cross-day decoding, the best of five independent searches reached a median accuracy of 0.71 across 11 held-out test days spanning 35 days; on a representative regression task it reached a median R squared of 0.80 across 13 test days spanning 88 days, with different searches converging on distinct module combinations. The LLM-driven mode replaces random sampling with a guided loop: the model reads a task specification, a description of every module, and a running history of searched architectures with their scores; it plans the search, assembles candidate decoders against a frozen shared encoder, and either commits passing designs to the history file or diagnoses failures and designs new modules. The loop stops under explicit caps: at most 3 new module designs, 25 LLM calls, 60 minutes, or 2 consecutive rounds with no gain.
The deployment path is where the engineering bites. BCIJelly fuses sequences of convolution, batch normalization, and leaky integrate-and-fire layers into single spike-convolution operators that target chips execute natively, then optimizes the computational graph placement, cutting the cores needed for one spiking model from 113 to 33, a more than threefold reduction, with no loss of accuracy. Measured on the MFSNN decoder, an NVIDIA RTX A6000 GPU drew 16.48 W versus 0.57 W on the Lynxi neuromorphic chip, about a 30-fold reduction, and an estimated 0.34 W on the TaiBai FPGA prototype, about a 50-fold reduction, with decoding accuracies of 98.41 percent on GPU, 82.54 percent on Lynxi, and 92.71 percent on TaiBai. Three further decoders deployed on Lynxi confirm the pattern: essentially identical accuracy to GPU with roughly 36-fold to 52-fold lower power, for example an LSTM at 96.10 percent accuracy on both platforms, dropping from 40.610 W to 1.033 W.
Where a skeptic should push
The most load-bearing assumption is that the power numbers represent deployable reality. They partially do not. The TaiBai results come from an FPGA prototype with power estimated from a behavioral chip simulator, because the ASIC is still in fabrication; the headline 50-fold figure is a projection, and the Lynxi figure carries a real deployment tax, with accuracy falling from 98.41 percent on GPU to 82.54 percent on chip for the same model. The authors are honest about a second gap: the LLM-searched encoder-decoder models were never deployed on neuromorphic hardware at all, because their Transformer-based shared encoder is not supported by current chip toolchains. So the full story, language model designs the decoder straight onto low-power silicon, is demonstrated only piecewise.
Third, the LLM comparison used the same shared encoder across all three backends, isolating decoder search but not testing independent full pipeline design, and multiple searches per backend were collapsed into performance summaries, so the reproducibility of a single search run is left unquantified. Fourth, all of this is validated on invasive in vivo recordings from humans, macaques, and mice; nothing here touches in vitro neural tissue, organoids, or multielectrode-array cultures, and the field knows that cross-domain decoding does not transfer for free. Finally, the visualization platform ships on Windows only. None of this invalidates the engineering; it bounds the marketing.
Benchmark ecosystems quietly govern access
Strip away the neuroscience and BCIJelly is an access play. Its real product is standardization: 18 datasets converted into AI-ready inputs, 15 decoders in one API, 80 modules with documented interfaces. Once a field's data flows through one loader and its claims are scored on one benchmark suite, the party that curates the suite sets the terms of what counts as capability. This is leaderboard governance, and it arrives without a statute, a regulator, or an appeal process. The direct read for living-tissue computing is uncomfortable: the moment organoid and culture recordings are formatted into ecosystems like this, the question "does this tissue compute?" stops being a scientific question and becomes "how does it score on the suite?", and the suite's authors will have already decided what the tasks are, how splits are drawn, and which baselines count. Platforms selling neural tissue as compute should expect their customers to grade them with someone else's benchmark, and should watch who writes it.
The LLM-in-the-loop is the second governance object, and the paper gives the template for auditing it whether the authors intended to or not. Their stopping conditions (at most 3 new module designs, 25 LLM calls, 60 minutes, 2 consecutive no-gain rounds) are an explicit, tunable verification budget, and their history file is a provenance record: every architecture the machine proposed, with its scores, is written down. That is exactly the artifact a regulator should someday demand for any system that designs closed-loop controllers attached to living tissue, where the question "why did the machine choose this decoder?" currently has no answer anywhere in the field. The threat is that the same pipeline makes it trivially cheap to generate and ship controllers nobody fully designed: the verification budget becomes the only thing standing between a search loop and a deployed intervention on neural tissue, and today that budget is a config parameter set by the vendor.
The power result is the third thread, and it is the one that most directly touches computing on living substrates. A 30-fold to 50-fold reduction in decoder power changes what can run at the tissue interface at all, because chronic interfacing with living neural tissue is governed by heat and energy budgets as much as by bandwidth; every watt dissipated near the culture is a constraint on how much living tissue can be instrumented, for how long, and at what density. Equally important is where the computation happens. An edge neuromorphic chip keeps neural data inside the rig; a GPU pipeline sends it elsewhere. BCIJelly's compile step is therefore also a data-governance decision implemented in hardware, and the cross-species joint-training modes are a second one: pooling human, macaque, and mouse recordings in a single training run quietly erases the consent and oversight seams that exist precisely because those categories are different. In vitro human neural data will enter such pools as a fourth category with no settled consent class, and it will enter through the data loader, not through a deliberative process. The open-source release cuts both ways: it democratizes access and inspection, and it hands every actor, including actors nobody would license, a single command that turns a trained model into a low-power closed-loop controller. The deployment toolchain is the chokepoint; it is where governance should sit.
The bottom line
Established: a competent, genuinely integrated open-source stack that standardizes BCI data, searches decoders competitively with pretrained models, and compiles spiking decoders onto neuromorphic hardware at large measured or estimated power savings. Not established: silicon-verified power for TaiBai, deployment of LLM-designed models on chips, reproducibility of single search runs, or any validation on in vitro neural tissue. What would confirm the platform's importance is third-party reproduction, real ASIC measurements, and an organoid or multielectrode-array culture dataset onboarded through the same pipeline. What would break its governance pretensions is the opposite: benchmark suites and deployment toolchains are already becoming infrastructure, and infrastructure that curates data, designs models, and decides where computation runs is infrastructure that governs, whether or not anyone votes on it.
Frequently asked questions
What is BCIJelly?
An open-source computational ecosystem for brain-computer interface research that integrates standardized datasets, 15 benchmark decoders, a library of 80 reusable modules, automated architecture search, LLM-guided decoder design, neuromorphic deployment, and a visualization platform.
What does the automated architecture search do?
It samples and assembles combinations from an 80-module library to build task-specific decoders without manual design. Its LLM-driven extension plans searches, diagnoses failed candidates, and designs new modules, stopping after at most 3 new designs, 25 LLM calls, 60 minutes, or 2 consecutive rounds without improvement.
How large are the power savings?
For one spiking decoder, a GPU drew 16.48 W versus 0.57 W on the Lynxi neuromorphic chip (about 30-fold less) and an estimated 0.34 W on the TaiBai FPGA prototype (about 50-fold less, with the ASIC still in fabrication). Three further decoders on Lynxi showed roughly 36-fold to 52-fold reductions with essentially unchanged accuracy.
Which species and paradigms were tested?
Five paradigms (motor, visual, speech, emotion, and auditory) using recordings from humans, macaques, and mice, including cross-species joint-training settings. No in vitro neural tissue or organoid data was used.
What are the platform's main limitations?
The headline 50-fold power figure rests on an FPGA prototype with simulator-estimated power; LLM-designed models could not be deployed on chips because their Transformer encoder is unsupported; search reproducibility across runs is unquantified; and the visualization software currently runs on Windows only.
Why does a BCI platform matter for living-tissue computing?
Because it shows where the access battle will be fought: curated benchmarks define what counts as capability, LLM search loops design the controllers with only a tunable verification budget as oversight, and single-command neuromorphic compilation plus cross-species data pooling quietly decide where neural data goes and who audits it.
References
- Han L, Yang X, Zheng T, Yang Q, Qin Y, Chen L, Wei Q, Hong B, Zhang X, Xiong R, Gu Y, Poo MM, Xu B, Li C, Zhang T. BCIJelly: An integrated ecosystem for brain-computer interface research. bioRxiv. 2026. doi:10.64898/2026.08.13.744531. https://www.biorxiv.org/content/10.64898/2026.08.13.744531. Code: github.com/LiyuanHan/BCIJelly. Accessed 2026-09-30.