Paper List

Computational Neuroscience

Translating Measures onto Mechanisms: The Cognitive Relevance of Higher-Order Information

2025-12-02

This review addresses the core challenge of translating abstract higher-order information theory metrics (e.g., synergy, redundancy) into defensible, ...
Artificial Intelligence

Emergent Bayesian Behaviour and Optimal Cue Combination in LLMs

2025-12-02

This paper addresses the critical gap in understanding whether LLMs spontaneously develop human-like Bayesian strategies for processing uncertain info...
Bioinformatics

Vessel Network Topology in Molecular Communication: Insights from Experiments and Theory

2025-12-02

This work addresses the critical lack of experimentally validated channel models for molecular communication within complex vessel networks, which is ...
Biophysics

Modulation of DNA rheology by a transcription factor that forms aging microgels

2025-12-02

This work addresses the fundamental question of how the transcription factor NANOG, essential for embryonic stem cell pluripotency, physically regulat...
Systems Biology

Imperfect molecular detection renormalizes apparent kinetic rates in stochastic gene regulatory networks

2025-12-02

This paper addresses the core challenge of distinguishing genuine stochastic dynamics of gene regulatory networks from artifacts introduced by imperfe...
Bioinformatics

PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-Cancer

2025-12-02

This paper addresses the dual challenge of achieving computational efficiency without sacrificing accuracy in whole-transcriptome single-cell represen...
Mathematical Biology

Beyond Bayesian Inference: The Correlation Integral Likelihood Framework and Gradient Flow Methods for Deterministic Sampling

2025-12-02

This paper addresses the core challenge of calibrating complex biological models (e.g., PDEs, agent-based models) with incomplete, noisy, or heterogen...
Bioinformatics

Contrastive Deep Learning for Variant Detection in Wastewater Genomic Sequencing

2025-12-02

This paper addresses the core challenge of detecting viral variants in wastewater sequencing data without reference genomes or labeled annotations, ov...

14 / 18

期刊: ArXiv Preprint

发布日期: 2026-03-13

Computational NeuroscienceBioinformatics

Towards unified brain-to-text decoding across speech production and perception

Zhejiang University | Chinese Academy of Sciences | Huashan Hospital, Fudan University

Yang Yang, Meng Li, Zhizhang Yuan, Gaorui Zhang, Baowen Cheng, Zehan Wu, Yuhao Xu, Xiaoying Liu, Liang Chen, Ying Mao

30秒速读

IN SHORT: This paper addresses the core challenge of developing a unified brain-to-text decoding framework that works across both speech production and perception modalities for Mandarin Chinese, overcoming limitations of single-modality approaches and alphabetic language systems.

核心创新

Methodology First unified brain-to-sentence decoding framework for both speech production and perception in Mandarin Chinese, enabling direct comparison of neural dynamics across modalities.
Methodology Three-stage post-training and two-stage inference framework for 7B-parameter LLM that outperforms larger commercial LLMs (hundreds of billions of parameters) in mapping toneless Pinyin syllables to Chinese sentences.
Biology Revealed neural characteristics of Mandarin speech: production engages broader cortical regions than perception; shared channels show similar patterns with perception delayed by ~106.5ms; comparable decoding performance across hemispheres.

主要结论

Achieved best-case Chinese character error rates of 14.71% for spoken sentences and 21.80% for heard sentences across 12 participants with depth electrodes (mean speaking CER = 31.52%, mean listening CER = 37.28%).
NeuroSketch (2D-CNN) achieved mean initial/final accuracies of 59.54%/50.17% for speaking and 58.92%/48.05% for listening, representing 394.9%/412.0% and 389.7%/406.6% improvements over chance respectively.
Speech production involved neural responses across broader cortical regions than auditory perception (p<0.05), with perception showing consistent temporal delay relative to production (mean = -106.5ms, 90% CI [-249.4, 23.05]).

研究空白： Previous brain-to-text decoding studies have largely focused on single modalities (either production OR perception) and alphabetic languages, with limited work on logosyllabic languages like Mandarin Chinese that present unique challenges due to character-syllable ambiguity.

摘要: Speech production and perception constitute two fundamental and distinct modes of human communication. Prior brain-to-text decoding studies have largely focused on a single modality and alphabetic languages. Here, we present a unified brain-to-sentence decoding framework for both speech production and perception in Mandarin Chinese. The framework exhibits strong generalization ability, enabling sentence-level decoding when trained only on single-character data and supporting characters and syllables unseen during training. In addition, it allows direct and controlled comparison of neural dynamics across modalities. We collected neural data from 12 participants implanted with depth electrodes and achieved full-sentence decoding across multiple participants, with best-case Chinese character error rates of 14.71% for spoken sentences and 21.80% for heard sentences. Mandarin speech is decoded by first classifying syllable components in Hanyu Pinyin, namely initials and finals, from neural signals, followed by a post-trained large language model (LLM) that maps sequences of toneless Pinyin syllables to Chinese sentences. To enhance LLM decoding, we designed a three-stage post-training and two-stage inference framework based on a 7-billion-parameter LLM, achieving overall performance that exceeds larger commercial LLMs with hundreds of billions of parameters or more. In addition, several characteristics were observed in Mandarin speech production and perception: speech production involved neural responses across broader cortical regions than auditory perception; channels responsive to both modalities exhibited similar activity patterns, with speech perception showing a temporal delay relative to production; and decoding performance was broadly comparable across hemispheres. Our work not only establishes the feasibility of a unified decoding framework but also provides insights into the neural characteristics of Mandarin speech production and perception. These advances contribute to brain-to-text decoding in logosyllabic languages and pave the way toward neural language decoding systems supporting multiple modalities.