← SSI archive

Dataset selection guide 10 additional paper and project entries

Silent speech datasets and code resources

Start with the signal you need: facial muscle activity, or images of the tongue and lips. The three data sources below include silent and voiced material, but they are not interchangeable benchmarks. Follow the source documentation before choosing an experiment.

EMG and TaL documentation checked on 5 October 2026; SSR7000 sources checked on 8 October 2026. This is a selection guide, not a claim that we downloaded or reproduced the datasets. Additional links from the silent-speech review corpus remain below; a paper, demo, or project link does not guarantee a downloadable dataset or model.

What it is
Three starting points for data-driven silent-speech experiments, followed by related paper and project links.
Who it’s for
Researchers and students looking for SSI datasets or code after a review, and readers who landed here from search and need to know this is an artifact list—not a new paper.
Verdict
Choose by input, speaking mode, and evaluation split—not by the largest headline score. Check each source’s access and reuse terms separately from the availability of a paper or demo.

Which silent speech dataset should I start with?

Facial muscle signals: Silent Speech EMG

Choose this starting point for electromyography (EMG): electrical activity recorded from facial muscles during silent and vocalized speech. The release includes raw eight-channel signals, audio, and prompt metadata. Reference samples marked with sentence_index: -1 are not ordinary utterance examples.

Dataset files and description on Zenodo · Author’s code and setup instructions

For the research task and its limits, read our reviews of Digital Voicing of Silent Speech and An Improved Model for Voicing Silent Speech. The official code has separate paths for speech-audio synthesis and direct text recognition; select the intended output first.

Tongue ultrasound and lip video: TaL

Choose the Tongue and Lips corpus for synchronized ultrasound, lip video, and audio. TaL1 records one speaker across six sessions; TaL80 records 81 speakers, each in one session. Its prompt tags distinguish audible speech (aud), silent speech (sil), and whispered speech (whi, TaL1 only). Shared prompts carry an x prefix. Do not assume every recording is silent.

Official TaL documentation, samples, and download instructions · Review: voiced versus silent recognition

SSR7000: which download works with the recognition recipe?

Choose SSR7000 for synchronized tongue ultrasound and lip video recorded without speaking aloud. The original corpus contains 7,484 utterances from one English speaker, not a multi-speaker benchmark. Read the LREC 2022 paper and corpus description for its recording conditions and evaluation split.

The author’s repository currently directs downloads to SSR7000 Lightweight Proxies Dataset v1. That release contains reduced-resolution lip and ultrasound videos; it does not include full-resolution raw data. Do not treat it as an unchanged copy of every artifact described in the original paper or README.

  1. Match the files. The release lists proxies/lip, proxies/uti, and splits/train/text/splits/valid/text. Match videos and transcripts by utt; inspect its inventory and checksums. Confirm any test split separately.
  2. Check the recipe input. The repository documents ESPnet, not ESPnet2. Its checked recipe version expects feature files named train, val and test with .ark/.scp files. Downloaded MP4 videos are not those feature files.
  3. Keep comparisons honest. Record the data version, preprocessing and split before comparing a result with the paper. This guide checks documentation, not archive contents or a successful training run. Confirm access and reuse terms at the linked sources before reuse.

Before downloading or reporting a result

  1. Identify the speaking mode. Removing microphone audio from a model’s input does not make a voiced recording an example of silent articulation.
  2. Choose a split that answers your question. Keep speaker, session, and prompt overlap explicit. TaL documentation warns that prompts overlap between TaL1 and TaL80; pooling them without checking can change what a held-out test means.
  3. Match the implementation to the paper. The EMG repository’s current model differs from earlier paper versions, and its default validation set is larger than the original EMNLP 2020 set. Record the commit and split rather than comparing unmatched scores.
  4. Inspect samples and terms first. TaL offers sample directories before its much larger core downloads. Check the current access, license, citation, and storage instructions at each official source. Code, trained weights, and data can have different conditions.

This existing list includes papers and demonstrations, not only downloadable datasets. Links are retained for discovery; their presence does not verify a reusable implementation.