1. Sign In

Voices Obscured in Complex Environmental Settings (VOiCES)

VOiCES is a speech corpus recorded in acoustically challenging settings, using distant microphone recording. Speech was recorded in real rooms with various acoustic features (reverb, echo, HVAC systems, outside noise, etc.). Adversarial noise, either television, music, or babble, was concurrently played with clean speech. Data was recorded using multiple microphones strategically placed throughout the room. The corpus includes audio recordings, orthographic transcriptions, and speaker labels.

The Data

ARN: arn:aws:s3:::lab41openaudiocorpus
Region: us-east-1

Tutorials

Tags

aws-pds machine learning automatic speech recognition speaker identification denoising speech processing