File size: 6,752 Bytes
f6f4b59 f5428ac f6f4b59 f5428ac f6f4b59 f5428ac f6f4b59 f5428ac | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 | ---
license: mit
language:
- en
pipeline_tag: automatic-speech-recognition
tags:
- speech-recognition
- whisper
- singapore-english
- sociolinguistics
- corpus-construction
- domain-adaptation
base_model:
- openai/whisper-large-v2
---
# Model Card for `whisper-large-v3-NSC`
## Model Details
### Model Description
`whisper-large-v2-NSC` is a domain-adapted automatic speech recognition (ASR) model for conversational Singapore English (SgE).
It is a fine-tuned version of OpenAI’s Whisper Large-V3 model trained on aligned conversational speech from Part 3 of the National Speech Corpus (NSC). The model is optimized to improve the recoverability of interactional, morphosyntactic, and discourse-pragmatic features characteristic of contemporary spoken Singapore English.
The model was developed as part of a workflow for constructing **YCSEP_v2 (YouTube Corpus of Singapore English Podcasts, version 2)**, demonstrating how targeted fine-tuning can improve linguistic fidelity in large-scale corpus creation.
- **Developed by:** Steven Coats, Carmelo Alessandro Basile, Cameron Morin, Robert Fuchs
- **Funded by:** EU NextGenerationEU / Research Council of Finland (grant 358720)
- **Model type:** Automatic Speech Recognition (seq2seq Transformer)
- **Language(s):** English (Singapore English)
- **License:** MIT
- **Finetuned from model:** `openai/whisper-large-v2`
### Model Sources
- **Related publication:** Coats et al. (2025), *The YouTube Corpus of Singapore English Podcasts*, English World-Wide
- **Corpus produced with the model:** YCSEP_v2
---
## Uses
### Direct Use
The model is intended for:
- Transcription of conversational Singapore English speech
- Linguistic corpus construction and annotation workflows
- Research in World Englishes, sociolinguistics, and interactional linguistics
- Recovering discourse particles and local morphosyntax often missed by general ASR systems
### Downstream Use
The model can be integrated into pipelines involving:
- Speaker diarization (e.g., WhisperX + pyannote)
- POS tagging and syntactic annotation (e.g., spaCy pipelines)
- Construction-grammar-based corpus analysis
- Speech-based sociolinguistic research
### Out-of-Scope Use
This model is **not optimized for**:
- Standard American/British broadcast speech
- Multilingual ASR outside the Singapore English domain
- Real-time or low-latency applications
- Legal, medical, or safety-critical transcription
---
## Bias, Risks, and Limitations
- The model is trained narrowly on conversational Singapore English and may underperform on other English varieties.
- Training data reflects the demographic distribution of NSC Part 3 and is not fully balanced sociolinguistically.
- Domain specialization may bias outputs toward informal conversational registers.
- The model is designed for research and corpus-building, not general-purpose ASR deployment.
### Recommendations
Users should evaluate performance carefully before applying the model outside conversational Singapore English contexts.
---
## How to Get Started with the Model
```python
from faster_whisper import WhisperModel
from huggingface_hub import snapshot_download
model_dir = snapshot_download("stcoats/whisper-large-v2-NSC")
model = WhisperModel(model_dir, device="cuda", compute_type="float16")
```
## Training Details
### Training Data
Training data was derived from **National Speech Corpus (NSC) Part 3**, consisting of same-room conversational recordings between two speakers (friends, partners, or family members).
Processing steps included:
- Alignment of WAV recordings with Praat TextGrid transcripts
- Removal of markup and non-speech symbols
- Merging adjacent utterances separated by ≤ 0.25 s pauses
- Segmentation into 10–30 second conversational chunks
- Export of aligned 16-bit PCM WAV segments with metadata (JSONL format)
**Dataset statistics:**
| Measure | Value |
|----------------|-------------|
| Speakers | 428 |
| Segments | 69,603 |
| Words | 4.57 million|
| Audio length | 458 hours |
**Data split:** 80% training / 20% evaluation
---
### Training Procedure
Six Whisper model sizes were fine-tuned; this released model corresponds to the best-performing configuration.
#### Training Hyperparameters
- Learning rate: 2.5e-6
- Weight decay: 0.02
- Warm-up steps: 300
- Effective batch size: 16
- Epochs: 8
- Logging interval: 200 steps
- Checkpoint interval: 1,000 steps
- Training regime: bf16 mixed precision
---
### Compute Infrastructure
Training was conducted on the **LUMI supercomputer (CSC Finland)** using:
- 16 × AMD Instinct MI250X GPUs (128 GB memory each)
---
## Evaluation
### Testing Data
Two evaluation datasets were used:
- MERaLiON NSC-derived test set
- Random 1,000-segment sample from held-out NSC data
### Metrics
- **WER (Word Error Rate)**
- **CER (Character Error Rate)**
These metrics assess transcription fidelity for conversational speech.
### Results
Fine-tuning produced substantial gains across all Whisper variants.
The best fine-tuned model slightly outperformed MERaLiON-2-10B-ASR despite being substantially smaller.
**Example comparison (NSC sample):**
| Model | WER | CER |
|-----------------------------|--------|--------|
| whisper-large-v3 (baseline) | 0.3302 | 0.2452 |
| whisper-large-v2-ft | 0.2245 | 0.1599 |
| MERaLiON-2-10B-ASR | 0.2553 | 0.1844 |
---
## Model Examination
The model improves recoverability of interactional and morphosyntactic structure, enabling more reliable extraction of constructions and discourse particles in YCSEP_v2 and related linguistic analyses.
---
## Environmental Impact
Training used shared HPC infrastructure (LUMI), which operates with a high proportion of renewable energy.
- **Hardware:** AMD MI250X GPU cluster
- **Training duration:** 8 epochs across 16 GPUs
- **Compute region:** Finland (CSC)
---
## Technical Specifications
### Architecture and Objective
Sequence-to-sequence Transformer ASR model based on the Whisper architecture, optimized through supervised fine-tuning on aligned conversational speech.
### Software
- PyTorch
- Hugging Face Transformers
- WhisperX + pyannote (for downstream corpus creation)
- spaCy 3.8 (linguistic annotation)
---
## Citation
**BibTeX**
```bibtex
@article{coats2025ycsep,
author = {Coats, Steven and Basile, Carmelo Alessandro and Morin, Cameron and Fuchs, Robert},
title = {The YouTube Corpus of Singapore English Podcasts},
journal = {English World-Wide},
year = {2025},
volume = {46},
number = {3},
pages = {274--298},
doi = {10.1075/eww.25018.coa}
} |