| --- |
| language: |
| - vi |
| license: mit |
| pipeline_tag: text2text-generation |
| tags: |
| - vit5 |
| - asr-post-processing |
| - error-correction |
| - phonetic-rescoring |
| - vietnamese-speech |
| --- |
| |
| # Vietnamese Phonetic-Weighted ECLM (Error Correction Language Model) |
|
|
| Mô hình **ViT5-Base** được huấn luyện tăng cường bằng cơ chế **Phonetic-Weighted Loss (Scale x3.5)** nhằm nắn chỉnh và chấm điểm các lỗi phát âm/đồng âm đặc thù trong tiếng Việt sau bước nhận dạng giọng nói ASR. |
|
|
| ## Đặc điểm nổi bật |
| - **Phân rã âm vị học toàn diện (Unicode NFD):** Tự động bao phủ các cặp âm đầu (`tr/ch`, `s/x`, `d/gi/r`, `l/n`), âm cuối (`c/t`, `n/ng`, `ch/t`), vần (`i/y`, `ươn/ương`) và thanh điệu (`hỏi/ngã`). |
| - **Phonetic Syllable Gating:** Thiết kế đặc thù cho bài toán Rescoring / Post-processing ASR, đảm bảo không gây suy giảm cấu trúc câu đúng (Zero Over-correction). |
|
|
| ## Cách sử dụng (Inference) |
|
|
| ```python |
| import torch |
| from transformers import AutoModelForSeq2SeqLM, AutoTokenizer |
| |
| model_name = "Toan2408/vit5-vietnamese-phonetic-eclm" |
| device = "cuda" if torch.cuda.is_available() else "cpu" |
| |
| tokenizer = AutoTokenizer.from_pretrained(model_name) |
| model = AutoModelForSeq2SeqLM.from_pretrained(model_name).to(device) |
| |
| def correct_text(text: str) -> str: |
| input_text = "sửa lỗi chính tả: " + text.strip().lower() |
| inputs = tokenizer(input_text, return_tensors="pt", max_length=64, truncation=True).to(device) |
| |
| with torch.no_grad(): |
| outputs = model.generate(**inputs, max_new_tokens=64, num_beams=3) |
| return tokenizer.decode(outputs[0], skip_special_tokens=True) |
| |
| # Ví dụ |
| raw_asr_text = "có thể giúp ca quên đi những nỗi u phiền" |
| print("Đầu ra ECLM:", correct_text(raw_asr_text)) |
| |