StyleTTS2 (LJSpeech)

yl4579's StyleTTS2 LJSpeech checkpoint, exported for loom.cpp. Takes phoneme ids, not text.

This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.

Original model

Exported from yl4579/StyleTTS2-LJSpeech. Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format.

License

mit, inherited from the base model above.

Language(s)

en

the HF repo carries no license:/language: tags; MIT per the upstream GitHub repo's LICENSE (github.com/yl4579/StyleTTS2)

Usage

Run it with loom-py -- loom-py-rt on PyPI:

pip install -U "loom-py-rt[hub,phonemes]"

NOTE: This is a work in progress. For now, in order to avoid license conflicts and keep dependencies at a minimum, we opted for orthography2ipa as our "swiss-knife" phonemizer. For deep-orthography languages like English, to get stressing rules and context-based phonemization, the phonemizer must register a reference lexicon (e.g., ipa-dict) to get highly accurate phonemization. Full quality can be achieved by phonemizing the text yourself using your engine of choice and feeding it to the model via the argument phonemes as in the example below.

import loom

model = loom.Model.from_pretrained("loom-ai-org/styletts2-ljspeech-loom")

# styletts2-ljspeech is trained on phonemes. Its symbol table ships in the GGUF, so the only piece that is not in
# the file is grapheme-to-phoneme -- a property of the language rather than of this checkpoint, which
# is why it is the `phonemes` extra above rather than part of the model.

# THE FULL-QUALITY PATH: phonemes you produced yourself, with whatever G2P you trust. The symbol table
# in the GGUF is what encodes them, so anything that emits IPA works.
audio = model.text2speech.infer(phonemes="həˈloʊ wˈɜːld", sample_rate=24000)
audio.save("out.wav")

# THE BUILT-IN PATH: text straight in, phonemized by the bundled rule-based G2P. Good enough for
# shallow orthographies; for English see the note above, and give it a lexicon so it has stress and
# real vowels to work with -- "time" is /tɪm/ without one.
#
# open-dict-data/ipa-dict (MIT) publishes ~65k-entry wordlists WITH stress for en_UK and en_US, in
# almost the right shape: its IPA is wrapped in slashes and a rare entry carries two comma-separated
# variants, both of which the loader rejects. One line converts a downloaded data/en_UK.txt:
#     sed 's:/::g; s/\t\([^,]*\),.*/\t\1/' en_UK.txt > en_UK.tsv
loom.phonemizers.set_lexicon("en_UK.tsv")    # a path, an http(s):// URL, or hf://<repo>/<path>

# sample_rate=24000: this checkpoint does not carry its own rate, so it is a value you have to
# know from the model's documentation and pass. It is used only if the GGUF declares none; a wrong rate
# does not fail, it plays the voice at the wrong speed.
audio = model.text2speech.infer("hello world", sample_rate=24000)
audio.save("out.wav")

The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...) passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name.

model.driver_source prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See loom-py for the API and loom.cpp for what the engine does between the two.

Files

  • styletts2-ljspeech.gguf -- the model, exported with loom-exporter.
Downloads last month
102
GGUF
Model size
0.1B params
Architecture
loom-loom-styletts2-mil
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for loom-ai-org/styletts2-ljspeech-loom

Quantized
(3)
this model

Collection including loom-ai-org/styletts2-ljspeech-loom