Instructions to use vanch007/AuK-Base-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use vanch007/AuK-Base-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir AuK-Base-MLX vanch007/AuK-Base-MLX
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
π AuK-Base-MLX (Native Apple Silicon Port)
π GitHub Source & Web Demo: vanch007/mlx-AuK Β· 8-Bit Quantized Base Weights Β· 4-Step AuK-Flash
This repository provides the native Apple Silicon MLX port of the full AuK Base foundation model, Tencent Hunyuan's 1.5B speech foundation model designed for high-fidelity speech generation, voice cloning, and zero-shot speech editing with configurable NFE steps and CFG guidance.
- Main GitHub Repository: https://github.com/vanch007/mlx-AuK
- Original Project: Tencent-Hunyuan/AuK
- Architecture: 10 Double-Stream MMDiT Blocks + 20 Single-Stream DiT Blocks + Causal BigVGAN-Flow-VAE + Qwen2.5-Omni Thinker
- Inference Mode: Configurable NFE (default 32 steps) & CFG strength (default 2.0)
- Sampling Rate: 24kHz Mono High-Fidelity Audio
Performance on Apple Silicon
| Metric | PyTorch MPS | Native MLX (This Model) | Native MLX 8-Bit |
|---|---|---|---|
| 32-Step Sampling Latency | 29.8s | 7.92s | 8.15s |
| Real-Time Factor (RTF) | 2.98 | 0.792 (1.26x real-time) | 0.815 |
| Backbone Memory | 5.70 GB | 5.70 GB | 0.56 GB |
Files Included
- : Full-precision MLX weights for the AuK Base Flux2Edit backbone (5.7GB).
- : Full-precision MLX weights for BigVGAN Flow VAE (608MB).
- : Architecture and sampling parameters.
Quickstart
- Downloads last month
- 3
Hardware compatibility
Log In to add your hardware
Quantized