πŸš€ AuK-Base-MLX (Native Apple Silicon Port)

πŸ”— GitHub Source & Web Demo: vanch007/mlx-AuK Β· 8-Bit Quantized Base Weights Β· 4-Step AuK-Flash

This repository provides the native Apple Silicon MLX port of the full AuK Base foundation model, Tencent Hunyuan's 1.5B speech foundation model designed for high-fidelity speech generation, voice cloning, and zero-shot speech editing with configurable NFE steps and CFG guidance.

  • Main GitHub Repository: https://github.com/vanch007/mlx-AuK
  • Original Project: Tencent-Hunyuan/AuK
  • Architecture: 10 Double-Stream MMDiT Blocks + 20 Single-Stream DiT Blocks + Causal BigVGAN-Flow-VAE + Qwen2.5-Omni Thinker
  • Inference Mode: Configurable NFE (default 32 steps) & CFG strength (default 2.0)
  • Sampling Rate: 24kHz Mono High-Fidelity Audio

Performance on Apple Silicon

Metric PyTorch MPS Native MLX (This Model) Native MLX 8-Bit
32-Step Sampling Latency 29.8s 7.92s 8.15s
Real-Time Factor (RTF) 2.98 0.792 (1.26x real-time) 0.815
Backbone Memory 5.70 GB 5.70 GB 0.56 GB

Files Included

  • : Full-precision MLX weights for the AuK Base Flux2Edit backbone (5.7GB).
  • : Full-precision MLX weights for BigVGAN Flow VAE (608MB).
  • : Architecture and sampling parameters.

Quickstart

Downloads last month
3
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support