πŸš€ AuK-Flash-MLX-8bit (Apple Silicon 8-Bit Quantized)

πŸ”— GitHub Source & Web Demo: vanch007/mlx-AuK Β· Full Precision (FP32/BF16) MLX Weights

This repository provides the 8-Bit affine quantized MLX port of AuK-Flash, Tencent Hunyuan's 1.5B speech foundation model for ultra-fast generation and zero-shot speech/lyric editing on Apple Silicon.

  • Main GitHub Repository: https://github.com/vanch007/mlx-AuK
  • Original Project: Tencent-Hunyuan/AuK
  • Quantization: 8-bit group-wise affine quantization ()
  • Weight Size: 0.56 GB (90.1% reduction from 5.70 GB)
  • Inference Speed: 1.022s for 10.0s audio on Apple Silicon (RTF 0.1022, ~10x real-time)

Performance Benchmark

Metric PyTorch MPS Native MLX FP32 Native MLX 8-Bit (This Model)
4-Step DiT Latency (10s Audio) 3.820s 0.992s 1.022s
Real-Time Factor (RTF) 0.3820 0.0992 0.1022 (9.79x real-time)
Backbone Memory 5.70 GB 5.70 GB 0.56 GB (90.1% saving)

Files Included

  • : 8-bit quantized MLX weights for the Flux2Edit diffusion transformer (569MB).
  • : MLX weights for the BigVGAN Flow VAE (608MB).
  • : Architecture and quantization hyperparameters.

Quickstart

For full inference workflows, CLI tools, and the interactive A/B comparison web demo, visit the project GitHub: πŸ‘‰ https://github.com/vanch007/mlx-AuK

Downloads last month
26
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support