--- title: Soprano TTS API emoji: 🎤 colorFrom: blue colorTo: purple sdk: docker pinned: false license: apache-2.0 app_port: 7860 --- # Soprano TTS API Space This Hugging Face Space provides a REST API for the Soprano Text-to-Speech model loaded from [Gaston895/aegis001](https://huggingface.co/Gaston895/aegis001). Soprano is an ultra-lightweight, on-device TTS model designed for expressive, high-fidelity speech synthesis. ## Model Source This space automatically loads the Soprano TTS model from the Hugging Face repository: - **Repository**: [Gaston895/aegis001](https://huggingface.co/Gaston895/aegis001) - **Model Type**: Qwen3ForCausalLM optimized for text-to-speech - **Parameters**: 80M parameters - **Architecture**: Ultra-lightweight design for CPU inference ## Features - **Ultra-fast generation**: Optimized for CPU inference on Hugging Face Spaces - **Low memory usage**: <1GB memory footprint - **High quality**: Crystal clear 32kHz audio generation - **REST API**: Easy integration with any application - **Automatic model loading**: Downloads model from HF repository on startup ## API Endpoints ### Health Check ```bash GET /health ``` ### Single Text Synthesis ```bash POST /synthesize Content-Type: application/json { "text": "Hello, this is Soprano TTS speaking!", "temperature": 0.7, "top_p": 0.9, "format": "wav" } ``` ### Batch Text Synthesis ```bash POST /batch_synthesize Content-Type: application/json { "texts": [ "First sentence to synthesize.", "Second sentence to synthesize." ], "temperature": 0.7, "top_p": 0.9 } ``` ## Parameters - **text** (required): Text to synthesize - **temperature** (optional): Controls randomness (0.1-2.0, default: 0.7) - **top_p** (optional): Controls diversity (0.1-1.0, default: 0.9) - **format** (optional): Output format - "wav" or "base64" (default: "wav") ## Response Formats ### WAV File Response Returns audio file directly for download. ### Base64 Response ```json { "success": true, "audio_base64": "UklGRiQAAABXQVZFZm10...", "sample_rate": 32000, "duration": 2.5, "text": "Input text" } ``` ## Usage Examples ### cURL ```bash # Synthesize text and save as WAV curl -X POST https://huggingface.co/spaces/Gaston895/aegis001/synthesize \ -H "Content-Type: application/json" \ -d '{"text": "Hello world!", "format": "wav"}' \ --output speech.wav # Get base64 encoded audio curl -X POST https://huggingface.co/spaces/Gaston895/aegis001/synthesize \ -H "Content-Type: application/json" \ -d '{"text": "Hello world!", "format": "base64"}' ``` ### Python ```python import requests import base64 # Synthesize speech response = requests.post( "https://huggingface.co/spaces/Gaston895/aegis001/synthesize", json={ "text": "Hello, this is Soprano TTS!", "temperature": 0.7, "format": "base64" } ) if response.status_code == 200: data = response.json() audio_data = base64.b64decode(data['audio_base64']) with open('output.wav', 'wb') as f: f.write(audio_data) ``` ### JavaScript ```javascript const synthesizeText = async (text) => { const response = await fetch('https://huggingface.co/spaces/Gaston895/aegis001/synthesize', { method: 'POST', headers: { 'Content-Type': 'application/json', }, body: JSON.stringify({ text: text, format: 'base64' }) }); const data = await response.json(); return data.audio_base64; }; ``` ## Model Information This space uses the Soprano TTS model from [Gaston895/aegis001](https://huggingface.co/Gaston895/aegis001), which features: - **Architecture**: Qwen3ForCausalLM - **Parameters**: 80M - **Context Length**: 1024 tokens - **Vocabulary Size**: 8192 - **Audio Quality**: 32kHz sampling rate - **Optimization**: CPU-optimized for Hugging Face Spaces ## Deployment The model is automatically downloaded from the Hugging Face repository when the space starts up. No manual model files are needed in the space repository. ## Limitations - English-only support - No voice cloning capabilities - Occasional mispronunciation of uncommon words - CPU-optimized for Hugging Face Spaces - Model loading time on first startup (~1-2 minutes) ## License This project is licensed under the Apache-2.0 license.