TripoSG
Create a textured 3D model from a single image
OmniParser, turn your LLM into GUI agent
Consistency generation of portrait and subject
High-quality speech synthesis powered by Kokoro TTS
FitDiT is a high-fidelity virtual try-on model.
Detect human poses in images and videos
Compare and rank TTS voices by listening and voting
Generate 3D models from images
Upscale images with control and customization
Image to Video Generation
Transcribe audio or YouTube videos into text
Execute dynamic Python scripts from environment variables