MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
Paper โข 2607.25948 โข Published โข 13
MODUS for any-to-any generation (15 aligned modalities) as presented in MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities.
Project page: https://modus-multimodal.epfl.ch/ Code: https://github.com/EPFL-VILAB/Modus
Inference config (important):
conf/modalities/instruction_16mod_stage2.yamlmodel.safetensors โ trained weights (bf16)ae.safetensors โ VAE (image decode)config.json / llm_config.json / vit_config.json โ architecture configvocab.json / merges.txt / tokenizer_config.json โ tokenizer