Any-to-Any
Safetensors
multimodal
modus

MODUS

MODUS for any-to-any generation (15 aligned modalities) as presented in MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities.

Project page: https://modus-multimodal.epfl.ch/ Code: https://github.com/EPFL-VILAB/Modus

Inference config (important):

  • modality config: conf/modalities/instruction_16mod_stage2.yaml
  • weights are bf16.

Files

  • model.safetensors โ€” trained weights (bf16)
  • ae.safetensors โ€” VAE (image decode)
  • config.json / llm_config.json / vit_config.json โ€” architecture config
  • vocab.json / merges.txt / tokenizer_config.json โ€” tokenizer
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
15B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using EPFL-VILAB/MODUS 1

Paper for EPFL-VILAB/MODUS