--- library_name: transformers base_model: deepseek-ai/deepseek-coder-1.3b-instruct pipeline_tag: text-generation tags: - reinforcement-learning - code - test-generation --- # DeepSeek-Coder-1.3B Coder6.7B Offline SFT + GRPO Unfiltered offline SFT from DeepSeek-Coder-6.7B teacher generations followed by two-epoch GRPO. Checkpoint revisions: `checkpoint-10`, `checkpoint-20`, `checkpoint-30`, `checkpoint-40`, `checkpoint-50`, `checkpoint-60`, `checkpoint-70`, `checkpoint-80`, `checkpoint-90`, `checkpoint-98`. This model is part of the RL4TG experimental model collection.