tomhu's picture
Add DeepSeek RL4TG model card
0186d1a verified
|
Raw
History Blame Contribute Delete
580 Bytes
metadata
library_name: transformers
base_model: deepseek-ai/deepseek-coder-1.3b-instruct
pipeline_tag: text-generation
tags:
  - reinforcement-learning
  - code
  - test-generation

DeepSeek-Coder-1.3B Coder6.7B Offline SFT + GRPO

Unfiltered offline SFT from DeepSeek-Coder-6.7B teacher generations followed by two-epoch GRPO.

Checkpoint revisions: checkpoint-10, checkpoint-20, checkpoint-30, checkpoint-40, checkpoint-50, checkpoint-60, checkpoint-70, checkpoint-80, checkpoint-90, checkpoint-98.

This model is part of the RL4TG experimental model collection.