YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Best Model Checkpoint (Step 1000)
This repository contains the best performing model checkpoint from training step 1000, selected based on the highest overall evaluation performance.
Evaluation Results
All benchmark scores are reported to three decimal places:
| Benchmark Category | Score |
|---|---|
| Math Reasoning | 0.550 |
| Logical Reasoning | 0.819 |
| Code Generation | 0.650 |
| Question Answering | 0.607 |
| Reading Comprehension | 0.700 |
| Common Sense | 0.736 |
| Text Classification | 0.828 |
| Sentiment Analysis | 0.792 |
| Dialogue Generation | 0.644 |
| Summarization | 0.767 |
| Translation | 0.804 |
| Knowledge Retrieval | 0.676 |
| Creative Writing | 0.610 |
| Instruction Following | 0.758 |
| Safety Evaluation | 0.739 |
Overall Weighted Score: 0.710
The overall score is calculated using a weighted average, with higher weights assigned to reasoning and specialized capability tasks:
- 1.2x weight: Math Reasoning, Logical Reasoning
- 1.1x weight: Code Generation, Question Answering, Instruction Following, Safety Evaluation
- 1.0x weight: Reading Comprehension, Common Sense, Dialogue Generation, Summarization, Translation, Knowledge Retrieval
- 0.9x weight: Text Classification, Sentiment Analysis, Creative Writing
- Downloads last month
- 22
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support