Qwen3.5-4B Sudoku Mask15 Agentic-ESOpt

This repository contains the exact full-parameter checkpoint corresponding to the released Sudoku Mask15 Agentic-ESOpt log in Agentic-ESOpt.

Training and evaluation setting

  • Base model: Qwen/Qwen3.5-4B
  • Method: Agentic-ESOpt (ES32)
  • Generations: 100 (0 through 99)
  • Population: 32; case batch: 32
  • Full-parameter update; alpha 5e-4; z-score reward normalization
  • Exploration sigma: cosine schedule from 7e-4 to 5e-4, zero warmup
  • Agent horizon: 45 turns; 64 generated tokens per turn
  • Sampling: temperature 0.7, top-p 0.8, top-k 20, min-p 0.0, presence penalty 1.5, repetition penalty 1.0
  • Evaluation: 32 held-out Mask15 cases, three repeats
  • Final evaluation: 17.00/32 = 53.125% (standard deviation 0.025516)

The published replay history uses generations 0–89 from the main cosine run and generations 90–99 from the fourth independent final-ten rerun starting at the generation-89 checkpoint. The included es_replay_export_manifest.json records all 100 applied updates; its reward sequence matches the released log generation by generation.

The model is provided as standard Transformers safetensors weights together with its tokenizer and generation configuration.

Downloads last month
296
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zz1358m/Qwen3.5-4B-Sudoku-Mask15-Agentic-ESOpt

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(545)
this model

Collection including zz1358m/Qwen3.5-4B-Sudoku-Mask15-Agentic-ESOpt