ishaanranjan's picture
Upload reproduced read_span_selector qwen2.5-1.5b specialist
b7552ad verified
|
Raw
History Blame Contribute Delete
3.1 kB
metadata
license: apache-2.0
language:
  - en
library_name: transformers
pipeline_tag: text-generation
base_model: Qwen/Qwen2.5-1.5B-Instruct
base_model_relation: finetune
tags:
  - code
  - coding-agent
  - tool-use
  - function-calling
  - small-language-model
  - full-parameter-finetuning
  - supervised-fine-tuning
  - deterministic-verification
  - safetensors
  - subroutine:read_span_selector
model-index:
  - name: Code Read-Span Selector (Qwen2.5 1.5B)
    results:
      - task:
          type: text-generation
          name: Code Read-Span Selector
        dataset:
          name: Held-out HTTPX and Jinja2 oracle benchmark
          type: custom
        metrics:
          - type: accuracy
            value: 0.832
            name: Success after one schema-feedback retry
          - type: accuracy
            value: 0.996
            name: First-pass schema validity

Code Read-Span Selector (Qwen2.5 1.5B)

This is a full-parameter supervised fine-tune of Qwen/Qwen2.5-1.5B-Instruct for one narrow, schema-bound developer-agent subroutine:

Report the exact line span of a symbol's definition in a window.

The model is one cell from the Parameter Floors for Developer-Agent Subroutines experiment. Labels are generated by deterministic oracles over real Python repositories; no teacher model or human judge labels the data.

Intended Use

Use this checkpoint inside the repository's verified subroutine harness, which renders the task-specific prompt, parses strict JSON, permits one localized schema-feedback retry, applies deterministic guards, and falls back to rules where appropriate. This is not a general coding assistant or chat model.

Evaluation

Evaluation uses up to 250 examples from HTTPX and Jinja2, both held out entirely from training. Decoding is greedy.

Metric Result
Success after one schema retry 83.2%
First-pass success 83.2%
First-pass schema validity 99.6%
Base instruct success after retry 10.4% for the base instruct model
Rules-only success 81.7%

Experiment verdict for this subroutine: unsolved (best 0.83 @ qwen2.5-1.5b).

Training

  • Training examples: 1772
  • Epochs: 2.0
  • Learning rate: 2e-05
  • Effective batch configuration: 2 per device x 8 gradient accumulation
  • Maximum sequence length: 2048
  • Seed: 0
  • Final training loss: 0.137933
  • Reproduction hardware: one NVIDIA A100 80GB PCIe
  • Source revision: d0fd7bf

The dataset was generated from pinned Flask, Click, and Rich repositories for training/validation. HTTPX and Jinja2 were reserved for testing.

Limitations

The checkpoint is specialized to one closed JSON schema and should not be expected to retain broad instruction-following ability. The experiment mixes two base-model families across its size sweep. Some subroutines are better served by deterministic rules; consult the verdict above before deployment.

License

Apache-2.0, following the base model. Experiment code is MIT licensed.