OpenGrad · Study 002 · in progress
Relabelling, on-policy distillation and speculative decoding.
Study 002 is everything after the Study 001 freeze. It has no results yet. Each experiment below is pre-registered before any GPU time is spent, and its results will be reported against the frozen Study 001 checkpoints, not in place of them.
SCOPE
Three lines of work, none yet run.
WHAT STUDY 002 STARTS FROM
- Study 001’s checkpoints, corpora and evaluations, frozen at the OpenGrad tag
study-001. None of them is modified, relabelled or replaced. - The promoted checkpoint refuses all 1,319 zero-shot GSM8K questions. The regression is associated with tool-policy post-training on When2Call-derived data; causation is not established.
- Cross-model replication and joint capability-efficiency studies stay on the roadmap, not assigned to Study 002.