This is the official implementation of CalibSFT, a plug-and-play supervised fine-tuning stage that shapes a calibrated, broadly supported verbalized-confidence prior before confidence-aware reinforcement learning. For each training question, CalibSFT samples
Please refer to docs/INSTALL.md to set up the environment.
We release the CalibSFT and CalibSFT → RLCR checkpoints for quick reproduction. Evaluate them on all 16 benchmarks (temperature 0.6; 4 responses per question, 32 on AIME):
Qwen3-8B
CKPT=Qwen/Qwen3-8B \
bash scripts/examples/eval/base.shCalibSFT
CKPT=SUSTech/Qwen3-8B-CalibSFT \
bash scripts/examples/eval/calib_sft.shCalibSFT → RLCR
CKPT=SUSTech/Qwen3-8B-CalibSFT-RLCR \
bash scripts/examples/eval/calib_sft_rlcr.shPlease refer to docs/RUN.md for CalibSFT data generation, CalibSFT training, confidence-aware RL, and evaluation of trained checkpoints.
@article{wang2026pitfalls,
title={On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models},
author={Wang, Shuoyuan and Luo, Beier and Zeng, Hao and Yu, Chengyao and Zhang, Songxin and Xie, Zejian and Jing, Bingyi and Wei, Hongxin},
journal={arXiv preprint arXiv:2609.32470},
year={2026}
}