Thanks again, Dipankar. Good catch on your own 27B row too. We get 0.049 to 0.061 on Nimble public HelpSteer2 as well.
I went back through your SummEval numbers and the mean confidence point against our records. Consistency and relevance, all three sizes, match to the third decimal. So does the confidence gap: at the train T the small models sit within 0.02 of their accuracy on both HelpSteer2 sets, and the hard T pushes that up by 0.06 to 0.07. Your flat mix result matches ours too. The loss roughly halves and never flips.
On your question. The HelpSteer2 part of the calibration split is 161 questions over 87 prompts, and on all of it the 4B gets 0.50 and the 9B gets 0.48.
That number is misleading though. 68 of those 161 are the other four HelpSteer2 attributes (coherence, complexity, verbosity, correctness), which Jevals and Nimble never ask, and coherence and complexity run around 0.7. On helpfulness alone, the question your two eval sets actually ask, it's 0.42 on the 4B and 0.40 on the 9B. Reweighted to a flat mix it's 0.44 and 0.41.
So it isn't well above your 0.43. It's about the same.
You're right about where it comes from. The calibration split is the 5% side of a 95/5 split over the training pool. The HelpSteer2 items in it come from HelpSteer2's train split, drawn evenly across the helpfulness levels. The split is by prompt, so both responses to a prompt land on the same side and no calibration prompt was ever trained on. Jevals and Nimble both sample HelpSteer2's validation split, and the pool builder throws out every validation row.
Here's the part I didn't expect. At the same accuracy, the model is less sure of itself on those calibration items than on the eval items. Fit on helpfulness alone, the calibration slice wants T 0.79 on the 4B and 0.83 on the 9B. A family bootstrap puts the 90% range at 0.64 to 0.99 and 0.67 to 1.04. Fit on Jevals and Nimble directly, the same models want 1.05 to 1.20.
So the calibration fit really is off for HelpSteer2 traffic. It just isn't off because the split is in distribution and flatters the accuracy. Something about that draw makes the model less confident than it is on the validation data, and the label mix is only part of it, same as you found.
The cheap test is sitting right there. HelpSteer2's validation split has 418 rows, over 209 prompts, that neither Jevals nor Nimble uses. They're at the natural label mix and never went near the training pool. We pull raw scores for those from the fleet, fit the score T on them, and see which side it lands on. If it comes out near 1.1, the calibration draw is the problem. If it comes out near 0.8, the two eval subsets are the odd ones. It's a few hundred requests per model, so we'll run it before we touch v3.
What goes on the v3 list: the calibration split draws HelpSteer2 at the natural label mix instead of evenly across levels, a held-out validation slice stays in the eval suite as a standing calibration check, and the calibration split gets bigger than 87 HelpSteer2 prompts.
I really appreciate you pushing on this. :)