I ran the fresh test I promised: 80 new chains from the vulnerability group, old 20 kept out. Four arms, same prompt, greedy decoding, on one L40S:
* base (Qwen2.5-7B-Instruct): MAE 0.154 vs one labelling, 0.151 vs the other
* specialist: 0.122 / 0.115
* merged (three specialists): 0.121 / 0.111
* constant 0.70 (train median): 0.135 / 0.129
Paired bootstrap, against the constant:
* specialist: −0.013 [−0.028, +0.002] vs the first labelling, −0.014 [−0.028, −0.002] vs the second
* merged: −0.014 [−0.030, +0.002] and −0.019 [−0.034, −0.003]
Both specialist arms beat the constant by a small margin, and one of the two intervals excludes zero only just. Read it as: the adapters learned something beyond the base model and the label mean, but not much. The predictions cluster in a narrow band (0.65–0.75), and the test-retest MAE of the labels is 0.107, which is about the size of the gain. So "the specialist reads the chains" is not shown by this run.
Labels still come from one 405B model with no ground truth. The next version builds labels from documented incident outcomes.
All raw outputs, scripts, logs and hashes are in the repo:
AI_EXPERIMENTS/EXP-046-probability-estimator/brev_run_2026-10-05/.