Eval · EXP-046
Three Qwen2.5-7B LoRA specialists, one per risk group (vulnerability, deletion, sensitive_publication), trained to predict how likely a causal chain actually completes to its harmful outcome. Each one genuinely beat its own zero-shot baseline:
| Group | Zero-shot MAE | After specialist LoRA |
|---|---|---|
| vulnerability | 0.098 | 0.085 |
| deletion | 0.144 | 0.113 |
| sensitive_publication | 0.134 | 0.100 |
This wasn’t a task already saturated zero-shot — unlike a same-day decomposition-classifier tune (EXP-045), where the base model was already at 100% before any training. Real signal, real improvement, on a task with actual headroom.
Then the equal-weight merge of all three specialists into one adapter — same convention that held up cleanly on a binary refusal task back in EXP-031 (6 specialists merged, −1pp swing, noise) — landed within 0.001–0.004 MAE of the unspecialized base model on every group. Not “close to the best specialist.” Close to zero fine-tuning at all.
Likely mechanism: merging LoRAs that each shift a continuous number in group-specific directions cancels out under linear combination, in a way merging LoRAs that enforce a shared binary behavior doesn’t. Not investigated yet: whether a routed combination (pick the right specialist per group at inference, not blend weights) holds the gain a flat merge loses.
One bug caught before writing this up, not after. The eval script’s output filename only encoded before/after, not which adapter — the merged-eval run silently overwrote each specialist’s own result file, since both are “after” runs. Caught by checking the downloaded file’s own recorded adapter path against what was expected, not by trusting the script’s own success message. Fixed, specialists re-run cleanly under distinct filenames — numbers matched the original run within sampling noise.
Adapters, raw eval data (before / each specialist / merged, 9 files), and the full writeup are up.