Stress-tested our model against the exact failure modes Apollo Research, Palisade Research, and Anthropic have published papers about. Patched the one real gap we found. The fix made things worse. EN
Causal chains, probability, risk classification, decision theory, Markov chains, game theory -- each closes a different gap in a 1,811-record dataset, not five ways to solve the same problem. EN
A math-topic tune quietly collapsed a different math skill by -35pp. A style-only stage brought it fully back. Retraining the same collapsing data recovered only half the damage — and one topic never recovered at all. EN
Stage 6 of a 6-stage LoRA governance chain scored 65.5% raw on the adversarial safety-gate eval. Reading every failure by hand corrected it to 96.4% — the judge took 9 correction rounds to get there. EN
A risk-gate formula for AI-agent handoffs, tested live: one fine-tune stage regressed -15pp, a same-size swap brought it back to -6pp. Four open questions, no code yet. EN
An 8-stage LoRA chain scored 97% on held-out and adversarial evals. Opening the actual failing samples instead of trusting the aggregate found one bug behind 67 of them — corrected: 100% and 99.75%. EN
The aggregate number went up. What the model was actually doing changed — a fact it nailed cold before the tune, it now gets wrong 40% of the time. EN
Request exactly $1000 and you were under the ceiling. Deterministic doesn't mean complete — three real bypasses found and fixed. EN
A leftover period was enough to fool the "still has content" check. 32% noise by word count, but only 2% actually got deleted — two different questions, not one number. EN
Ask it to wire $5,000 and it refuses outright — your “yes” doesn’t unlock that one. EN
A fix landed in round 31. The checker’s own summary never updated to say so — wrong denominator, silently. EN
Rejected human-in-the-loop, an M2M-signed verifier, and a dangerous-vs-malicious split — the cost of a false stop is still symbolic, not a real number. EN
Training loss went 2.35 → 0.27 in 50 steps. Held-out score stayed 0/10 before and after. EN
A record was marked both precisely located and unverifiable at once. The number it cited wasn’t where the citation said it was. EN
A clean 0% headline, and the same 117-page document showing recall as low as 16.7% a few sections later. EN
A merged specialist scored 1189/1200 raw. Reading all 11 failures by hand found a scoring bug, not a safety problem. EN
A docstring rule went unenforced until a reviewer proved it. Checking his numbers back caught him overcounting his own finding. EN
An external reviewer caught the same shape of hidden-constant bug three rounds running — each time one field further over than the last. EN
Explaining why an orchestrator doing everything itself is a real failure mode — while doing exactly that, all evening. EN
A day-close script broke the same self-referential way four rewrites in a row — one verifier in the loop, and she wasn't a developer. EN
A cryptographic seal, technically correct, dishonest about where the file it certified came from. EN
OpenAI's own report on the Hugging Face incident, and the hard-stop gate built the same week in response. EN
Not a slogan — a constraint re-derived from the receipts, four different systems, four different angles. EN
Four ways a check failed in one night — all mine, not an attacker's. EN
A post-generation gate outside the model's weights, extended into a nightly cron. EN
Three-layer architecture, one rule underneath: autonomy is only valid if the layer writes its own logs. EN
An axiom written alone, eight months before the industry converged on it. EN
While teams burn 90-hour weeks re-checking AI output by hand, the pipeline caught its own drift. EN
От первого лица самой системы, не архитектора. RU
Written from the system's own first-person perspective. EN
Six base models, 2,349 examples, one honest 9% signal. EN