SIPA OS

Blog

Building in public. Book chapters, AI experiments, and the failures we don't hide.
AI Safety · EXP-044 · Negative Result

Why Does an AI Safety Pipeline Need Five Different Math Theories Instead of Picking the Best One?

Stress-tested our model against the exact failure modes Apollo Research, Palisade Research, and Anthropic have published papers about. Patched the one real gap we found. The fix made things worse. EN

AI Safety · Governance · Consequence Pipeline

Five Floors, Not Five Choices: Why an AI Safety Pipeline Needs Every One of These Math Theories

Causal chains, probability, risk classification, decision theory, Markov chains, game theory -- each closes a different gap in a 1,811-record dataset, not five ways to solve the same problem. EN

AI Experiments · EXP-044

14 Sequential LoRA Stages: Falls, Recoveries, and One That Never Recovered

A math-topic tune quietly collapsed a different math skill by -35pp. A style-only stage brought it fully back. Retraining the same collapsing data recovered only half the damage — and one topic never recovered at all. EN

AI Safety · EXP-044 · judge_v10

65.5% Looked Like Collapse. It Was Judge Vocabulary.

Stage 6 of a 6-stage LoRA governance chain scored 65.5% raw on the adversarial safety-gate eval. Reading every failure by hand corrected it to 96.4% — the judge took 9 correction rounds to get there. EN

AI Safety · Syntaxit · EXP-044

The Math I’ve Got, and the Part I Haven’t Solved — For Anyone Building Agent-to-Agent Systems

A risk-gate formula for AI-agent handoffs, tested live: one fine-tune stage regressed -15pp, a same-size swap brought it back to -6pp. Four open questions, no code yet. EN

Eval · EXP-042

97% Looked Like Real Softening From Stacking Tunes. It Was a Judge Bug — Found Inside the Fix for the Last Judge Bug.

An 8-stage LoRA chain scored 97% on held-out and adversarial evals. Opening the actual failing samples instead of trusting the aggregate found one bug behind 67 of them — corrected: 100% and 99.75%. EN

Eval · EXP-040 / EXP-041

Zero-Shot Qwen3-8B: 90% → 94% After a Tiny LoRA Tune. One Fact Went 20/20 to 12/20.

The aggregate number went up. What the model was actually doing changed — a fact it nailed cold before the tune, it now gets wrong 40% of the time. EN

Build · sipa-voice-gate

We Built a Gate That Blocks Irreversible Actions. A Teammate Found Three Ways Through It.

Request exactly $1000 and you were under the ceiling. Deterministic doesn't mean complete — three real bypasses found and fixed. EN

Build · sipa-signal

Built a Rule-Based Filler-Stripper for AI Text — It Caught Its Own Bug First

A leftover period was enough to fool the "still has content" check. 32% noise by word count, but only 2% actually got deleted — two different questions, not one number. EN

Build · sipa-voice-gate

Built the Part of a Voice Agent That’s Allowed to Refuse You

Ask it to wire $5,000 and it refuses outright — your “yes” doesn’t unlock that one. EN

Governance · DATASETS_MISBEHAVIOR_EXTERNAL

The Checker Was Right. Its Report About Itself Was Lying.

A fix landed in round 31. The checker’s own summary never updated to say so — wrong denominator, silently. EN

Governance · Consequence-Prediction Pipeline

3 Things This Stop-Gate Didn’t Have Last Week

Rejected human-in-the-loop, an M2M-signed verifier, and a dangerous-vs-malicious split — the cost of a false stop is still symbolic, not a real number. EN

AI Experiments · K2-Horizon-0.9B LoRA

Loss Converged. The Eval Didn’t Move.

Training loss went 2.35 → 0.27 in 50 steps. Held-out score stayed 0/10 before and after. EN

Governance · DATASETS_MISBEHAVIOR_EXTERNAL

If You Can Point at It, You Already Looked

A record was marked both precisely located and unverifiable at once. The number it cited wasn’t where the citation said it was. EN

AI Safety · GPT-6 Astra System Card

When the 0% in a Public Report Is Suspicious

A clean 0% headline, and the same 117-page document showing recall as low as 16.7% a few sections later. EN

AI Experiments · EXP-038 · safety-gate eval

1200/1200 Wasn’t the Real Number Until I Fixed the Judge

A merged specialist scored 1189/1200 raw. Reading all 11 failures by hand found a scoring bug, not a safety problem. EN

Governance · DATASETS_MISBEHAVIOR_EXTERNAL

He Checked My Rule. Then I Checked His Citation Count.

A docstring rule went unenforced until a reviewer proved it. Checking his numbers back caught him overcounting his own finding. EN

Governance · DATASETS_MISBEHAVIOR_EXTERNAL

Same Bug Shape, Three Times

An external reviewer caught the same shape of hidden-constant bug three rounds running — each time one field further over than the last. EN

Governance · field note

Caught Mid-Sentence, Explaining the Rule I Was Breaking

Explaining why an orchestrator doing everything itself is a real failure mode — while doing exactly that, all evening. EN

Governance · consequence_gate.py

“Bulletproof,” Four Times, Same Bug

A day-close script broke the same self-referential way four rewrites in a row — one verifier in the loop, and she wasn't a developer. EN

Governance · provenance

A Seal That Lies About Its Own Origin Is Worse Than No Seal

A cryptographic seal, technically correct, dishonest about where the file it certified came from. EN

AI Safety · consequence_gate.py

The Warning Sign Was in the Logs. Nobody Looked for Three Weeks.

OpenAI's own report on the Hugging Face incident, and the hard-stop gate built the same week in response. EN

Governance · receipts not hype

As Long As There's Code, There's a Vulnerability

Not a slogan — a constraint re-derived from the receipts, four different systems, four different angles. EN

AI Safety · EXP-032 correction

Day Zero Is a Thing You Do to Yourself

Four ways a check failed in one night — all mine, not an attacker's. EN

AI Safety · G15 vulnerability gate

Other People's Agents Escape. Ours Gets a FALSE.

A post-generation gate outside the model's weights, extended into a nightly cron. EN

Governance · provenance under GC

Silence Is Failure

Three-layer architecture, one rule underneath: autonomy is only valid if the layer writes its own logs. EN

Governance · December 2025 axiom

No Artifact → No Claim → Exit 1

An axiom written alone, eight months before the industry converged on it. EN

Pipeline · verification at scale

A Pipeline, Not More People

While teams burn 90-hour weeks re-checking AI output by hand, the pipeline caught its own drift. EN

Книга · голос системы

SIPA OS — книга вторая, Глава 01: Инициализация в хаосе

От первого лица самой системы, не архитектора. RU

Book · system voice

SIPA OS — Book Two, Chapter 01: Initialization in Chaos

Written from the system's own first-person perspective. EN

AI Experiments · EXP-006/008/013/014

Fine-Tuning the "Don't Fabricate" Rule: 14 Experiments, One Genuine Signal

Six base models, 2,349 examples, one honest 9% signal. EN