One panic shouldn't kill the scan: isolating rules with catch_unwind
Cloud Cost Analyzer runs 92 rules against your cloud account in a single scan. A rule is a real unit of work: it pulls a resource's config and metrics, evaluates them against a cost model, and maybe emits a finding. They're independent, and every one of them leans on libraries I don't control — cloud SDKs, response parsers, date and math crates.
Any of that can meet an input it didn't anticipate and panic: an unwrap() that couldn't
fail until some resource shape I'd never seen, an index out of bounds deep in a
dependency, an overflow on a field that's normally small. Usually it isn't "someone wrote
bad code" — it's a library panicking on an edge case nobody dependency-audited for.
The question that matters isn't "will a rule panic." It's "when rule #57 panics, do I lose the other 91 and the whole scan?" A cost report that dies two-thirds through because one rule hit one weird EBS volume is worse than useless.
"So handle it with Result"
The obvious objection: this is Rust — model the failure as a Result and there's nothing
to catch. And the rules do that for every failure they can see coming. Can't reach the
metrics API? Result. Missing a field the check needs? Result. Expected failures travel
the normal error path, always have.
But a panic isn't the expected-failure path — it's the failure you didn't foresee. You
can't write ? for a bug you don't know exists, and you especially can't for one buried
three crates deep in a dependency that panics on input its author never tested. Result
handles the errors you anticipated; catch_unwind is the backstop for the ones you
couldn't. It's both, not either — and the second only exists because the first can't cover
code you didn't write.
std::panic::catch_unwind runs a closure and hands you back a Result instead of letting
the panic tear through your stack. Here's the real dispatch, one rule at a time:
let mut eval = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| {
rule.evaluate(resource, metrics, costs)
}))
.unwrap_or_else(|_| {
tracing::error!(
rule_id = rule_id,
resource_id = %resource.resource_id,
"rule panicked during evaluate; skipping its finding"
);
None
})?;
Rule panics, we log which rule and which resource, return None for that one finding,
and the loop moves on to the next rule. One panic costs you one finding, not the scan.
AssertUnwindSafe is a promise, not a bypass
That AssertUnwindSafe wrapper isn't decoration — the compiler demands it, and it's worth
understanding why instead of reflexively reaching for it.
catch_unwind only accepts a closure that's UnwindSafe. The idea: if a panic unwinds
through the middle of your closure, any state it shared with the outside could be left
half-updated, and UnwindSafe is the marker that says "there's no such state to observe."
My closure captures &rule, &resource, and friends by reference, so the compiler can't
prove that on its own and refuses to compile.
AssertUnwindSafe is me telling the compiler: I've checked, catching this panic won't let
anyone observe a broken invariant. And here that's true for a concrete reason — on
panic I throw the result away. I return None and never read a partially-built finding.
There's no torn state to see because I don't look. That's the bar for reaching for
AssertUnwindSafe: not "make the error go away," but "I can point at why the post-panic
state is unobservable."
The setting that silently turns all of this off
Here's the part people miss, and the reason catch_unwind is not a try/catch: it only
works if panics unwind. Rust has two panic strategies, and you choose per build:
# Cargo.toml — a size-optimized profile
# Inherits all release settings, adds panic=abort to remove unwind tables
panic = "abort"
Under panic = "abort", a panic doesn't unwind the stack — it calls the abort handler and
kills the process immediately. catch_unwind never gets a chance to run. Every one of
those careful per-rule guards becomes a no-op, and rule #57 takes the whole process down
with it — in the exact build you ship to production, because panic = "abort" is something
you turn on to drop unwind tables and shrink the binary.
So panic isolation isn't a library feature you add; it's a whole-binary decision. If
you want one bad rule to cost one finding, you have to keep the unwinding runtime — and
eat the binary size — on purpose. Nobody's assert or catch_unwind can buy that back for
you. It's the trade you make once, in the profile, and it's invisible in the code that
depends on it.
One more edge: poisoned locks
If a rule panics while holding a Mutex/RwLock, that lock is poisoned and the next
rule to touch it inherits the panic. The isolation above holds partly because the shared
data these rules read is immutable and lock-free — there's no lock to poison. If yours
mutates shared state under a lock, catching the panic isn't enough; you have to decide
what a poisoned lock means.
Where it landed
Rules and the second-stage analyzers both run inside that catch_unwind, under an
unwinding profile. A rule — or a library it leans on — can meet an input nobody
anticipated, and the cost you pay is one missing finding and a log line naming the culprit,
not a dead scan. The isolation is real because the binary is built to allow it, and it's
honest because I can say exactly why the post-panic state is safe to ignore: I never read
it.
This is part of a series on building Cloud Cost Analyzer in Rust.