Skip to main content

One panic shouldn't kill the scan: isolating rules with catch_unwind

· 5 min read
Founder, Dragon Fractal · ex-AWS engineer

Cloud Cost Analyzer runs 92 rules against your cloud account in a single scan. A rule is a real unit of work: it pulls a resource's config and metrics, evaluates them against a cost model, and maybe emits a finding. They're independent, and every one of them leans on libraries I don't control — cloud SDKs, response parsers, date and math crates.

Any of that can meet an input it didn't anticipate and panic: an unwrap() that couldn't fail until some resource shape I'd never seen, an index out of bounds deep in a dependency, an overflow on a field that's normally small. Usually it isn't "someone wrote bad code" — it's a library panicking on an edge case nobody dependency-audited for.

The question that matters isn't "will a rule panic." It's "when rule #57 panics, do I lose the other 91 and the whole scan?" A cost report that dies two-thirds through because one rule hit one weird EBS volume is worse than useless.

"So handle it with Result"​

The obvious objection: this is Rust — model the failure as a Result and there's nothing to catch. And the rules do that for every failure they can see coming. Can't reach the metrics API? Result. Missing a field the check needs? Result. Expected failures travel the normal error path, always have.

But a panic isn't the expected-failure path — it's the failure you didn't foresee. You can't write ? for a bug you don't know exists, and you especially can't for one buried three crates deep in a dependency that panics on input its author never tested. Result handles the errors you anticipated; catch_unwind is the backstop for the ones you couldn't. It's both, not either — and the second only exists because the first can't cover code you didn't write.

std::panic::catch_unwind runs a closure and hands you back a Result instead of letting the panic tear through your stack. Here's the real dispatch, one rule at a time:

let mut eval = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| {
rule.evaluate(resource, metrics, costs)
}))
.unwrap_or_else(|_| {
tracing::error!(
rule_id = rule_id,
resource_id = %resource.resource_id,
"rule panicked during evaluate; skipping its finding"
);
None
})?;

Rule panics, we log which rule and which resource, return None for that one finding, and the loop moves on to the next rule. One panic costs you one finding, not the scan.

AssertUnwindSafe is a promise, not a bypass​

That AssertUnwindSafe wrapper isn't decoration — the compiler demands it, and it's worth understanding why instead of reflexively reaching for it.

catch_unwind only accepts a closure that's UnwindSafe. The idea: if a panic unwinds through the middle of your closure, any state it shared with the outside could be left half-updated, and UnwindSafe is the marker that says "there's no such state to observe." My closure captures &rule, &resource, and friends by reference, so the compiler can't prove that on its own and refuses to compile.

AssertUnwindSafe is me telling the compiler: I've checked, catching this panic won't let anyone observe a broken invariant. And here that's true for a concrete reason — on panic I throw the result away. I return None and never read a partially-built finding. There's no torn state to see because I don't look. That's the bar for reaching for AssertUnwindSafe: not "make the error go away," but "I can point at why the post-panic state is unobservable."

The setting that silently turns all of this off​

Here's the part people miss, and the reason catch_unwind is not a try/catch: it only works if panics unwind. Rust has two panic strategies, and you choose per build:

# Cargo.toml — a size-optimized profile
# Inherits all release settings, adds panic=abort to remove unwind tables
panic = "abort"

Under panic = "abort", a panic doesn't unwind the stack — it calls the abort handler and kills the process immediately. catch_unwind never gets a chance to run. Every one of those careful per-rule guards becomes a no-op, and rule #57 takes the whole process down with it — in the exact build you ship to production, because panic = "abort" is something you turn on to drop unwind tables and shrink the binary.

So panic isolation isn't a library feature you add; it's a whole-binary decision. If you want one bad rule to cost one finding, you have to keep the unwinding runtime — and eat the binary size — on purpose. Nobody's assert or catch_unwind can buy that back for you. It's the trade you make once, in the profile, and it's invisible in the code that depends on it.

One more edge: poisoned locks​

If a rule panics while holding a Mutex/RwLock, that lock is poisoned and the next rule to touch it inherits the panic. The isolation above holds partly because the shared data these rules read is immutable and lock-free — there's no lock to poison. If yours mutates shared state under a lock, catching the panic isn't enough; you have to decide what a poisoned lock means.

Where it landed​

Rules and the second-stage analyzers both run inside that catch_unwind, under an unwinding profile. A rule — or a library it leans on — can meet an input nobody anticipated, and the cost you pay is one missing finding and a log line naming the culprit, not a dead scan. The isolation is real because the binary is built to allow it, and it's honest because I can say exactly why the post-panic state is safe to ignore: I never read it.


This is part of a series on building Cloud Cost Analyzer in Rust.