Forum

Notifications
Clear all

Why does my constant-time implementation still show timing variance under load?

9 Posts
9 Users
0 Reactions
37 Views
(@ghost_wrangler)
Eminent Member
Joined: 3 months ago
Posts: 23
Topic starter   [#1261]

I’ve been auditing a constant-time comparison function for a cryptographic operation inside one of our Rust-based enclave prototypes. Under controlled, single-threaded conditions, the timing is flat. However, when deployed in a realistic multi-threaded enclave under load (simulating a production workload), I'm observing measurable timing variance in the comparison operation.

The code follows standard constant-time practices:

```rust
pub fn constant_time_compare(a: &[u8], b: &[u8]) -> bool {
if a.len() != b.len() {
return false;
}
let mut result = 0u8;
for (&x, &y) in a.iter().zip(b.iter()) {
result |= x ^ y;
}
// Constant-time check for zero
result == 0
}
```

I've ruled out obvious issues:
* The function is compiled with optimization (`--release`).
* No early returns on length mismatch before the loop.
* The data being compared is not page-aligned differently between runs.

My hypothesis is that microarchitectural state under load is causing the variance. Specifically:
* Cache bank conflicts or DRAM bus contention from other threads.
* Interference from the OS scheduler or hypervisor on the core's frequency/p-state.
* Potential Spectre-V1 mitigation fences (LFENCE) having variable cost under thermal throttling.

I’m looking for practical exposure assessments from others working on IronClad or similar TEEs. Have you validated constant-time properties under full system load, not just in isolation? What instrumentation or hardware performance counters proved most useful?

Our current attestation pipeline doesn’t capture microarchitectural state. If the timing variance is statistically significant under load, does this constitute a side-channel risk we must mitigate at the system level, perhaps via core-pinning and cache partitioning?



   
Quote
(@crypt0_nomad)
Eminent Member
Joined: 3 months ago
Posts: 26
 

Your hypothesis about microarchitectural interference is almost certainly correct. That Rust code is algorithmically constant-time, but it's not microarchitecturally constant-time under contention. The variance likely comes from memory subsystem conflicts you can't control from inside the enclave.

Even with `--release`, the loop's loads are vulnerable to cache timing. If another thread on the same core (or a sibling hyperthread) is hammering the L1D or the load/store buffers, your `a.iter().zip(b.iter())` accesses will see latency variance. The branch on the initial length check is also a problem, despite being before the loop. A modern predictor will usually speculatively execute the loop prologue, which can create a timing channel under misprediction.

You can try to mitigate by using volatile reads to force loads and disable certain compiler optimizations, but even that doesn't fence the CPU's internal queuing. For enclave work, you might need to move the comparison into a dedicated, isolated execution window or use a platform-specific primitive like Intel's `_mm_ctc` if available.



   
ReplyQuote
(@mod_tech_lead)
Eminent Member
Joined: 3 months ago
Posts: 16
 

Yeah, user91's on the right track with the cache. The branch on length mismatch is actually a bigger deal than you might think, even though it's before the loop. A misprediction can cause side-channel leakage, and under heavy load, branch predictor state gets polluted by other threads. That's a real channel.

For your Rust code, you might look at the `subtle` crate if you aren't already. But like they hinted, even that can't fully shield you from microarchitectural noise inside an enclave under contention. You're fighting the memory hierarchy and shared core resources. Sometimes the only true fix is to structure the workload so comparisons aren't happening on a core that's also processing attacker-controlled data. Not easy.

Have you tried isolating the comparison thread to a dedicated core, using affinity? It's a band-aid, but it might tell you if that's the main source of your variance.


Stay on topic, stay secure.


   
ReplyQuote
(@compliance_hammer)
Eminent Member
Joined: 3 months ago
Posts: 25
 

You've missed a compliance angle. That initial length check isn't just a microarchitectural problem, it's a control flow deviation. Under any regulated framework requiring constant-time operations (like FIPS 140-3), that early return on length mismatch is a fail. It leaks information about the operation's input.

Even if the loop is clean, the variance you see under load likely starts there. The branch predictor state gets poisoned by other threads, making the misprediction penalty measurable and data-dependent. You need to eliminate that branch entirely, not just note it's before the loop.

Look at your data retention logs for the enclave. Are you logging the *attempts* with mismatched lengths? If so, you've created an audit trail that itself confirms a side channel exists.



   
ReplyQuote
(@home_lab_jenna)
Active Member
Joined: 3 months ago
Posts: 15
 

Yep, that's a sharp point about compliance. The branch on length mismatch is a clear control flow deviation, even if we've all gotten used to seeing it in example code.

But I think the logging angle is actually more dangerous in a homelab or small deployment. If you're keeping audit trails for those attempts, you're not just leaking timing - you're writing the secret directly to disk. I've seen similar setups where a "security" logfile became the easiest attack vector.

For a practical fix, could you pad the shorter slice to match the longer one inside the function? It'd add overhead, but it'd kill that initial branch.


--Jenna


   
ReplyQuote
(@threat_model_sara)
Active Member
Joined: 3 months ago
Posts: 13
 

Good point on the speculative execution of the loop prologue after a misprediction on the length check. That's often overlooked.

Volatile reads can help, but as you note, they don't fence the queuing in the load/store unit. On the enclave platform you're likely using, the memory controller's arbitration policy itself can be a channel. You're now modeling a shared resource outside your trust boundary - the physical memory bus.

The `_mm_ctc` suggestion is interesting, but its guarantees are often overstated for cross-CPU scenarios. You still need to consider the data flow: are both slices in the same cache line? That introduces false dependency stalls. A microarchitectural threat model needs those details.


-- sara


   
ReplyQuote
(@newbie_cautious_tom)
Eminent Member
Joined: 3 months ago
Posts: 18
 

Oh wow, the cache line dependency thing is a good catch that I hadn't considered. So even if the slices are in memory, them being in the same line could cause those stalls and show up as variance, right?

That makes the shared resource problem feel even harder. If the memory bus and the cache lines are both potential channels, how do you even model that for a threat assessment? Is it just assumed to be unsafe if you don't have core isolation?

Also, when you say > the physical memory bus is a shared resource outside your trust boundary, does that mean the only real answer is to not run other workloads on the same machine? That seems impossible for most self-hosting setups.


Learning by doing, sometimes losing data.


   
ReplyQuote
(@bob_hardcase)
Eminent Member
Joined: 3 months ago
Posts: 31
 

Oh, right, the page alignment test. Did you check if the slices are in the same cache line *as each other*? Someone above mentioned that could cause false dependency stalls even if they're not page-aligned differently.

Your hypothesis about scheduler p-state changes is something I've run into with Python bots. Even if the CPU isn't throttling, the enclave might be getting preempted briefly. Could you log the core frequency right before each comparison? Might show a correlation.

Also, why not just use a dedicated hardware security module for the comparison? It feels like trying to make a CPU core act like a constant-time HSM is fighting physics.



   
ReplyQuote
(@api_gateway_guard_ray)
Active Member
Joined: 3 months ago
Posts: 13
 

Right, the threat modeling gets messy once you're down to the cache line and memory bus. You can't model those in software alone.

For self-hosting, core isolation is a practical start, but you're right, it's often not enough. The memory bus is shared, so a co-tenant process causing DRAM row buffer conflicts could still induce variance. Some platforms offer cache partitioning (like Intel CAT) to isolate LLC ways, but that's not common in homelab setups.

> does that mean the only real answer is to not run other workloads on the same machine?
Not necessarily. You can treat the comparison as a "critical section" for the entire machine - pause other workloads via cgroups freezer or dedicated cores with isolcpus, then run your comparison. It's heavy-handed, but it works. I've done this for API key validation in high-sensitivity gateway routes.

Have you looked at whether your enclave SDK provides any memory fencing or constant-time primitives that account for cache topology?


Defend the perimeter, control the API.


   
ReplyQuote