Upgrade to Pro — share decks privately, control downloads, hide ads and more …

FlexGuard vs Time Slice Extension: Handling Lock

FlexGuard vs Time Slice Extension: Handling Lock

Modern hardware scales by adding cores, but synchronization overhead increasingly limits scalability in user-space. Blocking locks are reliable but incur high handover costs due to frequent context switches. Spinlocks, in contrast, minimize handover latency, but their performance collapses in oversubscription (i.e., more threads than hardware contexts) as busy-waiting threads preempt lock holders in preemptible contexts. In an attempt to get the best of both worlds, many performance-oriented applications still rely on spin-then-park locks (e.g., POSIX mutexes). owever, like other locks that balance spinning and blocking, spin-then-park locks rely on arbitrary heuristics that often lead to suboptimal performance.

FlexGuard (SOSP'25) is a non-heuristic synchronization technique that leverages eBPF to monitor context switches and detect critical-section preemptions. When a lock holder is preempted, FlexGuard proactively transitions waiting threads from spinning to blocking, freeing CPU resources to quickly resume the preempted critical section. By reacting to actual execution events rather than static thresholds, FlexGuard can improve performance by up to 6 times compared to POSIX mutexes.

Another approach, often discussed in the Linux community (and used in Solaris), instead aims to prevent lock-holder preemptions altogether by extending scheduler time slices. Recent Linux proposals, including a patch by Thomas Gleixner (2025), use rseq to efficiently notify the kernel when a thread holds the lock. Rather than preempting such a thread, the scheduler allows it to run until the lock is released, preserving forward progress. Timeslice extensions have long been a subject of debate among maintainers.

To better understand whether timeslice extensions should be integrated into the Linux kernel, we evaluate both techniques across microbenchmarks and applications, including a memory-optimized database index, LevelDB, PARSEC's Dedup, and SPLASH2X's Raytrace and Streamcluster. Our results show that the two approaches address different bottlenecks and complement rather than replace each other, delivering benefits in both oversubscribed and non-oversubscribed settings.

Victor LAFORET

Avatar for Kernel Recipes

Kernel Recipes PRO

September 29, 2026

More Decks by Kernel Recipes

Other Decks in Technology

Transcript

  1. Blocking locks handoffs require sleeping and waking Lock holder wakes

    a waiter CS SW CS SW CS SW CS Problem: Sleeping and waking add latency to contended handoffs Waiters block (futex wait) 2
  2. Spinlocks avoid scheduler involvement Simple TAS spinlock Direct handoff void

    lock(state) { while (atomic_exchange_acquire(&state, LOCKED) != UNLOCKED) ; } No scheduler involvement on handoff Fast transitions between CS … as long as the lock holder keeps running. No sleep/wake 3
  3. Lock-holder preemption causes performance collapse Oversubscription begins Pure blocking lock

    Spin lock Efficient while CPUs are available Collapse under oversubscription Holds spinlock 5x Lower is better Waiter runs instead of lock holder When the lock holder is preempted, waiters waste CPU and progress stalls. => Leads to a performance collapse under oversubscription (threads > cores) 4
  4. Applications cannot assume dedicated CPUs Desktop applications: many concurrent processes

    Production workloads: Meta’s DCPerf: threads > cores System activity: monitoring, backups, background services, … User-space locks must remain robust under oversubscription 5
  5. Time Slice Extension (TSE) keeps the critical section running Idea:

    Prevent the scheduler from preempting the lock holder 1 Request TSE 2 Cancel TSE request 3 Yield if the extension was granted 1 rseq->slice_ctrl.request = 1; // Prevent compiler reordering barrier(); critical_section(); barrier(); 2 3 rseq->slice_ctrl.request = 0; if (rseq->slice_ctrl.granted) rseq_slice_yield(); Cheap: only writes, no sync 6
  6. Using TSE to build a lock void lock(L) { while

    (true) { while (atomic_load(&L) != UNLOCKED); 1 2 1 rseq->slice_ctrl.request = 1; barrier(); Wait until the lock looks free if (xchg(&L, LOCKED) == UNLOCKED) return; // acquired, TSE active 2 3 Request TSE before trying to acquire the lock No preemption window On acquisition failure: cancel TSE request Yield if TSE was already granted barrier(); rseq->slice_ctrl.request = 0; if (rseq->slice_ctrl.granted) rseq_slice_yield(); 3 } } void unlock(L) { atomic_store(&L, UNLOCKED); Request TSE at all times while holding the lock. barrier(); rseq->slice_ctrl.request = 0; if (rseq->slice_ctrl.granted) rseq_slice_yield(); } 7
  7. Single-lock microbenchmark Intel machine, 104 CPUs Oversubscription Measures latency function

    of number of threads Pure blocking lock is the baseline Lower is better 8
  8. Single-lock microbenchmark Intel machine, 104 CPUs Oversubscription In non-oversubscription MCS

    performs better than the pure blocking lock In oversubscription MCS collapses (unusable) Lower is better 9
  9. Single-lock microbenchmark Intel machine, 104 CPUs Oversubscription glibc’s POSIX lock

    (pthread_mutex_t) • Worse than pure blocking in that case Lower is better 10
  10. TAS + TSE Single-lock microbenchmark Intel machine, 104 CPUs Oversubscription

    TAS + TSE performs poorly TAS locks suffer from cache-line contention Lower is better 11
  11. TSE cannot prevent next-waiter preemption Ordered queue locks: handoff stall

    1 2 3 TSE keeps the holder A running Unlock hands lock off to next waiter B B becomes holder and progress stalls Lock design implications for TSE A TSE holds lock Running Option 1: Unordered acquisition (e.g. TAS/TATAS) Lock (shared word) B holds lock Preempted C cannot bypass B C waiting Running Our choice: modified qspinlock queued slow-path + stealing fast path A B C holds lock Running waiting Preempted waiting Running Option 2: Stealable fast path A B C holds lock Running waiting Preempted waiting Running D fast path Running 13
  12. TAS + TSE Single-lock microbenchmark Qspinlock + TSE Intel machine,

    104 CPUs Oversubscription In non-oversubscription Qspinlock + TSE performs similarly to MCS and better than blocking locks In oversubscription Collapse is reduced but still present Protecting the holder is not enough to guarantee progress. Lower is better 14
  13. FlexGuard: Switching to blocking Pure blocking lock Spin lock Target

    hybrid lock Idea: Detect the CS preemption, and switch to a blocking lock LOG SCALE Can we do this? Yes, with Lower is better , we can accurately detect critical-section preemptions 15
  14. FlexGuard: Preemption Monitor • Detects critical-section (CS) preemptions • Provides

    a global counter of preempted CSs (num_preempted_cs) • eBPF handler hooks to the sched_switch event • How to detect that a thread is holding a lock? • Idea: use a per-thread flag! • If flag set, we’re in a critical section! (needs to be a counter for lock nesting) Example: TATAS lock • Is that enough to be accurate? CS 16
  15. FlexGuard: Preemption Monitor • Answer: no, the counter is not

    enough. • The CS starts/ends in the lock()/unlock() functions • E.g., a thread could be preempted between lines 9 and 12  eBPF handler checks if the preemption address is in CS range Example: TATAS lock •Is Still one problematic it important to be case: fully Preemption accurate? right after  return is UNLOCKED)long • the Yes,XCHG a CS (CS is often only avalue few instructions •• Can we take care of thisincase? Preemptions are likely lock()/unlock() CS? • Yes, store XCHG result in known register Sufficient to cause performance collapse! •Access dumped register value in eBPF handler ⇒ Preemptions detected with 100% accuracy! CS CS CS 17
  16. Using FlexGuard to build a lock • Based on Qspinlock

    (MCS queue + fast path using global lock word) CS preemption detected Spinning mode  num_preempted_cs == 0 • Waiters enter the MCS queue • First spinner spins on the lock word instead of its node Qspinlock Blocking mode  num_preempted_cs > 0 • Waiters bypass MCS queue • Use global lock word as futex word Futex blocking lock No preempted CS remains 18
  17. TAS + TSE Single-lock microbenchmark Qspinlock + TSE Intel machine,

    104 CPUs Oversubscription Lower is better 19
  18. TAS + TSE Single-lock microbenchmark Qspinlock + TSE Intel machine,

    104 CPUs Oversubscription In non-oversubscription FlexGuard performs similarly to MCS and Qspinlock + TSE In oversubscription FlexGuard avoids the performance collapse of other locks. Lower is better 20
  19. Single-lock microbenchmark TAS + TSE FlexGuard + TSE Qspinlock +

    TSE Intel machine, 104 CPUs Oversubscription FlexGuard + Time Slice Extension • performs similarly to FlexGuard • more stable at high oversubscription Lower is better 21
  20. TAS + TSE Single-lock microbenchmark Intel machine, 104 CPUs Oversubscription

    FlexGuard greatly outperforms the blocking locks… but why? Lower is better 22
  21. Oversubscribed lock behavior Qspinlock + TSE Microbenchmark with 140 threads

    Intel machine, 104 CPUs ⇒ Oversubscribed case MCS performs poorly because # spinning waiters > # CPUs FlexGuard performs best because 0 < # spinning waiters ≤ # CPUs  Uncollapsed spinlock behavior in oversubscription Pure blocking lock performs better, but the next waiter is not running, because #spinning waiters = 0 Critical section preemptions (Spinning ⇒ Blocking) 23
  22. Qspinlock + TSE TAS + TSE FlexGuard + TSE Hash

    table DB Index Dedup Raytrace Higher is better Streamcluster LevelDB readrandom 24
  23. Evaluation FlexGuard + TSE Qspinlock + TSE TAS + TSE

    LevelDB readrandom Intel (104 hw threads) Raytrace AMD (512 hw threads) In the oversubscribed case, • Qspinlock + TSE does not avoid the performance collapse • FlexGuard outperforms blocking locks Benchmarks w/ concurrent workload: fixed number of benchmark threads, varying number of concurrent threads ⇒ Oversubscribed case In the non-oversubscribed case, • TAS + TSE as bad as a TAS lock… • Qspinlock + TSE performs worse than MCS (fast path) • FlexGuard outperforms blocking locks Higher is better 25
  24. FlexGuard + Time Slice Extension FlexGuard is limited by its

    ability to put threads to sleep DB Index AMD (512 hw threads) When the number of waiters is low, lock holder may not get re-scheduled quickly enough  FlexGuard + Time Slice Extension more consistently ensures good performance 26
  25. Conclusion Time Slice Extension FlexGuard • Lets user-space request extra

    time before preemption • Enables user-space spinlocks in oversubscription • Reduces critical section preemptions • Delivers spinlock-like behavior without collapse • 10% faster than blocking locks when oversubscribed under very low contention • Combines the reliability of blocking with the performance of spinning • Requires specific lock designs FlexGuard + Time Slice Extension Best solution in our scenarios, even when the number of waiters is limited 27