Kernel Tuning
Design: PREEMPT_RT kernel + CPU isolation for inference threads, with api-server and cluster-agent pinned to dedicated cores and fallible I/O paths planned for io_uring. Aimed at measurable wins on tail latency.
- Design: PREEMPT_RT for bounded scheduling latency under load.
- Planned: isolcpus — cores 4–7 reserved for inference; kernel housekeeping cannot steal them.
- Planned: io_uring batched submission queues for KV cache spill + log flush.
- Target: cut p99 enqueue jitter toward <1ms on saturated nodes (design baseline ~8ms; lab validation in progress).