Skip to content
Platform · Performance Engineering

Performance Engineering on Jetson Orin NX

ThoxOS is under active performance engineering on Jetson Orin NX. Three concurrent engineering tracks target tail latency, UMA contention, and hot-path predictability: a PREEMPT_RT kernel sprint to bound scheduler jitter, a memory broker so Ollama, TensorRT-LLM, and cuQuantum cannot starve each other, and a Rust inference daemon that retires Ollama from the hot path.

Platform stack: ThoxOS (local ownership) → MeshStack (distributed) → ThoxWork (workspace). This page is the ThoxOS inference-predictability track on Orin NX.

What this page is

A public status board for ThoxOS inference-predictability workstreams on Orin NX, plus a qualitative R&D lab teaser. Numbers are labeled Target unless THOX-measured evidence has cleared publication.

What this page is not

Not a shipped benchmark sheet, not a live public Orin lab demo, and not a place for vendor marketing numbers, tok/s claims, or unreleased SKU pricing. Private ThoxOSv2 work continues offline.

Active engineering tracks

Three concurrent fronts, one objective: predictable inference.

Status as of Sep 2026
In development
Track 01

Kernel Tuning

Design: PREEMPT_RT kernel + CPU isolation for inference threads, with api-server and cluster-agent pinned to dedicated cores and fallible I/O paths planned for io_uring. Aimed at measurable wins on tail latency.

  • Design: PREEMPT_RT for bounded scheduling latency under load.
  • Planned: isolcpus — cores 4–7 reserved for inference; kernel housekeeping cannot steal them.
  • Planned: io_uring batched submission queues for KV cache spill + log flush.
  • Target: cut p99 enqueue jitter toward <1ms on saturated nodes (design baseline ~8ms; lab validation in progress).
In development
Track 02

UMA Discipline

Orin's unified memory pool is the biggest performance trap — every byte the GPU grabs is a byte the CPU loses. Managed aggregation is the right direction; what is still missing is a hard memory broker between Ollama, TensorRT-LLM, and cuQuantum so they cannot starve each other.

  • Design: hard per-runtime ceilings (Ollama 5.5GB, TensorRT-LLM 6GB, cuQuantum 2.5GB) enforced at allocation time.
  • Planned: broker daemon refuses allocations that would trip OOM and surfaces back-pressure to callers.
  • Planned: LRU model + KV-cache eviction with NVMe spillover when a higher-priority runtime needs headroom.
  • Planned: per-runtime telemetry in the GPU Dashboard with pre-ceiling alerts.
In development
Track 03

Rust Inference Daemon

Design goal: retire Go Ollama from the hot path by wrapping llama-cpp-2 in an Axum-based Rust HTTP server for deterministic memory behavior without touching the kernel.

  • Design: llama-cpp-2 — direct Rust bindings to llama.cpp; same GGUF models, no Go runtime in the hot path.
  • Planned: Axum HTTP surface aligned with cluster-agent and the inference gateway.
  • Planned: arena-allocated request buffers; tokenizer + KV cache in long-lived pools; no GC tracing.
  • Design goal: drop-in /api/generate + /v1/chat/completions so existing clients need zero changes.

Target outcomes

What the three engineering tracks are measured against. Every figure below carries a Target badge — these are not shipped benchmarks.

Methodology: public figures replace targets only after THOX-measured EVAL_SHEET rows exist and Tommy clears publication. Until then, treat every numeric claim on this page as a design target, not a result.

p99 enqueue jitter
Target
<1ms
Target — design baseline ~8ms; lab validation in progress
UMA contention OOMs
Target
0
Target — enforced at alloc time (design)
Hot-path GC pauses
Target
None
Target — Rust + arena allocation (design)
Client API breakage
Target
0
Design goal — drop-in surface
What we're working on

R&D lab — qualitative callouts only.

Public-safe engineering surfaces and framing. No vendor model numbers, no tok/s, no seed IDs, no POC links. In development stays In development; targets stay targets.

Status as of Sep 2026
Shipped on site
Lab 01

ThoxKey interactive surface

Live interactive WebGL for ThoxKey on the public site — engineering surface work for product fidelity, not a performance benchmark.

  • Interactive 3D viewer shipped on /thoxkey.
  • Plate-matched exterior for public product storytelling.
  • No latency or throughput claims attached to this surface.
R&D lab
Lab 02

Local-first control path

Qualitative R&D on on-device route discipline and refuse-when-uncertain behavior so ownership stays local. No vendor model names, no tok/s, no seed IDs.

  • On-device route discipline — prefer local answers when the device can own them.
  • Refuse-when-uncertain — decline rather than invent when confidence is low.
  • Targets stay targets until THOX-measured EVAL_SHEET rows clear publication.
R&D lab
Lab 03

Platform stack framing

ThoxOS (local ownership) → MeshStack (distributed) → ThoxWork (workspace). This page is the ThoxOS inference-predictability track on Orin NX — not a MeshStack sales surface.

  • ThoxOS: local ownership and on-device control.
  • MeshStack: distributed coordination across devices — framing only; no pricing or tok/s here.
  • ThoxWork: workspace layer on top of the stack.

Built into the platform, not bolted on.

Performance work is planned to land in ThoxOS builds and targets Nova devices and MagStack clusters when those tracks ship — not a claim that every device already runs these paths. Read the OS deep-dive or browse hardware specs for more context. Your AI. Your Data. Your Rules.™