The role
A useful local assistant has to respond quickly, keep up with a conversation, and leave room for the rest of the device. You’ll own the work that turns model capability into a smooth experience: understanding where time and memory go, choosing the right optimization, and proving that it helps on the hardware people actually use.
What you’ll work on
- Trace the path from a submitted prompt to the first token and through every decode step. Find bottlenecks in computation, memory bandwidth, transfers, and synchronization.
- Improve prefill and decode for everyday workloads, including long conversations, repeated instructions, and short interactive requests.
- Develop KV-cache and prefix-reuse strategies with clear rules for cache validity, eviction, and context changes. Reduce repeated work without introducing stale state or incorrect results.
- Evaluate quantization, attention implementations, speculative decoding, and prompt-processing strategies against latency, memory use, and output quality.
- Build reproducible benchmarks across Apple, Android, and Windows hardware. Measure cold starts, sustained performance, peak memory, battery use, and thermal behavior.
- Carry improvements into the product with the runtime team, including regression checks and device-specific fallbacks where an optimization does not hold up.
What you bring
- You have improved the performance of an ML system, inference engine, compiler, or other compute-intensive software and can explain the evidence behind the improvement.
- You understand how transformer inference uses compute and memory, including the different demands of prefill, decode, attention, and the KV cache.
- You can use profiling tools to follow a slow operation across CPU and accelerator work, and distinguish a faster benchmark from a better user experience.
- You are comfortable working in low-level code and reading unfamiliar implementations. Experience with Rust or C++ is useful, along with a willingness to work in our shared Rust core.
- You treat numerical accuracy, model quality, and reproducibility as part of performance engineering.
Experience that helps
- Hands-on work with local inference projects such as llama.cpp or MLX, or with mobile and desktop GPU kernels.
- Experience with Metal, Vulkan, SIMD, model compression, or performance under constrained memory and power budgets.
- An upstream contribution, benchmark investigation, or technical write-up that shows how you reason about a difficult performance problem.
You don’t need experience with every tool listed. Show us the problems you’ve solved and how you approach unfamiliar ones.
Join Loci
Apply for this role
Email careers@askloci.ai with your résumé or a link to your work. Include the role title in the subject.
Tell us about a performance problem you investigated: what you measured, what you changed, and how you checked the result. A short explanation or a link is enough.
Apply by email