The role
The runtime is where a model becomes an application people can depend on. You’ll build the foundation that lets Loci run across different chips, operating systems, and inference engines, with predictable resource use and responsive controls. Your work will connect our shared Rust core to the execution backends that make local AI possible.
What you’ll work on
- Design clear interfaces between the application and inference engines for model loading, prompt submission, streamed output, cancellation, and cleanup.
- Integrate model architectures and execution backends across macOS, iOS, Android, and Windows, keeping shared behavior in the Rust core and platform details in focused adapters.
- Own the model lifecycle: loading and unloading, switching models, handling interrupted work, and recovering cleanly when a device runs short of memory.
- Manage context and cache ownership so conversations, edits, and cancellations produce correct results without leaking resources or mixing state.
- Make CPU and accelerator execution work together reliably. Diagnose failures at the boundaries between native libraries, drivers, threads, and the application.
- Build compatibility and regression coverage for model formats, backend updates, and device differences. Turn hard-to-reproduce failures into small, repeatable cases.
- Work with performance engineers to ship optimizations safely, and contribute fixes upstream when the issue belongs in a dependency.
What you bring
- You have shipped systems software with meaningful concurrency, memory-management, or native-integration requirements.
- You can work confidently in Rust and navigate C or C++ at library boundaries, with careful attention to ownership, lifetimes, and error handling.
- You understand the stages of language-model inference and can reason about how model weights, context, and cached state consume device resources.
- You can debug across layers, from an unresponsive application to the runtime, operating system, or accelerator backend beneath it.
- You build explicit, testable interfaces and care about cancellation, recovery, and correctness as much as the successful path.
Experience that helps
- Experience integrating inference engines such as llama.cpp, MLX, LiteRT, or ONNX Runtime into an application.
- Familiarity with Apple or Android native development, Windows systems APIs, Metal, or Vulkan.
- Work on model packaging, hardware capability detection, cross-platform builds, or maintaining an open-source systems project.
You don’t need experience with every tool listed. Show us the problems you’ve solved and how you approach unfamiliar ones.
Join Loci
Apply for this role
Email careers@askloci.ai with your résumé or a link to your work. Include the role title in the subject.
Tell us about a system you built or a difficult runtime failure you diagnosed. We’re interested in the constraints, your decisions, and what happened after it shipped.
Apply by email