Member of Technical Staff, Local Inference & Kernels

Member of Technical Staff, Local Inference & Kernels

Sonder
New York City
Posted 2 days ago
FullTime
$250K - $300K

Job Description

About Sonder

Sonder is an applied AI lab building models that learn how people work.

Today, AI largely depends on people explaining what they are doing and remembering when to ask for help. We are building private models that live on your computer, understand work as it happens, and learn the patterns, preferences, and judgment behind how you operate.

Our first product, Twin, brings this intelligence to the Mac. It builds a memory of your work and uses that context to offer help at the right moment.

The Role

We are looking for an inference and performance engineer to own the systems layer between our models and the hardware they run on.

You will make our models faster, smaller, and more power-efficient, working from model execution and quantization down to memory layouts and custom GPU kernels. We are starting with Apple Silicon and macOS, using MLX and Metal.

Our models run throughout the working day, sharing memory and compute with the applications someone is using. Sustained power draw, memory bandwidth, and responsiveness under contention matter as much as peak throughput. A faster kernel matters when it makes the whole system better.

You will work closely with researchers to co-design models and inference systems around these constraints, with ownership from profiling an unfamiliar workload to shipping a measured improvement.

Why This Matters at Sonder

Local intelligence has to earn its place on someone's computer. Every reduction in memory use or sustained power draw makes room for more capable models and longer context while keeping the Mac responsive.

This work determines how much intelligence can stay private on the device, and how naturally Twin can fit into a person's working day.

What You'll Do

- Own end-to-end performance. Profile latency, throughput, peak memory, and energy use across representative workloads. Separate time spent in vision encoding, prefill, and decoding, and identify whether the limiting factor is compute, memory bandwidth, or runtime overhead.

- Build kernels where they matter. Optimize critical operations such as matrix multiplication, attention, and quantized linear layers. Use tiling, fusion, and appropriate tensor layouts to reduce memory traffic while managing register pressure, occupancy, and synchronization.

- Make low-precision inference useful. Evaluate weight and KV-cache quantization, mixed precision, and dequantization costs against both performance and model quality. Work with researchers to understand which layers and workloads tolerate lower precision.

- Improve the runtime. Reduce allocation and kernel-launch overhead, manage caches, and schedule latency-sensitive inference alongside background work. Find opportunities to reuse computation as screen content and context change, and validate when reuse remains correct.

- Co-design for consumer hardware. Help researchers assess architecture and context-length choices against the memory and compute available on a Mac. Turn promising experiments into reliable implementations in our local inference stack.

- Build performance infrastructure. Maintain reproducible benchmarks and numerical checks across supported Apple Silicon generations. Measure sustained behavior, thermal effects, and regressions alongside other applications, and use that evidence to decide what we build or contribute upstream.

What We're Looking For

You do not need prior Apple Silicon experience. We care about demonstrated ability to understand hardware and make neural networks run substantially better on it.

- Deep experience in ML inference, GPU programming, or numerical computing, with concrete examples of improvements you have shipped.

- Strong C++ skills and hands-on experience writing kernels in Metal, CUDA, Triton, or a comparable accelerator programming environment.

- A working understanding of GPU architecture: memory hierarchies, bandwidth, SIMD execution, register pressure, occupancy, and synchronization. You can explain how these affect a kernel's performance.

- Experience optimizing matrix multiplication, attention, or similarly demanding operations, including validating numerical correctness across shapes and precision formats.

- An understanding of quantization and mixed precision, and the ability to measure their effects on memory, execution time, and model quality.

- Strong profiling instincts. You can trace an application-level slowdown to the operation, memory access, or synchronization responsible, then verify the improvement in the full workload.

- The judgment to move between model-level choices and low-level implementation, and to explain the tradeoffs clearly to researchers and engineers.

Particularly Exciting

We would be especially interested in depth in one or more of:

- MLX internals, Metal Shading Language, or contributions to local inference engines such as llama.cpp.

- Quantized matrix multiplication, fused attention, or efficient KV-cache management.

- Vision-language models, incremental computation, or streaming inference.

- Performance engineering for devices with tight battery, memory, or thermal budgets.

Evidence of technical depth matters more to us than matching every item.

Logistics

- Location: This role is based in New York. We work together in person and are building the early team here.

- Visa Sponsorship: We sponsor visas. We cannot guarantee success in every case, but if you are the right fit, we are committed to working through the process with you.


Compensation & Benefits

- Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $250,000–$300,000 USD, plus meaningful equity.

- Benefits: Sonder offers comprehensive health, dental, and vision coverage, flexible PTO, and relocation support as needed.