
Member of Technical Staff, Data and RL
Job Description
About Sonder
Sonder is an applied AI lab building models that learn how people work.
We bring research, engineering, and design together to build useful personal AI, with privacy and efficiency at the core. We are building our early team in New York.
The Role
We are looking for a research engineer to improve model capabilities through better data, training environments, and reinforcement learning.
You will work across data generation, environment design, and post-training, with ownership from an initial hypothesis to a measured improvement in model behavior. The work calls for strong engineering, careful experimentation, and judgment about what is worth pursuing.
You will work directly with researchers and engineers to identify capability gaps, design experiments, and bring successful results into our models.
What You'll Do
- Build high-quality data pipelines. Develop systems for generating, curating, and validating training data. Understand how coverage, quality, and data mixtures affect learning.
- Develop training environments. Build reliable environments for multi-step tasks, with reproducible execution and useful feedback for training and evaluation.
- Improve post-training. Establish supervised baselines, run RL experiments, and investigate the effects of data, rewards, and optimisation choices.
- Make evaluations trustworthy. Design graders and held-out evaluations, investigate reward hacking and data contamination, and distinguish capability gains from benchmark artifacts.
- Close the experimental loop. Inspect model failures, form testable hypotheses, and use controlled ablations to guide the next experiment. Build tooling that makes results reproducible and easy to inspect.
What We're Looking For
- Strong Python engineering skills and hands-on experience training or post-training language or multimodal models with PyTorch, JAX, or a comparable framework.
- Technical depth in reinforcement learning, synthetic data, or interactive environments, and the ability to work across the broader training pipeline.
- A practical understanding of policy optimization, reward design, and credit assignment in multi-step tasks.
- Evidence of rigorous experimental judgment: credible baselines, informative ablations, and clear explanations of what a result does and does not establish.
- Experience building dependable data or experiment infrastructure and debugging failures across the learning pipeline.
Particularly Exciting
You have built RL environments used by other researchers, improved model capabilities through data quality, or taken a post-training experiment from prototype to a reliable system.
We care about the quality of your work and the depth of your understanding. A strong project, research result, or open-source contribution can demonstrate both.
Logistics
- Location: This role is based in New York. We work together in person and are building the early team here.
- Visa Sponsorship: We sponsor visas. We cannot guarantee success in every case, but if you are the right fit, we are committed to working through the process with you.
Compensation & Benefits
- Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $250,000–$300,000 USD, plus meaningful equity.
- Benefits: Sonder offers comprehensive health, dental, and vision coverage, flexible PTO, and relocation support as needed.

