Senior Machine Learning Engineer, AI Systems
COMPANY OVERVIEW:
There are over 5 billion users using basic applications today such email, notes, tasks that are not AI-native. This company’s mission is to build a proactive smart assistant for everyday users to bring intelligence to conversations, errands, organizing and workflows, with minimal prompting.
Their product focuses on achieving high reliability for long-running workflows, persistent context, and real-world task completion. The system must handle multi-step reasoning, interact with external tools, and remain reliable despite non-deterministic model behavior. The objective is to help users complete tasks daily enjoyable with over ~90%* reduced time.
The best products today in the world were built by small, world class teams. This is a high talent density and hands-on team. They make decisions collectively, move at rapid speed, striking a balance between shipping high quality work and learning. Joining the team requires the ability to bring structure, exercise judgment, and execute independently. The goal is to put in hands of users a truly magical product.
POSITION OVERVIEW:
As a Senior Member of Technical Staff, Machine Learning, you are an independent owner of critical ML subsystems in production. You take ambiguous problems, design practical solutions, and ship systems that operate reliably at scale.
This is a hands-on, high-impact role focused on depth.
RESPONSIBILITIES:
- Build core ML systems that power a proactive, long-horizon AI product.
- Own work end-to-end: data preparation, training, evaluation, inference, and iteration.
- Turn research ideas into working systems that run reliably in production.
- Debug model failures and system issues using real production signals.
- Iterate quickly: ship, measure outcomes, refine, and repeat.
- Collaborate closely with research, product, and engineering to deliver real user impact.
- Mentor and review work from other ML engineers through example and technical judgment.
- Work under real production constraints: latency, cost, reliability, and safety
PREFERRED PROFILE:
- You have built and shipped ML systems used by real users.
- You understand how modern ML models behave — and misbehave — in production.
- You write strong, production-quality code and think in systems, not scripts.
- You take ownership, work independently, and push work across the finish line.
- You learn fast, communicate clearly, and improve through iteration.
TECH STACK:
- Python
- PyTorch / JAX
- GPU-based training and inference systems
DESIRED OUTCOMES:
- ML models and systems in production consistently meet accuracy, latency, reliability, and efficiency targets.
- Complex production issues are monitored, debugged, and resolved with minimal disruption.
- Training, inference, and data pipelines are robust, scalable, and maintainable over time.
- Drives measurable improvements in ML systems based on real-world signals and user feedback.
- Provides mentorship and technical guidance to peers, raising the overall ML engineering standard.
- Collaborates cross-functionally to ensure ML features integrate seamlessly into products and meet business goals.
LOCATION: Remote
Job ID# 3635093
Artemis invites you to subscribe to our free Job Alerts and “The Hunt” Blog for free insights on hiring and career development.
Artemis Referral Bonus – $500! If you know someone for this job, please join our Referral Bonus Program.