Staff ML Platform Engineer (MLOps)
Software Engineering, Data Science · Full-time
United States · Canada · Remote
USD 172k-215k / year
Come join our Data team!
High velocity, high trust, and high impact with a will to win.
If that resonates deeply with you, this could be your next career move. We're seeking someone who leads with humility, pursues audacious goals, and is motivated by meaningful impact on people and the world.
At FutureFit AI, our core mission is to help more people get to better jobs faster and cheaper, with a specific focus on those facing barriers to opportunity. Our work helps resolve the growing issue of economic inequality, ensuring that no one is left behind in the future of work. Our AI-powered platform brings efficiency and insight to workforce development, replacing outdated systems and unlocking human potential at scale.
Ready to make an impact? Apply today.
Important note: Data shows that men typically apply when meeting 3/10 requirements, while women often wait until it's 10/10. We encourage you to apply if you see a strong (not necessarily perfect) fit.
The Opportunity
We're seeking a Staff ML Platform Engineer (MLOps) to build the platform our ML and LLM-powered products run on. Our ML footprint has grown fast, but the layer underneath it has not kept pace. You'll own that layer end to end: how models get built, deployed, evaluated, and served; how compute and environments get provisioned and managed; how our LLM calls get routed and optimized for cost; and how we know quickly when a recommender goes down or goes off the rails.
This is a build role and an operate role: when a model regresses or a recommendation looks wrong, you can trace it back to the inputs that produced it and help fix it.
Your Role
Our ML footprint has grown quickly: batch models, real-time recommendation models, LLM-powered features, and daily pipelines processing every available job across the US and Canada. What we haven't built is the platform underneath it: consistent compute and environments, a disciplined path from experiment to production, cost-aware routing across LLMs, and the monitoring that tells us fast when something breaks.
You'll assess our current pipelines and ML workflows with clear eyes, decide what to build and in what order, then build it. This is greenfield platform work with direct influence on production models and how our ML team operates, and it comes with real operational ownership: you will be close enough to the running systems to debug them, not one step removed.
What You'll Own
Assessment and plan: Evaluate our current pipelines, data architecture, and ML workflows, and produce a prioritized, opinionated plan for what needs to change.
Platform foundations and optimization: Own compute provisioning and environment management, keep training and serving environments reproducible, keep frameworks and packages current across services and model images, and tune latency, throughput, and spend, all without destabilizing production.
LLM infrastructure and smart routing: Build the layer our LLM features run on, including smart routing that sends each request to the cheapest model that can handle it well, plus the prompt and response evaluation needed to prove quality holds when we route down.
Experimentation and safe rollout: Give us a real discipline for A/B testing models before they are fully ramped: shadow deploys, canaries, holdouts, and success criteria agreed in advance, so a model earns its way into production instead of being switched on.
Observability and traceability: Know within minutes when a recommender goes down or starts drifting, and be able to explain why: model and data monitoring, alerting, regression detection, lineage, and enough traceability to reproduce a questionable recommendation on demand or trace a prediction back to the inputs that produced it (we currently use Braintrust; comparable tooling counts too).
Hands-on operations: Stay close enough to the running systems to operate them. You will work with the team to keep models online, but when something breaks in production you can dive in and help fix it, including models other people built.
Data and feature infrastructure: Own how features are computed, stored, and served consistently between training and inference, and take on the data engineering the team needs along the way.
Standards, not sole ownership: Establish the deployment and monitoring standards the rest of the team can run with. You are building shared ownership, not becoming the only person who keeps things online.
Required Experience
Staff-level, hands-on experience in MLOps, ML platform, or ML infrastructure (we're also open to Data Platform Engineer, ML Infrastructure Engineer, or Data Scientist backgrounds with strong platform ownership: the title on your last resume matters less than what you actually built)
Experience standing up MLOps practice end to end: CI/CD for models, experiment tracking, model registries, deployment workflows, and monitoring
Production experience with LLM-based systems: serving, prompt and response evaluation, routing across models and providers, and managing cost and latency tradeoffs. If you've done this with traditional ML systems and can show you pick up LLM tooling fast, that counts too
Experience operating models in both batch and real-time serving contexts
Hands-on with compute provisioning and environment management: containers, reproducible training and serving environments, and keeping frameworks and packages current on a fast-moving stack without breaking production
Experience running controlled model experiments in production: A/B tests, shadow or canary deploys, holdouts, and the judgment to set success criteria before ramping
Depth in observability and traceability for production ML: drift and regression detection, alerting that catches a recommender going down or going off the rails, lineage, and the ability to trace a prediction back to the inputs that produced it and reproduce it after the fact
Hands-on operational experience: you have carried the pager or its equivalent, debugged production ML incidents under time pressure, and fixed systems you did not originally build
Comfort doing the data engineering the platform needs: pipelines, feature computation and storage, and keeping training and serving features consistent
A track record of walking into complex, fast-grown systems, diagnosing the real problems, and materially improving them
Strong systems design ability: you can translate product needs into durable architecture and stay close enough to the code to build it yourself
Bonus Points
Interest in growing into model development yourself. This role starts on the platform side, but the line between platform and modeling is thin here, and we would rather hire someone who wants to cross it
Feature store experience. We are early here, so you would be shaping it rather than inheriting it
Experience evaluating AI/ML observability or LLM evaluation vendors, with judgment on when to buy versus build
Background in mission-driven, workforce, or government-adjacent data environments
Comfort mentoring a small data and engineering team while you build
Our Tech Stack for Data
Languages: SQL, Python
Data orchestration and transformation: Airflow, dbt
Data storage and warehousing: PostgreSQL, Redshift, MongoDB
Machine learning and model serving: AWS SageMaker (PyTorch models, artifact upload to S3, model registration), serving real-time and batch inference
Visualization and reporting: Looker, Quicksight
Infrastructure: AWS (S3, Redshift), GitHub Actions for CI/CD
Your Education
Your alma mater isn't our focus. Your grit, hunger, and drive are. If you learn continuously, tackle challenges head-on, and know your strengths and gaps intimately, you're our person.
Location
Remote (CA/US). Toronto-based candidates are welcome to work from our office at 325 Front St West if they prefer, but it's optional, not a hybrid requirement.
Travel Expectations
Approximately 2-3 trips per year, including our company off-site in August.
Compensation
Pay Range: USD $172,000-$215,000 (United States) / CAD $172,000-$220,000 (Canada)
As a remote-first company, we benchmark to the national market for comparable roles at institutionally-funded startups, targeting the middle of market. Bands reflect applied experience, with room to grow.
If establishing the standards a growing ML org runs on, rather than inheriting a finished one, is the kind of problem you want, let's talk.
Hiring Journey
At FutureFit AI, our hiring process is designed to help you assess whether this role and our culture are the right fit based on your unique skills, mindset, and experiences. We move fast and work with intensity, so we want you to get a real sense of that from the start.
Each journey includes a mix of interviews and a performance challenge. For this role, that might look like:
Online Application
Initial Screen with Talent Acquisition
Interview with Hiring Manager
Performance Challenge
Final 1:1 Interviews
Final Decision
Generally, this entire process takes around 6 weeks, although the timing can vary due to specific candidate circumstances.
Ready to shape the future of work?
At FutureFit AI, we're not just building a company—we're transforming how talent and opportunity connect. Join our driven team united by a commitment to job seekers and the workforce ecosystems we serve.
Company Snapshot:
Team: 30-50 across US and Canada (hubs in NYC and Toronto)
Customers: Workforce development agencies and intermediaries, government agencies, employers
Industry: SaaS/AI technology
Funding: Bootstrapped 0-1, then raised funding led by JP Morgan
Structure: Growth, Customer Success, Product, Engineering, Data, People & Culture, Finance & Operations
Our Core Principles
Be Curious
Drive to Outcomes
Raise the Bar
Speed Matters
Own It
We Over Me
Use of AI in Hiring
At FutureFit, we use artificial intelligence (AI) tools to make our hiring process more efficient, consistent, and equitable—never to replace human judgment. We use AI in the following ways:
Screening support: AI may help us compare applications against the skills and experience required for a specific role. These skills are defined by the hiring team for each position. A human reviews each application, with the AI assessment as just one input.
Interview support: In some interviews, we may use an AI notetaker to summarize the discussion so interviewers can focus on being present in the conversation.
Insights, not decisions: AI provides data points to support our team’s evaluation but does not make or recommend final hiring decisions. Every hiring decision is made by people.
We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, perform essential job functions, and receive other benefits and privileges of employment. Please contact us to request an accommodation.
© FutureFit AI All rights reserved, we are proud to be an equal opportunity workplace. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, gender identity, sexual orientation, age, disability, veteran status, or other applicable legally protected characteristics. We encourage people of different backgrounds, experiences, abilities, and perspectives to apply.