About the Job
Buildkite's CI platform is trusted by the world's leading engineering teams, shipping software to over 1,000,000,000 daily users.
We're hiring a Staff Engineer (ML) to join our Test Engine team. In this role, you'll define and lead the technical strategy for machine learning within Test Engine — specifically, building the models and infrastructure behind predictive test selection: using code changes to determine which tests actually need to ru...
Staff Engineers at Buildkite are hands-on technical leaders. You'll influence how we design, build, and scale systems while supporting other engineers to deliver their best work. You'll be the most senior ML practitioner in the company, setting the technical direction for how we approach test selection and establish...
Key Responsibilities
Lead and define the ML strategy for predictive test selection — from early experimentation through to models running reliably in production at scale
Lead the technical investigation into how we build a generalised test selection model, and shape the approach based on what the data tells you
Lead the design of the ML architecture end-to-end: feature engineering from code changes and test history, model training and evaluation, serving infrastructure, and feedback loops for continuous improvement
Drive key decisions around model operationalisation — latency constraints (test selection has to be fast enough to sit in the critical path), prediction accuracy trade-offs, and graceful degradation when confidence is low
Shape how ML capabilities integrate with Test Engine's existing data infrastructure — billions of ingested test runs, test-to-code mapping, and the intelligent splitting engine
Build the ML platform layer so that getting a model into production is fast and repeatable
Design, build, and maintain the data pipelines that feed ML workloads — connecting code change signals with test execution history at scale
Train, evaluate, and deploy models, taking ownership through to monitoring and retraining in production
Instrument production models with observability metrics: prediction accuracy, latency, coverage, false negative rates, and drift detection
Solve the hardest technical challenges at the intersection of code analysis and test data — feature extraction from diffs, generalisation across languages and frameworks, and handling the cold-start problem for new tests and repositories
Investigate and resolve complex performance and reliability issues across the data and ML stack
Share knowledge and drive engineering best practices across teams through documentation, mentorship, and pairing
Support the wider engineering organisation by contributing to cross-team tooling, infrastructure, and frameworks
Communicate trade-offs effectively and build alignment around technical decisions
Work closely with customers to understand how test selection fits into their development workflows, and ensure the product delivers real impact
Required Skills & Abilities
Deep proficiency in Python, with strong experience building production ML systems end-to-end
Proven experience designing and operating ML infrastructure at scale — model registries, feature stores, serving layers, experiment tracking, or similar
Strong experience with data processing at scale — whether batch or streaming frameworks (Spark, Flink, or similar)
Deep proficiency in SQL
Comfort working in cloud environments (AWS) and with containerised workloads (Docker, Kubernetes)
In short, we'd expect equal comfort and high level capability in the end to end process from designing and building models through to deploying them.
Hands-on experience training, evaluating, and deploying ML models in production — you're a practitioner, not only an infrastructure builder
Experience with classification, ranking, or prediction problems where the signal-to-noise ratio is challenging — test selection shares characteristics with anomaly detection, change-point detection, and predictive filtering
Track record of building ML capabilities that scaled beyond a single use case — not just one-off models but repeatable, generalised approaches
Experience with feature engineering from structured and semi-structured data (code diffs, execution logs, dependency graphs, or similar)
Experience instrumenting production models with observability: accuracy, latency, coverage, drift
Excellent written and verbal communication skills, especially in a remote-first environment
Ability to distil complex technical concepts into clear explanations for diverse audiences
A collaborative, pragmatic mindset — balancing technical quality with business context
Comfortable mentoring engineers and leading technical discussions across teams
Proven ability to build alignment across teams and influence technical direction without authority
Experience with code analysis, static analysis tools, or building features from source code structure
Familiarity with CI/CD systems, developer tooling, or test infrastructure
Experience with Ruby on Rails, React, GraphQL, or Go
Background in search ranking, recommendation systems, or other domains where you're predicting relevance from sparse signals
Experience working with test frameworks or test execution data