About the Job
We’re Fundraise Up - a global fundraising platform built to make donating to nonprofits fast, seamless, and accessible to all. Every month, our technology powers tens of millions of dollars in donations across the globe. We focus on innovation that directly impacts results: faster load times, higher conversion rates...
Our platform is trusted by many of the world’s leading nonprofits, including UNICEF, the Alzheimer’s Association, and a wide range of global NGOs. With a 4.9/5 rating across top software review platforms, we’re recognized not just for our impact - but for the quality of the product we deliver.
We operate in the enterprise segment, serving nonprofit organizations across North America, the United Kingdom, Australia, and Europe.
Key Responsibilities
You will join the DevOps team responsible for the platforms our engineering teams rely on daily: CI/CD, observability, logging, and developer tooling. Our philosophy is simple: the team owns its systems end-to-end, and every engineer should be able to diagnose and fix issues in their area of responsibility.
This is a senior-only role. We are looking for an engineer we can hand an entire area — observability or CI/CD — and trust to run it: from requirements and technical design through rollout, operations, and mentoring others. You will be the go-to technical reference for your area and the senior escalation point for complex incidents in it.
Own one of our core platform areas end-to-end: observability (VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, alerting) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus, build agents) — you drive its architecture, reliability, and roadmap.
Drive technical initiatives end-to-end: gather requirements, write the design doc, decompose into tasks, implement, deliver to production, and own the operational health afterwards.
Drive clarity in ambiguous situations by defining requirements, assumptions, and next steps.
Design for reliability and scale: evolve the architecture of our platforms — topology, integration points, scaling approach, and reliability model.
Support developers: deploy and monitor applications on both on-premise servers and Kubernetes (Helm), troubleshoot builds and deploys, help teams with metrics, alerts, and logs; participate in chat duty in developer support channels.
Automate away toil: repetitive operations, provisioning, and maintenance should be codified, not performed by hand.
Investigate production incidents as the senior escalation point for your area: drive resolution, lead post-mortems, implement systemic fixes. Participate in on-call rotations and raise the bar for how on-call works.
Mentor less experienced engineers through design discussions, reviews, and pairing; catch debt-inducing shortcuts at the review stage.
Use AI in all aspects of day-to-day work: researching, troubleshooting, developing.
Required Skills & Abilities
6+ years as a DevOps Engineer / SRE (or very close responsibilities).
Track record of owning technical initiatives end-to-end — from requirements and technical design through production delivery. You can showcase initiatives that were yours, not just tasks you completed.
Confident Linux skills (we use Ubuntu).
Working knowledge of the Prometheus stack: metric types, exporters, and how alerting works — enough to navigate and extend an existing setup.
Hands-on experience with CI/CD: pipeline design, build orchestration, artifact delivery.
Containers: Docker, image building, registries.
Ansible.
Git.
Experience with Bash or Python scripting for automation and observability (writing exporters, eliminating routine work).
Production/on-call experience: diagnosing incidents, restoring service, leading post-mortems.
Experience mentoring less experienced engineers.
Ownership and attention to detail. Downtime is expensive: during busy events 10 minutes of downtime can cost us around $500k.
We understand it’s impossible to be an expert in everything, but it’s important to have solid hands-on experience in two or more of the areas below:
VictoriaMetrics / Prometheus stack at scale: architecture, cardinality control, exporters, alerting infrastructure.
Log pipelines at scale: Graylog / VictoriaLogs / ELK — collection (fluent bit or similar), retention, sharding, performance.
Jenkins scripted pipelines: shared libraries, pipeline infrastructure, build agent fleets.
Container registries and artifact management: Harbor, Nexus, base images, image policies.
Operating applications on Kubernetes: Helm, workload monitoring and log delivery, deploy troubleshooting.
Grafana: dashboards as code, alerting, performance at scale.
Great if you’ve worked with any of the following:
Analytics & DS platforms: JupyterHub, Airflow, Tableau, MLflow, Airbyte — deployment, maintenance, resource limits. Building platform around these tools to improve Quality of Life for Analytics.
Remote development environments and AI agent execution environments. E.g. Coder/Telepresence.
Bare-metal Kubernetes: provisioning, networking, scaling.
Flux and GitOps.
Terraform.
Sentry on-premise: operating self-hosted error tracking.
ClickHouse, MongoDB.
Qualifications
Experience:
10 years experience