Notifications

Loading notifications...
vCluster Labs Cover

AI Infrastructure Engineer

vCluster Labs
Worldwide Full Time Negotiable 19 days ago

About the Job

As vCluster’s AI Infrastructure Specialist, you will work directly with customers at the earliest and most critical stage of their journey: from bare metal GPU nodes through to a production-ready deployment. This is not a traditional professional services role; you operate pre-sale as part of a proof of value engage...

vCluster is gaining rapid traction with GPU AI Clouds and enterprises building AI Factories: organizations that need to offer Kubernetes as a managed service on bare metal GPU infrastructure, and need to do it fast. This role exists to make that happen.

As an AI Infrastructure Engineer, your role will include:

Key Responsibilities

Lead Technical Deployments: Drive end-to-end technical deployments for GPU neocloud and AI Factory customers, from initial bare metal configuration to a validated vCluster environment.
Infrastructure Optimization: Configure and troubleshoot bare metal GPU node infrastructure, including CNI configuration, GPU Operator setup, distributed storage backends, and RDMA/InfiniBand.
Validation: Deploy and validate Kubernetes and vCluster to provide GPU-powered managed K8s.
Knowledge Transfer: Work alongside customer teams to build self-sufficiency, ensuring they can operate and grow the platform independently.
Scaling through Documentation: Document reusable playbooks and deployment architectures so your learnings become the next customer's head start.
Feedback Loop: Collaborate with Engineering and Product to surface recurring infrastructure challenges, acting as a direct feedback loop from the field into the roadmap.
Strategic Partnering: Join Sales in the pre-sales process where deep infrastructure work is required to achieve a meaningful proof of value.
Production K8s Mastery: 5+ years of experience deploying and operating Kubernetes in production, ideally on bare metal or in high-complexity environments.
GPU Fluency: Practical knowledge of NVIDIA GPU Operators, CUDA tooling, and systems-level configuration for GPU nodes.
Networking Fundamentals: Deep understanding of CNI plugins, overlay networks, load balancing, and connectivity diagnosis in layered environments.
Storage Expertise: Experience with persistent volume configuration, CSI drivers, and distributed systems like Ceph, Rook, Weka, or Longhorn.
Operational Agility: Comfort operating in ambiguous, fast-moving environments where you are often writing the playbook in real time.
Modern Tech Mindset: You thrive in environments that reject legacy tech and prefer a modern stack where you can solve a variety of problems from pipelines to internal services.
Automation Skills: Experience writing automation scripts with Bash, Python, or Go.
Kubernetes Depth: Relevant certifications such as CKA (Certified Kubernetes Administrator) or experience writing Kubernetes Operators.
AI/ML Familiarity: Experience with inference serving, GPU scheduling, and the tooling around LLM deployment.
Documentation: Experience building AI Automation in documentation to contribute to a shared knowledge base.
We are a venture-backed tech startup and the company pioneering Kubernetes virtualization for the AI era. We raised +$30M from top-tier VCs such as Khosla Ventures (first investor in OpenAI, GitLab, Stripe, Doordash) and are in a hyper-growth phase looking for motivated people to complement our team. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe and we have a remote-first work culture.
We are the leading platform for operating GPU infrastructure, enabling AI Cloud providers to deliver a hyperscaler-like experience to their customers and AI factories that need to build that same experience for their internal teams. Our platform delivers the full operational stack operators need to run their GPU data centers — managed Kubernetes, fast isolated tenant provisioning, and automated node provisioning and lifecycle management — enabling them to accelerate time to value, reduce operational burden, and maximize the ROI of every GPU.
We're the company behind vCluster, an open-source technology for virtualizing Kubernetes (10k+ GitHub stars, 40M+ virtual clusters created since 2021). Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI — a Kubernetes-native framework purpose-built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.

Qualifications

Experience: 5 years experience

Apply now

Please let vCluster Labs know you found this job on Job Vista. This helps us grow!