Notifications

Loading notifications...
Firmus Cover

Site Reliability Engineer, AI Infrastructure

Firmus
Melbourne, AU Full Time Negotiable Today

About the Job

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. Fo...

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.

Key Responsibilities

Support in the deployment, configuration, and maintenance of various high-end GPU servers, storage servers, networking equipment and software components in highly secure environments. 
Perform hardware diagnostics, systems functionality and firmware updates as required. 
Collaborate with engineering teams to assist in tailored customer environments deployment (eg: bare-metal systems, HPC Clusters, Kubernetes, Slurm etc). 
Serve as first line of engineering support for onsite operational issues, including troubleshooting hardware, network and software problems, and firmware compliance. 
Troubleshoot incidents, escalate critical issues and provide feedback to appropriate teams for improvements. 
Participate in an on-call rotation to ensure 24/7 availability and responsiveness to critical issues. 
Provide technical support to the GOC Support Specialist team in troubleshooting compute infrastructure related problems. 
Document incident details, resolutions, and lessons learned to enhance future problem-solving. 
Maintain clear, accurate, and up-to-date documentation to promote effective knowledge sharing across the team. 
Communicate effectively with GOC, HPC Engineers, internal teams, stakeholders, and end-users to ensure alignment on issue resolution. 
Take part in team meetings and knowledge-sharing sessions to foster collaboration and continuous learning. 

 

Required Skills & Abilities

Bachelor’s degree in computer engineering, computer science, or a related technical field.  
5+ years of experience in field service technical areas. 
Strong understanding of server hardware technology, firmware lifecycle, Linux environments and troubleshooting hardware problems, with adherence to physical and system-level security standards. 
Experience with scripting languages (eg: Bash, Python) 
Familiarity with using configuration management, CICD tools, workload manager and cluster softwares (eg: Slurm, Kubernetes, Nvidia BCM) and Observability tools (eg: Prometheus, Grafana, ELK, etc) 
Excellent problem-solving and analytical skills.  
Ability to work independently and as part of a team.  
Strong communication skills, both written and verbal. 

 

Qualifications

Experience: 5 years experience

Apply now

Please let Firmus know you found this job on Job Vista. This helps us grow!