Remote Senior Kubernetes Operations Engineer at Lambda

Summary

The job is for an experienced operations engineer to remotely manage Kubernetes clusters. The role involves daily operations, incident response, improving tooling, assisting customers, initial cluster build-outs, mentoring, and product direction. The ideal candidate should have deep knowledge of Linux clusters, bare-metal, containers, virtualization, Kubernetes, on-call environments, and incident post-mortems.

Requirements

Are an experienced operations engineer, SRE, sysadmin or similar with a deep knowledge of running Linux clusters and systems
Are very familiar with running on bare-metal (including knowledge of BMCs, kernel drivers, PXE, RAID, VLANs, hypervisors)
Have a good understanding of containers, virtualization, and the mechanisms underpinning them
Have a good understanding of daily operation, bug-fixing, and maintenance of Kubernetes
Have experience in an on-call environment and with incident response
Can perform incident post-mortems and develop procedures and tooling to prevent root causes from reoccurring
Have an excellent ability to learn on-the-fly and adapt to solve problems
Are able to work either independently with limited direction, or as part of a team
Are able to work with customers during incidents either via tickets, live messaging, or as part of a larger call

Responsibilities

Remotely install, upgrade, operate, and maintain bare-metal Kubernetes clusters (up to thousands of nodes each)
Handle cluster degradation, recovery, and resizing using our fleet management tooling
Perform out-of-hours on-call response for critical incidents as part of a well-balanced on-call rotation
Work on improving our tooling, automation, and processes, for both daily operations, alerting, and incident response
Dive into systems at a low level to solve unique cluster problems and write up your findings
Assist customers with high-level Kubernetes questions and integration with applications, storage, and authentication
Assist with initial cluster build-outs and validation to help identify failed hardware before customer delivery
Work closely with our HPC Ops and Datacenter Ops teams on issues that require lower-level expertise or cross-functional solutions
Mentor and assist less-experienced team members

Preferred Qualifications

Deep Kubernetes experience
Experience with user-level restrictions and hardening (e.g. AppArmor)
Experience with network engineering
Experience with HPC clusters, environments & tooling
Experience with large-scale AI/ML training clusters
Experience with machine learning/AI frameworks

Benefits

Generous cash & equity compensation
Investors include Gradient Ventures, Google’s AI-focused venture fund
We are experiencing extremely high demand for our systems, with quarter over quarter, year over year profitability
Our research papers have been accepted into top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
We have a wildly talented team of 250, and growing fast
Health, dental, and vision coverage for you and your dependents
Commuter/Work from home stipends for select roles
401k Plan with 2% company match
Flexible Paid Time Off Plan that we all actually use

Remote Senior Kubernetes Operations Engineer

Lambda

Job highlights

Summary

Requirements

Responsibilities

Preferred Qualifications

Benefits

Remote

DevOps

Senior

Similar Remote Jobs

Senior Operations Engineer

InfluxData

Remote

DevOps

Senior

Senior Operations Engineer

InfluxData

Remote

DevOps

Senior

Senior Staff Operations Engineer

Airbnb

Remote

DevOps

Senior

Senior Technical Operations Engineer

LetsGetChecked

Remote

DevOps

Senior

Senior Machine Learning Operations Engineer

Omnidian

Remote

DevOps

Senior

Software Engineer, Senior

SumerSports

Remote

Software Development

Mid-level

Senior Engineer

Veritone

Remote

Software Development

Mid-level

Platform Engineer (Kubernetes)

steercom - Key Message. Delivered.

Remote

Software Development

Mid-level

Platform Engineer (Kubernetes)

steercom - Key Message. Delivered.

Remote

Software Development

Mid-level

Platform Engineer (Kubernetes)

steercom - Key Message. Delivered.

Remote

Software Development

Mid-level