← Serch more jobs

Senior Customer Reliability Engineer - Weights & Biases

LinkedIn Weights & Biases Bellevue, WA
Not Applicable Posted March 27, 2026 Job link
Thinking about this job
Not Met Priorities
What still needs stronger evidence
Requirements
  • 5+ years of experience in technical support, customer engineering, production engineering, reliability engineering, or a similar role supporting enterprise or strategic accounts.
  • Expert in Python, with strong debugging, profiling, and production-grade development skills.
  • Strong background in computer science or software engineering (B.S. in CS or equivalent experience).
  • You have strong experience running or supporting large-scale, high-availability systems (Kubernetes/GKE, cloud services, distributed systems, or similar).
  • Deep familiarity with the AI/ML ecosystem: training frameworks (PyTorch, TensorFlow), generative AI stack (Hugging Face, LangChain, vector databases), and modern experimentation workflows.
  • Skilled at diagnosing distributed systems, APIs, containerized environments, and multi-tenant cloud architectures.
  • Exceptional communication skills, with the ability to interface effectively with customer engineering teams, executives, and internal stakeholders.
  • Demonstrated success partnering with product and engineering teams to drive reliability improvements and influence roadmap priorities.
  • Self-driven, customer-obsessed, and passionate about building reliable, scalable systems and great customer experiences.
  • Proficient with monitoring and observability tools (Datadog, Prometheus/Grafana, OpenTelemetry, etc.) for debugging production environments.
Preferred Skills
  • Experience with Docker, Kubernetes, and cloud platforms (AWS, GCP, Azure).
  • Familiarity with GPU compute environments and distributed model training pipelines.
  • Previous experience in SRE, incident management, or cloud platform reliability roles.
  • Experience owning reliability for a specific major customer or “tenant” (e.g., dedicated instances, VPC deployments, or on-prem/isolated environments).
  • Experience participating in or running on-call rotations and incident management processes.
Education
  • (Not required) – Strong background in computer science or software engineering (B.S. in CS or equivalent experience).