Cirrascale

Senior Cloud Engineer - Storage – AI & HPC Systems (Ceph / WEKA Expert)

San Diego, CAITFull-time
Apply Now

About Cirrascale

Cirrascale is at the forefront of providing cloud solutions tailored for AI and HPC workloads. Our mission is to create scalable and resilient storage infrastructures that empower businesses to achieve unprecedented performance and data integrity. We pride ourselves on fostering a collaborative and dynamic work environment.

Position Summary

We are seeking a Senior Cloud Engineer with deep expertise in distributed storage systems, specifically Ceph and WEKA, to architect, deploy, and maintain scalable storage infrastructures supporting AI and HPC workloads. This role is critical to ensure performance, resiliency, and data integrity across our customer environments. You will play a key role in supporting large-scale GPU infrastructure deployments, collaborating with engineering and operations teams to deliver best-in-class storage solutions tailored to the demanding requirements of AI and ML workloads.

Key Responsibilities

  • Architect, deploy, and manage high-performance, scalable storage solutions based on Ceph and WEKA for AI and deep learning workloads. - Design and implement storage infrastructures that support AI and deep learning workloads, ensuring high performance and scalability.
  • Optimize IOPS and throughput across distributed systems supporting hundreds of GPUs per cluster. - Enhance input/output operations per second and data throughput in distributed systems to efficiently support large GPU clusters.
  • Develop and maintain infrastructure-as-code templates for automated storage deployments. - Create and update code templates that automate the deployment of storage systems, ensuring consistency and efficiency.
  • Monitor system performance and implement improvements to ensure low latency, high bandwidth, and data integrity. - Regularly check system performance metrics and make necessary adjustments to maintain optimal latency, bandwidth, and data integrity.
  • Lead incident response and root cause analysis for storage-related issues across production environments. - Direct the response to storage incidents and conduct thorough analyses to identify and resolve root causes in production settings.
  • Collaborate with system engineers, network teams, and customer success to tailor storage performance to specific workload needs. - Work closely with various teams to customize storage solutions that meet the unique requirements of different workloads.

Requirements

  • 7+ years of experience managing and scaling enterprise storage systems.
  • 3+ years of hands-on experience with Ceph and/or WEKA in production environments.
  • Deep knowledge of storage architectures: object, block, and parallel file systems.
  • Strong understanding of RDMA, InfiniBand, NVMe-oF, and distributed metadata systems.
  • Proficiency in Linux (preferably Ubuntu/CentOS) and scripting (Bash, Python).
  • Experience with performance tuning for AI/ML workloads using storage-intensive frameworks like TensorFlow, PyTorch, etc.
  • Familiarity with containerized and virtualized environments: Docker, Kubernetes, KVM, etc.
  • Strong troubleshooting and diagnostic skills in large-scale, multi-tenant environments.

Additional Requirements

  • Availability for after-hours support as needed for system emergencies.
  • Willingness to participate in on-call rotations for production support.

Skills

Must Have

Expertise in Ceph for architecting scalable storage infrastructures.Proficiency in WEKA for deploying scalable storage solutions.Proficiency in Linux for managing enterprise storage systems.Experience with Python for scripting and automation.Strong understanding of RDMA for high-performance networking.Expertise in InfiniBand for high-speed data transfer.Knowledge of NVMe-oF for efficient storage networking.Familiarity with Docker for containerized environments.Experience with Kubernetes for managing containerized applications.Experience with TensorFlow for AI/ML workloads.Experience with infrastructure monitoring using Prometheus to ensure system reliability.Proficiency in API integration to facilitate seamless data exchange with cloud services.Experience in network management protocols to optimize data transfer processes.

Benefits

comprehensive medicaldentalvision coverage401(k) with company matchgenerous paid time-offprofessional development support
Apply Now