Title: Site Reliability Engineering Manager (OpenSearch)
Bangalore, Karnataka, IN
Job Summary
NetApp is looking for Technical Operations Manager to join our growing NetApp Instaclustr team in India. NetApp’s Instaclustr offering provides open source as-a-service company, delivering reliability at scale. We manage cutting edge open-source technologies (OpenSearch, Kafka, Cassandra, Kafka, PostgreSQL, Valkey, Postgres, ClickHouse and Cadence) for our customers around the world.
NetApp Instaclustr makes it easy for our customers to run powerful open-source applications at the highest levels of scale. We have developed a platform that takes care of the whole lifecycle: provisioning infrastructure, installing applications and, most importantly, keeping the applications running reliably in production. Since being founded in 2013, Instaclustr has grown strongly, with over 300 customers worldwide, and over 30,000 nodes under management.
Our Technical Operations Engineers are the frontline team keeping our large fleet of cloud-hosted open-source clusters up and running. Your work will ensure the security, reliability and performance of world-class systems and databases. You will collaborate with our customer’s technical teams, from globally recognised companies in the gaming, banking and logistics industry sectors, ranging from big multinationals to emerging start-ups.[
Job Requirements
As TechOps Manager, you will be a leader in our TechOps team who maintain and support our large fleet of cloud-hosted Apache OpenSearch and Kafka open-source technology clusters. Every day will raise interesting and challenging problems ranging across many different areas such as fleet management, customer support, incident response and team development. Based from our Bangalore, office, you will be working closely with other teams, including Product Management, Development Teams and Customer Success to help drive and shape the way we support our customers.
Duties and Activities:
- Maintain an understanding of our product offerings (OpenSearch and Apache Kafka) as well as working knowledge of Instaclustr systems.
- Lead a team of talented engineers to deliver a very high quality of customer service and outcome. Including performance reviews and other people management activities.
- Reviewing the teams work demands and support requests daily, prioritizing them and ensuring timely responses are provided to the customer
- Ensuring that customer issues are responded to within SLA timeframes through staff training, ongoing queue management, allocation of appropriate resources.
- Project management of major operational tasks such as cluster migrations and fleetwide upgrades. This includes resource allocation and scheduling.
- Participate customer calls in different geographies related to the management of major operations
- Be aware of conflicting fleetwide operations and manage these appropriately to eliminate risk (e.g. fleet upgrades targeted to clusters under migration).
- Ensuring that defined processes and procedures are adhered to.
- Manage Shift Rostering
- Participate in our Level 3 on-call roster (Major Incident Manager) which includes rotating weekend coverage and contribute to post incident activities.
- Provide training to new team members and customers in TechOps operational procedures, and ensure these processes are adhered to.
- Be a proactive, reliable and supportive member of the TechOps team.
- Identify and drive the implementation of opportunities to continuously improve TechOps activities through the development and
- Maintenance of tools and procedures and/or working closely with the product engineering team.
Education
- We require you to have the right skills and maintain professional qualifications for the role covering:
- Proven experience in a Technical Lead capacity, with a clear interest in pursuing a Technical Manager career path, or recent experience in a Technical Manager role. Should have 8+ years of overall experience.
- Exceptional ability to communicate clearly and professionally in written and verbal English is essential as this role will involve customer interaction.
- Ability to perform the role as Major Incident Manager
- Demonstrated excellent experience with managing teams in challenging environments and take measures to keep the team morale high, drive positive work culture
- Previous system administration and/or programming experience is preferred
- Demonstrated ability to multitask, this is a critical attribute to have in this role
Job Segment:
Open Source, Engineering Manager, Engineer, Technology, Engineering, Marketing, Customer Service