Mongodb

Cloud Operations Engineer

Save to Kiter
What Mongodb is looking for in applicants

MongoDB Atlas is the premier multi-cloud database as a service built and operated by the makers of MongoDB.

The Cloud Operations Engineering team at MongoDB is a worldwide team responsible for the consistent operational success of every MongoDB Atlas customer. As a Cloud Operations Engineer, you will help ensure the success of our Atlas customers, whether they are early startups or large multinational companies, cloud native or just getting started with a digital transformation to the cloud. You are excited about the core mission of MongoDB, and the opportunity to join the team responsible for operating Atlas, the fastest growing Multi cloud database as a service in the world. You are prepared to be one of the founding members of 24/7/365 global cloud operations team.

Cloud Operations Engineers will be responsible for day to day duties such as creating and monitoring systems alert dashboards, reviewing critical event and system logs, accessing customer instances that underpin their production databases and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace.

At MongoDB you will grow your career and skills, wear multiple hats, and be part of an operations team that works at the frontier of Cloud services and database systems.

Responsibilities

  • Troubleshoot, maintain and support the Atlas platform and associated server hardware, software and MongoDB databases ensuring optimum system integrity and performance for customers
  • Successfully coordinate with a global team of Cloud Operations Engineers who are tasked with ensuring our uptime guarantees to our Atlas customer base
  • Help scale the worldwide Cloud Operations Engineering team with the strategic implementation of new processes and tools
  • Assist in scoping, designing and deploying systems that reduce Mean Time to Resolve for customer incidents, including installation of diagnostic system or database tools
  • Monitor and detect emerging customer facing incidents on the Atlas platform; assist in their proactive resolution
  • Automate routine monitoring and troubleshooting tasks
  • Diagnose live incidents, differentiate between platform issues versus usage issues, and take the next steps toward resolution
  • Cooperate with our product management and cloud engineering organizations by identifying areas for improvement in the management applications powering the Atlas infrastructure
  • Inform executive leadership and escalation management personnel of major outages
  • Plan, coordinate and participate in a weekly on call rotation, where you will handle short term customer incidents (from direct surveillance or through alerts via our Technical Services Engineers)
  • Handle remediation activities where database corruption or hardware failures have occurred

Requirements

  • Experience with being an oncall DevOps, SRE, or Cloud Operations engineer (at least 4 years)
  • Expertise with Linux system administration and networking technologies like DNS,
  • TCP/IP, etc.
  • Knowledge of database installation, operations and disaster recovery,
  • including HA concepts like sharding and replication
  • Knowledgeable about a wide range of web and internet technologies
  • Familiarity with Amazon Web Services and other Cloud infrastructure platforms (e.g.GCP, Azure)
  • Experience in monitoring, system performance data collection and analysis, and reporting
  • Capability to write small programs/scripts to solve both short term systems problems
  • A CS/CE degree or equivalent experience
  • At least 1 of the following programming languages: Java, Go, Python, Javascript
  • A keen interest in learning new things

Nice to haves

  • MongoDB
  • Splunk
  • Kubernetes

To drive the personal growth and business impact of our employees, we’re committed to developing a supportive and enriching culture for everyone. From employee affinity groups, to fertility assistance and a generous parental leave policy, we value our employees’ wellbeing and want to support them along every step of their professional and personal journeys. Learn more about what it’s like to work at MongoDB, and help us make an impact on the world!

MongoDB is committed to providing any necessary accommodations for individuals with disabilities within our application and interview process. To request an accommodation due to a disability, please inform your recruiter.

MongoDB is an equal opportunities employer

 

Position: Cloud Operations Engineer

Department: Engineering, Support

Reports to: Karol Zapolski, Manager, Cloud Operations Engineering

Status: Full time, Permanent 

Want some tips on how to get an interview at Mongodb?

What is Mongodb looking for?
If this role looks interesting to you, a great first step is to understand what excites you about the team, product or mission. Take your time thinking about this and then tell the team! Get in touch and communicate that passion.
What are interviews for Cloud Operations Engineer like?
Interview processes vary by company, role and team. The best plan is to see what others have experienced and then plan accordingly.
How to land an interview at Cloud Operations Engineer?
A great first step is organizing your path to an offer. Check out Kiter for tools to get started!