Nvidia Corporation
Software Engineering Manager - Cloud Infrastructure Services, DGX Cloud (Finance)
As a Site Reliability Engineering leader you will manage the operations of our observability platform focused on multi-colo distributed NVIDIA GPU cloud clusters. You will be the leader for all aspects of cluster operational excellence planning and grow your team. You thrive in a fast-paced iterative engineering environment and have experience delivering scalable distributed systems. Most importantly, you will have a track record of having past teams respect you as both a technical leader and manager. NVIDIA DGX Cloud Computing team is responsible to work all across the company, in areas such as information retrieval, artificial intelligence, natural language processing, distributed computing, large-scale system design, Life science, Image Processing; the list goes on and is growing every day in Machine Learning. Operating with scale and speed, our world-class software engineers are just getting started -- and as a manager, you guide the way to solve reliability both our internally critical and our externally-visible systems.
What you'll be doing:
The base salary range is 200,000 USD - 385,250 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.
You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.
NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.