Metal Toad
DevOps Engineer
About this role
Metal Toad is a strategic AI partner. We help our customers to use AI to build the engines of their future growth, aligning AI’s capabilities with their most important business opportunities. We are an AWS, Anthropic, and OpenAI partner. Our team includes consultants, software engineers, devops engineers, UX designers, project managers, marketers, and supporting staff. It should go without saying, but everyone working at Metal Toad should be interested in the impact of AI, and believe in a human-first, AI-enabled future. Metal Toad is a fully remote company, offering all team members the ability to work from home. Although Metal Toad is a remote company, we are currently seeking contractors residing in Brazil for this position. Compensation will be made in Brazilian Reais (R$), in accordance with the local currency of the country where the position is located. Due to legal limitations, we are unable to sponsor any type of visa. To apply for this position, please submit your resume in English as a PDF. Job Description The Cloud Engineer position at Metal Toad requires experience in designing and maintaining infrastructure for high-availability, scalable, enterprise-grade applications. You will be part of a talented team that works on mission-critical applications. Responsibilities Planning
- Analyzing customer requirements for software components, system availability, security, and performance.
- Designing and documenting complete cloud hosting systems, including capacity planning software and instance type selection, allocation, and network design.
- Estimating the costs of the recommended system design.
- Building systems by executing installation, configuration, and testing of cloud resources.
- Using automation and configuration management to ensure repeatability and traceability of changes. Managed Services
- Troubleshooting system hardware, software, networks, and operating systems.
- Protecting the integrity and security of systems through proper use of controls and monitoring tools, and providing written evaluations and recommendations for ongoing improvement.
- Maintaining system performance through system monitoring and analysis, performance tuning, and planning for future growth.
- Designing and running load and stress tests, documenting outcomes, debugging infrastructure issues, and escalating documented application problems to the development team.
- Maintaining internal systems and customer deployment documentation.
- Partnering with project managers, technical consultants, software architects, and developers to validate infrastructure deliverables against the requirements and document all technical hand-offs.
- Experience with Amazon Web Services (AWS).
- Responding to support tickets and incidents in a timely manner that corresponds to SLA commitments. Expertise