Cloud Operations Architect - AWS

This listing is synced directly from the company ATS.

Role Overview

This senior-level role leads IT operations and cloud architecture for a remote company, focusing on AWS infrastructure, CI/CD pipelines, and edge/GPU workloads. The architect will mentor a small team, implement operational processes, and ensure reliable, scalable, and cost-effective systems that enable engineering velocity. Impact includes improved incident response, security, and infrastructure maturity.

Perks & Benefits

This is a fully remote role open to candidates in Latin America, offering the flexibility to work from anywhere in the region. The company emphasizes professional growth through mentorship and development of cloud engineering skills. Benefits likely include a collaborative culture, opportunities to work with cutting-edge AI/ML and edge computing technologies, and a focus on work-life balance typical of remote tech roles.

Full Job Description

Cloud Operations Architect - AWS

Location: Remote (Latam)

About the Role

We're seeking a process-oriented Cloud Operations Architect - AWS to elevate our cloud, GPU compute, and edge device infrastructure maturity and lead our IT operations team. You'll bring operational rigor to our AWS environment, establish best practices for CI/CD pipelines, and mentor our IT team to deliver reliable, scalable infrastructure that enables engineering velocity.

Key Responsibilities

  • Team Leadership: Manage and mentor IT operations team member(s), providing technical guidance, establishing clear processes, and developing their cloud engineering capabilities through hands-on coaching

  • Cloud Architecture: Own and grow our AWS environment, implementing proper architecture patterns, cost optimization, security hardening, and disaster recovery procedures across web, AI/ML, and edge computing workloads

  • CI/CD Operations: Design and maintain automated CI/CD pipelines for infrastructure and application services, including support for containerized, hardware-dependent, and ML-related deployments

  • Process & Standards: Build operational maturity through documentation, runbooks, change management processes, incident response procedures, and knowledge transfer protocols across cloud and edge systems

  • Engineering Support: Partner with engineering teams to provide infrastructure that enables fast, safe deployments while maintaining system reliability and security

  • Incident Management: Lead incident response, conduct blameless post-mortems, and implement preventive measures to reduce recurring issues

Required Experience & Skills

  • Cloud Expertise: 5+ years working with AWS (EC2, RDS, S3, ECS, Fargate, IAM, IaC, networking, security groups) with hands-on architecture and troubleshooting experience

  • DevOps/CI-CD: Strong background with CI/CD tools (GitHub Actions, Jenkins, GitLab CI, CircleCI) including pipeline design, testing strategies, and deployment automation

  • People Management: Experience managing and developing technical team members, with patience and skill in coaching less experienced engineers through complex technical concepts

  • Process Orientation: Track record of implementing operational processes that stick - documentation standards, change management, on-call rotations, incident response

  • Security & Compliance: Understanding of cloud security best practices, IAM policies, SOC 2 considerations, and infrastructure-as-code for audit trails

  • Communication: Ability to translate technical complexity into clear explanations for both engineering teams and non-technical stakeholders

Nice to Have

  • Optimizing costs for ML/GPU workloads

  • CDK, Terraform or CloudFormation infrastructure-as-code experience

  • Kubernetes/container orchestration knowledge

  • Experience with monitoring/observability tools (DataDog, CloudWatch, Grafana, PagerDuty)

  • Background in SRE (Site Reliability Engineering) practices

  • AWS certifications (Solutions Architect, SysOps Administrator)

  • Experience with SOC 2 compliance and security audits

  • Scripting skills in Python, Bash, or (bonus points) Ruby for automation

What Success Looks Like

  • IT operations mistakes decrease significantly through better processes and mentorship

  • AWS infrastructure follows architecture best practices with documented patterns

  • CI/CD pipelines are reliable, fast, and trusted by engineering teams

  • Team member(s) develop stronger cloud engineering skills and confidence

  • Incident response times improve with clear runbooks and escalation procedures

  • Infrastructure costs are optimized without sacrificing reliability or performance

Similar jobs

Found 6 similar jobs

Browse more jobs in: