Flexible work model
Work from home, from the office, or in a hybrid format that supports focus and collaboration.
ABOUT THE ROLE
In this role, you will own the reliability, evolution, and automation of cloud networking foundations supporting internal services and engineering productivity for a fast-growing US-based cloud data platform company. You will provide senior technical ownership for cloud networking and site reliability, making connectivity predictable, observable, recoverable, secure, and easier to operate.
RESPONSIBILITIES
Define cloud networking standards across cloud providers, covering routing, segmentation, firewalls, NAT, DNS, ingress, egress, and private connectivity
Own the technical direction of VPN and private-access connectivity, improving the reliability and operability of critical network services
Collaborate with Cloud Platform, DevPlat, Systems Engineering, IT, and InfoSec teams on connectivity architecture and technical decisions
Establish SLOs for critical network services and implement actionable monitoring and alerting using Grafana, Mimir, Loki, Pingdom, and PagerDuty
Lead diagnosis and mitigation of significant connectivity incidents, driving permanent remediation and maintaining runbooks, failure-mode documentation, and escalation paths
Develop infrastructure automation using Terraform or OpenTofu, Terragrunt, Python, and Go, with safe validation, deployment, and rollback practices
Guide VPN, bastion, DNS, device-trust, IPAM, and cloud-to-data-center connectivity patterns, including technologies such as Tailscale, Headscale, and NetBox
Mentor engineers and distribute operational knowledge through standards, architecture decisions, and operational practices that strengthen coverage across time zones
REQUIREMENTS
Significant experience with designing and operating production cloud networks across both GCP and AWS
Deep practical knowledge of routing, segmentation, firewalls, DNS, VPNs, bastion access, private connectivity, and cloud-network troubleshooting
Strong SRE experience with SLOs, observability, incident response, post-incident remediation, capacity planning, and disaster recovery
Hands-on experience with Infrastructure as Code using Terraform or OpenTofu and Terragrunt, plus infrastructure automation using Python, Go, or a comparable language
Practical experience with Kubernetes networking or adjacent platform infrastructure
Hands-on experience with FreeIPA and Keycloak, including practical understanding of cloud identity and access dependencies; familiarity with OIDC, SAML, SCIM, or JIT access is beneficial
Solid understanding of security best practices for cloud infrastructure, including safe infrastructure changes and operational guardrails
Hands-on experience with AI-driven infrastructure and workflows; experience with Tailscale, Headscale, WireGuard, NetBox, or mixed cloud and data-center environments is a plus
Proven ability to lead cross-team technical decisions, document architecture and operational standards, and collaborate without relying on formal authority
Demonstrated B2+ English proficiency, both written and spoken, for daily communication with the team and a US-based client, including technical ownership during US business hours and critical-incident escalation
SoftServe is an equal opportunity employer. Qualified applicants will receive consideration regardless of race, color, ancestry, ethnicity, national origin, religion, sex, sexual orientation, gender identity or expression, age, citizenship, disability, health condition, marital or family status, veteran status, or any other characteristic protected by applicable law.
Colombia, Chile
Remote/Office
Engineering & Technology
Clouds & DevOps
Senior
Fill out the form, and we'll be in touch shortly.
We are a digital engineering and technology consulting company where expertise grows alongside people. For more than 30 years, we have been elevating technology: helping organizations navigate complex business challenges by combining deep engineering knowledge with thoughtful, research-backed innovation. Our teams work across key areas: digital engineering, data and analytics, Сloud, and AI/ML. In each, we deliver practical, scalable solutions rooted in real business needs and measurable human impact.
You bring your perspective and ambition. We create an environment where your work meets clarity, confidence, and purpose.
Work from home, from the office, or in a hybrid format that supports focus and collaboration.
Competitive, market-based pay, benchmarked by role and location — plus health coverage, paid time off, wellness support, and learning opportunities.
Approachable leaders who communicate openly, keep teams close to the strategy, and support long-term planning.
Stay close to AI/ML, Cloud, Quantum Computing, IoT, and Robotics communities, with projects built on modern frameworks.