Architecting high performance, highly-available (HA) and scalable technical solutions on Amazon Web Services.
Provide an efficient way to provision and modifying existing infrastructure using code.
Allow developers to create and deploy applications faster on various environments.
Provide an automated way to build, test & deploy apps to the cloud.
Provide least privilege access and shut down unnecessary services to reduce surface area for attack.
Choosing the right system to manage the data and provide optimum performance using various storage technology.
Create dashboards to understand current systems health and performance.
* Platform Ownership & Design
Authority: Serve as ultimate
technical authority for global HPC
Cloud infrastructure, owning design
decisions, technical standards, and
long-term platform evolution.
* Full-Stack Infrastructure: Oversee
end-to-end design and operation of
HPC web portals, usage metering,
telemetry, and high-throughput SLURM
job scheduling systems.
* Modernization & Resiliency:
Partner with HPC IT leadership to
audit legacy SOPs, drive
observability, and modernize
bare-metal automation, ensuring high
availability, security, and
compliance.
* Toil Reduction & Automation:
Standardize automated bare-metal
provisioning and network
orchestration to eliminate manual
operations and boost platform
stability.
* Data-Driven Performance: Analyse
telemetry across large-scale server
clusters to optimize workloads and
enforce SLAs, SLOs, and SLIs.
* Team Building & Talent
Acquisition: Build, retain, and
directly manage high-performing
systems software engineering teams;
partner with HR to recruit top-tier
infrastructure talent.
* Management Pipeline & Succession:
Actively coach senior individual
contributors into Engineering
Managers, building their
capabilities in project ownership,
delegation, and personnel
management.
* Career Growth & Mentorship:
Structure individual development
plans, deliver continuous feedback,
and establish clear technical and
managerial career progression
paths.
* Operational Leadership: Lead
day-to-day team operations, delegate
critical infrastructure projects,
and foster an engineering culture
centred on accountability,
mentorship, and high performance.
* Architecting high performance,
highly-available (HA) and scalable
technical solutions on AWS.
* Identify suitable technology
stacks and approaches to be used.
* Review existing application and
infrastructure with the intention to
improve existing system.
* Build tools to automate
operations, enhance productivity,
maintain and improve CI/CD
pipelines.
* Advocate and mentor team members
on DevOps best practices and
methodologies as well as collaborate
with development teams to adopt
DevOps practices.
* Act as a Subject Matter Expert to
the organization’s cloud end-to-end
architecture including networking,
monitoring, governance, BCDR, and
management.
* Update and maintain Jenkins
pipeline. Troubleshoot the pipeline
whenever there is issue.
* Update and maintain Terraform code
to ensure AWS resources are deployed
automatically via infra-as-code.
* Update and maintain Ansible code
to ensure application deployment on
EC2 instances can be done
automatically.
* Monitor systems performance and
issues via Instana. Solve any issues
detected during the monitoring
process.
* Participate in root cause analysis
whenever there is downtime on the
systems.
* Performs senior-level
responsibilities for the overall
system architecture, design,
installation, configuration,
technical support, and maintenance
of system mainly hosted in Major
Cloud Providers.
* Works with limited supervision to
establish, monitor and maintain
cloud environment, systems hardware,
operating systems, and related
network and security infrastructure
to ensure reliable operations.
* Monitors cloud systems for optimal
performance and establishes and
monitors best practices, policies,
and procedures.
* Optimize network infrastructure
for IaaS, SaaS, PaaS and other cloud
applications.
* Recommend and plan for future
growth of systems taking into
consideration capacity planning,
monitoring, Disaster Recovery, and
Business Continuity for operating
infrastructure.
* Research connectivity, performance
and related security issues to
determine root cause and implement a
plan of action to resolve these
issues.
* Lead and design technical
infrastructure & cloud processes,
integrate solutions into existing
infrastructure, consult on
development projects, help deploy
solutions that meet business and
technical requirements.
* Provide technical leadership to
teammates through coaching and
mentorship.
* Maintain the stability and
performance of the cloud platform.
Additionally, provide operational
support on all cloud solutions.
* Design, develop and maintain
Infrastructure-as-Code (IAC) that
automates and orchestrates
Continuous Integration and
Continuous Development (CI/CD) which
enable agile team to deliver quality
application.
* Design, develop and implement
monitoring system by using tools
such as Cloudwatch, NewRelic, Wormly
and PagerDuty alerts to ensure high
availability of the system.
Troubleshoot incidents and provide
Post Incident Report for future
system improvement.
* Drive projects assigned by
management to improve cost
efficiency, security and implement
industry best practices.
* Collaborate with agile team to
improve engineering tools, systems
and procedures.
* Mentor other engineers about
DevOps/SRE practices and help build
a fast growing team.
Kevin See — [email protected] — (+60) 16 225 1805