Lead Site Reliability Engineer – MSSQL (Sri Lanka)
About Platned
Platned is a rapidly expanding Enterprise Services company headquartered in the UK with a growing global presence. We are business transformation specialists, helping organisations maximise their investment in the world-class IFS product suite through comprehensive implementation, consultancy, and managed services.
Role Overview
The successful candidate will act as a technical lead for Microsoft SQL Server environments supporting highly available and business-critical workloads.
This role requires significantly more than general SQL Server administration. Candidates must have strong, demonstrable hands-on experience designing, configuring, troubleshooting and operating complex SQL Server High Availability and Disaster Recovery architectures.
You should be comfortable working with and explaining multi-node SQL Server architectures, including environments with primary and secondary replicas across multiple nodes, and understand how clustering, Failover Cluster Instances and Availability Groups interact at both SQL Server and Windows infrastructure levels.
You will also be expected to lead complex database performance investigations, planned and unplanned failover activities, production incidents, database validation and recovery activities.
Key Responsibilities
- Design, implement, administer and troubleshoot highly available Microsoft SQL Server environments supporting business-critical applications.
- Design and support SQL Server High Availability and Disaster Recovery solutions, including WSFC, Failover Cluster Instances (FCI) and Always On Availability Groups (AG).
- Manage multi-node SQL Server architectures, including primary, secondary and Disaster Recovery environments, and ensure database replication and synchronisation meet defined RPO and RTO requirements.
- Plan and execute SQL Server failovers, recovery activities and Disaster Recovery exercises, ensuring minimal business impact and maintaining application availability.
- Monitor and maintain SQL Server health, availability, performance, database synchronisation, backups, storage, transaction logs and overall operational reliability.
- Lead complex SQL Server performance investigations, using execution plans, Query Store, DMVs, wait statistics, Extended Events and performance monitoring tools to identify and resolve performance bottlenecks.
- Troubleshoot and optimise SQL Server parallelism, including MAXDOP, Cost Threshold for Parallelism, CXPACKET and CXCONSUMER waits, taking into account server hardware, NUMA architecture and workload characteristics.
- Identify and resolve issues related to blocking, locking, deadlocks, long-running queries, CPU, memory, storage and I/O performance.
- Install, configure, patch, upgrade and maintain Microsoft SQL Server environments, including applying security updates and cumulative updates.
- Manage SQL Server Agent jobs, maintenance activities, database housekeeping, backup and recovery processes, including database restores and point-in-time recovery.
- Support database deployments, schema changes, release validation and rollback activities while ensuring database integrity and availability.
- Act as a senior technical escalation point for complex SQL Server incidents, outages, performance issues and HA/DR failures.
- Lead root cause analysis and implement permanent solutions to improve database reliability, resilience and operational stability.
- Troubleshoot issues across SQL Server, Windows Server, clustering, storage, networking, firewalls, Active Directory and application connectivity, working closely with relevant infrastructure and application teams.
- Maintain database architecture, operational procedures, recovery documentation and technical standards, ensuring compliance with security, governance, change management and business continuity requirements.
- Conduct regular database health assessments, capacity planning and reliability reviews to identify and address potential risks before they impact production systems.
Knowledge, Skills & Abilities
- Strong hands-on expertise in Microsoft SQL Server administration, with experience supporting complex enterprise and production environments.
- In-depth knowledge of SQL Server High Availability and Disaster Recovery, including Always On Availability Groups, Failover Cluster Instances (FCI) and Windows Server Failover Clustering (WSFC).
- Strong understanding of SQL Server architecture, clustering, replication, failover, database synchronisation and recovery, including RPO/RTO requirements.
- Advanced SQL Server performance tuning and troubleshooting skills, including execution plans, Query Store, DMVs, wait statistics, blocking, locking, deadlocks, CPU, memory, storage and I/O analysis.
- Strong knowledge of SQL Server query optimisation and parallelism, including MAXDOP, Cost Threshold for Parallelism, CXPACKET and CXCONSUMER.
- Strong T-SQL and database administration skills, including backup and recovery, database maintenance, patching, upgrades and deployment activities.
- Good knowledge of Windows Server, networking, storage, Active Directory and infrastructure components relevant to SQL Server environments.
- Strong analytical and problem-solving skills, with the ability to diagnose complex technical issues and identify permanent solutions.
- Ability to lead critical production incidents, failover and Disaster Recovery activities and make sound technical decisions under pressure.
- Strong communication and stakeholder management skills, with the ability to explain complex technical concepts to both technical and non-technical audiences.
- Ability to review architectures, identify availability, performance and reliability risks, and recommend appropriate technical solutions.
- Strong documentation and technical writing skills, with the ability to maintain operational procedures, recovery plans and technical documentation.
- Ability to mentor and support engineers, share technical knowledge and provide technical leadership within the team.
- Relevant Microsoft database certification is required; additional Azure, Windows Server, ITIL or infrastructure certifications will be an advantage.
- Experience with IFS, ERP platforms, Microsoft Azure, PowerShell, database monitoring and automation tools will be considered an advantage.
What We Offer
- High-Growth Environment: Join a company scaling rapidly across global markets.
- IFS-Focused: Work exclusively with best-in-class IFS technology.
- Competitive Package: Attractive salary, commission structure & benefits.
- Career Development: Clear progression as we expand our global footprint.
How to Apply
If you are an experienced Microsoft SQL Server professional with deep hands-on expertise in Always On Availability Groups, Failover Cluster Instances, Windows Server Failover Clustering, SQL Server performance engineering and enterprise HA/DR and ready to make a meaningful impact in a high-growth global company, we would love to hear from you. Submit your CV via the Platned Careers Portal or email directly to careers@platned.com.
Due to our continued success and rapid growth, interviews will be conducted promptly and decisions made within a short timeframe.
Join our team
Let’s talk next steps towards your accelerrated growth.