We are seeking a Principal Server Engineer to lead the development and maintenance of backend systems that support our live games, currently serving over 1 million daily active users. This role demands ownership of performance, scalability, and reliability at the core of our infrastructure. You will collaborate closely with the engineering team on various aspects, including schema design, incident response, and system optimization to ensure seamless gameplay experiences.
Key Responsibilities
- Design, build, and maintain backend services using Go and Python for high-traffic live game titles.
- Own database schema design and optimize query performance within PostgreSQL environments.
- Develop and implement horizontal scaling strategies focused on optimizing CPU and memory usage while minimizing lock contention, favoring horizontal scaling over vertical.
- Profile and troubleshoot production issues such as memory leaks, CPU spikes, and deadlocks.
- Lead root-cause investigations for performance bottlenecks and outages, managing the process from detection to resolution.
- Architect and maintain in-memory caching and state layers using technologies like Redis, KeyDB, or Memcached.
- Design and execute load tests to identify system breaking points prior to release.
Required Qualifications
- Extensive expertise in Go and Python, including deep understanding of language internals, execution models, and performance tuning techniques such as goroutine management, garbage collection tuning, and zero-allocation strategies.
- Strong proficiency in PostgreSQL, with experience in scalable schema design, query plan analysis, and index optimization, supported by real-world examples of resolving performance bottlenecks through data layout or query architecture improvements.
- Proven track record of scaling backend systems to support at least 350,000 daily active users, emphasizing horizontal scaling and efficient serialization/deserialization methods like Protobuf or MsgPack.
- Hands-on experience with profiling tools to identify and resolve memory leaks, CPU spikes, and deadlocks in production environments.
- Demonstrated ability to analyze and resolve complex production performance issues or outages, with at least two detailed examples of root cause analysis and resolution.
- Experience as a core backend engineer on at least two released production projects handling live traffic.
Preferred Qualifications and Benefits
- Experience with caching and state management solutions such as Redis, KeyDB, or Memcached, including techniques for cache stampede mitigation, serialization overhead reduction, and eviction policy design.
- Strong judgment in test-driven and behavior-driven development (TDD/BDD) practices, along with practical experience in load testing tools like Artillery.
- Familiarity with the Nakama framework, including custom Go/Lua modules, matchmakers, storage engines, and real-time multiplayer socket management.
- Practical experience deploying, monitoring, and scaling serverless or containerized services on Google Cloud Platform (GCP) or Google App Engine.
Benefits
- Competitive, tax-free USD salary packages.
- Private health insurance coverage.
- Paid time off to support work-life balance.
- Performance bonuses aligned with individual and company success.
- Annual performance reviews to foster career growth and development.
This role offers an exciting opportunity to impact millions of players worldwide by ensuring the backend infrastructure is robust, scalable, and efficient. If you have a passion for backend engineering and live game systems, this position will allow you to leverage your expertise in a dynamic and fast-paced environment.