Lead Distributed Systems Software Engineer π₯οΈβ‘
Lead Developer required
Role Overview π
The Lead Distributed Systems Software Engineer architects, builds, and optimises large-scale backend systems that support millions of concurrent operations with reliability, speed, and elegance. You will guide a team of engineers developing cloud-native services, real-time data pipelines, and fault-tolerant distributed architectures that underpin mission-critical products.
This role is perfect for someone who thrives on designing systems that remain calm under extreme load, who treats latency like an arch-nemesis, and who enjoys turning complex distributed problems into beautifully engineered solutions.
Key Responsibilities π 1. System Architecture & Engineering Leadership
This role belongs to someone who enjoys the dance between chaos and order in distributed systems. Someone who can look at logs, traces, packet timings, or flame graphs and feel the pulse of the machine. Someone who takes delight in shaving off milliseconds, eliminating single points of failure, and building systems that keep running even when the universe conspires otherwise.
The Lead Distributed Systems Software Engineer architects, builds, and optimises large-scale backend systems that support millions of concurrent operations with reliability, speed, and elegance. You will guide a team of engineers developing cloud-native services, real-time data pipelines, and fault-tolerant distributed architectures that underpin mission-critical products.
This role is perfect for someone who thrives on designing systems that remain calm under extreme load, who treats latency like an arch-nemesis, and who enjoys turning complex distributed problems into beautifully engineered solutions.
Key Responsibilities π 1. System Architecture & Engineering Leadership
- Design high-availability distributed systems that scale horizontally across global infrastructure.
- Lead development of backend services using microservice patterns, RPC frameworks, event streaming, and container orchestration.
- Establish architectural standards for resilience, observability, testing, and documentation across engineering teams.
- Collaborate with platform and SRE groups to ensure systems are robust under heavy traffic, fail gracefully, and recover autonomously.
- Write high-performance backend code in languages such as Go, Rust, Java, or Python.
- Optimise memory, CPU usage, and network behaviour for distributed workloads.
- Build internal libraries, frameworks, and tooling that elevate engineering velocity and reliability.
- Own end-to-end delivery of featuresβfrom design docs to rollout strategies.
- Architect real-time streaming systems supporting telemetry, analytics, ML features, and user-facing experiences.
- Work with technologies like Kafka, Pulsar, gRPC, Envoy, Redis, Bigtable-like databases, or custom sharded stores.
- Ensure correctness, ordering guarantees, and performance for high-volume events.
- Implement metrics, tracing, and logging systems that provide deep visibility into complex distributed behaviour.
- Lead incident reviews and design fault-tolerant mechanisms such as circuit breakers, retries with backoff, load-shedding, and graceful degradation.
- Define SLOs and drive efforts to consistently meet or exceed reliability targets.
- Partner with product managers to translate business needs into technical designs.
- Mentor engineers at all levels, reviewing code, writing design documents, and modelling best practices.
- Help grow a healthy engineering culture emphasising curiosity, craftsmanship, and measured risk-taking.
- 6+ years of professional experience building backend or distributed systems.
- Strong fluency in at least one systems-level programming language.
- Solid understanding of concurrency, network protocols, distributed consensus, and data consistency models.
- Experience with cloud platforms, containerisation (Docker, etc.), orchestration (Kubernetes), and CI/CD pipelines.
- Proven track record designing high-scale, low-latency systems.
- Experience with service mesh architectures or zero-trust networking.
- Familiarity with raft/paxos, vector clocks, CRDTs, or conflict-resolution strategies.
- Background in real-time analytics or stream-processing frameworks.
- Contributions to open-source distributed systems projects.
- Ability to break down complex problems using clear, thoughtful design documents.
- Core distributed services are faster, more reliable, easier to debug, and simpler for other teams to build on.
- System latency decreases across multiple critical pathways.
- A sustainable architecture emerges, reducing complexity while enabling new product capabilities.
- Junior engineers evolve into confident system-thinkers under your guidance.
- Major incidents drop due to better observability, improved failover logic, and more resilient designs.
This role belongs to someone who enjoys the dance between chaos and order in distributed systems. Someone who can look at logs, traces, packet timings, or flame graphs and feel the pulse of the machine. Someone who takes delight in shaving off milliseconds, eliminating single points of failure, and building systems that keep running even when the universe conspires otherwise.
Referral reward: $750
IT & Telecomms > IT & Telecomms > Software - Developer