Lead Distributed Systems Software Engineer πŸ–₯️⚑

Permanent / Full Time

Lead Developer required

Role Overview 🌍
The Lead Distributed Systems Software Engineer architects, builds, and optimises large-scale backend systems that support millions of concurrent operations with reliability, speed, and elegance. You will guide a team of engineers developing cloud-native services, real-time data pipelines, and fault-tolerant distributed architectures that underpin mission-critical products.
This role is perfect for someone who thrives on designing systems that remain calm under extreme load, who treats latency like an arch-nemesis, and who enjoys turning complex distributed problems into beautifully engineered solutions.
Key Responsibilities πŸš€ 1. System Architecture & Engineering Leadership
  • Design high-availability distributed systems that scale horizontally across global infrastructure.
  • Lead development of backend services using microservice patterns, RPC frameworks, event streaming, and container orchestration.
  • Establish architectural standards for resilience, observability, testing, and documentation across engineering teams.
  • Collaborate with platform and SRE groups to ensure systems are robust under heavy traffic, fail gracefully, and recover autonomously.
2. Core Software Development 🧩
  • Write high-performance backend code in languages such as Go, Rust, Java, or Python.
  • Optimise memory, CPU usage, and network behaviour for distributed workloads.
  • Build internal libraries, frameworks, and tooling that elevate engineering velocity and reliability.
  • Own end-to-end delivery of featuresβ€”from design docs to rollout strategies.
3. Data Pipelines & Real-Time Processing πŸ“‘
  • Architect real-time streaming systems supporting telemetry, analytics, ML features, and user-facing experiences.
  • Work with technologies like Kafka, Pulsar, gRPC, Envoy, Redis, Bigtable-like databases, or custom sharded stores.
  • Ensure correctness, ordering guarantees, and performance for high-volume events.
4. Observability, Monitoring & Reliability πŸ› οΈ
  • Implement metrics, tracing, and logging systems that provide deep visibility into complex distributed behaviour.
  • Lead incident reviews and design fault-tolerant mechanisms such as circuit breakers, retries with backoff, load-shedding, and graceful degradation.
  • Define SLOs and drive efforts to consistently meet or exceed reliability targets.
5. Collaboration & Mentorship 🀝🌟
  • Partner with product managers to translate business needs into technical designs.
  • Mentor engineers at all levels, reviewing code, writing design documents, and modelling best practices.
  • Help grow a healthy engineering culture emphasising curiosity, craftsmanship, and measured risk-taking.
Required Qualifications πŸŽ“
  • 6+ years of professional experience building backend or distributed systems.
  • Strong fluency in at least one systems-level programming language.
  • Solid understanding of concurrency, network protocols, distributed consensus, and data consistency models.
  • Experience with cloud platforms, containerisation (Docker, etc.), orchestration (Kubernetes), and CI/CD pipelines.
  • Proven track record designing high-scale, low-latency systems.
Preferred Qualifications ⭐
  • Experience with service mesh architectures or zero-trust networking.
  • Familiarity with raft/paxos, vector clocks, CRDTs, or conflict-resolution strategies.
  • Background in real-time analytics or stream-processing frameworks.
  • Contributions to open-source distributed systems projects.
  • Ability to break down complex problems using clear, thoughtful design documents.
What Success Looks Like in 12 Months πŸ†
  • Core distributed services are faster, more reliable, easier to debug, and simpler for other teams to build on.
  • System latency decreases across multiple critical pathways.
  • A sustainable architecture emerges, reducing complexity while enabling new product capabilities.
  • Junior engineers evolve into confident system-thinkers under your guidance.
  • Major incidents drop due to better observability, improved failover logic, and more resilient designs.
Team Culture & Identity 🌱
This role belongs to someone who enjoys the dance between chaos and order in distributed systems. Someone who can look at logs, traces, packet timings, or flame graphs and feel the pulse of the machine. Someone who takes delight in shaving off milliseconds, eliminating single points of failure, and building systems that keep running even when the universe conspires otherwise.
  • Bigtable
  • Docker
  • Envoy
  • gRPC
  • Java
  • Kafka
  • Kubernetes
  • Pulsar
  • Python
  • Redis
  • REST
  • Rust
  • OnSite
  • Director
  • Distributed Systems
  • Observability
  • CI/CD

Referral reward: $750

IT & Telecomms > IT & Telecomms > Software - Developer

Back to Jobs