Engineering Group Lead – Reliability & SRE

We're a recruitment agency partnering with a global technology company building a leading AI-powered work platform used by teams worldwide. On their behalf we're looking for an Engineering Group Lead to own platform reliability, observability and incident response.

The role is based in Warsaw — a growing engineering hub working on large-scale product and infrastructure challenges. Hybrid model: 3 days a week from the office, 2 days from home.

You'll work in a modern, AI-assisted engineering environment: AI-powered IDEs, customisable agent rules, prompt engineering tooling and AI-infused CI/CD pipelines.

Your role

The group's mission is simple and critical: guarantee that the platform is always up and running. With hundreds of thousands of customers and continuous deployments happening daily, the platform faces constant pressure — and when things break, they are restored as fast as possible. The group is advancing agentic incident response and AI-SRE practices across three areas:

  • Traffic shaping — tooling that lets service owners contain unusual usage patterns and traffic spikes that could destabilise the platform
  • Observability — alerting when things misbehave, and accelerating inspection, troubleshooting and incident resolution
  • Reliability engineering — tooling and practices that combine software engineering and DevOps expertise

As Group Lead you will:

  • Own the mission and outcomes of the group — platform availability, incident response speed, and reducing the frequency of production incidents
  • Identify bottlenecks across the platform and lead the group to solve them systematically, e.g. improving observability coverage and detection rates
  • Work with the Group Tech Lead to define technical strategy and translate it into actionable scopes of work for each team
  • Manage and develop Team Leaders, supporting their growth and ensuring cross-team alignment
  • Understand the platform's architecture, usage patterns and product context deeply enough to identify the next set of problems worth solving
  • Own and improve incident response practices — tooling, escalation paths, post-mortems and preventative mechanisms

Offer

  • Total target monthly compensation: 73,000–80,000 PLN — base salary + bonus target + Restricted Stock Units (RSUs). Bonus and RSU grants are discretionary and depend on individual and company performance.
  • Employment contract (B2B is not an option)
  • Hybrid work: 3 days a week from the Warsaw office, 2 days from home
  • Free breakfast and lunch in the office from Monday to Wednesday
  • Comprehensive private medical care, life insurance and a Multisport card
  • Mental health support, including access to a mindfulness app
  • Discounts on partner products and services
  • Regular team get-togethers and events, plus gifts for birthdays and work anniversaries
  • A dedicated learning & development team: workshops, new skills and AI tooling
  • An equity incentive programme (for eligible roles)

Recruitment process

20–30 min intro call with the Talent Acquisition Partner → technical stages: system design and end-to-end interviews (1 h each) → final stages on-site in Warsaw: Management Interview and HR Interview (approx. 1 h each) → offer

Requirements

  • Strong distributed systems understanding — you can navigate a microservices environment, understand infrastructure and application-level reliability, and assess trade-offs across observability, safety mechanisms and operational tooling
  • Deep background in SRE or production engineering, and excitement about where the field is heading: agentic SRE, agentic incident response, AI-assisted operations
  • Senior engineering management experience — you've managed Team Leaders and Senior Technical Leaders and know how to develop people who run complex technical domains
  • A problem-solver who understands reality: the system, organisational constraints, and engineering-creative approaches that actually ship
  • Familiarity with observability best practices, traffic management and production reliability at scale
  • Experience owning end-to-end reliability programmes — defining what needs to be built and why, not just running team ceremonies
ID: 1388 job_post.published_on: 14/07/2026
announcement.apply