


Team Lead - Site Reliability
Our client is an international producer of a patented software platform predicted to become number 1 technology trend in 2022 (according to Gartner), and one of the key disruptors in the Enterprise space.
The company has a proven track record of 10+ years in the enterprise data integration and management space, growing consistently year on year and supporting Fortune 1000 customers making sense of the data (among them AT&T, Verizon, Vodafone, GlobalPayments, etc).
The company's platform (Fabric) processes data from multiple systems, enriches it with real-time insights and transforms it into a patented Micro-Database - one for every customer, based on particular requirements.
To maximize performance, scale and security, every Micro DB is compressed and individually encrypted. It is then delivered in milliseconds to fuel quick, effective and pleasing customer interactions.
We're looking for a Lead SRE who will coordinate a multicultural SRE team (Romania & Israel, covering 24x7 managed platform services) - approx. 15 members.
This role best fits those candidates who are ready for a full time people management role, diving into people processes and supporting the team grow on multiple levels.
The SRE team focuses on securing high availability, reliability and performance of the system within a managed PaaS environment, powered by all of the public cloud solutions.
Tech stack: Linux, Prometheus, ELK, Grafana, Honeycomb, PostgreSQL, Docker, Kubernetes, AWS, Azure, GCP, Oracle Cloud, etc.
As a Team Lead - SRE, your role will be to:
Build and lead the SRE team (10-20 engineers), mentoring the more junior members of the team.
Work with and introduce the latest monitoring technologies - Prometheus, Honeycomb, Grafana, ELK, etc.
As a senior member of the team, input into technical discussions, decisions and strategy;
Drive quality accountability within the organization with well defined processes, metrics and goals for process quality;
Work with clients, stakeholders and senior executives in the team to showcase progress on strategic initiatives;
Support new installations of Managed PaaS for new customers;
Drive various production investigations covering technologies such as Cassandra, Golden Gate & Kafka.
To thrive in this role, you'll need to demonstrate:
Leadership skills and genuine interest in working with people and growing talents (we don't require experience in a people management role, as long as you're up for the challenge);
Good experience with Linux (including scripting);
Knowledge of DevOps environment / Containerisation (Docker, Kubernetes);
Some experience working with large amounts of data and databases such as PostgreSQL, SQLite is nice to have, definitely not mandatory;
Experience with monitoring tools (Grafana/Data dog);
Ability to analyze/debug large and complicated systems;
Project and process management skills, with a focus on process improvement;
A passion for performance excellence, robustness and engineering mindset;
Advanced level of English.
Why you should consider this role?
You'll be part of one of the key disruptors in the market, working on innovative solutions and shaping the future of data;
Your package will include: attractive salary (according to expertise), stock option plans, meal vouchers, gym subscription, medical services.
You'll work in a hybrid model, with offices located downtown at Piata Victoriei.
If you feel you still need to fill some gaps for this position don't give up, let's talk first.
If interested: send your resume to recruitment@itworx.ro and we'll reach out to you shortly.
Else: you can refer an interested friend or acquaintance.