Sergei Iakovlev
Principal SRE — Ethereum staking infrastructure and platform engineering
- selfuryon@pm.me
- github.com/selfuryon
- linkedin.com/in/sergei-iakovlev
- Paphos, Cyprus · Remote
Summary
Principal SRE and tech lead of the Ethereum unit at P2P.org. Responsible for the reliability of 32,000+ mainnet validators (1.1M+ ETH) on bare-metal Kubernetes: 99.998% attestation participation, zero slashing on threshold-signed validators, #1 Lido operator by performance on Rated Network. Managed a team of up to four SREs as SRE manager in 2024–2026. 15 years in infrastructure: data-center networks first, Kubernetes platforms since 2019.
Experience
Principal SRE Engineer, Tech Lead · P2P.org 2026-02 – now
Non-custodial staking provider, #1 Lido operator by performance on Rated Network · Remote
Tech lead of the Ethereum unit: reliability, security and technical direction.
- Own the reliability of 32,000+ active mainnet validators (1.1M+ ETH) with zero slashing on threshold-signed validators: Vouch with 2-of-3 Dirk signing, distributed validators on SSV and Obol, 63 node pairs across four European regions.
- Keep the fleet at the top of the field: 99.998% attestation participation and 99.57% correct head votes against 98.7% for the network (Jul–Oct 2026). On Rated Network as of October 2026: #1 Lido operator over 7 and 30 days, #1 of all operators network-wide over 30 days and #3 over 7 days.
- Run five execution and four consensus clients in many different pairings, so a single client bug cannot take the fleet down.
- Own the Kubernetes reference architecture and security baseline used across all teams, agreed through RFCs; act as the technology advisor and technical interviewer for every engineering team.
- Lead DevSecOps practice: keyless cosign signatures and SLSA provenance for every image, binary and deployment artifact; SHA-pinned, audited CI workflows; reproducible builds with Nix.
- Write the unit’s services in Rust: validator inventory, exits via threshold BLS and Safe role policies, Lido reward claims with safety checks.
- Built network analytics on Xatu, Kafka and ClickHouse, and exposed it and Grafana to LLM agents over MCP, now used company-wide.
SRE Manager (Senior → Principal) · P2P.org 2024-06 – 2026-02
Led the Ethereum SRE team of up to four engineers, staying hands-on.
- Ran the team: hiring, 1:1s, growth plans and performance reviews.
- Organised on-call and incident management for the unit: rotations, escalation, P1–P5 severities by network; ran blameless postmortems with tracked action items.
- Built the reliability model: about 60 alert rules with severity by network, missed-duty detection from ClickHouse, daily reports on attestation correctness (rated.network methodology) and APR against Lido peers.
DevOps Engineer · P2P.org 2021-07 – 2024-06
Moved the Ethereum infrastructure off cloud VMs onto company-owned bare metal.
- Launched the company’s first bare-metal Kubernetes cluster and grew it into the internal platform used by every team. Moving workloads off cloud VMs cut infrastructure cost by more than 70%.
- Designed the delivery model for Ethereum workloads: CUE and Nix, Crossplane and Terraform on bare metal and three clouds, ArgoCD from OCI.
- Designed disaster recovery for signing: 2-of-3 Dirk threshold across three regions with slashing-protection databases, and wallet stores replicated to versioned, KMS-encrypted object storage in a separate cloud.
Senior DevOps Engineer · SoftPro 2019-09 – 2021-07
Software company · Moscow, Russia
- Led the move from VMs with apt packages and docker-compose to bare-metal Kubernetes: 7+ clusters of about 30 nodes across several data centers.
- Introduced infrastructure as code, CI/CD and observability as standard practice, wrote internal Terraform providers; cut deploy time from 40 minutes to 4.
Network Architect, Team Lead · HOST 2011-04 – 2019-09
IT integrator · Perm, Russia · Network Engineer (2011) → Senior (2014) → Architect, Team Lead (2017)
- Led the network engineering group: staffing across projects, workload planning, training; owned pre-sales for the group, including its budget.
- Designed multi-DC infrastructure for a bank (150+ servers, 1,000+ VMs) and replaced two data centers’ network for Russia’s largest reseller with no maintenance window over one hour.
Skills
- Ethereum
- Geth, Nethermind, Besu, Reth, Lighthouse, Prysm, Teku, Nimbus; Vouch, Dirk, SSV, Obol; commit-boost (MEV-Boost); Lido, StakeWise
- Platform
- Kubernetes on bare metal and in the cloud, Talos, Cilium, ArgoCD
- IaC
- Nix, CUE, Terraform, Crossplane
- Observability
- Prometheus, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, Alertmanager
- Data
- ClickHouse, Kafka, PostgreSQL
- Security
- HCP Vault, SLSA, Tailscale, CI hardening
- Programming
- Rust, Go, Python
Education
Computer Security, 5.5-year specialist degree · Perm State University 2006 – 2012
Faculty of Mechanics and Mathematics · Perm, Russia
Other
- At work
- 900+ merged PRs in 21 repositories and ~1,800 reviews over the last year
- Open source
- Maintainer of ethereum.nix; 130+ merged PRs to 22 projects since 2018 (Lighthouse, Dirk, Commit-Boost, Lido, nixpkgs); author of dkc and netdev
- Certifications
- CKA (2021); Cisco CCSP, Check Point CCSE, Palo Alto ACE, Huawei HCNA
- Languages
- Russian (native), English (professional working proficiency)