INFRASTRUCTURE ENGINEER

  • Full-time
  • Infrastructure
  • Remote

ABOUT BETBY

As a fast-growing, award-winning company, Betby powers the industry with our premium sportsbook, featuring world-class risk management and seamless omni-channel support, reaching millions of players across countless markets.


With offices in Latvia, Malta, Spain, Greece and Montenegro, we offer a vibrant work culture, relocation opportunities, and full support for remote talent across the globe.


Join Betby and unlock endless opportunities for growth, success, and making a real impact in the world of iGaming!

SHORT DESCRIPTION

At BETBY, our team is responsible for building, managing and maintaining secure, robust and scalable infrastructure to run our prod and beta/test/dev environments and core services using DevOps principles. This includes designing, deploying, operating, scaling, automating, and standardizing monitoring, logging, and alerting systems across on-premise and Kubernetes environments, as well as cloud infrastructure hosted within AWS.


We are looking for a strong, experienced technical person to join our team to help take observability maturity to the next level using the right tools. You will make infrastructure and business-critical services easier to understand and operate through reliable metrics and log pipelines, useful dashboards, actionable alerts, and clear operational standards. If you are a hardworking and result-oriented person, please apply now!

RESPONSIBILITIES

  • Designing, deploying, configuring, and maintaining scalable monitoring, logging, and alerting platforms for production and beta/test/dev environments;
  • Operating Prometheus, Alertmanager, Grafana, Fluent Bit, Kafka, Fluentd, OpenSearch, and OpenSearch Dashboards, including upgrades, reliability, availability, capacity, and retention planning;
  • Building and maintaining reliable metrics and log collection pipelines for infrastructure and business-critical services;
  • Creating dashboards that provide clear, useful visibility into service health, performance, capacity, and operational risks;
  • Designing, tuning, and maintaining actionable alert rules and notification routing; reducing alert noise and improving incident response;
  • Monitoring infrastructure and application metrics and logs, troubleshooting issues, and improving stability and performance under heavy loads;
  • Managing metric cardinality, log volume, retention, storage consumption, and query performance to keep observability platforms scalable and cost-effective;
  • Establishing high-availability and recovery approaches for observability services and validating operational readiness;
  • Automating configuration management and standardizing observability configuration through Ansible, Terraform, Python, and bash;
  • Developing self-service observability patterns, reusable dashboards, alert templates, and documentation for engineering teams;
  • Supporting production incidents, investigating root causes with telemetry, and improving dashboards, alerts, and runbooks after incidents;
  • Evaluating new technologies and their implementation in existing infrastructure;
  • Working with Kubernetes, Linux systems, networking, databases, and message brokers to ensure meaningful observability coverage;
  • Maintaining and writing documentation of observability architecture, configurations, standards, and operational procedures.

REQUIREMENTS

  • Minimum 3 years of experience with administering Linux systems and operating monitoring, logging, or observability systems;
  • Experience with Debian-based systems;
  • Experience with Docker and Kubernetes;
  • Hands-on experience with Prometheus, Alertmanager, and Grafana, including metric collection, alert rules, routing, and dashboards; experience with VictoriaMetrics would be a plus;
  • Experience with Fluent Bit, Kafka, Fluentd, OpenSearch, and OpenSearch Dashboards for log collection, transport, processing, storage, search, and visualization;
  • Understanding of metrics and log pipeline design, including reliability, scalability, data retention, capacity planning, and cardinality management;
  • Experience designing actionable alerts, reducing alert noise, and troubleshooting infrastructure and application issues using metrics and logs;
  • Proficiency in shell command line usage, scripting, and automation tools like Ansible/Terraform;
  • Python and bash scripting skills;
  • Understanding of networking concepts, including TCP/IP, DNS, VPN, Firewalls and the ability to configure and troubleshoot network settings;
  • Experience operating highly available services and planning capacity for production workloads;
  • Experience with configuration-as-code, Git-based workflows, and enabling self-service observability for engineering teams;


OUR STACK

  • OS: Debian;
  • Virtualization: KVM, Proxmox;
  • Storages: Ceph, S3;
  • Networking: IPsec, Open vSwitch, iptables, VRRP, OpenVPN;
  • DBMS: MongoDB, PostgreSQL, ClickHouse;
  • Message brokers: Kafka, RabbitMQ;
  • Service orchestration: Kubernetes;
  • Monitoring systems: VictoriaMetrics, Prometheus, Alertmanager, Grafana;
  • Logging pipeline: Fluent Bit, Kafka, Fluentd, OpenSearch, OpenSearch Dashboards;
  • Revision control and CI/CD tools: GitLab;
  • Cloud services: Amazon Cloudfront/WAF/S3/EC2/EKS/ELB;
  • Web services: Nginx, HAproxy;
  • Configuration management: Ansible, Terraform;
  • Scripting: Python, bash.

PERKS AND BENEFITS

  • Comprehensive health insurance with coverage for your well-being
  • Paid sick leave up to 10 days without medical certificate
  • 20 days of paid vacation plus additional leave for important life events
  • Learning and growth opportunities with support for professional development
  • Language learning support for multilingual collaboration
  • Modern hardware provided for your work
  • International team environment across multiple countries
  • Corporate events and team activities
  • Welfare support program for critical situations
  • Gifts and support for major life milestones 

TAGS

#level middle

#relocation yes

#category infrastructure