Technology Strategy · December 2025

Where We Are &
Where We're Going

An honest look at our current challenges, what we're doing about them, and what it will take to get to a better place.

Context

How We Got Here

The platform we've built over the past 9 years has served us well. It's taken us from zero to 450k+ PAPs and £120M+ in revenue without requiring significant investment in technology or platform infrastructure. But we're now at a point where the gaps and debt we've accumulated need deliberate investment to address.

What follows is an honest breakdown of where we are, what we're doing, and what we still need.

For context: We had 25 incidents in Q4 (Sept–Nov). Roughly half were related to external providers (game providers, payment providers, Cloudflare). The rest were internal — and many of those map directly to the gaps we're addressing.

Overview

The Challenges We're Addressing

1. Infrastructure

Migrate from Layershift to AWS — resolve network issues, enable infrastructure-as-code

2. Build & Deployment

Replace ageing build system with modern tooling fit for today's world

3. Architecture Ownership

No one owns the high-level architecture view across all services and squads — issues growing unchecked

4. Observability

Centralise logs and metrics, enable proactive alerting and faster diagnosis

5. Platform Engineering

Load testing, performance profiling, staging environment — the foundations

6. QA & Test Automation

1:1 squad parity, E2E automation on critical flows, free QAs for exploratory testing

7. Cyber Security

No one owns this today — hire specialist, assess posture, establish governance

Context: Product Function

Lack of product function has stretched tech leadership and impacted squad direction

Q1 Priority

1. Infrastructure: Layershift → AWS

Layershift has been good value as we've grown, but we're seeing intermittent network issues, deployment inconsistencies, and we can't achieve infrastructure-as-code — everything is manual click-ops.

What We're Doing

This isn't a lift-and-shift. We're using the migration as an opportunity to completely rework our build and delivery pipelines to be dynamic, resilient, and flexible.

Timeline

  • Q1: All MrQ services migrated to AWS
  • Q2+: PQ modernisation work continues (more complex, longer timeline)

⚠️ Capacity required: 1–2 backend engineers per squad for a full month in Q1, focused solely on migration work

📊 Q4 impact: 5+ infrastructure incidents in Q4 were directly related to Layershift limitations — connectivity issues, network failures, socket timeouts, service unavailability.

Q2 Initiative

2. Build & Deployment Pipeline

Once on AWS, our new infrastructure-as-code pipelines will still be running through our current build system — an ageing setup that doesn't fit our needs in an AI-powered world where build and deployment tooling has evolved significantly.

The Plan

After the AWS migration completes, Platform Squad will evaluate modern alternatives (GitHub Actions, GitLab CI, etc.) and select a replacement that better supports our ways of working.

Timeline

  • Q1: Investigation and tool selection
  • Q2: Rollout to all squads
Needs Hire — Q1

3. Architecture Ownership

The gap: I used to own architecture across all services. When I stepped back 2 years ago, no one filled that gap. Issues have grown unchecked without anyone owning the high-level view across all microservices and squads.

What we've done: In the last 2–3 months, I've formed an architecture function with three IC architects — Shawn (MrQ Backend), Gonzalo (Frontend), Diego (Games). They're my "Avengers" — I can deploy any of them to zoom into a problem and work with squads to get it done. The impact has already been significant. Imagine if we had more of them.

What we need: I'm still directing them personally, which isn't sustainable. We need a Head of Architecture who can own this full-time: understand our current state, surface the unknown unknowns, identify risks and opportunities, and chart where we need to go. This person would lead the architecture function and free me to focus on broader technology strategy.

📅 Timeline: Hire in Q1. First 90 days focused on assessment and building the architecture roadmap.

Q1 Rock

4. Observability

The Challenge

Our metrics, dashboards, logs, and alerts are scattered across multiple tools with no unified view. No anomaly detection. When something goes wrong, we're often reactive rather than proactive, and diagnosis depends too heavily on individual knowledge. We've outgrown our current tooling.

Where We Need to Be

Comprehensive observability means: centralised logs and metrics, proactive alerts that catch issues before customers do, anomaly detection, and dashboards that let any engineer diagnose problems quickly — not just the ones who built the system.

📅 Timeline: Q1 is foundation — adopt a mature observability platform (under evaluation). Then 2–3 quarters to achieve comprehensive coverage across all critical player flows. This will require ongoing Platform Squad capacity and squad involvement to instrument their services properly.

📊 Q4 impact: Many Q4 incidents had extended impact because we detected issues slowly or couldn't diagnose root causes quickly. Better observability means faster detection, faster resolution, and smaller blast radius.

Q1 Rock

5. Platform Engineering

Why it matters: A mature Platform Squad provides the shared tooling, infrastructure, and practices that let product squads move fast without reinventing the wheel. They own developer experience, reliability engineering, and the foundations everyone else builds on.

Where we are: Our Platform Squad is only 2–3 quarters old. They're already making significant impact, but we have gaps:

Load Testing

We've only just started. Our current framework tests endpoints in isolation — we need to simulate real player behaviour at scale, especially before peak events.

Performance Profiling

The architects have pushed this at squad level and it's helped catch bottlenecks early. But it's manual and inconsistent. We need automated profiling on every service, mandated for all squads on the critical path.

Staging Environment

Our staging environment doesn't reflect production well enough. We need staging to be much closer to prod so we can properly replicate and test performance issues before they hit customers.

📅 Timeline: Q1 establishes the foundations (framework, tooling, initial automation). Q2–Q3 for mature load testing and profiling across critical paths. Staging-like-prod is a longer effort. Platform Squad leads, but squads will need to invest time instrumenting and testing their own services.

📊 Q4 impact: 3 scaling/performance incidents in Q4 — including critical issues during peak weekender traffic — would have been caught or prevented with proper load testing and performance profiling.

Q1 Rock

6. QA

Current Challenges

  • Until now, a manual QA was shared across 2 squads — we underestimated the pressure this created
  • Limited E2E automation coverage on critical player flows
  • Manual QA as a bottleneck slows down deployments
  • QAs stretched thin, unable to do proper exploratory testing

The Goal

  • 1:1 parity between QA and squads (final SDET joins January)
  • Close to 100% automation coverage on core player flows
  • Engineers can deploy confidently without manual QA bottleneck
  • Manual QAs freed for true exploratory testing and edge cases

📅 Timeline: Q1 focuses on achieving 1:1 parity, tooling selection, and beginning E2E automation on critical flows. Q2–Q3 expands automation coverage progressively. This is a multi-quarter investment — there are no shortcuts to comprehensive test coverage.

📊 Q4 impact: 6 frontend incidents in Q4 — broken landing pages, CVV fields not loading, affiliate URLs returning 500s, cookie issues — would have been caught by better QA automation before reaching production.

Needs Hire — Q1

7. Cyber Security

The reality: No one owns cyber security today. We have no dedicated security function, no clear governance, and critically — we don't know what we don't know. This is a significant risk at our scale and in a regulated industry.

What we need: A Cyber Security Specialist in IT Ops. Their first job will be to assess where we actually are, surface the unknowns, and establish proper security governance.

📅 Timeline: Hire in Q1. First 90 days focused on security assessment and identifying gaps. Findings may drive additional work for Platform Squad and other teams in Q2+, depending on what's discovered.

Important Context

The Product Function Gap

I think we're severely underestimating the impact that not having a proper product function has had on Engineering this year.

It put enormous strain on Tech Leadership (2.5 people) to take on product responsibilities while delivering against targets. We could only focus on a subset of priorities, and even then it felt like we were doing many things in a mediocre way rather than mastering a few.

At times, squads were left without clear direction, trying to deliver against rocks without the product guidance they needed.

This isn't a technology problem to solve — it's context that helps explain why some foundations have slipped, and why Tech Leadership capacity has been stretched thin.

Summary

What We're Doing in Q1

AWS Migration

Complete migration of all MrQ services. New pipelines live. Requires 1–2 BE per squad for a full month.

Observability Foundation

Evaluate and adopt mature observability platform. Begin centralising logs and metrics.

Platform Engineering

Establish load testing framework, performance profiling baseline, begin staging improvements.

QA & Test Automation

Achieve 1:1 squad parity (final SDET joins Jan). Select tooling. Begin E2E automation on critical player flows.

Head of Architecture

Hire and onboard. First 90 days: assess current state, build architecture roadmap.

Cyber Security Specialist

Hire and onboard. First 90 days: security assessment, identify gaps, establish governance.

The Ask

What We Need to Succeed

Two Key Hires in Q1

Head of Architecture — to own architecture full-time, lead the architect function, and surface the unknowns across our platform.

Cyber Security Specialist — to own security governance and assess our current posture.

Protected Capacity

Platform work needs to be protected in quarterly planning alongside product delivery.

AWS migration alone requires 1–2 backend engineers per squad for a full month in Q1.

Sustained Focus

Q1 is foundations. We need 3 more quarters of consistent execution on these priorities to get to a genuinely good place.

Platform Squad in particular needs to stay focused on platform work — they're the engine driving infrastructure, observability, and tooling improvements.

Product Function

Tech leadership cannot continue to carry product responsibilities at this level. We need a proper product function to partner with.

Roadmap

Realistic Timeline

Q1 2026

Jan – Mar

AWS migration (MrQ). Jenkins investigation. Foundations laid across observability, platform, QA. Key hires made.

Q2 2026

Apr – Jun

New build system rollout. Observability coverage expanding. Load testing maturing. QA automation growing. PQ modernisation continues.

Q3 2026

Jul – Sep

Comprehensive observability. Performance testing at scale. Security governance established. Architecture roadmap executing.

Q4 2026

Oct – Dec

Platform maturity achieved. Proactive monitoring. High QA coverage. Clear architecture ownership. Staging closer to prod.

Bottom line: Q1 is foundations. We need 3 more quarters of sustained focus to get to a genuinely good place. By end of 2026, we should have a mature, observable, well-architected platform that can scale confidently.