Infraspec
Back to Case Studies
ObservabilityReliability EngineeringPlatform Engineering

Infrastructure modernisation for a fast-growing hyperlocal platform

A large hyperlocal grocery delivery company

A large hyperlocal grocery delivery company in India was handling roughly 250,000 orders a day. With new funding and ambitious expansion plans, its CTO and architect wanted to know which parts of the platform would struggle as the business grew, before an incident forced the question.

Infraspec assessed the setup, built a prioritised roadmap and embedded with the platform team to work through the highest-impact initiatives.

58%

lower mean time to detect

67%

shorter lead time to change

$161k

in infrastructure savings

Why modernise when nothing was broken yet?

The next phase of growth would put pressure on delivery speed and reliability. Without an active incident, the team had room to assess the platform before those pressures became harder to manage.

We spent a week alongside the teams reviewing applications and infrastructure and gathering stakeholder input. The findings became an action plan with clear priorities.

How did the team decide what to tackle first?

The roadmap weighed impact, return on investment and ownership. Some initiatives addressed reliability and visibility; others addressed delivery speed, infrastructure cost or operational risk. Each priority needed a rationale stakeholders could support and a path the internal platform team could continue to own.

Infraspec embedded with the platform organisation to execute those priorities transparently alongside the team.

What ended up on the roadmap?

From March 2023 to April 2024, work included a centralised observability platform, branch-based continuous delivery, FinOps adoption, API Gateway changes, Karpenter-based autoscaling and canary deployments with ArgoCD.

API Gateway reduced the public API surface by roughly 60%. The newer autoscaling setup reduced recovery risk and helped optimise cloud spend.

What changed with observability?

A centralised observability platform was introduced between March and June 2023. Mean time to detect dropped by 58%, while telemetry gaps fell by 70%.

What changed in the release process?

Branch-based continuous delivery cut lead time to change by 67%. Canary deployments with ArgoCD halved the change failure rate, with a typical release affecting only about 5% of users.

What changed about infrastructure cost?

FinOps made infrastructure cost a shared responsibility across the organisation and contributed $161,000 in savings. Even as the business scaled, infrastructure spend grew more slowly than business volume.

How did the changes stick?

We remained embedded with the platform team throughout execution. Documentation, infrastructure-as-code and repeatable runbooks reduced dependency on any one person's knowledge and helped the internal team continue to own the work.

The company kept growing around 10% month on month without reliability becoming a bottleneck. Developer productivity and release velocity improved, while automation and simpler workflows freed up engineering time for higher-value work. Starting before a major outage gave the team room to make lasting changes.

More case studies
→

Want to make your infrastructure work harder for your team?

Tell us what your engineers are wrestling with.

Talk to us