What made observability difficult as the platform grew?
Before the change, an online wellness platform relied on a combination of vendor-specific APM tools and separate CloudWatch log streams. Each tool offered part of the picture, but none of them showed how the different parts of the system connected.
This created blind spots during incidents. Engineers could see individual logs or metrics, but it was difficult to understand how one event led to another across services, queues, databases, and background processes.
Even a seemingly simple issue, such as a notification that never reached a user, could require a long manual investigation. Engineers often had to move between multiple CloudWatch streams and search for gaps between log entries. It was not always clear whether a delay was caused by slow processing, a timeout, a failed request, or a retry that happened silently in the background.
Why did troubleshooting take so long?
The fragmented setup made everyday maintenance harder than it needed to be.
Instead of following a request through the system, engineers had to piece together its story from disconnected sources. Troubleshooting became a process of searching, comparing, and making educated guesses.
Over time, the team found themselves spending too much energy responding to problems and not enough time improving the product. The lack of a unified view made it harder to identify patterns, understand the impact of infrastructure issues, and decide where improvements would have the greatest effect.
For a platform built around consistency and daily habits, reliable reminders and timely sessions matter. Improving visibility into the system was therefore not just an engineering concern. It had a direct connection to the experience an online wellness platform wanted to provide its users.
What did the platform need from a new observability setup?
The platform needed an approach that could bring together signals from across its infrastructure without locking the team into a single monitoring vendor.
The new setup had to make it easier to trace requests across services, connect logs with metrics and traces, and investigate failures without manually searching through multiple systems. It also needed to work reliably with AWS Lambda and remain cost-effective as the volume of telemetry increased.
To address these needs, the platform partnered with Infraspec to redesign its architecture around an open source observability stack.
How did the open source observability stack fit into the architecture?
The team introduced a vendor-neutral open source observability stack for collecting, processing and exploring telemetry.
The pipeline used two collector patterns. Each pattern served a different purpose within the infrastructure.
How did agent mode help?
In agent mode, Collector instances run as daemons on each EC2 host. They collect local signals, including container logs and host metrics, close to where those signals are produced.
This keeps collection efficient and helps reduce the delay between an event occurring and the engineering team being able to see it.
How did gateway mode help?
In gateway mode, a centralized Collector runs on Amazon ECS and acts as the main ingestion point for telemetry from AWS-managed services.
This includes data from Lambda, Kinesis Firehose, RDS, and other parts of the AWS environment. Bringing those signals through a common gateway made it easier to apply consistent processing and routing rules across the platform.
Together, the two patterns allowed the team to collect signals from both individual hosts and managed AWS services through a more consistent architecture.
How did the team handle the challenges of AWS Lambda?
AWS Lambda was an important part of the platform's architecture, but it also introduced some of the most difficult observability challenges.
Standard, off-the-shelf approaches did not provide the environment isolation or debugging reliability the team needed. The team therefore implemented a custom solution designed around the way its Lambda workloads behaved.
The implementation addressed several issues. It managed module state persistence so that telemetry remained reliable across Lambda execution environments. It also introduced parallel flushing to reduce telemetry loss when functions completed.
In addition, Axios interceptors were used to capture metrics for outbound HTTP requests. This gave the team better visibility into calls made to external services and internal APIs.
These details made a meaningful difference. Instead of seeing only that a Lambda function had failed or taken too long, engineers could better understand what the function was doing and where time was being spent.
How did the team manage the volume and cost of telemetry?
The platform was processing approximately 82.9 million metric samples every day. At that scale, collecting everything without filtering would have increased storage and query costs while making the resulting data harder to use.
The team chose to manage this at the collector layer rather than adding filtering logic throughout the application code. This kept the instrumentation cleaner and allowed observability rules to evolve without requiring changes across multiple services.
They introduced severity-based filtering for logs so that the most useful events received the appropriate level of attention. Raw URL paths were replaced with templated routes, which reduced metric cardinality and prevented every unique path from creating a separate time series.
For traces, the team used tail sampling. This allowed the system to retain traces for errors and slow requests while sampling more routine traffic.
The result was a more focused dataset that preserved the information engineers were most likely to need during an investigation, without creating unnecessary costs.
What changed after the new observability platform was introduced?
The new open source observability stack gave the team a much clearer view of its entire platform.
Telemetry from Kafka, Postgres, ElastiCache, SQS, Elasticsearch, CloudWatch, and other parts of the stack can now be brought together in one observability environment. The team also retired its fragmented legacy vendor tools and rebuilt its dashboards in the new stack.
Troubleshooting has become much more direct. Engineers can filter traces by service name and error state, then follow a request through its complete lifecycle. They can inspect individual spans, review outbound call durations, and see related logs in the same view instead of jumping between several disconnected systems.
How has the change affected the engineering team?
The transformation has changed more than the team's tooling. It has changed how the team works.
When an issue occurs, engineers can rely on correlated metrics, logs, and traces to understand what happened instead of starting with a collection of isolated clues. They have fewer blind spots during incidents and can spend less time on manual investigation.
The broader shift has been from reactive troubleshooting to more proactive engineering. With better data and clearer context, the team can make more informed decisions, improve reliability, and focus more of its time on building features that help the platform’s users maintain healthier daily routines.
