> For the complete documentation index, see [llms.txt](https://stephen-tsoi.gitbook.io/stephen-tsoi-docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://stephen-tsoi.gitbook.io/stephen-tsoi-docs/opentelemetry/from-missing-events-to-complete-visibility-tracing-enterprise-transactions-with-opentelemetry/day-32-34-designing-a-resilient-and-scalable-collector-architecture.md).

# Day 32/34 Designing a Resilient and Scalable Collector Architecture

In my earlier post “8. OTel Collector” (<https://lnkd.in/g4VMivVC>), I introduced the role of the OpenTelemetry (OTel) Collector.\
The collector is responsible for:\
Receiving OTel traces, metrics, and logs from source systems\
Filtering and processing telemetry data\
Sending data in batches to the target SaaS APM backend\
\
The OTel Collector can be deployed on-host (alongside the application) or off-host, depending on architectural and operational requirements.\
\
In posts #24 (<https://lnkd.in/gNh5mWqj>) and #25 (<https://lnkd.in/gqFSnaWm>), I shared a two-tier collector strategy to optimize data volume and ensure observability quality:\
Success traces are sent based on N% sampling\
Failure traces are captured and sent at 100%\
Timeout traces are also captured at 100%\
\
Designing a Resilient and Scalable Collector Infrastructure\
The next step is to design an effective collector infrastructure that supports resilience, scalability, and operational simplicity.<br>

## **Tier 1 Collector**

Deploy the collector on the same host as the application, or share it with the APM agent\
Deploy additional collectors for appliances or third-party products\
Forward traces to a specific Tier 2 collector using the formula:\
\[Trace ID] % \[Number of Active Tier 2 Collectors]\
Send metrics and logs directly to the SaaS APM backend<br>

## **Tier 2 Collector**

Deploy Tier 2 collectors within the same network segment\
Create multiple collector sets, with each set configured in an active–standby model to enhance availability\
DMZ Design\
Establish a DMZ between Tier 1 / Tier 2 collectors and the Internet\
Deploy a proxy server to route traces, logs, and metrics to the SaaS APM backend\
The proxy simplifies firewall rules and collector configurations\
It avoids the need for each collector to directly terminate Internet connections\
Internet Client Devices\
Send 100% of trace data directly to the SaaS APM backend

## This layered collector design helps achieve:

✅ High availability and fault tolerance\
✅ Controlled telemetry volume\
✅ Simplified network and security management\
✅ Better support for large-scale and hybrid environments

<figure><img src="https://2617374589-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FcQN1DZY6gJZPxlsQf9Re%2Fuploads%2FEeUSfdqPOqzjhAZ3I2bP%2Fimage.png?alt=media&amp;token=83bcbab0-9681-434c-b5df-7dc1ce49fc1f" alt=""><figcaption></figcaption></figure>

## 🚀 Let's Connect Beyond GitBook!

If you found this article helpful, you can find more of my technical insights, daily discussions, and deep dives across these platforms:

* **Read more of my work:** Check out my articles on [dev.to](https://dev.to/stephen_tsoi_5b2c4055f3a9) and [Hashnode](https://stephentsoi.hashnode.dev/).
* **Join the daily conversation:** Connect with me directly on [LinkedIn](https://www.linkedin.com/in/stephen-tsoi-16309730/).

***

#### 📬 Stay Ahead of the Curve

Enjoyed this piece? I break down complex technical topics into bite-sized, actionable insights every week.

👉 **Subscribe to my** [**LinkedIn Newsletter**](https://www.linkedin.com/build-relation/newsletter-follow?entityUrn=7487299517642612736) to never miss an update and get the latest articles delivered straight to your feed!


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://stephen-tsoi.gitbook.io/stephen-tsoi-docs/opentelemetry/from-missing-events-to-complete-visibility-tracing-enterprise-transactions-with-opentelemetry/day-32-34-designing-a-resilient-and-scalable-collector-architecture.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `build a script that syncs our docs to a CMS` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
