> For the complete documentation index, see [llms.txt](https://stephen-tsoi.gitbook.io/stephen-tsoi-docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://stephen-tsoi.gitbook.io/stephen-tsoi-docs/opentelemetry/from-missing-events-to-complete-visibility-tracing-enterprise-transactions-with-opentelemetry/day-19-34-tail-based-sampling-enhance-two-tier-collector-design-by-eda-messaging-bus.md).

# Day 19/34 Tail-based Sampling Enhance Two Tier Collector Design by EDA (Messaging Bus)

In my previous post, we discussed the two‑tier collector architecture and its major challenges:

* High memory consumption
* Data loss during failover or scaling
* Limited scalability due to point‑to‑point integration<br>

At the root of these issues is a tight coupling between collectors.\
As part of this series, where we use OpenTelemetry to trace EDA data flows across multiple system components, a natural question comes up:

**👉 Can a messaging bus help solve these problems?**\
Yes — absolutely.

Let’s look at a few design options.\
1️⃣ Non‑exclusive queue\
First‑layer collectors publish spans to topics (keyed by Trace ID). Second‑layer collectors scale horizontally, consume via selectors, and batch‑ack spans per trace.\
2️⃣ Partitioned queue\
Spans are published with a hash key (Trace ID or modulo). This ensures all spans of a trace land on the same partition and are processed and acknowledged together.\
3️⃣ Error‑first processing\
Span status and Trace ID are added to headers. Second‑layer collectors prioritize error traces, fetch all related spans, send them to APM, and batch‑ack. Successful traces can be sampled.

**Key Advantages of Using a Messaging Bus**\
✅Reduced collector memory usage\
Spans are temporarily stored in persistent message queues, not in collector memory.\
✅Auto‑scaling support\
Non‑exclusive queues allow second‑layer collectors to scale based on queue depth.\
✅Guaranteed trace consistency\
Partitioned queues and selectors ensure spans for the same trace are handled by the same collector.\
✅Faster failure identification\
Status codes in message headers allow error traces to be prioritized.\
✅Improved reliability\
Persistent queues + manual batch acknowledgment prevent data loss during crashes or restarts.

**Disadvantages / Trade‑offs**\
⚠️Need to enhance the collector to support the messaging protocol\
⚠️Need to implement the new logic to handle trace log based on topic name or header value.

## 🚀 Conclusion

By decoupling collectors with a messaging bus, we gain scalability, resilience, and trace‑aware processing—at the cost of slightly higher implementation complexity, which is often a worthwhile trade‑off for large‑scale EDA observability.

**💬 Thoughts or experiences with similar patterns? Let’s discuss. I will cover the detailed design in my upcoming article.**

<figure><img src="https://2617374589-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FcQN1DZY6gJZPxlsQf9Re%2Fuploads%2F2PaTn0T1HaEgVTHgj7ex%2Fimage.png?alt=media&amp;token=e5dc366d-9acf-479e-ace9-706e93e34697" alt=""><figcaption></figcaption></figure>


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://stephen-tsoi.gitbook.io/stephen-tsoi-docs/opentelemetry/from-missing-events-to-complete-visibility-tracing-enterprise-transactions-with-opentelemetry/day-19-34-tail-based-sampling-enhance-two-tier-collector-design-by-eda-messaging-bus.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `build a script that syncs our docs to a CMS` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
