Datadog developed a distributed observability system to verify data completeness across billions of ingestion payloads per second. By modeling their ingestion pipelines as segments and tracking telemetry through each stage, the team created a framework that identifies where and why data drops occur. This approach enables real-time diagnostic capabilities and provides reliable signals for both human operators and automated decision-making agents.
Key points
Tracking completeness requires measuring data flow per tenant to account for high-cardinality routing paths.
Decoupling the monitoring system from the ingestion pipeline ensures visibility even during service-level failures.
Complex end-to-end problems can be simplified by modeling pipelines as discrete segments and monitoring progress through each stage.
At extreme scale, prioritize reliable signals with bounded computational costs over perfect, high-precision metrics.
Contextualizing incomplete data is more critical than just detecting the loss; identifying the 'where' and 'why' is essential for actionable remediation.