Managing the 2026 observability transition
Security vulnerabilities like CVE-2026-15815 in Grafana plugins pose risks for Prometheus-native stacks. Organizations must weigh Datadog's modular pricing against New Relic's usage-based model when scaling observability.
Security vulnerabilities in Grafana plugins present direct risks for teams using Prometheus-native stacks. CVE-2026-15815 allows attackers to escape plugin installation directories and execute arbitrary backend binaries with server privileges. A crafted plugin archive can chain relative symbolic link entries to escape the plugin installation directory, writing arbitrary files and an executable backend binary outside that directory to achieve remote code execution. This vulnerability carries a CVSS base score of 8.8 and affects Grafana OSS and Enterprise versions ranging from 11.6.0 to 13.2.1. Other recent security issues include CVE-2026-42504, which enables denial of service via malicious MIME headers, and CVE-2026-33818, which uses excessive recursion in Go encoding/asn1 to cause denial of service. CVE-2026-56853 involves unencrypted HTTP/2 connections vulnerable to denial of service, while CVE-2026-56858 involves cross-site scripting via pathological input. CVE-2026-33819 involves excessive recursion in Unmarshal. Managing these open-source tools requires senior engineers to handle scaling, retention, storage, and alert hygiene. Organizations should audit contributors and community activity to assess project health. Even established projects face risks, much like the Log4j situation where security breaches necessitated patches that introduced new vulnerabilities. The FTC has even stepped in to issue fines regarding such security failures.
Datadog and New Relic define the two primary paths for teams outgrowing self-managed Prometheus. Datadog provides high-level UX and tight correlation between metrics, logs, and traces through a modular, per-host, per-product pricing model. This model creates a risk where a 100-host environment using APM, infrastructure, logs, RUM, and synthetics costs between $60,000 and $120,000 annually. Datadog charges $1.50 per 1,000 RUM sessions and $5 per 10,000 synthetic API test runs. New Relic offers a different economic reality through usage-based pricing that focuses on data ingestion and per-user seats. New Relic provides a 100 GB/month free tier that allows small teams to operate without a budget approval cycle. New Relic also offers a Data Plus tier at $0.55 per GB for those requiring longer retention and governance. You know that managing a massive Kubernetes estate with hundreds of pods makes per-host pricing dangerous.
| Feature | Datadog Price (Approx.) | New Relic Price (Approx.) |
|---|---|---|
| Infrastructure | $15-$23 per host/month | Included in bundle |
| APM Pro | $31 per host/month | Included in bundle |
| Log Ingestion | $1.27 per GB | $0.30 per GB (Standard) |
| 100-Host Fleet | $60,000-$120,000/year | $15,000-$30,000/year |
| Full-Platform User | N/A | $99/month (Standard) |
Datadog users face unpredictable bills if they do not monitor custom metric cardinality. Datadog’s modularity means every capability, from infrastructure to CI Visibility, acts as a separate SKU. New Relic instead bundles APM, infrastructure, logs, and tracing into a single license.
Data quality failures jeopardize the reliability of machine learning models in production environments. If new data differs significantly from historical training sets, models fail to predict outcomes accurately. This phenomenon occurred during the COVID-19 pandemic when demand patterns for delivery services shifted higher and more volatile than historical data predicted. Data failures include significant fluctuations in ingestion rates, changing schemas, or shifts in the relationships between features. Kubernetes monitoring must cover cluster health, workload health, resource usage, application telemetry, and cost awareness to prevent these failures. Effective monitoring covers node readiness, API server health, scheduler signals, pod restarts, and container throttling. OpenTelemetry collectors receive, process, and export telemetry to multiple backends, preventing reliance on a single vendor’s agent. Grafana Cloud also imposes limits on ingestion rates, such as a 10,000 samples per second limit for series and a 200,000 burst size. Grafana Cloud also limits traces to 5MB by default. Will engineers spend more time tuning Prometheus cardinality or managing New Relic ingestion limits? The decision between Datadog and New Relic rests on the trade-off between feature breadth and cost predictability.