Beyond “Supports OpenTelemetry”: A maturity model for Cloud Native Observability

A presentation at ContainerDays Hamburg in September 2026 in Hamburg, Germany by Kasper Borg Nissen

Slide 1

Slide 1

Beyond “Supports OpenTelemetry”: A maturity model for Cloud Native Observability Kasper KasperBorg BorgNissen, Nissen,Principal Director Developer of Developer Advocate Relations at Dash0 kaspernissen.xyz /in/kaspernissen kaspernissen

Slide 2

Slide 2

Who? Director of Developer Relations at Dash0 KubeCon+CloudNativeCon EU/NA 24/25 Co-Chair (former) Author of OpenTelemetry for Dummies CNCF, AAIF, MergeForward Ambassador Golden Kubestronaut CNCG Aarhus, KCD Denmark Organizer Co-founder & Community Lead Cloud Native Nordics

Slide 3

Slide 3

Scattered signals. Growing haystacks. Every signal its own tab. Every tab its own query language or filter. And one engineer trying to correlate it all - in their head. THREE BROWSER TABS OF OBSERVABILITY

Slide 4

Slide 4

Stamped on every project page. Same two words on READMEs and release notes. The ink is everywhere — what’s underneath varies wildly.

Slide 5

Slide 5

“Supports OpenTelemetry” has become the new “Supports Kubernetes”. A phrase that means everything - and nothing - at the same time

Slide 6

Slide 6

OpenTelemetry: a 1000-mile view. Instrumentation OTel API & SDK Telemetry Backends The OpenTelemetry Collector auto-instrumentation Time-series database … Log database Receive Process Analysis Tools Export Trace database Infrastructure … Kubernetes … Generate and Emit transmit Collect, Convert, Process, Route, Export transmit Store & Analyze

Slide 7

Slide 7

Three signals. One context. The building blocks behind the diagram Produce & Move Make it meaningful SDKs v Auto & manual instrumentation - in apps and on the request path. Semantic Conventions v Agreed names & shapes - so data mean the same thing across tools. OTLP The wire protocol. One pipe, three signals - spoken by SDK, Collector, and backends. Resource attributes The identity on every signal - who emitted this. Collector Shared plumbing: filter, batch, sample, route outside your apps. → Context How trace context & baggage flow through sync and async hops.

Slide 8

Slide 8

Telemetry without context is just data

Slide 9

Slide 9

What are we looking at?

Slide 10

Slide 10

What are we looking at? Awww… Adorable! Cute Cuteness Pretty Normal Unfortunate Creepy Reddit /r/funny, “Cuteness Vs Number of legs” (circa 2010) Gaah! Kill it! Kill it! 0 1 2 3 4 5 Number of Legs 6 7 8

Slide 11

Slide 11

How we talk about system context Organization (By whom) 1 Architecture (What / Why) Which service / system component is this? 2 Compute (How/2) 3 Platform (How) Kubernetes? Which cluster / namespace / deployment / cronjob / job / pod? AWS ECS? Which cluster / service / task? … Which team owns it? “Who you gonna call?” .. 4 Which container? Which process? Pid? Startup args? Which runtime is it? Node.js? JVM? .NET? Which build? Which version? … Infrastructure (Where) 5 Which datacenter / Cloud region / availability zone / account does it run in? …

Slide 12

Slide 12

OpenTelemetry semantic conventions to context layers 1 Organization 😢 Architecture Service (stable) and (experimental) Deployment Environment 2 Compute 3 Platform Kubernetes Cloud (cloud.platform specifically) Cloud-provider specific 4 COM NOT PRE A HE LIST NSIVE ! Telemetry SDK (stable) and (experimental) Compute Unit and Instance Operating System Process & Process Runtimes Device, Browser, Webengine, … … 5 Infrastructure Cloud (general stuff)

Slide 13

Slide 13

Resource is the identity layer - and it belongs at the source. Every span, metric, and log is emitted by something. Resource describes who - the hinge on which correlation, routing, and ownership all turn. Resource, concretely A stable bundle of attributes attached once, at process start - then carried on every signal the process emits.

Emitted with every span, metric, log resource: service.name: api-gateway service.namespace: payments service.version: 2.14.0 deployment.environment: production k8s.cluster.name: eu-west-1-prod k8s.namespace.name: payments k8s.pod.name: api-gateway-7b9f-x2k

Source > Pipeline ◗ At the source, the component knows what it is deployment metadata, config, versioning. ◗ In the pipeline, the Collector is guessing - joining on pod IPs, racing pod lifecycle, correlating by convention. ◗ Drift is silent. Reconstructed downstream, signals can quietly disagree about who - correlation breaks.

Slide 14

Slide 14

1/1/20251/1/2026 Commits: 37.959 PRs+Issues: 46.709 Commits: 53.495 PRs+Issues: 40.597 Source: CNCF Velocity Report

Slide 15

Slide 15

49% of respondents using OpenTelemetry in production. 26% of respondents evaluating OpenTelemetry. Source: https://www.cncf.io/wp-content/uploads/2026/01/CNCF_Annual_Survey_Report_final.pdf

Slide 16

Slide 16

Foundational infrastructure raises expectations on the ecosystem. A natural consequence… 01 Plug into existing pipelines INTEGRATION Connect to an OpenTelemetry Collector without custom adapters, sidecars, or per-project flags. 02 Speak a shared language SEMANTICS Consistent attributes across traces, metrics, and logs - so dashboards and alerts transfer. 03 Stable, predictable, documented IDENTITY & STABILITY Resource identity you can trust across environments. Telemetry treated as a long-lived contract. Users don’t just want OTLP ingest - they want a component that behaves like a first-class citizen of their observability platform.

Slide 17

Slide 17

Ecosystem support is scattered. In practice… WHAT MOST PROJECTS SHIP WHAT PLATFORM TEAMS NEED Supports OTLP Supports OpenTelemetry → a wire protocol ◗ ◗ v Can export to a Collector Metric/log/trace bytes on the wire v → a platform capability ≠ ◗ ◗ ◗ ◗ Semantic correctness & resource identity Trace modeling & propagation Multi-signal correlation Stable, documented configuration surface. “Supports OpenTelemetry” is often interpreted as “can export OTLP”. That framing hides major differences in how well projects actually integrate.

Slide 18

Slide 18

The Collector becomes a compensation layer. Every project that ships half-baked resource identity adds another processor to someone else’s pipeline. 01 02 03 04 05 k8sattributes THE SHIFT resource / transform Identity missing at the source becomes config sprawl in the platform v reach back into the API server to discover what pods this even is. v rewrite service.name , synthesise deployment.environment.name. attribute processor v paper over wrong SemConv names from individual projects. filter / redact v strip PII the component leaked into attribute values. .. & a per-project branch for every new onboarding v Pipeline grows linearly with the estate. Unowned. Untested.

Slide 19

Slide 19

Correlation isn’t emitted. You manufacture it. What “supports OpenTelemetry” looked like in practice

Slide 20

Slide 20

Where the bill is paid When projects ship weak OpenTelemetry support, platform teams absorb the cost. 01 02 03 04 Collector YAML sprawl Silos & stitching Fragile dashboards MTTR flatlines Per-component transform rules, schema mappings, ad-hoc enrichment. Logs, tracing, and metrics correlate only after pipeline massaging. A refactor upstream silently breaks alerts and runbooks downstream. More tooling. Same debugging. Same outcomes. Platform components are long-lived dependencies. Weak OpenTelemetry support limits the value platforms can extract - from reliable debugging to automation and AI-assisted analysis.

Slide 21

Slide 21

We have discovery. We have runtime quality. DISCOVERY RUNTIME QUALITY Ecosystem Explorer Instrumentation Score A registry. Tells you what exists - binary inclusion, no maturity dimension. Rule-based checks on live OTLP. Tells you how clean the emitted data is. explorer.opentelemetry.io instrumentation-score.com

Slide 22

Slide 22

JANUARY 2026 - DRAFT PROPOSAL OpenTelemetry Support Maturity Model. A shared vocabulary for evaluating how intentionally a project supports OpenTelemetry - not whether it can speak the protocol.

Slide 23

Slide 23

Four levels - from instrumented to optimized. LEVEL 0 LEVEL 1 LEVEL 2 LEVEL 3 Instrumented OTel–Aligned OTel–Native OTel–Optimized Observability as implementation detail Incremental migration Telemetry exists mostly to serve internal debugging. OpenTelemetry is not yet a design concern. OTel supported explicitly - often alongside legacy exporters. Works for common cases; legacy assumptions still shape design. OTel shapes architecture OpenTelemetry is the integration surface. Telemetry is designed intentionally for correlation and UX. Telemetry as contract Telemetry as a long-lived product surface. Continuously refined for scale, cost, quality and stability.

Slide 24

Slide 24

Semantic Conventions How consistently telemetry meaning aligns with OpenTelemetry semantic conventions, and how domain-specific meaning is introduced when needed. Resource Attributes & Config How identity, scope, and configuration are handled across environments, including correct use of resource attributes and standard OpenTelemetry configuration mechanisms. Trace Modeling & Context Propagation Integration Surface How traces are structured and how context flows through synchronous and asynchronous execution paths. How users connect a project to their observability pipelines and how strongly telemetry is coupled to specific tools or vendors. Dimensions Stability & Change Management Multi-Signal Observability How telemetry evolves over time and how changes are communicated and managed once users depend on it. How traces, metrics, and logs are supported together and correlated to form a coherent observability experience. Audience & Signal Quality Who telemetry is designed for, how noisy it is by default, and how well it communicates meaningful system behavior.

Slide 25

Slide 25

Resources - where identity is set decides who pays LEVEL 0 Accidental service.name derived from binary name, or absent entirely. No k8s attrs. No environment service.name = “./gateway” LEVEL 1 LEVEL 2 LEVEL 3 Configurable,partial k8s-aware (source) OTel–Optimized A flag exists for service.name. Everything else - namespace, pod, env - still patched in Collector. + k8sattributesprocessor Downward API wired in by default. Pod, namespace, node, cluster emitted as resource. k8s.* set a process start Versioned, documented. Invariants hold across rollouts. Identity never drifts between signals. Telemetry as API TELLS FROM THE OUTSIDE ◗ Does it emit k8s.* attrs without the Collector k8sattributesprocessor? ◗ Is service.name first-class config - or reverse-engineered from a flag?

Slide 26

Slide 26

Semantic Conventions - the vocabulary that makes data transferable. LEVEL 0 LEVEL 1 LEVEL 2 LEVEL 3 Project-native names Mixed vocabulary Tracks SemConv Contrib to SemConv upstream_addr, envoy_cluster - whatever the project calls it internally. Standard attrs emitted. Keeps up with SemConv migrations. Domain extensions marked clearly. Some OTel attrs (http., net.) alongside legacy names. Platform teams rename in the Collector Upstream it’s domain model. Proposes new conventions. Telemetry meaning is shared, not invented. TELLS FROM THE OUTSIDE ◗ Spans use http.route, url.path, server.address - or project-local names? ◗ Do docs reference semconv by version? ◗ Did the http → client/server SemConv migration land?

Slide 27

Slide 27

I tested the model against 5 ingress controllers Ingress is the edge - every request passes through it. If any component should be OpenTelemetry-native Traefik Istio Gateway kgateway Contour Emissary Ingress OTel-first design Envoy-based, battle-tested Envoy + Gateway API. Envoy-native, CNCF Envoy, API-gateway lineage link link link link link

Slide 28

Slide 28

More telemetry ≠ better telemetry. Default metrics exposed by each ingress controller — same job, the edge of your cluster. A 44× spread for the same job - no two projects agree on what “enough” looks like. More isn’t better, fewer isn’t better either. What counts is whether they’re the right signals, follow semantic conventions, and let you answer it. Count is not quality.

Slide 29

Slide 29

Same box ticked. Very different shapes. 01 Only one is OTel-native end‑to‑end Traefik. Everyone else is OpenTelemetry‑aligned at best - legacy assumptions still shape the design. 02 Metrics still reflect Prometheus gravity Traces go out via OTLP; metrics still come from a /metrics scrape. The Collector bridges the two. Hybrid by default. 03 Resource identity leaks into your pipeline Most controllers expect you to enrich service.name, workload, and environment downstream in the Collector. 04 Correlation is not free Trace‑to‑log correlation emerges in the pipeline, not at the source. You assemble it; the component doesn’t hand it to you. 05 Stability is rarely discussed Telemetry changes show up in refactors, not release notes. Dashboards break quietly. 06 “Supports OpenTelemetry” hides all of it Every project ticks the same box. The shape tells a completely different story.

Slide 30

Slide 30

The model didn’t survive review. community#3435 is now “OpenTelemetry Support Self-Assessment and Maintainer Guidance, co-led with Graziano Casto. WHERE IT STARTED A comparative maturity score — levels 0–3 on every dimension, one number-shaped verdict per project. → WHERE IT’S HEADING Opt-in tooling maintainers run on their own telemetry, plus guides per project type. No score. Results belong to whoever ran it.

Slide 31

Slide 31

This only works in the open. Next steps & how to help. 01 Find it a home Land the model in an OpenTelemetry SIG so it has owners and a process — not a doc that rots. 02 Build the self-assessment tooling Deterministic, objective, run locally by the maintainer, no score in the output. Reuse existing projects, like the semantic-conventions-conformance. 03 Write the maintainer guides Libraries, then services and infrastructure - databases, brokers, proxies, gateways, controllers. 04 We need more help! Tooling contributors, guide authors, and projects willing to run the prototype against their own telemetry and tell us if the report is useful.

Slide 32

Slide 32

OpenTelemetry is the foundation for much faster incident resolution Lower MTTR requires more than data - it requires telemetry that means the same thing across every component on the request path. WITHOUT SHARED SEMANTIC CONVENTIONS WITH SHARED SEMANTIC CONVENTIONS Stitching by hand in the war room Symptom → cause without friction Engineers navigate multiple tools, align timestamps, and guess which attribute means what. MTTR stagnates despite “advanced” tooling. Start from a metric, drill into a trace, land on correlated logs. Automation & AI-assisted analysis become viable because the data has meaning.

Slide 33

Slide 33

Platform Engineers are the leverage point. You pick the components that every team inherits - the ingress, the service mesh, the message bus, the CI runner. Your selection is the contract your organization runs on, and it quietly decides how much observability tax every product team pays for the next five years. CONSUME PROVIDE ADVOCATE Select with intent Observability as a contract Push upstream Use the model in technology radars and golden-path decisions. Prefer components whose shape matches your signal priorities. Guarantee baseline telemetry on the paved road. Not opt-in heroics - a provision every service gets for free. File the issue. Ask the maintainer. “Which dimensions are you working on?” is a better question than “Does it support OTel?”.

Slide 34

Slide 34

AI doesn’t replace platform thinking. It increases the need for it. → → GARBAGE IN Inconsistent, fragmented, uncorrelated telemetry AI/LLM Doesn’t invent clarity. Doesn’t reason. Just summarizes. GARBAGE OUT Automated confusion. Faster. At scale. 34

Slide 35

Slide 35

Structure it to the semantic conventions. Conventions are what turn telemetry into something an LLM understands. OpenTelemetry Semantic Conventions THE SHARED VOCABULARY Your telemetry → → CONFORMING TO SEMCONV YOUR DATA Conform to them, and the magic is real. AI/LLM Now it has structure to reason over. UNICORNS AND RAINBOWS Correlation. Root cause. Answers you can trust. 35

Slide 36

Slide 36

Without conventions, correlation fails. Without correlation, AI guesses. With structure and context, AI reasons. 36

Slide 37

Slide 37

One vocabulary. Shared vocabulary for GenAI workloads STABILIZED CONCEPTS

span.attributes gen_ai.operation.name = “chat” gen_ai.system = “openai” gen_ai.request.model = “gpt-4o” gen_ai.request.temperature = 0.2 gen_ai.request.max_tokens = 2048 gen_ai.usage.input_tokens = 1283 gen_ai.usage.output_tokens = 412 gen_ai.response.finish_reasons = [“stop”] # tool call gen_ai.tool.name gen_ai.tool.call.id

= “kubernetes.list_pods” = “call_84ax2”

resource (process) service.name deployment.environment.name

= “agent-sre” = “production” ◗ ◗ ◗ ◗ ◗ Operation name & system Model + parameters Token usage Finish reasons Tool calls 37

Slide 38

Slide 38

One vocabulary? Not yet. (picture from a couple of months ago) Five conventions for the same span OTel GenAI SemConv OpenInference OpenLLMetry LangSmith Langfuse OPTIMIZES FOR VENDOR NEUTRALITY OPTIMIZES FOR EVALUATION WORKFLOWS OPTIMIZES FOR DEVELOPER ERGONOMICS OPTIMIZES FOR THE LANGCHAIN ECOSYSTEM OPTIMIZES FOR IT’S OWN PLATFORM SCHEMA gen_ai.request.model llm.model_name gen_ai.request.model Framework-native taxonomy langfuse.observation.m odel.name Still some traceloop leftovers These are legitimate differences. Not naming preferences. They are not going away soon. 38

Slide 39

Slide 39

One vocabulary? Not yet. Now Five conventions for the same span OTel GenAI SemConv OpenInference OpenLLMetry LangSmith Langfuse OPTIMIZES FOR VENDOR NEUTRALITY OPTIMIZES FOR EVALUATION DONATED TO WORKFLOWS OPENTELEMTRY DONATED TO OPTIMIZES FOR DEVELOPER OPENTELEMTRY ERGONOMICS REJECTED OPTIMIZES FOR THE LANGCHAIN ECOSYSTEM OPTIMIZES FOR IT’S OWN PLATFORM SCHEMA gen_ai.request.model llm.model_name aligns with the gen_ai.* Framework-native taxonomy langfuse.observation.m odel.name These are legitimate differences. Not naming preferences. They are not going away soon. 39

Slide 40

Slide 40

The platform answer. Normalize at the edge. OTel Collector processor Spring AI genainormalizer Arconia v v Rewrites OpenInference and OpenLLMetry spans into GenAI semconv — in the pipeline, not in the apps. All five conventions behind one config property. Switch flavors without touching application code. OpenTelemetry By Thomas Vitale 40

Slide 41

Slide 41

Conformance for OpenTelemetry? WE’VE DONE THIS BEFORE TWO COMPLEMENTARY LAYERS CERTIFIED KUBERNETES THE TRAJECTORY Objective conformance tests Self Assessment — descriptive guidance Run the suite (Sonobuoy, originally Heptio), pass, get the mark. Machine‑checkable. “Certified” means something. How intentionally support evolves. A shape, not a verdict. sits on top of PROMETHEUS CONFORMANCE THE FLOOR “Prometheus‑compatible” Conformance — pass / fail baseline A program that verifies PromQL & remote‑write compliance — so the claim is testable, not marketing. Does it emit (and ingest) OpenTelemetry correctly? Testable. The thing the TC actually asked for.

Slide 42

Slide 42

semantic-conventions-conformance “Does an instrumentation actually emit what the semantic conventions say it should?” Same answer, every library, every language. 01 Exercise Run a small program against the library. A mock LLM server keeps it deterministic. 02 Collect Capture what it actually emitted, through Weaver live-check 03 Check Compare against expectations. Pass or fail

Slide 43

Slide 43

WANT TO GO DEEPER? Free copy of OpenTelemetry for Dummies. The Dash0 Special Edition - by Ayooluwa Isaiah and Kasper Nissen. Unified telemetry, correlation with context, and scalable observability, in plain language. DOWNLOAD: dash0.com/lp/opentelemetry-for-dummies

Slide 44

Slide 44

Follow us on Merge Forward (CNCF) Join Merge Forward! Building a stronger open source future together! Learn more #merge-forward on Slack! community.cncf.io/merge-forward

Slide 45

Slide 45

Thank you! Kasper KasperBorg BorgNissen, Nissen,Principal Director Developer of Developer Advocate Relations at Dash0 kaspernissen.xyz /in/kaspernissen kaspernissen

Slide 46

Slide 46

Get in touch! kaspernissen.xyz /in/kaspernissen kaspernissen