Getting Real Observability Into Spring Boot With OpenTelemetry and Dynatrace
Why OpenTelemetry and Not Just the Dynatrace Agent Honest answer: the Dynatrace OneAgent is impressive. It does a lot automatically. But on a few services where we couldn't use the agent approach (containerized workloads with locked-down base images, mostly), we needed something we could instrument programmatically. And I didn't want to write Dynatrace-specific code throughout the codebase. If we ever switched backends, we'd be in a real mess. OpenTelemetry gives you a vendor-neutral instrumentation API. You instrument once, and you point the exporter wherever you want. Dynatrace supports the OTLP (OpenTelemetry Protocol) ingestion endpoint natively, so the two actually play together quite well. Not ideal to have to wire it up yourself. But worth it for the flexibility. The Spring Boot Setup We were on Spring Boot 3.5 with Java 24, Spring Actuator, and Micrometer already in place. The starting point for adding OpenTelemetry was the opentelemetry-spring-boot-starter , which handles a lot of the boilerplate stuff automatically. <!-- pom.xml --> <properties> <java.version>24</java.version> </properties> <dependencies> <dependency> <groupId>io.opentelemetry.instrumentation</groupId> <artifactId>opentelemetry-spring-boot-starter</artifactId> <version>2.12.0</version> </dependency> <!-- Bridges Micrometer traces through the OTel pipeline --> <dependency> <groupId>io.micrometer</groupId> <artifactId>micrometer-tracing-bridge-otel</artifactId> </dependency> </dependencies> This starter auto-instruments a bunch of things out of the box: incoming HTTP requests via Spring MVC (or WebFlux), outgoing calls with RestClient (the cleaner replacement for RestTemplate that Spring Boot 3.x brought in), JDBC queries, and more. You don't have to write a single line of custom span code to get useful traces flowing. The exporter config goes in application.properties : otel.service.name=order-management-service otel.exporter.otlp.endpoint=https://<your-environment>.live.dynatrace.com/api/v2/otlp otel.exporter.otlp.headers=Authorization=Api-Token <your-dt-api-token> otel.exporter.otlp.protocol=http/protobuf otel.traces.exporter=otlp otel.metrics.exporter=otlp otel.logs.exporter=otlp You generate the Dynatrace API token with the openTelemetryTrace.ingest , metrics.ingest , and logs.ingest scopes. Don't forget logs.ingest . I missed it the first time and spent twenty minutes wondering why traces were showing up but logs weren't. I've defaulted to http/protobuf rather than gRPC here, because gRPC has a habit of silently failing behind proxies. More on that later. Custom Spans When You Need Them Auto-instrumentation gets you pretty far. But sometimes you need to trace a specific business operation, not just the HTTP boundary. On the order service, we had a pricing calculation step that was sometimes slow, and we wanted to see it as its own span in the trace. Java 24 doesn't change the OpenTelemetry API itself, but the code reads a lot cleaner with records and the newer concise syntax. Here's what the manual span approach looks like: import io.opentelemetry.api.trace.Span; import io.opentelemetry.api.trace.StatusCode; import io.opentelemetry.api.trace.Tracer; import io.opentelemetry.context.Scope; import org.springframework.stereotype.Service; import java.math.BigDecimal; import java.util.List; // Java 16+ record, works perfectly in Java 24 public record Order(String id, List<OrderItem> items) {} @Service public class PricingService { private final Tracer tracer; public PricingService(Tracer tracer) { this.tracer = tracer; } public BigDecimal calculatePrice(Order order) { Span span = tracer.spanBuilder("pricing.calculate") .setAttribute("order.id", order.id()) .setAttribute("order.item_count", order.items().size()) .startSpan(); try (Scope scope = span.makeCurrent()) { return doCalculation(order); } catch (Exception e) { span.recordException(e); span.setStatus(StatusCode.ERROR, e.getMessage()); throw e; } finally { span.end(); } } } Worth calling out: span.setStatus(StatusCode.ERROR, ...) alongside recordException is important. Without the explicit status, Dynatrace won't mark the span as errored in the UI even though the exception is attached. Took me a while to notice that, and it meant some real errors were invisible in our error rate charts. The Tracer bean is registered automatically by the starter, so you just inject it. Being deliberate about span attributes early pays off fast. The order.id attribute saved us multiple times when filtering traces in Dynatrace for a specific order a customer reported as broken. If you prefer annotations and want less ceremony, @WithSpan works too: import io.opentelemetry.instrumentation.annotations.SpanAttribute; import io.opentelemetry.instrumentation.annotations.WithSpan; @WithSpan("pricing.calculate") public BigDecimal calculatePrice(@SpanAttribute("order.id") String orderId) { // pricing logic here } I use @WithSpan for simple cases and the manual API when I need exception status recording or dynamic attributes. Pick whichever one doesn't make you cringe when you read it back. Virtual Threads and Context Propagation This one is worth its own section if you're on Java 24, because it's a genuine gotcha and I've seen it trip up multiple teams. Spring Boot 3.2+ supports Project Loom virtual threads, and enabling them is a single property: spring.threads.virtual.enabled=true The throughput improvement on I/O-heavy services is real. But OpenTelemetry's context propagation relies on ThreadLocal storage internally, and virtual threads can be remounted across carrier threads in ways that break context in some edge cases. In practice, the standard auto-instrumentation handles this fine. Where it breaks is when you're manually spawning threads with Thread.ofVirtual() or using @Async . The fix is wrapping your executor: import io.opentelemetry.context.Context; import java.util.concurrent.ExecutorService; import java.util.concurrent.Executors; // Context.taskWrapping ensures OTel context flows into virtual threads ExecutorService executor = Context.taskWrapping( Executors.newVirtualThreadPerTaskExecutor() ); executor.submit(() -> { // trace context is correctly propagated here doSomeAsyncWork(); }); Our DevOps lead hit this without the wrapping and ended up with disconnected spans in Dynatrace, showing up with no parent trace. The traces weren't wrong exactly, just useless for following a request flow. Easy fix once you know about it. Not obvious at all when you don't. Connecting Traces to Logs This is the part that made the biggest practical difference for us, more than any dashboard or metric. Before we had this wired up, finding a slow trace in Dynatrace meant you'd identified the problem but still had no easy way to pull the exact log lines for that specific request. You'd be grepping for timestamps and hoping. Not great. OpenTelemetry automatically injects trace_id and span_id into the MDC (Mapped Diagnostic Context) when you're using Logback or Log4j2. Every log line gets tagged with the current trace context, as long as your log pattern includes those MDC fields. For Logback, the logback-spring.xml config looks like this: <configuration> <appender name="CONSOLE" class="ch.qos.logback.core.ConsoleAppender"> <encoder> <pattern> %d{ISO8601} [%thread] %-5level %logger{36} traceId=%X{trace_id} spanId=%X{span_id} - %msg%n </pattern> </encoder> </appender> <root level="INFO"> <appender-ref ref="CONSOLE"/> </root> </configuration> Ship those logs to Dynatrace via the OTLP log exporter or the Dynatrace log forwarder, and it correlates them with the corresponding trace automat...