sequenced.ai
Articles/Data & analytics/Blueprint///8 min read

Datadog connects AI investigation with the telemetry behind production software

Datadog brings Bits agents and Agent Observability into its monitoring platform. Evaluate incident evidence and distinguish AI credits from trace billing.

By Sequenced deskAI-assisted, source-led · how we work
Visit Datadog website ↗
Bits InvestigationIncident analysisInvestigate alerts using operational context.
Bits ChatConversational accessAsk about telemetry and monitoring objects.
Agent ObservabilityAI application insightTrace, evaluate and improve AI workflows.
AI CreditsBits consumptionSeparate from Agent Observability span pricing.
Datadog mark
Datadogdatadoghq.com · independent research

Represent this company? Verify your work email to access its workspace, or send the desk a factual correction.

Datadog combines infrastructure and application monitoring with AI tools that investigate operational signals and help teams understand AI applications themselves. Bits Investigation works on incidents; Agent Observability follows model calls and agent behavior. Keeping those jobs separate makes the product easier to assess and prevents two different forms of AI consumption from being mistaken for one bill.

In brief
  1. 01The offer Observability data, AI operational agents and tooling for evaluating AI applications.
  2. 02The fit Engineering teams with maintained telemetry and a recurring investigation or agent-quality problem.
  3. 03The scope Public-source research and a proposed incident pilot; no authenticated monitoring access or remediation tests.

01 / ProductTwo AI jobs sit inside the Datadog platform

The Bits AI overview describes Bits Chat, Investigation, Code, Security Analyst and Agent Builder. This article concentrates on investigation and conversational access. They use operational context to help engineering teams understand signals and decide what to do next. The wider family includes action-oriented capabilities that deserve separate access and change-management decisions.

Bits Investigation is the current product name for the AI incident-investigation offer. Its job is to inspect alerts and related telemetry, develop possible explanations and present findings. The presence of evidence links matters: a team must be able to inspect the data behind an explanation before using it to justify a production change.

Agent Observability addresses a different reader need. It supports tracing, evaluations, experiments and datasets for AI applications. Datadog positions it alongside backend and user-experience context. A team building an agent can use that view to investigate model behavior and service performance together, rather than treating every slow or incorrect response as a model problem.

02 / AudienceA maintained telemetry estate creates the starting point

An on-call engineering team already using Datadog has a concrete reason to investigate Bits. It may receive alerts with enough data to diagnose the issue, yet still spend time locating the relevant logs, traces, service dependencies and recent changes. AI can be evaluated against that search and synthesis work without assuming that it can safely own every remediation.

A team with incomplete instrumentation faces a different problem. If the relevant service has no useful traces or its logs omit the identifier needed to connect requests, an AI explanation may remain speculative. Adding an assistant can expose missing context, but it cannot make unrecorded events observable. Instrumentation and service ownership may need attention before an incident pilot produces meaningful results.

The Elastic blueprint examines search and retrieval infrastructure that can support an AI application. The Arize blueprint focuses attention on AI evaluation and application behavior. These cover adjacent layers: retrieving evidence, assessing an agent's output and investigating the production service that runs it. Choose the comparison according to the specific operational question your team needs to answer.

03 / WorkflowA proposed investigation for a rising checkout error rate

Consider a proposed pilot around a checkout service that experiences a higher error rate after a release. Use a historical or controlled incident that the engineering team already understands. This is a workflow design, not a test of Datadog. The aim is to learn whether Bits makes the investigation easier to inspect and whether it distinguishes supported conclusions from unresolved possibilities.

Begin with the alert's service, environment, time window and customer impact. Check that the monitor message points to useful context rather than a vague instruction to investigate everything. A good starting record identifies what changed in the observed behavior and which operational owner should assess it. That helps both human responders and an AI investigator avoid pursuing unrelated noise.

The Bits Investigation launch article describes gathering monitor context, consulting runbooks and exploring hypotheses through telemetry queries. It describes hypotheses as validated, invalidated or inconclusive. In the pilot, inspect those categories against the underlying evidence. A plausible timing correlation is not automatically a demonstrated cause, especially when several services changed during the same deployment window.

Ask the engineer to follow one promising hypothesis through a trace, its related logs and the relevant deployment information. The point is to understand whether the AI selected the right time interval and service boundary. A database slowdown and a client retry storm can appear together while requiring different interventions. The team should be able to explain which observation supports each step of the proposed causal story.

Use Bits Chat to ask follow-up questions about the telemetry and monitoring objects. A useful follow-up might narrow the affected region or compare successful and failed requests. Keep the result anchored to the same incident identifiers and time window so a conversational answer does not drift into a different population of requests.

Preserve human review before any production change in this initial pilot. A suggested code fix, restart or configuration adjustment is a proposed intervention with its own consequences. The engineer should identify the expected effect and verify the actual state afterward. If the error rate improves, check whether the affected customer operation recovered, not merely whether one dashboard stopped showing an alert.

End with a short incident account: initial signal, inspected evidence, rejected explanations, accepted action and unresolved questions. Compare that account with the known historical outcome. Measure useful investigative work and reviewer corrections, including cases where the AI appropriately remained uncertain. A fast but poorly supported diagnosis can increase operational work if responders must later undo a change based on it.

04 / PricingAI Credits and LLM spans measure different work

OfferCommercial basisReader implication
AI Credits$500 per 500 credits/month, billed annuallyOn-demand displayed at $1.30 per credit; consumption varies.
Agent Observability Free40,000 LLM spans/month15-day retention shown; separate from Bits usage.
Agent Observability Pro$160/month annually, 100,000 LLM spans includedAdditional annual spans displayed at $3.50 per 10,000.

Commercial model from Datadog pricing, consulted 22 September 2026. Confirm the applicable account, region and contract before purchase.

The Datadog pricing page lists AI Credits for the Bits family separately from Agent Observability. Its illustrative consumption figures vary by task complexity and context; they are not fixed prices for an incident. The table captures selected displayed USD rates. Other monitoring products and the account's commitments remain separate parts of the overall bill.

For Bits, the public page says unused monthly credits do not roll over and usage above the purchased bundle is billed at the on-demand rate. An incident-heavy month can therefore have a different cost pattern from a quiet one. Estimate consumption using the investigations the team actually runs, including unsuccessful or inconclusive work, rather than assuming every use reaches a complete resolution.

For Agent Observability, distinguish an LLM span from the whole workflow. A multi-step agent can call a model several times while also invoking tools or retrieval steps. The product's commercial unit follows model calls. A longer answer is not necessarily a separate trace, and a single customer request can contain multiple billed spans. Reconcile the instrumentation with the pricing definition before forecasting usage.

05 / DistinctionsShared operational context is the useful distinction

Datadog's platform approach is attractive when a team needs to connect an AI explanation to the same telemetry used by its ordinary incident process. Moving between an alert, a trace and a supporting log can make a diagnosis inspectable. The relevant advantage is continuity of evidence, rather than the claim that an agent reasons like a senior engineer.

For AI applications, Agent Observability can connect output evaluation with system behavior. An apparently poor response may come from an empty retrieval result, a failed tool or a model decision. A trace that preserves those relationships helps the team choose the right fix. Changing the prompt will not repair a broken upstream API, just as adding infrastructure capacity will not correct a mistaken business rule.

The breadth of the product family can also create confusing evaluation scope. A team may want a read-oriented incident assistant while a demonstration includes autonomous coding or custom agents. Name the exact features in the pilot and decide which actions they may take. Treat any expansion into production mutations as a separate operational change with its own evidence.

06 / QuestionsCheck missing signals, permissions and site availability

Which data does the investigator actually see? Ask the platform owner to show the services, logs, traces and documents available to the selected user and feature. A coherent explanation can still reflect a narrow view of the incident. Include a case where the evidence lies outside the connected source set and check that the resulting uncertainty remains visible.

How will the team handle repeated alerts from the same incident? The investigation may produce useful work each time, but repeated triggering can affect both attention and consumption. Examine the monitor configuration and actual usage records before expanding the pilot. The appropriate fix may be better alert design rather than a larger AI credit purchase.

Where is the feature available? The public pricing page marks AI Credits and Agent Observability unavailable on the US-FED site. Confirm the selected Datadog site and account entitlement before depending on either product. The general marketing page is not sufficient evidence that an organization's regional or regulated environment exposes the same feature set.

07 / DecisionChoose a diagnosis the on-call team can defend

Datadog is a strong company to examine when AI work needs to meet production evidence. Begin with a known incident and require the investigator's claims to survive inspection by the engineer responsible for the service. That approach reveals useful synthesis, missing telemetry and misleading inferences without confusing a demonstration with dependable operations.

If the main problem is an AI application's own behavior, begin with Agent Observability and a representative set of traces instead. The common decision criterion is whether the team can find the reason for a failure and verify a correction. Expand the product scope when that evidence supports it and when the commercial units match the team's measured usage.

01

Existing Datadog incident process

Compare Bits findings with a known incident and inspect every important evidence link.

Start with investigation
02

AI application quality problem

Trace model calls, tools and backend failures with Agent Observability.

Evaluate the application
03

Sparse or unclear telemetry

Improve instrumentation and ownership before relying on generated diagnoses.

Prepare the evidence
What should we explore next?

A business worth understanding.

Suggest your business or one you find interesting. Tell us what you want to understand about its product, positioning, design or workflows.

Suggestions are free. Selection and publication stay with the desk.

Sources
Filed under Data & analyticsCompany DatadogNot affiliated with DatadogRequest a correctionRequest a refresh by email

Continue reading

All in this category