LaunchDarkly helps teams decide which users receive a software behavior and when that behavior changes. AgentControl applies the same operational idea to AI prompts and model settings. The product is useful when changing a prompt has become a production decision that needs versioning, targeted exposure and evidence, rather than a text edit deployed to every user at once.
- 01The offer Runtime configuration and release controls for software features and AI behavior.
- 02The fit Teams operating AI features that need controlled prompt or model changes across user groups.
- 03The boundary Public-source research with a proposed rollout experiment; no feature or model configuration was deployed.
01 / ProductAgentControl makes an AI configuration a managed release object
The AgentControl documentation defines a configuration as a resource containing variations, model settings and either messages or instructions. Completion mode supports single-step responses; agent mode supports structured instructions for multi-step workflows. The application and SDK implementation still determine how tools execute, so choosing a mode does not automatically grant a model new operational abilities.
AgentControl was previously called AI Configs. The API reference explains that the product name changed while existing API paths and operation identifiers remain compatible. This is useful for teams encountering older examples: an AI Configs endpoint can still be relevant, but its naming should not be mistaken for a separate current product.
The quickstart shows the application retrieving a resolved configuration and using its own model-provider credentials. That divides responsibilities clearly. LaunchDarkly supplies configuration and targeting; the application must still call the provider, handle the response and enforce its business rules. A successful configuration evaluation does not establish that the subsequent model request succeeded.
02 / AudienceThe audience owns an AI feature that already has users
A concrete audience is a product team improving an assistant inside an existing application. It may want a new prompt for one customer segment, a different model for an internal pilot or a quick return to the previous configuration after a regression. LaunchDarkly is relevant when those changes need to be managed independently of the application's release cycle.
The Datadog blueprint is a useful comparison when the immediate need is observing model and application behavior. The Cloudflare blueprint covers a broader infrastructure route for building and operating applications. LaunchDarkly's particular role is deciding which configured behavior a context receives and connecting that decision to release evidence.
A team still exploring whether an AI feature is useful may need a smaller local evaluation process first. A controlled rollout cannot compensate for an unclear task or missing quality criteria. Define what a useful answer looks like and how a user can recover from a bad one. Those product decisions determine whether targeting and experiments will reveal meaningful differences.
03 / WorkflowA proposed prompt change should have a baseline and an exit path
Use a proposed pilot around an internal knowledge assistant that returns short answers with references. Preserve the current prompt as the baseline and create one candidate that changes how the assistant handles missing evidence. Keep the model fixed initially. Changing the prompt and provider simultaneously would make it difficult to explain whether a measured difference came from instructions or model behavior.
Build a small dataset covering ordinary questions, conflicting source material and questions the system cannot answer. Label the expected behavior, including when the assistant should decline to infer a fact. A correct reference is more valuable than a fluent answer with an unsupported claim. Human reviewers should agree on those expectations before looking at candidate outputs.
Integrate the configuration through the SDK and provide a stable context representing the intended user or organization. The targeting guide explains individual, segment and custom-rule targeting, with environment-specific rules and an off variation. Use a dedicated test environment first and verify that the baseline remains the default for everyone outside the pilot segment.
Run the candidate against the dataset and inspect the failures. LaunchDarkly's online evaluation documentation describes judges and tracked evaluation metrics; offline evaluation is a separate product surface. Use model-based judging as one signal, then check exact source references and important factual claims directly. A judge can share the answer model's blind spots.
Release the candidate to a small, explicitly selected internal group. Record the configuration version with the result and capture relevant behavior such as unsupported-answer reports, request failures and time to response. An aggregate score without version attribution makes a rollback harder to justify. Keep the proposed pilot's data collection limited to what is needed to evaluate that change.
Test the off variation and return to the baseline before expanding. Confirm what the application does when configuration retrieval fails and which default it uses. The rollback should restore an understood behavior, not merely disable a dashboard switch while the application continues to serve a cached or independently hard-coded prompt. Inspect the actual next request after the change.
Compare the complete result with the baseline, including user corrections and operational cost. If the candidate improves refusal behavior but makes ordinary questions much less useful, the team needs to decide whether that tradeoff fits the feature. The proposed workflow treats rollout as a product decision informed by evidence, not an automatic reward for a higher model-generated score.
04 / PricingAI runs are separate from seats and model-provider tokens
| Plan or unit | Published basis | Qualification |
|---|---|---|
| Developer | Free plan; 5,000 AI runs/month advertised | Confirm AgentControl access for the account |
| Foundation | 5,000 AI runs/month; $5/additional 1,000 | Displayed annual-billing view; other usage billed separately |
| Enterprise | Custom pricing and AI run allocation | Confirm retention and advanced controls |
| AI run | Tracked LLM request or online evaluation | Not equivalent to one conversation or model token |
AI usage summary from LaunchDarkly pricing, consulted 3 October 2026. Confirm eligibility against the AgentControl overview.
The current pricing page lists Developer, Foundation and Enterprise plans. It advertises 5,000 AI runs per month on the first two, with Foundation charging US $5 per additional thousand runs in the displayed annual-billing view. An AI run is generated by tracking an LLM request or running an online evaluation, so it is not necessarily one end-user conversation.
That definition matters when estimating an evaluation rollout. A user interaction can involve several model calls, and online judging can add activity. Count the tracked operations in a representative workflow before multiplying by the expected audience. Model-provider charges remain a separate part of the application bill because production calls use the application's own provider access.
There is an availability discrepancy in the current public material. Pricing advertises included AgentControl usage, while the overview documentation calls AgentControl an add-on and says access depends on the organization. Verify that the feature appears in the intended account and obtain confirmation of its entitlement before planning a rollout around the advertised allowance.
Other LaunchDarkly usage has its own basis, including service connections, client-side monthly active users and observability data. The table intentionally summarizes the AI portion and the entitlement question. A platform estimate should include the software-release features the team will actually use, with retention and advanced controls confirmed for the selected plan.
05 / DistinctionsConfiguration changes can follow a release discipline of their own
The important distinction is that prompts and model settings become versioned operational objects rather than incidental strings in a codebase. This can shorten the path from a discovered behavior problem to a targeted correction. It also changes who can affect production behavior, so the permission to edit a configuration deserves the same attention as permission to change an application setting.
Targeting adds a concrete way to separate evaluation from full exposure. An internal group can receive a candidate while the wider audience remains on the baseline. That is useful when behavior varies by use case or customer context. It does not prove the candidate is safe for every untested group; the expansion decision still depends on whether the pilot represents the broader audience.
The configuration layer also gives model changes a visible identity. A team can explain which prompt and settings produced an answer and compare them with a previous version. This is most valuable when the surrounding application records the version and relevant outcome consistently. Without that connection, a sophisticated control panel can still leave operators guessing about what users actually received.
06 / QuestionsData paths differ between application calls and hosted evaluation
The privacy documentation makes an important distinction. Production model calls and online evaluations occur within the application using provider credentials. Playgrounds and offline evaluations can instead send prompts, variables and evaluation criteria through LaunchDarkly to a configured provider. Review the exact surface used in the pilot rather than applying one broad no-proxy statement to the entire product.
Context attributes need the same care. Targeting can use information about users or organizations, and prompt variables may insert that information into model input. Decide which attributes are necessary and configure privacy handling deliberately. A setting that limits analytics exposure does not automatically make it appropriate to send the same value to a model provider through a prompt.
Finally, evaluation metrics require interpretation. A judge's quality score can support a release decision but does not certify factual correctness or permission to act. Keep deterministic checks for exact requirements and human review for important edge cases. Automated rollback also depends on the purchased capability and configured metric; a general product promise should not replace testing the behavior in the actual account.
07 / DecisionUse AgentControl when AI changes need a controlled production path
LaunchDarkly is worth evaluating when prompt and model changes already affect real users and the team needs a clearer release process. Start with one candidate configuration, a stable baseline and a meaningful internal segment. Demonstrate that the application receives the intended variation and can return to its prior behavior when the candidate fails.
Broader adoption should follow evidence that the configuration process improves decisions and recovery. Confirm entitlement, count the operations that produce AI runs and map the data path for each evaluation surface. The result should be an AI feature whose changes can be explained to engineers and product owners, with authority and rollback behavior that remain understandable.
AI feature already serving users
Preserve a baseline and pilot one prompt change with an internal segment.
Unclear answer-quality criteria
Create representative cases and agree on acceptable outcomes before an experiment.
Sensitive prompts or variables
Map application calls, playgrounds and offline evaluations separately.
A business worth understanding.
Suggest your business or one you find interesting. Tell us what you want to understand about its product, positioning, design or workflows.
Suggestions are free. Selection and publication stay with the desk.
- LaunchDarkly AgentControl overviewConsulted
- AgentControl API namingConsulted
- AgentControl quickstartConsulted
- AgentControl targetingConsulted
- AgentControl online evaluationsConsulted
- LaunchDarkly pricingConsulted
- AgentControl privacyConsulted

