Use Agentic SRE for operations and support in ITRS Analytics

Beta

The ITRS Agentic SRE app is currently available as a beta release. Features, behavior, and documented procedures are subject to change without prior notice during ongoing development. Starting with Agentic SRE app version 0.6.0, your license must include the Agentic SRE entitlement, and the app must be enabled in the KOTS Admin Console. For details, see Enable Agentic SRE with a license entitlement. For information about major and minor updates, see the Agentic SRE app release notes.

Important

Starting with Agentic SRE app version 0.5.0, provider API keys and tokens are expected to be stored in the Secrets app instead of directly in the LLM configuration. If you upgrade from an earlier beta release, review any previously configured LLMs and recreate or update them to use API Key (Secret).

About the ITRS Agentic SRE app Copied

The ITRS Agentic SRE app is an AI-assisted capability within the ITRS Analytics Web Console. It helps users investigate operational conditions, ask product-related questions, and work from monitored entity context within a single interface.

The app provides an AI chat experience for operational and support workflows, together with an Admin > AI area where administrators can configure LLMs, test model connections, review service readiness, and review evaluation results before making models available for wider use.

Agentic SRE app overview

The Agentic SRE app provides the following capabilities for day-to-day use:

Note

Available agent modes and admin features can vary by deployment, user permissions, and enabled capabilities.

Prerequisites Copied

Before using the Agentic SRE app, ensure that the following requirements are met:

If no LLM is configured, start by adding your chosen AI provider in the LLMs section before using the AI chat.

Use case scenarios Copied

Use the Agentic SRE app when you want faster guidance without switching between multiple tools or documentation sources.

Note

Use SRE when you want a general entry point for the Agentic SRE app. The SRE agent acts as an orchestrator that can handle different types of requests and route them to the appropriate specialist workflow when needed.

Investigate what is happening in your environment Copied

Use Insights when you want a quick operational summary for one or more selected entities. To request a platform-wide summary without entity context, state explicitly that you want an overview of the system or environment.

Typical tasks include:

If Insights cannot resolve the entity named in a request, refine the entity reference or try the request again instead of relying on a platform-wide summary.

In most cases, you can send all tasks through SRE and let it direct the request to the most relevant workflow. This is useful when you need to move from an alert to a working hypothesis quickly, or when you are not sure which specialist agent to choose first.

Perform root cause analysis Copied

Use RCA when you already have an incident, service, or entity in scope and want a more structured investigation.

Typical tasks include:

For example, you might ask, “What changed in the last 30 minutes for this service?” after selecting the affected entity in Entity Viewer.

Value Entity

Get product and setup guidance Copied

Use Support when you need help with product features, setup tasks, or product documentation.

Typical tasks include:

For example, you can ask for help creating a configuration template or for guided steps to complete a setup task.

You can also start these requests in SRE and allow it to route the query to the appropriate support-oriented workflow.

Create an XML template

Create dashboards from AI results Copied

Use the app when you want to create a dashboard for a selected entity.

Add the entity to the prompt and ask the AI chat to create an entity dashboard. When the dashboard is created, the response includes a link that you can use to open it for further analysis.

Entity cards in a response provide separate actions for opening the related entity context or adding the entity to a prompt.

How to use the AI chat in the ITRS Analytics Web Console Copied

Follow these steps to start using the AI chat in the Web Console:

  1. In the ITRS Analytics Web Console, open the AI chat from the header.
  2. In the prompt toolbar, choose the agent mode that best matches your task.
  3. Select the LLM that you want to use.
  4. Enter your prompt.
  5. If you are already working in Entity Viewer, select the relevant entity to make it available among the recently visited entities in AI chat.
  6. Add each relevant entity to the prompt before sending it.

Use the agent modes as follows:

Add entity context to a prompt Copied

Add entity references when you want an answer to focus on specific monitored entities. Entity references appear as chips within the prompt, so you can review the scope before sending the request.

To add entity context:

  1. Open or select an entity in the Web Console. Recently visited entities appear above the prompt.

  2. Add an entity by using one of the following methods:

    • select Add to chat on the entity chip
    • open Contexts to apply and add one or all recently visited entities
    • enter @ in the prompt and select an entity from the filtered list
    • select Add to chat on an entity card in an AI response Add entity context
  3. Repeat the process to include more than one entity.

  4. Review the entity chips within the prompt, then send the request.

You can select an entity chip to open the related entity.

Review the AI response Copied

After you send a prompt, review the response details before acting on the recommendation:

Provide feedback on an AI response Copied

If response feedback is enabled, use the thumbs-up or thumbs-down controls below an AI response to record whether it was useful.

Submitted feedback remains associated with the response and is available to administrators in Audit.

Use chat history if it is enabled Copied

If chat history is enabled in your deployment, you can reopen earlier conversations from the history panel.

Use chat history to:

Configure AI settings in Admin Copied

Use Admin > AI to prepare the Agentic SRE app for operational users. From this area, you can configure LLMs, review service readiness and evaluation reports, and view AI audit activity.

Admin access

The procedures in this section are intended for users with access to Admin > AI. Available sections, such as Eval and Audit, can vary by deployment and assigned permissions.

Add an LLM Copied

Use the LLMs section to add, test, and manage the language models that the Agentic SRE app can use for chat, retrieval, and evaluation.

Important

Starting with Agentic SRE app version 0.5.0, provider API keys and tokens are expected to be stored in the Secrets app instead of directly in the LLM configuration. If you upgrade from an earlier beta release, review any previously configured LLMs and recreate or update them to use API Key (Secret).

Use this section when you need to:

Use this procedure to add a model and make it available in chat:

  1. In the Web Console, go to Admin > AI > LLMs.

  2. Click Add LLM to create a new entry, or copy an existing LLM configuration.

  3. Select the provider that you want to use.

  4. Enter the required provider settings, such as model name, endpoint or base URL, credentials, and any provider-specific options.

  5. In API Key (Secret), select the secret that stores the provider credential.

  6. If no secrets are available, create the required secret first in the Secrets app, then return to the LLM configuration and select it.

  7. Use the masked credential fields and reveal control only when you need to verify an entered value.

  8. Click Test to verify that the model connection works as expected. The control shows progress followed by a success or failure icon. Point to the control to review the response or failure reason, and click it again when you want to rerun the test.

    Add LLM in Admin UI

  9. Save the LLM.

  10. Open the AI chat and select the new LLM from the prompt toolbar.

Before you save a model, confirm the following:

Depending on your deployment, available providers can include managed services and self-hosted options.

Configure a Google GenAI model Copied

Google GenAI is available as an LLM provider. When adding a Google GenAI model, configure the following required settings:

Add LLM Google GenAI

You can also configure Temperature, Max Tokens, Top P, and Top K. Leave optional settings unchanged to use the provider defaults.

For supported Gemini models, use either Thinking Level or Thinking Budget to limit how much reasoning the model performs. Do not configure both settings for the same LLM. Leave both settings unset to use the model default.

Configure reasoning for an OpenAI model Copied

For OpenAI models that support reasoning, use Reasoning Effort to control how much reasoning the model performs. The available values depend on the selected model. Models that do not support this setting do not show the field.

Review service readiness in Status Copied

If Status is enabled in your deployment, use Admin > AI > Status to review whether the services required by the Agentic SRE app are ready.

The Status page provides live information for:

Service readiness in Status

The overall status can be Operational, Degraded, or Unavailable. Review the individual cards when the app reports reduced capability or an unavailable component.

The AI chat header also shows a status indicator. Administrators can select this indicator to open the Status page.

Review AI activity in Audit Copied

If Audit is enabled in your deployment, use it to review recorded AI conversations and related activity.

Use Audit and chat history for different purposes:

Chat history contains your own AI chat conversations. Audit provides administrators with tenant-wide activity, including conversations from all users and activity that is not added to user chat history, such as Eval runs and LLM connection tests.

  1. Go to Admin > AI > Audit.
  2. Review the list of recorded conversations. Each entry can show details such as Title, LLMs Used, Exchanges, Tokens, Created, and Status.
  3. Use the available filters to narrow the list when you need to find a specific conversation or time range.
  4. Use Refresh to reload the Audit list.
  5. Open a conversation from the list to review the recorded exchange in more detail.
  6. In the conversation detail view, review the summary information, including the LLMs used, exchange count, average latency, total tokens, user and team information, and conversation status.
  7. Review the conversation replay to inspect the recorded prompts and responses.
  8. If turn replay is available, select Inspect turn on a turn entry to open the turn inspection drawer.
  9. In Turn inspection, review the event timeline for that turn. This view can help you examine the recorded sequence of events for the selected exchange in greater detail.
  10. Review the feedback indicator for a turn to see whether the user approved or disapproved the response. Point to the indicator to review an optional feedback comment.
  11. Export the conversation record in the available format when you need to keep or share the audit output.

Use conversation details as follows:

If turn replay is available, each turn can also show:

Use Inspect turn when you need a more detailed view of a single exchange. In the turn inspection drawer, review the event timeline and use the available categories to focus on the part of the turn that you want to inspect, such as LLM activity, tools, retrieval, or routing.

For example, if you need to review a production support conversation, open the audit entry and confirm which LLM answered, how many exchanges occurred, how many tokens were used, and whether the conversation completed successfully. If a specific turn needs closer review, use Inspect turn to open the detailed event timeline for that exchange.

Run an Eval campaign and review the results Copied

Use Eval to compare outputs, validate model behavior, and review answer quality before wider rollout.

Use Eval when you want to:

  1. Go to Admin > AI > Eval.

  2. Click Add Eval to open the Add New Eval wizard.

    Run Eval

  3. In Configuration, select the target LLM, eval type, and number of iterations. Choose the model explicitly rather than relying on a default selection.

  4. In Graders, select one or more graders for the run.

  5. In Tests, load tests from one or more test sets, or define custom tests for your scenario.

  6. If your test sets use tags, filter the loaded questions by tag to narrow the scope of the run.

  7. For string-content checks, choose whether every expected string must match or the best-matching string determines the score.

  8. Select the tests that you want to include, then click Run Eval.

  9. Monitor the run status and wait for completion. Click Cancel Evaluation if you need to stop an active run.

  10. Open the generated report to compare scores, inspect test results, filter results by tag, and review answer quality.

  11. Repeat the run with adjusted settings if you want to compare models or tune the configuration.

Understand Eval results Copied

After an Eval run completes, open the generated report to review the results in detail.

Example use case:

Use Eval to validate Support answers before rolling out a new LLM.

Example test prompt:

Can Geneos monitor Linux and Windows hosts?

Example expected answer:

Yes. Geneos can monitor both Linux and Windows hosts. The answer should clearly mention both operating systems and provide a concise explanation.

Example expected strings:

This type of test checks two things at the same time:

Use the report to:

When you open a test answer, the view can include:

Use the grader tabs as follows:

Available Eval types and graders depend on the enabled capabilities. By default, production deployments provide Support evals. Agent-targeted evals for SRE, Insights, and RCA are available only when enabled for the deployment.

If a grader does not apply to a test, the result is shown as N/A instead of a percentage. An N/A result does not indicate a failed test.

For example, if a Support eval answer shows LLM 68% and String Content 100%, you can interpret the result as follows:

Use this combination of scores and feedback to decide whether you should adjust the prompt, refine the test, or compare the same test against another LLM.

Upload and review Eval reports Copied

Use this procedure when you want to review an existing Eval report in the Web Console.

  1. Go to Admin > AI > Eval.
  2. In the reports view, upload a saved eval report file.
  3. You can upload supported report files directly or drag and drop them into the reports area.
  4. Use Refresh to reload the report list after an upload or after another user has added reports.
  5. Open the uploaded report to review test results, grader output, and score ranges.
  6. Filter the report by tags if you want to focus on a subset of tests.
  7. Download a report when you want to retain a local copy for comparison or sharing.

For example, you can upload a previously saved report from another environment, refresh the reports list, and then review the same test set results in one place before deciding whether the target LLM is ready for wider use.

Use the Secrets app with the Agentic SRE app Copied

Beta

The Secrets app and the Agentic SRE app are both currently available as beta releases.

If your deployment includes the Secrets app, use it to store provider credentials and reference them from the Agentic SRE app.

Follow these steps:

  1. In the Secrets app, create a secret for the provider credential that the LLM will use.
  2. Assign an owner role that matches the users who are allowed to use that secret.
  3. Save the secret.
  4. In Admin > AI > LLMs, open the LLM configuration.
  5. In API Key (Secret), select the stored secret instead of entering the credential value directly. Agentic SRE API Key (Secret)
  6. Save the LLM configuration and test the connection.

This integration helps you keep provider credentials in one managed location and avoids storing provider API keys directly in day-to-day LLM configuration.

["ITRS Analytics"] ["ITRS Analytics > Agentic SRE"] ["User Guide"]

Was this topic helpful?