Ticket triage · Production

incident.io builds Workbench, an internal AI evaluation suite for their incident investigation agent

The problem

As incident.io moved from tightly focused first-generation AI features to a complex AI agent for incident investigation, triage, and resolution, their existing lightweight tooling was insufficient — it lacked eval suites, graders, and scorecards needed to ensure quality at that scale.

First attempt

Off-the-shelf AI tooling options existed but were rejected because relying on vendor marketing rather than first-hand experience risked adopting a product built for a different team context, which would have caused the team to skip learning AI engineering from first principles.

Workflow diagram · grounded in source
1
@incident interaction trigger
Trigger
Someone interacts via @incident, initiating LLM prompts to classify and score the interaction.
source quote
“When someone interacts with us via @incident, we run some further LLM prompts to classify and score the interaction.”
2
LLM classification and scoring
Ai action
LLM prompts classify and score the incoming interaction.
source quote
“we run some further LLM prompts to classify and score the interaction”
3
Investigations agent analyzes incident
Ai action
The Investigations agent analyzes source code to pinpoint what broke and reaches into Grafana to interpret telemetry data.
source quote
“The agent analyzes source code to pinpoint what broke, reaches into Grafana to interpret telemetry data, and connects dots across disparate systems”
4
Engineer review in Workbench
Human review
Engineers review each interaction in Workbench alongside its classification, scorecard, and contextual user feedback from the thread.
source quote
“We can then review each interaction in workbench, alongside our classification and scorecard. We also include extra context around the interaction so we can see if the user provided any relevant feedback in the thread (e.g. 'thank-you' or 'argh that's …”
5
Trace view debugging
Validation
A trace view shows the full prompt tree for each interaction, with errors highlighted in red.
source quote
“We use this trace view to show us the series of prompts that were run to power a given interaction, enabling us to debug more complex examples: you can easily see what path our interaction took through our prompt tree, …”
6
Reproducibility testing
Validation
Engineers re-run specific LLM requests multiple times to assess whether failure cases are easily reproducible or unpredictable.
source quote
“it's useful to be able to re-run it a few times to see if it's an easily reproducible failure case, or an unpredictable one”
7
AI performance dashboards
Feedback loop
Dashboards join product data with AI-specific metadata including scorecards, costs, and latency to enable spend analysis, forecasting, and aggregated performance reporting over time.
source quote
“we can easily build dashboards that join product data with our AI-specific metadata (e.g. scorecards, costs and latency). That helps us do spend analysis and forecasting, as well as being able to report on aggregated scorecards over time”
Reported outcome

incident.io built Workbench, a bespoke internal AI evaluation suite that enabled rapid iteration, a single pane of glass for debugging LLM interactions, and privacy-preserving performance analysis of their Investigations agent without exposing customer data to staff.

Reported metrics
Latency saved via speculative tool callingabout 2s
Idea-to-production cycle timein production by lunchtime
engineer deep-work time on AI improvementspend more time thinking deeply about how to improve the product
Reported stack
WorkbenchLLMGrafanaSonnet 3.7Slack
◆ Does this fit your context?

Compare to your context

Tell us your scale, team, and constraints. We'll show what changes at your size, what fails at your scale, and whether this case is a fit, needs adaptation, or won't scale to you. Free demo, no signup.

Compare to your context →
~30 seconds · free
Source
https://incident.io/building-with-ai/built-our-own-ai-tooling
Read source ↗

Frequently asked questions

What did this team achieve with this AI workflow?

incident.io built Workbench, a bespoke internal AI evaluation suite that enabled rapid iteration, a single pane of glass for debugging LLM interactions, and privacy-preserving performance analysis of their Investigati…

What tools did this team use?

Workbench, LLM, Grafana, Sonnet 3.7, Slack.

What results were reported?

Latency saved via speculative tool calling: about 2s; Idea-to-production cycle time: in production by lunchtime; engineer deep-work time on AI improvement: spend more time thinking deeply about how to improve the product (source-reported, not independently verified).

What failed first in this deployment?

Off-the-shelf AI tooling options existed but were rejected because relying on vendor marketing rather than first-hand experience risked adopting a product built for a different team context, which would have caused th…

How is this ticket triage AI workflow structured?

@incident interaction trigger → LLM classification and scoring → Investigations agent analyzes incident → Engineer review in Workbench → Trace view debugging → Reproducibility testing → AI performance dashboards.

WHAT TO DO WITH THIS

Now compare it to your context

This case is one data point. Whether its pattern fits you depends on your volumes, your stack, and your exception load — that comparison is the step no case study can do for you.