Decision Lab

Decision model versus LLM

Who should route your tickets: a decision model or an LLM?

Decision Lab sends one ticket to a decision model and to your LLMs at the same time. You see what each one decided, how sure it was, how long it took and what it cost.

Incoming ticket

This isn't an outage, the status page is green and everything loads for us. We were billed twice for the annual plan though.

Every model answers three questions

Which team should pick this ticket up first?
Billing, infra, account or security, with a probability for each.
Does this ticket need human attention right now?
A probability.
How severe is the impact on the customer?
0, 1 or 2.

Then plain code picks the queue

  • Human review when confidence is below your threshold (60% by default)
  • Page on-call when it is urgent and severe
  • The team's queue otherwise

What you can do in the lab

Connect your contenders

Paste your TypeSafe key for JEV, then choose where the LLMs run (OpenAI, or your own LiteLLM proxy) and add as many models as you want to compare. They load from your account, and prices from a public list.

Bring your own tickets

Type one ticket or paste a batch, each with the correct team. Or start from one of the use cases below and edit it.

Read who was right

Accuracy against your answers, median latency and cost for each contender, ticket by ticket. Your keys stay in your browser tab.

Results

JEV and LLMs on the same tickets

What each one decided on a labeled set of support tickets and on three situations where the answer is not obvious. These are real answers, recorded on a date, with the models named below.

Use cases

Ticket by ticket

Each case is a handful of realistic tickets built to be hard for routing by keywords. Stamps show what each contender decided. Open any case in the lab, edit the tickets and the answers you expect, and run them with your own models.

Method

What is measured, and what is not

In the lab

For every ticket: the team, the probabilities, the routed queue, and whether it matches the answer you gave. For the set: accuracy, paging accuracy, median latency and cost.

Not measured here

Calibration (when it says 80%, is it right 80% of the time), consistency over repeated runs and injection rates. The hostile-input case shows single examples, not rates.

Limits to keep in mind

JEV returns a probability per option; an LLM states its own confidence. They are not the same quantity. Latency depends on where each model runs: one you host yourself is limited by your own hardware. Forty tickets is a demo, not a benchmark: measure on your own domain before deciding anything.

Your keys, your tokens

Tick the free keyword baseline to try the lab without a key. Paste a key, or add the address of your own LiteLLM proxy, and your contenders unlock. Keys stay in your browser tab unless you choose to remember them, are sent only with your requests, and are never stored or logged by the server.

Open the lab