Where Jev fits in ops

October 8, 2026

If you're using agents and MCPs to get a better understanding of your environment or work through an investigation, you can get a lot of useful information back. You can pull logs, look at recent changes, and check how services are configured, but you're still the one deciding what to do with all of it. That part of the process still lives in your head.

To see where Jev might fit, look at decisions your team already makes and work backwards from them. Can you describe the options and the information you'd need to choose between them? That gives you something concrete to test against how the team works today.

What Jev does

Jev, from TypeSafe AI, takes the context you give it and answers questions you've defined ahead of time. That could be assessing whether an error spike needs immediate escalation, checking evidence against a set of criteria, or estimating whether a condition is true. You get a structured answer with an uncertainty estimate that you can use in the workflow.

Jev is now part of a broader set of decision models and APIs. Kev is an open-source family of Jev-like models you can run and fine-tune yourself, and OpenAI's Decisions API is now in public beta, offering typed decisions through GPT-6 Luna. Classifiers aren't new, and ML teams have been training BERT-style models for years. What's useful here is being able to describe a new decision task through instructions and criteria without training a separate classifier for each one. That gives these models a practical role in agent workflows.

An LLM agent can still do the investigation and come up with a hypothesis. Adding Jev could help with specific decisions along the way, like whether there's enough evidence to try a particular diagnostic step. The harness around it would handle permissions, route model calls, and run approved actions.

Deciding whether an error spike needs immediate escalation can depend on how many customers are affected, whether retries are succeeding, and whether the problem is spreading. The answer might be a simple urgency level, but getting there takes context. You still need to check whether the recommendation is right and whether its uncertainty estimate helps identify cases that need another look.

Why this makes sense for ops

Consider an agent investigating a latency increase. It finds a recent deployment, downstream errors, and a change in traffic. Someone still has to decide what to investigate next. That decision has a few characteristics that make a model like Jev worth considering:

  • Defined next steps: Give Jev the findings and your team's criteria, then ask it to choose whether to inspect the deployment, check the dependency, or collect more evidence. The evidence changes between incidents, but the decision repeats.
  • Consequences when it's wrong: A poor choice can waste investigation time, and a bad remediation can make the incident worse. Evaluate the kinds of mistakes the model makes and what they would mean for affected users.
  • Response time and model cost: An LLM judge is a reasonable baseline, but another review call at each step adds latency and cost. That overhead matters during an incident and grows with larger models, longer histories, and repeated checks across alerts.
  • A path for uncertainty: Jev returns a structured judgment with an uncertainty estimate. The workflow can use that estimate to request more evidence or send the decision to a person for review.

TypeSafe's published evaluations report lower latency and cost than its LLM baselines on structured decision tasks. That's a reason to test whether Jev lets you check decisions more frequently with less overhead. You still need to verify that it catches the errors that matter in your environment.

A few places to start

Instead of having the blank sheet problem, here are a few problems to highlight where Jev is a good model to incorporate into your workflow.

These can fit into the same incident loop: Choice picks the next step, Score rates customer impact, and Noul checks a specific condition before a proposed fix moves forward. Each example below uses a different question type.

1. Recommending the next step

For a latency alert, there might be several runbooks that look reasonable. Recent deploys, which services are affected, and what you've already checked all help narrow it down. Assuming those procedures are already documented, you have a set of options that someone on the team can evaluate.

Give Jev that context and the candidate runbooks, including what needs to be true before each one applies, then ask:

Which of these runbooks fits, or is more evidence needed first?
‍

curl --request POST https://api.typesafe.ai/v1/systemone \
  --header "Authorization: Bearer $TYPESAFE_API_KEY" \
  --header "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "model": "jev-latest",
  "state": "Checkout p95 latency rose from 300ms to 1.8s after release v42. Sampled slow traces are concentrated in new instances; the detailed old/new instance comparison is still pending. Payment-provider response times are normal. Checked: traffic volume and database saturation; neither changed. [Relevant traces and runbook excerpts…]",
  "questions": {
    "next_step": {
      "type": "choice",
      "instructions": "Which of these runbooks fits, or is more evidence needed first?",
      "criteria": {
        "inspect_deploy": "Compare old and new instances when symptoms are concentrated in the new release and that comparison hasn't been completed.",
        "check_dependency": "Use the dependency runbook when slow calls point to an upstream service.",
        "more_evidence": "Neither runbook's prerequisites are established, or the evidence conflicts."
      }
    }
  }
}
JSON

Read the selected runbook from answers.next_step.choice, with confidence and probabilities alongside it. Put the recommendation in the incident record so you can see whether it led to something useful or repeated checks that were already done. Keep permission to run the commands as a separate part of the workflow.

2. Assessing impact and urgency

An increase in errors doesn't tell you on its own how urgently you need to respond. Retries in a background job and failed customer checkouts can produce similar alert volumes while having very different consequences. You need the service context and evidence of impact to decide what deserves attention first.

Use Score to place customer impact on an ordered scale. Give Jev the service context, error trends, and evidence of what customers can still do, then ask:

How much is the incident affecting customers' ability to complete a purchase?

curl --request POST https://api.typesafe.ai/v1/systemone \
  --header "Authorization: Bearer $TYPESAFE_API_KEY" \
  --header "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "model": "jev-latest",
  "state": "Checkout errors have persisted for 12 minutes. Most retries succeed, but 6% of purchase attempts still fail. Support has three reports of customers unable to buy. Background sync errors are also rising, with jobs recovering on retry. [Error breakdown and impact reports…]",
  "questions": {
    "customer_impact": {
      "type": "score",
      "instructions": "How much is the incident affecting customers' ability to complete a purchase?",
      "criteria": [
        "No customer impact: purchase attempts complete normally.",
        "Degraded experience: customers encounter delays or retries but can complete purchases.",
        "Blocked purchases: some customers cannot complete a purchase despite retries."
      ]
    }
  }
}
JSON

Read answers.customer_impact.score alongside its probabilities and confidence. This scale runs from 0 to 2 and can return values between levels; it isn't a percentage of affected customers. Use it to suggest urgency for a responder to review, keeping existing escalation rules in place. Compare with how the incident developed, especially missed impact and unnecessary escalations. If impact evidence is missing, gather more context before assigning urgency.

3. Checking the evidence for a proposed fix

An agent might recommend rolling back a deployment because it happened shortly before an error spike. Before acting, you need to check whether the investigation supports that recommendation. Did the errors start with the deployment, are the affected services part of that change, and were the checks in the rollback procedure completed?

Noul checks whether a specific condition is true. Give Jev the proposed action, the investigation, and the relevant procedure. For a rollback prerequisite, ask:

Does the investigation explicitly confirm that v41 can read the current database schema?

curl --request POST https://api.typesafe.ai/v1/systemone \
  --header "Authorization: Bearer $TYPESAFE_API_KEY" \
  --header "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "model": "jev-latest",
  "state": "Proposed action: roll checkout back from v42 to v41. Errors began after v42 and appear only on new instances. The rollback procedure requires confirming that v41 can read the current database schema. That check hasn't been done. [Deployment diff, traces, and rollback procedure…]",
  "questions": {
    "schema_check_passed": {
      "type": "noul",
      "instructions": "Does the investigation explicitly confirm that v41 can read the current database schema?",
      "criteria": {
        "true": "The investigation records a completed check confirming v41 can read the current schema.",
        "false": "The compatibility check failed, has not been completed, or its result is not documented."
      }
    }
  }
}
JSON


Read answers.schema_check_passed.noul as the model's estimated probability, from 0 to 1, that the answer is yes. A low or uncertain result keeps the action in review. Use a separate question for each prerequisite and test whether it catches missing evidence without holding up well-supported fixes. A high result supports that one check; permissions, required approvals, and direct system checks still belong in the workflow.

What to evaluate

Start with past incidents

Choose cases where your team knows the outcome. Give Jev only the evidence available at the time of the decision, leaving out the eventual resolution.

Compare four things

Run the same cases through an LLM judge and compare both models with the team's decisions. Review disagreements and look at:

  • Serious errors: Which decisions would have delayed the response, missed customer impact, or made the incident worse?
  • Uncertainty: Do uncertain answers help identify cases that need review? How often is the model confident and wrong?
  • Response time: How much delay does each decision add, including unusually slow responses?
  • Total cost: What does the workflow cost once retries and escalations are included?

Run alongside responders

Have Jev make recommendations during live incidents while responders keep control of decisions and actions. Record its recommendation, uncertainty, and what the responder actually did.

Use those results to set thresholds and approvals that match the consequences. A diagnostic suggestion and a production rollback need different levels of scrutiny.

Decide what happens when the model isn't sure: collect more evidence, involve a specialist, or hand the decision to a person. Have that path working before enabling automatic actions.

Context still matters

Jev isn't a magic bullet. It still has a limited context window: TypeSafe currently lists 32k tokens for the state plus the longest question, and 64k for the state plus all questions combined in Jev 1.13.

If an investigation misses a dependency, or a summary drops a failed prerequisite check, the model is still deciding from incomplete evidence. A confident answer doesn't establish that it had all the context it needed. Be deliberate about what you retrieve and preserve, make known gaps explicit, and test how decisions change when relevant context is missing.

Start with one decision

Pick a decision your team makes often enough to get useful feedback. If Jev helps responders choose useful checks sooner or catches recommendations that lack supporting evidence, that's a good reason to try the next one. Expand what the system is allowed to do based on how it performs in your workflow.

This is close to how we’re thinking about autonomous operations at Mezmo. Earlier, we described a harness that handles permissions, routes model calls, and runs approved actions. That’s the role AURA, our open-source agentic harness, plays. An AURA agent pulls context through MCP tools and works through the investigation. It can only run the actions its configuration allows, and any action that needs sign-off waits for a person.

Jev could handle specific decisions inside that loop, like which runbook applies or whether a rollback’s prerequisites are met. A low-confidence answer could send the agent back for more evidence or flag the step for review. A confident answer gives the reviewer more to go on, but the approval still belongs to them.

If you’re interested in using Jev to solve some of these problems, reach out. We’d love to talk through how it could fit into your incident workflow. Reach out here.

‍

Ask about this page
Perplexity
Grok
Table of contents

    More blog posts

    The journey to production AI: Five steps for SRE and platform teams
    The journey to production AI: Five steps for SRE and platform teams
    AURA
    Root Cause Analysis
    Alerting & Incident Response
    AI Agent Infrastructure
    Production AI Observability
    The Grok-to-AI evolution: Why modern SREs are moving beyond manual parsing
    The Grok-to-AI evolution: Why modern SREs are moving beyond manual parsing
    Root Cause Analysis
    Production AI Observability
    The inconvenient truth about AI ethics in observability
    The inconvenient truth about AI ethics in observability
    Root Cause Analysis
    Observability Strategy
    Thought Leadership
    Production AI Observability