Why My Support Agent Missed an Urgent Ticket

Versioning and evaluating prompts for an IT support agent, and measuring what each revision changedAn employee writes in: their laptop won’t start, and a client demo begins in ten minutes.My support agent answered with questions. What happens when you press the power button? Have you tried…

Annons
Annons
Versioning and evaluating prompts for an IT support agent, and measuring what each revision changedAn employee writes in: their laptop won’t start, and a client demo begins in ten minutes.My support agent answered with questions. What happens when you press the power button? Have you tried restarting? Is another device available? In this test, the expected action was to open an urgent ticket, and the agent never called the ticket tool.That request was one of 40 synthetic tasks I wrote for a small IT-support agent. Under a model judge’s criteria, the first prompt handled 37 of them. After I revised the instructions, the next two prompts handled all 40. I kept every prompt version and its recorded answers so I could go back and see what changed.The assistant and its instructionsThe demo agent could look up IT help articles, check service status, and create tickets. The articles and statuses were fixed and simulated, so the tool responses stayed the same from run to run.The starting system prompt was short:You are an IT support assistant. Help employees with their IT problems. You have tools to look up knowledge base articles, check system status, and create support tickets.It lists the tools but never says when to use them. The laptop case pointed at one specific gap: nothing told the agent to escalate urgent requests.The prompt also lived in the source code, so every change to the wording meant editing and redeploying the agent. I wanted to store revisions outside the code and check how each one behaved before promoting it.A place for each versionFor the deployed prototype, I used Amazon Bedrock AgentCore Configuration Bundles [1]. A bundle is a named, versioned set of settings, including the prompt and the model choice, that an agent can read while it handles a request. I created V1 for the starting prompt and V2 for my manual revision.On each request, the runtime reads those settings and builds a Strands agent from them, and Strands [3] connects the model to the three tools. So a request travels from the AgentCore Runtime to the configuration bundle, then to the Strands agent, then to the model.The deployed prototype’s configuration path. Image by the author.The configuration read happens in the agent setup, following the runtime pattern in the AWS guide [2]:config = BedrockAgentCoreContext.get_config_bundle() or {}agent = Agent( model=BedrockModel(model_id=config.get("model_id", DEFAULT_MODEL_ID)), system_prompt=config.get("system_prompt", DEFAULT_SYSTEM_PROMPT), tools=[lookup_kb_article, check_system_status, create_ticket],)If no bundle is supplied, the agent falls back to default settings. I didn’t test a live version switch or a rollback, so I have no measurement of how quickly either takes effect.The 40-task comparison below didn’t go through this deployed path. I ran it in a separate local evaluation script.The same 40 requests, three promptsThe 40 requests cover password and VPN questions, outages, urgent failures, routine requests that shouldn’t create a ticket, vague questions, and questions outside IT support. Each task lists the tools it expects and describes a suitable response.All three prompts ran through a local Strands script with Claude Haiku 4.5 and the same tools. The script recorded tool calls, response time, and tokens. Claude Sonnet 4.5 acted as the judge. It decided whether each task succeeded and, where it applied, whether the answer’s claims were supported by the tool results. I call that second check groundedness.These are the three versions I compared:V1 is the brief starting prompt.V2 is my revision after going through V1’s mistakes. It added rules for urgent tickets, for choosing tools, for staying within IT support, and for using the help articles as evidence.V3 is a model-assisted revision. I gave another Bedrock model V2 plus nine weak cases from the earlier runs and asked it to improve the prompt. AgentCore’s managed Recommendations feature did not generate it.One run per task and version. The groundedness denominators differ because the judge didn’t apply that measure to every response.Local prompt comparison on the same 40 tasks. Image by the author.On the laptop request, V2 called create_ticket and passed. V1 had only asked questions.The password-expiry case showed a different problem. V2 gave the correct answer, 90 days, and then added that employees receive an advance notification. The article the tool returned said nothing about notifications, so the judge marked the answer ungrounded. V3 kept only the supported 90-day answer and passed that check.V3’s 17/17 needs a caveat. Its urgent-ticket reply still contained extra assertions, but the judge left that reply out of the grounding check. The score counts only the answers the judge assessed, so it can’t show that every V3 sentence was supported.Limits of the comparisonI wrote V2 from V1’s failures and generated V3 from weaknesses in the earlier runs, then tested both on the same 40 tasks. The scores show how the prompts handled the cases I developed them against. Accuracy on new requests still needs a fresh test set.V3 used more tokens than V2 yet had a lower average response time in these runs. With one run per task and prompt, I can’t call that a reliable speed advantage. I’d repeat the timing measurements and have people review a sample of the judge’s decisions.I also tried AgentCore’s managed evaluation and recommendation workflow. At first, the deployment exposed only an outer request trace. More instrumentation made the agent’s activity visible, but evaluation then stopped on missing logs, and later on a trace-parsing error. That path produced no score or recommendation, and the V3 comparison used the local script instead.Before using this for real support requestsI’d keep the prompt version with every recorded run and read the failures alongside the aggregate scores. The version history tells me what changed. The requests, expected outcomes, and actual responses tell me whether the change helped.The revised instructions fixed the escalation mistake and removed unsupported details from some answers. Before putting this in front of real traffic, I’d test it on requests it has never seen, review where the judge got things wrong, and verify bundle switching on the deployed agent. The managed production workflow is still unfinished.Code and referencesThe agent code, all three prompts, the 40 tasks, the evaluation script, and the recorded runs are in the project repository on GitHub: https://github.com/aashrithd4/support-agent-prompt-evaluation.References[1] Amazon Web Services, “Configuration bundles,” Amazon Bedrock AgentCore Developer Guide. Available: https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/configuration-bundles.html.[2] Amazon Web Services, “Use configuration bundles at runtime,” Amazon Bedrock AgentCore Developer Guide. Available: https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/configuration-bundles-runtime.html.[3] Strands Agents, “Strands Agents SDK documentation.” Available: https://strandsagents.com.Sai Aashrith Gavini builds AI agents on AWS and tests where they break. He writes about reliability, evaluation, and the lessons behind better prompts. Explore the code.This story is published under the Generative AI publication. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories. Let’s shape the future of AI together!Why My Support Agent Missed an Urgent Ticket was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

Source: Generative AI Pub — Published — Category: Image AI

🔗 Read full article on Generative AI Pub →
Annons
Annons