Home/AI integration/Custom AI agent development: tools, limits and evaluation
AI integration

Custom AI agent development: tools, limits and evaluation

A generic agent product will not know your order statuses, approval limits or naming conventions. Vascoh develops custom agents around your own tools and data, with test sets and logs so you can tell whether it works.

19.8%

Share of US businesses using AI in at least one business function as of May 3, 2026.

Source: U.S. Census Bureau, Business Trends and Outlook Survey (2026)
41%

Share of workers who reported using generative AI for job-related tasks as of November 2025.

Source: Federal Reserve Board, FEDS Notes: Monitoring AI Adoption in the U.S. Economy (2026)
Jan 26, 2023

Release date of version 1.0 of the NIST AI Risk Management Framework, organized around Govern, Map, Measure and Manage.

Source: NIST, AI Risk Management Framework

What goes into a custom agent

A custom agent has four parts: a model, a set of tools, a policy that says when to use them, and an evaluation harness. The tools are functions your developers write, such as get_order(order_id), create_ticket(summary, priority) or draft_refund(order_id, reason). Each has typed inputs, validation and a permission level.

Narrow tools beat broad ones. A tool that runs arbitrary SQL gives the model too much room. A tool that fetches an order by ID gives it exactly what it needs.

State management is another design choice. Short tasks can run in a single request. Longer ones, such as chasing a missing document across several days, need a stored state, a scheduler and idempotent steps so a restart resumes in the right place.

Planning is where agents differ. Some tasks need a fixed pipeline with a model at certain points, which is cheaper and easier to test. Others need the model to pick its own sequence of tools. Vascoh starts with the fixed pipeline wherever the process allows it, because fewer free choices means fewer unexpected paths.

Designing tools the model can use

Tool names and descriptions are part of the prompt, so vague descriptions produce wrong calls. Return structured errors the model can act on, such as a message saying the order ID was not found, instead of a stack trace. Separate read tools from write tools, and require confirmation tokens for writes that cost money.

Test each tool without the model first. A tool with unit tests and clear failure messages makes the agent easier to debug, because you can tell whether the model chose badly or the tool misbehaved.

Memory is a design question. Passing a long history on every call raises cost and can confuse the model, while passing nothing loses context. Store structured facts, such as the customer ID and the current case status, and retrieve only what the next step needs.

  • Typed arguments validated before execution
  • Read and write tools kept separate
  • Per-run caps on steps and spend
  • Human approval on irreversible actions

Evaluating before launch

Build an evaluation set from real cases: the input, the expected tool calls, and the expected outcome. Run it whenever the prompt, tools or model version changes. Without it, every edit is a guess. Add adversarial cases too, such as a customer pasting instructions into an email body that tell the agent to ignore its rules.

Keep a regression log of every real failure. Each incident becomes a new evaluation case, so the same mistake is caught before the next release.

Usage context

Federal Reserve survey work reports 41% of workers used generative AI for job tasks as of November 2025, so your staff probably already use chat tools informally. A custom agent brings that habit inside controlled access and logging.

Build a kill switch. A single configuration flag should stop the agent from taking actions while keeping read-only behavior, so an incident can be contained in minutes rather than requiring a deployment.

Monitoring in production

Log every model request, tool call, result and final action with a run ID. Track cost per run, tool error rate, escalation rate and how often staff edit the agent's output. Treat a rising edit rate as a signal that inputs have changed or a model update shifted behavior. NIST's AI Risk Management Framework, released January 26, 2023, is a reasonable checklist for documenting those measures.

Handover includes the evaluation set, the tool definitions, the prompt files, the deployment scripts and a runbook for incidents, so your team or another developer can maintain the agent.

How a project runs

From first call to working system.

Step 01

Define tools and boundaries

Vascoh specifies each tool, its permissions and the actions that always need approval.

Step 02

Build with an evaluation set

The agent is developed against real historical cases and scored on tool-call accuracy and outcomes.

Step 03

Deploy with logging

The agent goes live in suggest mode, with dashboards on cost, errors and edits.

Questions

Common questions

How is a custom agent different from a no-code agent builder?

A custom agent is written against your systems with your validation, permissions and logging. No-code builders suit simple flows but offer less control over error handling.

Which model should the agent use?

It depends on task difficulty, latency and data terms. Test candidates on your evaluation set and keep the design provider-neutral.

How do you stop an agent from taking the wrong action?

Narrow tools, validation, step limits, approval on risky actions and tests on adversarial inputs.

Can the agent use my internal documents?

Yes, through retrieval over indexed documents with access control matching your existing permissions.

Contact

Tell us what needs to talk to what.

Describe the systems and the manual work, and we will tell you what is realistic to build and what is not.

We reply within one business day. Your details are used only to answer this enquiry.