AI solution

AI Agent Development

An AI agent is software that decides which tool to use, and in what order, to reach the goal you give it: it reads a record, fills a gap from another system, drafts the output and checks the result. We do not set agents up as open ended assistants but like a new team member with a written job description, a permission list and clear stop rules.

Task and stop rulesLeast privilegeHuman approval gateAPI and MCP connectionsStep by step logging
  • Google Partner
  • Talha Aslan and team
  • English, German, Turkish

In short

AI agent development means giving a large language model access to your business tools and letting it plan a multi step task on its own. The agent reads a CRM record, interprets an email or document, fills in missing details and picks the next step. A well built agent reaches only the tools it needs, waits for a person before sending, paying or deleting, stays within set limits and logs every step with its reasoning.

Talha Aslan and teamLast updated:

When you need one

Which jobs genuinely call for an AI agent?

Most processes run cheaper and more predictably as a fixed workflow. An agent adds value where the steps change with each request and the answer depends on combining data from several systems. If a few of the situations below sound familiar, an agent is worth discussing.

Every request takes a different path

One customer question means checking the order system, then the courier, sometimes accounts as well. Whenever you try to draw a fixed flow, the exceptions outnumber the rules.

Your team ferries data between tabs

Staff copy and paste between the CRM, email, spreadsheets and the ERP. The work is slow not because it is hard but because the information is scattered, and most errors creep in during that shuffling.

Nobody has time for groundwork

Researching a new lead's company, finding past correspondence and listing what is missing before a quote is valuable but slow. So it gets skipped or stays superficial.

An agent you tried went off the rails

An off the shelf agent with broad access emailed the wrong person, repeated the same action in a loop or ran up an unexpected usage bill. Now the team no longer trusts automation.

Sources: Anthropic: Building effective agents

Our approach

An agent with a clear task, narrow access and a traceable trail

The first question is whether the job really needs an agent. We write out your process step by step using real cases, then assign the parts where the steps never change to a plain workflow and the parts where the decision depends on the request to the agent. Most projects end up as a mix: the agent only steps in where judgment is needed, and rules handle the rest.

Each tool the agent may use is defined separately, such as searching the CRM, reading an order status or creating an email draft. Tools are connected through an API, a webhook or a Model Context Protocol (MCP) server and start with read access only. Write access is granted in stages as test results come in. Build, monitoring and upkeep run as part of our AI automation services.

If the agent's only job is talking to customers, you most likely need a chatbot rather than an agent; see our AI chatbot development page. Where the target system has no API, or you need a new dashboard to review the agent's work, we plan that part through custom software development.

  • We first test whether a fixed workflow can do the job
  • The agent reaches only the tools its task requires
  • Sending, paying and deleting go through human approval
  • Caps on steps, actions and usage
  • Every tool call logged with its reasoning
Anatomy of an AI agent
  1. Task definitionGoal, scope and finish condition
  2. Tool listEach tool does one job, with written permissions
  3. Step logTool called, input, output and reasoning
  4. Approval gateSending, paying and deleting need a person
  5. Stop rulesHalts when a cap is hit or it is unsure
  6. HandoverUnfinished task passed on with a summary

What the agent must not do is written down as clearly as what it may do; those limits live in code level permissions, not in the model's goodwill.

Which agent?

The job decides the type of agent

The same foundation needs three different permission sets for three different jobs, so we first agree on whom the agent serves and what it does.

Sales

Lead research agent

Looks into a new lead's company and past correspondence, completes the CRM record and prepares a briefing note for your sales team.

  • Summary from the company website and CRM history
  • Suggested field updates, saved after approval
  • Draft reply ready, sending stays with sales

Operations

Back office agent

Compares supplier emails with order and invoice records, spots mismatches and opens them as tasks for the right person.

  • Field extraction from emails and PDFs
  • Matching orders with invoices
  • Task and alert when something does not match

Internal

Agent that supports your staff

Searches several systems for the answer to a colleague's question, combines the figures and replies with the sources it used.

  • Role based, read only access
  • Summaries from reports and spreadsheets
  • List of sources behind each answer

Essentials

Which rules keep an AI agent safe?

An agent does not just talk, it acts, so the risks are different; we write these rules down before any build starts.

Least privilege, narrow tools

OWASP's risk list for large language model applications treats systems given more functionality, permissions or autonomy than they need as a separate category, called excessive agency. So we give the agent single purpose tools and only the access its task requires.

Approval for high impact actions

Steps that are hard to undo, such as messages to customers, payments, deleting records or changing prices, wait for a person to approve them. The approval screen shows what the agent wants to do and why, side by side.

Outside text is never an instruction

An email, web page or document the agent reads may hide commands meant to steer it; this is called prompt injection. External content is treated purely as data, and any step that needs permissions is checked by rules that do not depend on that content.

Automated decisions about people

Article 22 of the GDPR gives people the right not to be subject to a decision based solely on automated processing that has legal or similarly significant effects on them. For screening applicants or assessing creditworthiness, the agent prepares a recommendation and a person decides.

An agent that discloses what it is

When the agent writes emails or messages to outside people, they know they are dealing with an AI system. Article 50 of the EU AI Act calls for this transparency, and we apply the same rule to recipients in every market.

Usage caps and rollback

An agent stuck in a loop can pile up both actions and model fees. Each task gets a cap on steps, daily actions and usage; every write action the agent takes is logged and can be reversed where needed.

Sources: OWASP Top 10 for LLM Applications 2025 · OWASP LLM06:2025 Excessive Agency · General Data Protection Regulation (2016/679), Article 22 and Chapter V, EUR-Lex · EU AI Act (Regulation 2024/1689), Article 50, EUR-Lex

Comparison

Fixed workflow or AI agent?

TopicFixed workflowAI agent
Who sets the stepsA flow drawn in advanceThe agent, based on the goal
Good fitProcesses that run in the same order every timeMulti system work whose path changes per request
Running costLow and predictableEvery decision is another model call
TestingSame input, same outputChecked repeatedly against a scenario set
How it failsThe flow stops and raises an alertIt may pick the wrong tool or the wrong order
Right choice forMost jobs whose rules can be written downOnly the steps that truly need judgment

Quick check

AI agent feature list

Must haves: is your process ready?

0 of 6 in place Tick the boxes to see how ready you are for automation.

Added as needed

  • Tool connections through an MCP server
  • Answers grounded in your documents (RAG)
  • Several agents splitting the work
  • Scheduled tasks
  • Assigning tasks from Slack, Teams or email
  • Monthly cost and success report

We choose which of these you need together during the first call.

Let us pick the first task to hand to an agent

Describe a job your team repeats every week that runs a little differently each time, plus the systems you use; we will tell you whether it needs an agent or a workflow and send the scope and a written quote.

Process

From discovery to launch in four steps

  1. First call and discovery

    We listen to your processes in a free 15-minute call. Then discovery maps your tools and tasks, scores the opportunities and ends with a written scope and fee for your approval.

  2. Build and test

    We build the first workflow in your accounts and test it with real but masked examples. Approval steps, error scenarios and alerts go in before anything reaches a customer.

  3. Go live and tune

    We switch the workflow on step by step, watch the logs and adjust thresholds with your team. You get documentation and a short training session.

  4. Monitor and expand

    On the monthly plan, we monitor running workflows, adapt them to model and API changes and add new workflows from the priority list, with a monthly report.

Free tools

Prepare your agent project with free tools

See which systems your website runs on, check authentication for the domain your agent will send email from, work out the time spent by hand and compare your pilot results.

Analysis

Website Technology Checker

Detect a website's CMS, e-commerce platform, server, and tracking tags such as GA4, GTM, Google Ads and Meta Pixel.

E-mail

SPF, DKIM & DMARC Checker

Why do your emails land in spam? Check a domain's SPF, DKIM and DMARC records, find the errors and get a corrected record to copy.

Work

Working Hours Calculator

Calculate daily and weekly working hours after breaks, in hours and decimals, and check legal breaks and rest periods for the UK and EU.

Conversion

A/B Test Calculator

Check whether your A/B test result is statistically significant and calculate the sample size and test duration you need.

Conversion

Conversion Rate Calculator

Calculate conversion rate, CPA and revenue per visitor, and plan how much traffic you need to hit your goal.

All free tools

How we work

The agent starts as an observer, then becomes a helper

We do not yet have a live client AI agent project we can show as a reference, so instead of claiming results we describe our method. You can see our other work, including publishing workflow automation, on the references page.

Shadow mode first

In the first weeks the agent only makes suggestions while your team keeps doing the job by hand, and the two results are compared side by side.

Scenario set with traps

We pick easy, hard and deliberately misleading cases from your past work; no prompt or tool change goes live without passing that set.

Permissions in stages

Read, draft and write access are unlocked in turn, and each step up depends on measured results.

You own it

Model, automation and cloud accounts are opened in your company's name; tool definitions, prompts and documentation stay with you at handover.

All references

FAQ

Questions about AI agent development

If your question is not here, write to us; we will send you an answer and a written quote.

Next step

Let us define your first agent's task together

Tell us which job you want to hand over and which systems it touches; after a free 15 minute call we will send the split between workflow and agent, the scope and a written quote.

In-depth guide

AI Agent Development: Task, Tool, Permission and Oversight Decisions

Talha Aslan and teamLast updated: 16 min read

In AI agent development, quality depends less on the model than on the rules written before it runs: which task the agent owns, which tools it may touch, when it stops and asks a person, and how its actions are recorded. This guide covers those decisions in the order a business owner makes them.

Technical terms are explained briefly where they first appear, so you can speak the same language as any vendor and judge a pilot with your own data.

Test the idea with a one page task card

A single page task card tells you whether a job suits an agent at all. It names the event that starts the job, the goal, the condition that proves the job is done, the systems involved and the person who would notice if something went wrong. If you cannot fill in those five lines, the problem is the process definition, not the technology, and AI agent development should wait until that definition is clear.

Once the card exists, make the call with four questions:

  • Path variety: List your last twenty requests; if most follow the same order, a rule based workflow automation is enough.
  • Checkability: Someone on your team must be able to verify the agent's output in a few minutes; if only a specialist can check it after long effort, the pilot cannot be measured.
  • Cost of a mistake: Can a wrong step be reversed? Actions that cannot be undone may only sit behind an approval gate.
  • Volume: A job done a few times a week rarely repays the build and upkeep.

If more than two answers are negative, postpone the agent. The card later becomes the first draft of the test set, so the work is never wasted.

First agent tasks by type of company

A first agent task should be narrow, frequent and easy to check. The examples below are not client results; they are typical starting points that meet those criteria in different kinds of companies.

  • Online store: An agent that reads a return request, compares the order date, the shipping history and the return policy, and leaves a ready decision note for the support rep, who still approves the refund.
  • Manufacturer or distributor: An agent that reads the technical specification attached to a quote request and lists matching catalog items plus the details still missing.
  • Accounting or consulting firm: An agent that finds, at month end, which client still owes which document and drafts the reminders.
  • Appointment based service business: An agent that matches a canceled slot with suitable people on the waiting list and proposes who to contact.
  • B2B sales team: An agent that pulls next steps out of meeting notes for the CRM record; fixing the record rules first through CRM automation makes this much easier.

Answering internal questions usually calls for a search assistant that cites its sources rather than an agent. If nothing gets written to any system, a company knowledge assistant is simpler and cheaper. Start with the job that is easiest to measure, not the most ambitious one, and bring the employee who owns that job into the project from day one; they are the right person to write the correct answers in the test set.

Plain language architecture: loop, state and orchestrator

Technically, an agent is a loop: the model reads the current state, decides to call a tool, sees the result and repeats until it reaches the goal or hits a stop rule. The part that runs this loop is ordinary code called the orchestrator. It takes the model's proposal, checks the permission list and decides whether the tool actually runs.

Pin down three concepts when you discuss architecture:

  • State: Which step the task is on, what each tool returned and which limits remain. This lives in a database, not in the model's memory, so an interrupted task resumes where it stopped.
  • Context: The amount of text a model can see at once is limited. Long threads and large spreadsheets are not passed in whole; only the relevant slice is.
  • Memory: Notes made during a task are discarded when it ends. Facts the agent must remember long term, such as a customer's delivery preference, sit in a separate record that is reviewed.

Several agents handing work to each other look attractive but are hard to debug, so we start with one and add a second only when a subtask needs separate permissions or another model. Fixed steps run outside the loop as plain code, so the model is involved only where judgment is required. Ask vendors for a sketch of this loop showing which decisions sit in the model and which in code.

Tool design: small parts that each do one job

Most of the effort in AI agent development goes into tool design, because the agent's quality often depends more on its tools than on the model. Each tool has a name, a short description of what it does, the parameters it accepts and the format of what it returns. The model picks tools by reading those descriptions, so write them as carefully as instructions for a new hire.

  • One job each: Instead of a broad "manage the CRM" tool, define separate ones such as "find a customer by email address" and "add a note to a record", and grant access per tool.
  • Safe retries: Write tools carry an idempotency key, meaning a repeated request does not create a second record; a retry after a network error never produces a duplicate invoice.
  • Readable errors: A failing tool returns a message the model can act on, such as "no permission" or "record not found", so it can change course.
  • Narrow output: Results are paged and limited to the fields the task needs, not thousands of raw rows.
  • Dry run mode: Write tools first run in a version that only reports what they would have done; the real version runs after approval.

Keep the tool list short. Many similar tools raise the odds of a wrong pick; as the task grows, split it rather than adding tools.

Data sources, credentials and connection paths

For each system the agent connects to, settle the access path first, then the credentials, then data quality. The access path is usually the system's API, a webhook that fires when something happens, or an MCP server that exposes tools in a standard form. MCP, the Model Context Protocol, is an open protocol that lets a model call tools across different systems in the same way.

  • Dedicated service account: The agent connects with its own account and only the scopes it needs, never with an employee's login, so the logs show clearly who did what.
  • No secrets in prompts: API keys stay in a secrets vault and are injected when a tool runs; the model never sees them.
  • Rate limits: Provider limits are read in advance, and the agent knows how to wait instead of hammering the system.

Unstructured documents such as invoices, contracts and scanned forms should not reach the agent raw. Splitting them into fields first through AI document processing makes the results far easier to audit.

An older program without an API needs an export, an intermediate database or a thin layer of custom software. Screen clicking bots break with every interface change, so keep them as a last resort. Data quality matters as much as the connection: if one customer appears under two names in two systems, the agent will multiply that confusion rather than fix it.

Model choice, hosting and version pinning

Choose the model by how often it gets the hardest step of your task right on your own test set, not by public leaderboards. For agent work the deciding factors are how reliably it calls tools with correct parameters, how well it follows instructions over long context, and response time.

  • Split roles: A more capable model for planning and decisions, and a faster, lighter one for simple steps such as field extraction or classification, keeps costs balanced.
  • Hosting: Weigh the provider's own API, cloud platforms that let you choose a region, and open weight models on your own servers against data location, operating effort and quality.
  • Version pinning: Behavior can shift when providers update a model, so production uses a fixed version and a new one goes live only after passing the test set.
  • Outage plan: When the provider does not respond, the agent either waits or hands the unfinished task to a person; nobody assumes the same prompt works unchanged on another model.

Self hosting keeps data in house, but hardware, updates and security become your job; let the volume of personal data and your capacity to run servers decide.

Designing the approval queue and the step log

An approval gate that a reviewer cannot understand in a few seconds is either ignored or rubber stamped. The screen should show, at a glance, the action the agent wants to take, the records it relied on, the before and after values of every field that will change, and its reasoning.

  • Three choices: Approve, edit and approve, or reject. A rejection asks for a short reason code, and those codes feed the next round of improvements.
  • Named owner: Every action type has an approver and a backup for when that person is away.
  • Timeout: If approval does not arrive within the set time, the action never runs by default; a reminder goes out or the task closes.

The step log answers, months later, why the agent did something. Each entry holds the task ID, timestamp, tool called, inputs, result, model version and the approver.

Retention periods for entries with personal data are set up front, access is limited by role, and the log records who viewed which entry.

Agent security: outside content, secrets and outbound limits

An agent is most dangerous when it can reach private data, reads untrusted outside content and can send data out. When those three abilities meet in one task, a single email carrying hidden commands can steer the agent into leaking data; the design must break at least one leg of that triangle.

  • Outbound allowlist: The agent may write only to listed domains and listed recipients; any message to an unlisted address lands in the approval queue.
  • Separated context: Information extracted from outside content goes into a marked data field, never into the same place as the agent's instructions.
  • Sending channel: If the agent sends email from your domain, SPF, DKIM and DMARC records must be correct; our SPF, DKIM and DMARC checker shows the status in a couple of minutes.
  • Isolated runtime: Agents that run code or process files work in a separate environment with restricted network access.

Security probes belong in the test set. A document with planted instructions, a fake request from a manager, or a customer message asking for an action outside the agent's permissions is tried on purpose.

They run again after every prompt or tool change, and when one succeeds, the fix goes into the permission layer first, because lasting protection comes from limits in code.

GDPR, transfers and the EU AI Act

Because an agent acts, write down separately what data it reads, where it sends it and what it produces about people. The points below are not legal advice; they are a checklist to work through with your legal adviser before the build.

  • Processor contracts: Under Article 28 of the GDPR, any provider processing personal data on your behalf, including the model and automation platforms, needs a data processing agreement; read their terms on data use, retention and training.
  • Transfers: If a provider sits outside the EU or EEA, Chapter V of the GDPR applies and the legal basis for the transfer must be documented.
  • Decisions about people: Article 22 restricts decisions based solely on automated processing with legal or similarly significant effects, so the final say on applicants, credit or contracts stays with a person.
  • Transparency: Article 50 of the EU AI Act requires that people know they are interacting with an AI system; we apply that to every recipient, inside the EU or not.
  • Other markets: The UK has its own version of the GDPR, and US state privacy laws differ; confirm which rules apply to the people your agent contacts.

A practical rule: mask fields the agent does not need at the tool level. If a task does not require a national ID number, the agent never sees it and so can never write it anywhere by mistake.

Rollout steps from pilot to full use

Widening permissions in measured steps, instead of switching the agent on all at once, lowers both the risk and the team's resistance. This is the sequence we follow:

  1. Collect cases: Gather task examples from past weeks along with how the team resolved them, masking personal data for testing.
  2. Build the test set: Label the cases easy, hard and misleading, and have the team write the correct outcome for each.
  3. Shadow run: The agent processes real requests but writes to no system; its suggestions are compared with the team's decisions.
  4. Draft mode: The agent prepares write actions, and each one lands in the approval queue.
  5. Limited writes: Low risk, reversible action types run without approval; everything else stays in the queue.
  6. Regular review: Rejection reasons, cost and error logs are reviewed monthly, and permissions widen or narrow based on them.

Each step has a pass criterion written in advance, such as the rejection rate staying below an agreed threshold for two weeks. Falling back a step is not a failure; it shows the controls work.

AI agent development does not end at launch. As connected systems, models and business rules change, the agent needs upkeep: monitoring, test set updates and small fixes run as part of our AI automation services.

The metrics that show whether the agent pays off

An agent's value is measured by what a task costs end to end and how much human intervention it needs; the number of completed tasks on its own is misleading. Before measuring, you need to know how long the same job takes by hand; our working hours calculator helps you estimate that from the team's weekly records.

  • Intervention rate: The share of tasks where a person had to correct, redirect or take over.
  • Steps per task: If the step count rises for the same task type, the agent is wandering; review tool descriptions or context layout.
  • Approval wait time: If the agent is fast but approvals sit for hours, the bottleneck is human, and approval ownership needs rework.
  • Total cost: Model usage, platform fees and approval time, counted together.
  • Escaped errors: Mistakes that passed approval and surfaced later are counted separately; they show whether the approval screen is clear enough.

Track these side by side in a monthly table. If value does not show, narrowing the task or moving that part to a fixed workflow is usually wiser than expanding the agent. Read the numbers with the person who owns the job; the table alone cannot tell you why interventions rose.

Limits, risks and criteria for stopping

Agents are probabilistic: the same input may not produce the same output, and a model can invent a wrong tool parameter with full confidence. In long chains, small errors compound, so checkpoints should increase as the number of steps grows.

  • Silent drift: When field names change in a connected system, the agent may read the wrong field without raising an error; a regularly run test set catches this.
  • Cost swings: Hard requests mean more steps and more model calls, so a monthly budget cap and alerts are set from the start.
  • Latency: Multi step tasks can take minutes rather than seconds; if a customer expects an instant reply, the agent belongs in the background.
  • Provider lock in: Without documented prompts and tool definitions, changing models means rewriting the project.

Write stop criteria at the outset too. If the intervention rate does not fall by the end of the pilot, if total cost exceeds the manual process, or if the team cannot keep up with the approval queue, the agent is switched off and the task returns to a fixed workflow or to people.

Common mistakes in AI agent development

These recurring patterns explain why agent projects lose trust; each comes with a healthier alternative.

  • Deciding from a demo: Instead of launching an agent that shines on a few hand picked cases, decide with a test set built from past work that includes misleading examples.
  • Starting with full access: Instead of connecting the agent with an admin account, start read only and open write access one action type at a time.
  • Handing everything to the model: Instead of letting the agent run fixed steps too, move every step whose rule can be written into plain code and use the agent only where judgment is needed.
  • Editing prompts off the record: Instead of changing prompts without version control, log every change and run it through the test set before release.
  • Leaving approval unowned: Instead of a queue nobody is responsible for, assign an owner and a backup to every action type.
  • Watching cost after the fact: Instead of discovering the usage bill at month end, set per task caps, a daily budget and alerts from day one.

All of them treat the agent like a feature you toggle; managed as a process with rules, owners and limits, most never appear.

Choosing a partner and the next step

Choose your AI agent development partner less by what their agent can do in a demo and more by how clearly they explain what it is prevented from doing. Ask for direct answers to these questions:

  • Which of our jobs would you solve with a fixed workflow instead of an agent, and why?
  • Will the tool list and each tool's permissions be handed over to us in writing?
  • In whose name are the model, automation and cloud accounts opened, and what can we take with us if we part ways?
  • How is the test set built, and how are prompt changes tested before release?
  • By which criterion will we call the pilot a success or a failure?
  • When the agent stops or fails, who is notified, through which channel and with what information?

Vague answers signal a vague scope. A written task card and test plan let you compare proposals on the same scale.

To start with us, briefly describe the job you want to hand over, the systems involved and how the work is done today; you can reach us through the contact page. We fill in the task card together and send, in writing, which part becomes an agent, which part a fixed workflow, and the scope.

Starter options are listed under AI automation pricing; model and platform usage fees are paid separately, from your own account, to the providers. If we conclude you do not need an agent, we will say so and suggest the simpler route.