Artificial Intelligence

What Is Function Calling? How AI Models Use Tools Step by Step

Talha Aslan 20 min read 2 views

What is function calling?

Function calling is the ability of an AI model to decide which of your predefined functions to call and with which arguments. The model does not run the function itself. It returns a structured call request, your application executes it, and the result goes back to the model.

For example, think of a restaurant. The customer tells the waiter what they want, and the waiter writes an order slip the kitchen understands. The waiter does not cook, because the kitchen does. In function calling, the model is the waiter and your code is the kitchen.

This analogy also shows a risk. Whatever the waiter writes on the slip gets cooked, so someone should check the slip first. In the same way, you should validate every call the model proposes before you run it.

So this article covers this one term only. We mention neighboring concepts briefly and link to our guides for the details. For exact parameter names, the provider's official documentation always wins.

What is function calling in practice: how does the flow work step by step?

The flow has four steps, and the official docs of the major providers all follow the same skeleton. However, the details differ from API to API while the idea stays the same.

  1. First, you send the model a list of available functions, including names, descriptions and parameters.
  2. The model reads the user's request and decides whether a function is needed.
  3. If so, the model returns a call proposal with the function name and arguments. Then your application runs that function in its own environment.
  4. Finally, you send the result back, and the model turns it into a natural answer for the user.

In other words, there is a back and forth between the model and your application. The model only produces text and structured proposals. Also, everything that touches the outside world stays on your side.

OpenAI's docs give each call an ID and ask you to match the result to it. So the model knows which result belongs to which call. Likewise, Anthropic's docs use the same pairing idea.

Does the model actually run the function itself?

No. Google's Gemini docs say it plainly: the model does not execute the function, and your application extracts the name and arguments and runs it. Meanwhile, the OpenAI and Anthropic docs describe the same flow.

This split matters a lot in practice. Because of this, authorization, logging, error handling and security decisions all live in your code. The model does not connect to a database, take a payment or send an email by itself.

Still, there is one exception. Some providers offer ready-made tools that run on their own infrastructure, such as web search. These are called server tools because the provider runs them. The flow in this article is about the functions you define yourself.

Knowing who runs what also settles who is responsible. So if your own function fails, the failure is yours. So plan testing, logging and monitoring from day one.

For example, when a customer asks "where is my parcel?" in a support bot, the model only writes the lookup request. The code that connects to the shipping system, checks identity and returns the answer is yours. Therefore, that boundary lets you catch mistakes on your side.

What are the parts of a function definition?

A function definition has three main parts. In practice, the more carefully you write them, the better the model chooses.

  • Name: a short, meaningful name that describes one job, such as get_order_status.
  • Description: plain sentences that say what the function does, when to use it and when not to.
  • Parameter schema: the name, type and required status of each argument. Providers use JSON Schema for this.

The description is the most underrated part. In fact, the model mostly relies on it when it picks a function, so wording matters. A vague description leads to wrong picks, and a clear one leads to right picks.

Some providers also offer a strict mode. In that mode the arguments the model produces match your schema exactly. Check the provider's official docs for current options and setting names.

Also keep free-text fields to a minimum. Define fixed option lists when you can, because models make fewer mistakes with limited choices.

Separate required fields from optional ones on purpose. For example, when a required field is missing, the model should ask the user. When optional fields pile up, the model tends to guess.

How do you write a function description so the model picks correctly?

Treat the description like a job posting for the model. Say what the function does, when it applies and what it does not touch. A vague posting attracts the wrong candidate.

  • State the job in one sentence: "Returns the current status of an order by order number."
  • State the limit: "Does not start returns or exchanges, it only reads the status."
  • Show the argument format: "The order number is a short code made of letters and digits."
  • If two functions look alike, explain the difference clearly, because this is where models slip most.
  • Do not over-explain. Every extra sentence also adds to token cost.

Also use a consistent naming pattern. Start every name with a verb and include the object. That way both the model and your teammates find what they need quickly.

How does the model decide which function to call and when?

The model decides based on the user message, the conversation history and your function descriptions. Anthropic's docs summarize it this way: if the request matches a tool's described capability and the answer is not already in context, the model calls the tool. For stable knowledge, creative tasks and small talk, it answers directly instead.

You steer this behavior in three ways. First, sharpen the descriptions. Second, add a rule to the system instructions, such as "verify facts you do not know with a tool". Third, set the tool choice mode.

Choice modeWhat it doesWhen you use it
AutoThe model decides whether to call anything.General chat assistants.
RequiredThe model must call some function.Flows where every message becomes an action.
Specific functionThe model may call only the one you pick.Single-step, predictable jobs.
NoneThe model never calls a function.Moments when you only want a text answer.

Mode names differ by provider. So check the official docs for the exact values.

What is function calling's relationship to tool use?

In short, they are largely the same idea. Anthropic's docs use the term tool use and say it is also called function calling. OpenAI and Google prefer the name function calling. Put simply, all three describe the same pattern.

There is a subtle difference, though. However, "tool" is a broader umbrella. For instance, a tool can be a function you wrote. It can also be a web search or code execution tool that the provider runs.

In short, the two terms are interchangeable in daily talk. For architecture decisions, ask one question: who executes the call, you or the provider?

When you research, searches for "tool calling", "tool use" and "function calling" mostly lead to the same pages. So search for all three while reading docs.

Can a model call several functions at once?

Yes, models allow it. Google's docs separate two cases: parallel calling for independent tasks, and sequential or compositional calling for tasks that depend on each other.

For example, a user asks "what is my order status and what are the return terms?" The order lookup and the policy search do not need each other, so the model can propose both in one turn. In a booking flow, you first need to query free slots and only then save the chosen one. Those calls run one after another, so order matters.

Parallel calls save time, but each call can fail on its own. However, each call can fail on its own, so you must send every result back with the right pairing.

The OpenAI and Anthropic docs also offer a setting that allows at most one call per turn. For critical actions, turning it on makes your life easier, because you inspect each step one by one.

How do OpenAI, Anthropic and Google approaches compare?

All three providers use the same core flow, but terms and settings differ. The table below summarizes only the conceptual differences we saw in the official docs. Above all, verify setting names and current behavior in the docs.

TopicOpenAIAnthropicGoogle Gemini
Name usedFunction calling.Tool use, also called function calling.Function calling.
Definition formatFunction definition with a JSON schema.Tool definition with an input schema.Function declaration with a JSON schema.
Returning the resultAn output item matched by call ID.A result block matched by tool use ID.A function response sent back.
Forcing a callAuto, required or a specific function.Auto, any or a specific tool.Auto, any, none or a validated mode.
Parallel callsSupported, can be switched off.Supported, can be switched off.Parallel and compositional calls.

So when you switch providers, you can carry your function logic over, but you rewrite details in the execution layer. Keeping that layer provider-neutral gives you flexibility later.

Which business scenarios benefit from function calling?

First, a model cannot reach live data or take actions on its own. Second, function calling fills both gaps. That is why it pays off wherever you want to connect a conversation to real work.

  • Order and shipment status lookups.
  • Creating, changing and canceling appointments.
  • Stock and price checks.
  • Finding a customer record in the CRM and adding a note.
  • Pulling data for a report and summarizing it.
  • Opening a support ticket and routing it to the right team.

They all share one pattern: the model understands the intent, and the code does the work. So if your product has an assistant that can chat but cannot act, function calling is the missing link.

The approach also helps marketing and sales teams. For instance, a lead form bot can understand a visitor's interest and open the lead record in your CRM. What matters is keeping its permission to create that record narrow.

What does the flow look like in an appointment assistant example?

Note that this is an example scenario. For instance, picture a consulting business. A customer writes: "Can we schedule a call on Friday afternoon?"

  1. The model proposes calling a get_available_slots function and passes the day as an argument.
  2. Next, your application queries the calendar system and returns the list to the model.
  3. Then the model offers the options, and the customer picks a time.
  4. This time the model proposes a create_appointment function. If something is missing, such as a phone number, it asks the customer first.
  5. Your application creates the record and sends the result back. The model writes the confirmation.

Notice that creating the record is a write action, so it needs care. So asking for the user's explicit approval makes sense. We return to this later, in the security sections.

Also think about what happens if the calendar system fails in step two. If your application reports the error to the model, it gives the customer an honest answer. Otherwise, if it stays silent, the model may invent a time.

Finally, ask for a summary at the end. When the model says "I booked a call for Friday", that sentence must rest on a real, successful record. In other words, if no record exists, the confirmation must not appear.

How does function calling differ from MCP?

Put simply, function calling is a model's ability to propose a tool call. MCP is an open protocol that standardizes how tools are offered to the model. One is a capability, and the other is a connection standard.

As a result, even with MCP the model still proposes a call behind the scenes. The difference is that you fetch tool definitions from a shared server instead of hand-writing them in every app. For details, see our guide on what MCP is.

In practice, plain function calling is usually enough in small, single-app projects. However, if you will use the same tools across several apps, MCP is worth considering.

Use this rule of thumb: when you have few tools that live in one app, keep it simple. When many tools are shared across teams, a common standard cuts maintenance work. Besides, the two do not exclude each other.

How does function calling differ from structured output?

On the other hand, with structured output, the goal is to get the model's answer as data that fits your schema. With function calling, however, the goal is to trigger an action. Both use schemas, but their purposes differ.

For example, if you want to extract an order number and an amount from an email, structured output is enough. If you want to look up the customer's order in your system, you need function calling. Sometimes you use both: the model first calls the function, then writes its answer in a schema-valid shape.

We cover the topic in detail in our article on structured outputs. So here we mention it only for comparison.

How do function calling, MCP, structured output and RAG compare?

These four concepts get mixed up often, so a side-by-side view helps. The table below summarizes the core job of each and who runs the work.

ConceptCore jobWho runs itWhen you choose it
Function callingThe model proposes an action or a query.Your application.To connect chat to live data and actions.
MCPOffer tools through a shared protocol.An MCP server and the client app.To share the same tools across many apps.
Structured outputGet the answer as schema-valid data.Nobody, the output is data.To extract fields and classify text.
RAGGround the answer in retrieved documents.Your retrieval layer.To answer from company documents with evidence.

We cover retrieval in our guide on what RAG is. In practice, an agent can even call document search as a function. So these concepts complement each other instead of competing.

How is function calling related to AI agents?

An AI agent is often function calling repeated in a loop. The model proposes a function, the application runs it, and the model looks at the result and decides the next step. The loop continues until the model says it is done.

So function calling is the building block, and the agent is the larger concept. An agent sets a goal, plans steps and uses tools in sequence. Still, a single call does not make an agent.

Two safety rules matter once you build a loop. First, cap the number of steps. Second, repeat the authorization check at every step. Otherwise the model could call tools forever or widen its own scope.

For agent architecture, see the links at the end of this article. Here we only draw the connection.

What security risks does function calling bring?

Giving a model the power to act enlarges the attack surface. Because the model is influenced by user messages and by the content it reads, it can be steered. For instance, a malicious piece of text can steer it toward the wrong function or the wrong argument.

  • Excess permission: if a function runs with broader rights than it needs, a small mistake causes big damage.
  • Prompt injection: an instruction hidden in a document or web page can mislead the model.
  • Wrong arguments: the model may guess missing information and invent a value.
  • Data leakage: because the function result goes back to the model, sensitive fields appear there too.

We cover this topic in a separate article on prompt injection. In short, schema validation is not security. A schema checks format, but it does not check permission.

Also treat function results as untrusted input. Text that comes from a web page or an email can look like a new instruction to the model.

How do you set approval and permission limits?

First, the core principle is least privilege. Run every function with only the permissions its job needs, because extra rights add risk. Split read and write actions into separate functions.

  • For read actions such as lookups and listings, automatic execution can be enough.
  • For write actions such as saving, canceling and sending, ask the user for explicit approval.
  • Keep irreversible actions such as deleting and paying away from the model, or tie them to human approval.
  • Take the user's identity from the session. Never leave identity to a model argument, because the model can be tricked.
  • Log every call: which function, which arguments, for whom and with what result.

So put the permission check inside the function, and do not trust the model. Even if the model suggests something wrong, your code must be able to say "this user cannot do that".

Also prepare for repeated calls. For example, after a network drop, the same write action can arrive twice. Design it so that running the same call a second time does no harm.

What happens if the model produces a wrong argument?

However, the model sometimes proposes a missing or wrong argument. When a user leaves out a detail, some models may guess a plausible value. Anthropic's docs state clearly that this behavior is not guaranteed.

That is why you validate every call before running it. Check type, range, required fields and business rules. For example, reject a call that books an appointment in the past.

When you reject a call, send the error message back to the model. Then the model usually retries by asking the user the right question or fixing the argument. However, set a cap on retries to prevent endless loops.

Strict mode reduces this risk but does not remove it. Even if the format is right, however, the value can be wrong. A date that fits the schema can still be the wrong date.

How does function calling affect cost and latency?

First, every function definition travels to the model as part of the request. Anthropic's docs say that tool names, descriptions and schemas count as input tokens. So many long definitions raise the cost.

The flow also means at least two model calls: one for the proposal and one to interpret the result. That lengthens response time. Check the provider's official pages for current prices and limits.

To lower cost, limit the number of functions to what you really need and keep descriptions short. We explain token logic in our guide on what a token is.

When you ask what is function calling worth to your business, include cost in the answer, because every extra definition and every extra turn has a price. Done right, though, it also cuts manual workload.

One tip for latency: run independent lookups in parallel. As a result, the user waits once, not four times.

What are common mistakes with function calling?

These are the traps teams fall into most often in first attempts. Fortunately, all of them are avoidable.

  • Offering too many functions. The more options, the more likely a wrong pick.
  • Loading several jobs into one function. Single-purpose functions produce fewer errors.
  • Taking identity and permission from arguments. The user identity must always come from the session.
  • Hiding error states from the model. In silence, the model also guesses.
  • Skipping the approval step. Write actions need user or human approval.
  • Not keeping logs. Otherwise, when something breaks, you cannot see what was called.
  • Returning raw results. Strip unnecessary and sensitive fields first.

Most of these mistakes come from the code around the model, not from the model. In other words, success depends on your engineering as much as on the model's intelligence.

How do you test a function calling flow?

First, start with fake functions before you connect real systems. A fake function returns fixed, predictable answers. So you measure only the model's decision quality.

  1. First, prepare a scenario list from real user messages. Mix clear, vague and incomplete examples.
  2. For each scenario, write the expected function and expected arguments.
  3. Record whether the model picked the right function and filled the arguments correctly.
  4. Try hostile inputs separately: asking to leave scope, asking about someone else's data, text with hidden instructions.
  5. Rerun the same list after every change, so you catch regressions.

Also watch real calls in your logs. After launch, add the new situations you meet to your scenario list. As the list grows, your quality also grows.

What is function calling's opposite: when do you not need it?

However, not every job needs function calling. If a process has fixed steps and clear rules, plain code or a form is often cheaper, faster and safer.

  • When the information to collect is known, a form is more predictable than a model.
  • When answers come only from a fixed document, RAG or simple search can be enough.
  • For actions with zero error tolerance, do not make the model the decision maker.
  • Do not automate irreversible actions without human approval.

So ask yourself: is understanding the intent really hard here? If the answer is no, then classic software wins. Function calling makes sense where intent arrives in free language and the job touches your systems.

Also run a small value check before you build. How many conversations will you have each month, and how many will turn into real actions? Calculate this number from your own data, because example numbers vary from business to business.

What should a business check when asking what is function calling?

Answer the items below before you start. Otherwise, every unanswered item comes back later as a bug or a security hole.

  • Which jobs will you really turn into functions, and which stay with people?
  • Does each function do one job, with a clear description?
  • Are read and write actions separate?
  • Do write actions require user approval?
  • Do identity and permission checks live in code or in the model?
  • Are argument validation and error feedback defined?
  • Did you set a cap on retries?
  • Are you logging every call?
  • Do you have a test environment with fake data before you touch real data?

Keep in mind that this list is a general starting point, not a legal or security audit. If you process personal data, ask a qualified expert about the rules that apply to you.

Where should you start with a function calling project?

First, start with the smallest, lowest-risk and most repetitive job. A read-only order lookup is a good first step. Then, once it works, add write actions with an approval step.

When you pick a provider, compare how each documents the flow, because the logic is similar but details differ. For API basics, our guide on how to use the OpenAI API is a good start. If you want to turn functions into an agent, our posts on AI agents for marketing and on running an agent on your own server are the next stops.

If you are planning an assistant that connects to your systems, you can talk to our team through our AI agent development service.

For official sources, read the providers' docs: the OpenAI function calling guide, the Anthropic tool use overview and the Google Gemini function calling docs. Docs change often, so verify setting names there.

Frequently Asked Questions

What is function calling in simple terms?
Function calling is a way for an AI model to ask your application to run one of your own functions. The model reads the request, picks a function and fills in its arguments. Your code runs it and returns the result. The model then writes the final answer, so chat connects to real data and actions.
Is function calling the same as tool use?
In everyday use, yes. Anthropic calls the capability tool use and notes that it is also called function calling, while OpenAI and Google use the name function calling. The word tool is broader, because it also covers provider-run tools such as web search. For architecture decisions, ask who executes the call.
Is function calling safe for production use?
It can be, but only if you design the safety layer yourself. The model only suggests calls, so your code decides what actually runs. Use least privilege, user approval for write actions, argument validation and logging. Schema validation checks format, not permission, so always put authorization checks inside the function.
Do I need to code to use function calling?
To connect it to your own systems, yes. Someone has to write the functions and the layer that runs them. Some no-code platforms hide this behind a visual builder, but you give up flexibility and control. A good first step is one read-only function, such as an order status lookup, tested with fake data.
Can function calling and MCP work together?
Yes. MCP is an open protocol that standardizes how tools reach the model, while function calling is the model's ability to propose a call. With MCP, the model still suggests a call, but the tool definitions come from a shared server. For a small, single-app project, plain function calling is often enough.
Does function calling increase cost and latency?
Yes, a little. Tool names, descriptions and schemas count as input tokens, and the flow needs at least two model calls: one to propose the call and one to read the result. Keep the number of functions small and descriptions short. Check the provider's official pricing page for current rates.
  • function calling
  • tool use
  • AI agents
  • LLM
  • MCP
  • structured output
  • AI integration
Share:
Talha Aslan

Google Partner digital marketing expert. Hands-on with SEO, Google Ads, web design and e-commerce projects since 2012; every post here comes from that experience.

Next project

Let's talk about your project.

Your brief goes straight to Talha Aslan and team: strategy led by Talha, delivery by an experienced team. The first consultation is free; we listen and come back with a clear roadmap.