What Is Structured Output? Getting Reliable JSON From AI

What is structured output?
Structured output is a method that makes an AI model shape its answer to a schema you define in advance. The schema names the fields, sets their data types and marks which ones are required, so the model returns predictable JSON that your software can read directly instead of free-form prose.
Think of a paper form. When you contact support, you can write a long letter, or you can fill in a form with separate boxes for name, phone number and request type. You cannot write outside the boxes. Likewise, structured output hands the model exactly that kind of form.
In short, what is structured output really about? It is about the shape of the answer, not about its wisdom. Because the shape stays fixed, the program that reads the answer never meets a surprise.
This guide explains the term at a conceptual level. Therefore, you will not see code blocks. The goal is to help you and your developers ask the right questions.
Why does free-form AI text break automation?
In a chat window, a person reads the model's answer. Meanwhile, people forgive a missing comma or a friendly extra sentence. Software does not. Also, an automation step expects a specific field, and when that field is missing, the whole flow stops.
For example, imagine you ask a model for the customer name and total of an order as JSON. Most of the time the answer is perfect. Sometimes, however, the model adds an intro such as "Here is the data you asked for:" or forgets the closing brace. The parsing step fails, and the process halts.
Each failure looks small. At scale, though, they pile up. In a flow that handles hundreds of records a day, a rare broken output still means manual fixes every week.
Older workarounds existed. Instead, teams cleaned the text with regular expressions or added "only write JSON" to the prompt. These tricks help, but they stay fragile, and that is the problem. When the model changes its writing habits slightly, the rules break, and maintenance grows.
Structured output closes that gap. Once the model follows a fixed form, the output format becomes predictable.
Where does structured output make sense in practice?
The term sounds abstract, but the use cases are concrete. In practice, structured output helps wherever a model must turn loose text into tidy data.
- Data extraction: Pull dates, amounts and parties out of an invoice, a contract or a form message into separate fields.
- Classification: Tag an incoming email with topic, urgency and the team that should own it.
- Summaries: Split meeting notes into decisions, owners and deadlines.
- Interface generation: Feed a screen with clean data for cards and lists.
- Agent workflows: Keep the intermediate results that an AI agent passes between steps in a stable format.
If you want to see how agents fit into marketing work, read our guide on AI agents for marketing. Here we only stress one point: agents need a schema whenever they hand data from one step to the next.
As a rule of thumb, if the first reader of the model's output is software, think about a schema.
What is a schema and how does it describe the output?
In short, a schema is a written description of the fields your data will contain. First, for each field, you state the name, the type and whether it is required. In the JSON world, an open standard called JSON Schema does this job, and you can find its rules on the official site at json-schema.org.
Conceptually, a schema carries these pieces of information:
- The overall type of the data, such as an object or a list.
- The name and type of each field, such as text, whole number, decimal or yes/no.
- Which fields you require.
- Fixed values a field may take, such as "low", "medium" or "high".
- A short description that explains what each field means.
The last item matters more than it looks. The model also reads field descriptions. So writing "delivery date, day and month only" instead of "date" raises the chance that the model picks the right value.
Why do field descriptions and fixed options matter?
Put simply, a model does not read a schema as a bare list of types. It also picks up hints from field names and descriptions. A well-written description therefore works like an extra instruction.
Consider a field called "status". For example, without a description, the model might write "resolved", "closed" or "done" on different days. If you limit the field to "open, waiting, closed", then you receive one of those three words every time. Reporting and filtering then become easy.
When you write descriptions, follow these habits:
- Explain in one sentence what the field means.
- State the format you expect, for example the order of day and month in a date.
- Say what to do when the information is missing.
- Add a note that separates two similar fields.
Also keep fixed options small. More than seven or eight categories can lower the model's decision quality and make your own analysis harder. If you need more, consider a two-step flow: ask for the main category first, then for the subcategory in a separate call.
How does structured output work?
The official documentation of the major providers describes a similar approach. First, you send the schema to the API inside your request, and the model builds its answer to match that schema. Conceptually, you can think of the process as constrained generation.
A model writes text piece by piece, token by token. So with constrained generation, the system filters out schema-breaking pieces at every step. For example, if a required field must end with a comma, the model cannot pick a different character there. As a result, the final text matches the schema in format.
If tokens are new to you, start with our article what is a token. For general API habits, our guide to the OpenAI API is a good first stop.
Details vary by provider. Therefore, read the current documentation of the provider you use before you build anything.
What is the difference between JSON mode and schema enforcement?
Many people mix up these two ideas, so let us separate them. JSON mode aims only for a response that is valid JSON. Schema enforcement aims for a response that also follows your fields and types. Provider documentation draws this same line and recommends the schema approach when it is available.
Here is an example. In JSON mode, the model returns valid JSON, but it may invent a field called "total" when you wanted "amount". The syntax is right. However, the structure is not the one you planned. With schema enforcement, the field names and types come from your schema.
In short, JSON mode promises that your parser can read the text. Schema enforcement promises that the text also has the shape you asked for. Neither promises that the values are correct, and we cover that point next.
What is structured output actually promising?
To set the right expectations, you need to know where the guarantee ends. Schema enforcement is mostly a format guarantee. In other words, fields exist, types match and required fields do not go missing.
Several things fall outside that promise:
- The value being true. A model can write a wrong amount in a perfectly valid shape.
- Consistency with the source. Inventing a fact that the document does not contain is not a schema error.
- Your business rules. A rule such as "the end date must come after the start date" is not the schema's job.
- A complete answer. If the output hits a length limit, it can stop halfway.
So structured output is not a quality fix on its own. Think of it as a reliability layer. For accuracy, though, you still add validation and checks.
Why can schema-valid data still be wrong?
Google's official documentation says this clearly: the output can be syntactically correct without being semantically correct. In addition, the documentation advises you to validate values in your application every time.
Here is a simple example scenario. For instance, a model fills an "urgency" field from a customer email. Your schema limits the field to "low, medium, high". The model writes "high", and the format is flawless. However, the email is actually a routine question. The format is right, and the judgment is wrong.
So you must answer two separate questions. First, is the output in the shape I expected? Second, is the information inside it correct and reasonable? Structured output helps with the first question only. For the second, you use rules, sample checks and, when needed, human review.
How do you build validation and retries?
In practice, a solid flow combines trust with checking. Even if you use schema enforcement, a second validation layer in your own application is a good habit. Conceptually, the flow looks like this:
- Ask the model for a schema-conforming answer.
- Run the answer through a validator in your own application.
- Once the format passes, check your business rules, for example whether the amount is above zero and the email field looks valid.
- When a check fails, retry once or twice and include the error details.
- After that, if it still fails, move the record to a human review queue.
Keep the retry count low. Endless retries raise both cost and delay. Also, log every failed attempt. That way you see which fields keep failing, and you can improve the schema or the descriptions.
This layer also protects you when the schema changes. Add a new field as optional first, then make it required once the receiving system is ready. Deleting or renaming a field is riskier, so produce the old and new fields side by side for a while.
What should you send back on a retry?
When validation fails, repeating the same request blindly usually produces the same mistake. Instead, a better approach is to tell the model briefly what went wrong.
For example, add a note such as "the phone field contains no digits" or "the amount cannot be negative". The model then reads the note, fixes the field and mostly keeps the rest. As a result, the second attempt tends to be more accurate than the first.
Keep these points in mind:
- Write the error note short and concrete.
- Send the original raw text together with the previous answer.
- Limit the retries to one or two.
- Log the result of every attempt.
If failures continue, the problem often sits in the source text, not in the model. For example, the form may contain no phone number at all. In that case, sending the record to a person is cheaper and safer than trying again and again.
How do you handle refusals and truncated replies?
Official documentation points to two edge cases. First, the model can refuse a request for safety reasons. The reply may then break your schema, and providers flag the refusal with a separate, detectable marker. Your application should check that marker and never treat a refusal as normal data.
Second, the output can hit a length limit. Therefore, a cut-off reply leaves the JSON incomplete, and it fails the schema. The usual fix is to raise the output limit or to shrink the schema.
For example, if a customer message trips a safety filter, the model may return a refusal note instead of your fields. If your code reads that as an empty record, it writes meaningless data into the CRM. Check for refusals before you parse.
Add separate branches for both cases:
- On refusal: flag the record and show a suitable message, or send it to human review.
- On truncation: raise the limit and retry once.
- In both cases: write the event to your logs.
These details look minor. Still, in production, they cause the most trouble.
What is the difference between function calling and structured output?
Both use schemas, but they serve different goals. First, function calling lets the model decide to call a function or tool that you defined. Second, structured output fixes the shape of the answer that goes back to the user or to your software. According to provider documentation, you reach for function calling when you connect the model to tools, functions or data. You choose a response format when you want to structure the answer itself.
We cover the first topic in a separate article, what is function calling, so we do not repeat it here. In practice, you often use both. The model first calls a tool, and then it gives its final answer in a fixed structure.
Here is a simple rule. If the model must do something in the outside world, such as booking an appointment, think about a tool call. If the model only produces an answer or a data package, such as splitting a form into fields, a response schema is enough. Most real systems use both side by side, because the two jobs often meet.
Some providers also offer strict schema checking for tool parameters. That way you secure both the call parameters and the final answer.
How do similar approaches compare side by side?
The table below helps you separate methods that look alike. Details vary by provider, so treat it as a conceptual comparison.
| Approach | What it does | Format guarantee | Typical use | Main risk |
|---|---|---|---|---|
| Prompting "give me JSON" | Asks the model for JSON in plain words | None, the model usually complies | Quick experiments | Intro sentences, missing braces |
| JSON mode | Aims for a valid JSON response | Syntax level | Simple, flexible shapes | Field names and types may drift |
| Schema enforcement (structured output) | Shapes the reply to your schema | Field and type level | Extraction, classification | Values can still be wrong |
| Function calling | Lets the model call a tool | For tool parameters | Assistants that take actions | Choosing the wrong tool |
| Parsing afterward | Pulls data from free text with rules | Only as good as your rules | Legacy systems | Fragile, hard to maintain |
The "format guarantee" column matters most. For an automation in production, you usually want a schema-level guarantee first and your own validation on top.
How do you move form data into your system with structured output?
Here is an example scenario. A consulting company wants to move requests from the free-text contact form on its website into a CRM. Customers describe their request in their own words. The sales team, however, wants tidy fields such as name, company, request type and urgency.
Conceptually, you build the flow like this:
- Define the schema: name, company, phone, request type (fixed options), urgency (fixed options) and a short summary.
- Then, for each submission, send the text to the model together with the schema.
- Allow empty values in the schema, so the model can leave out fields it cannot find.
- Next, validate the reply in your application and check the phone and email format.
- Write passing records to the CRM, and send the others to a review queue.
The sales team now works with tidy records instead of reading raw text. If you want help with the workflow side, see our CRM automation page.
How do you design a good schema?
Schema quality, therefore, directly affects output quality. These are the practical principles our team follows when it designs one:
- Keep the field count low. Ask only for fields you will actually use.
- Choose clear field names. Write "delivery_date" instead of "date1".
- Limit categories to a fixed list instead of free labels.
- Decide what happens when the data is missing. If you forbid empty values, the model may feel pushed to invent something.
- Write a short, clear description for every field.
- Avoid deeply nested structures without a reason.
The third and fourth items reduce the risk of wrong data the most. A schema that lets the model say "unknown" does not force it to guess. Instead, it keeps the data honest. Descriptions and examples also work together with prompt design, so our article on prompt engineering complements this one.
What limits come with structured output?
Structured output is powerful, but it has limits. Provider documentation highlights a few common ones, so let us go through them.
First, providers do not support the whole JSON Schema standard. They accept only a subset. Google and Anthropic say this openly. Some rules, such as numeric ranges or text length, may not work with certain providers. In that case, you enforce those rules in your own validator.
Second, very large or deeply nested schemas can get rejected. Providers set complexity limits. Check the current limits in the provider's official documentation.
Third, the first request can add some latency. The provider prepares the schema and can reuse that work later. Changing the schema often weakens this benefit.
Fourth, the schema adds tokens to each call, which can raise cost slightly. For the cost side, see our article on prompt caching.
What is structured output not the right tool for?
However, not every AI output needs structure. For a blog draft, a warm reply to a customer or a brainstorm of ideas, a schema is often an unnecessary restriction.
Structured output is probably unnecessary in these cases:
- Only a person will read the output.
- The fields change every time, and you cannot describe a stable shape.
- The task is a small, one-time experiment.
On the other hand, it brings real value for document extraction, email triage, form import and any step that feeds another program. In short, if the next stop of the output is a program, a schema deserves a look.
When the job also needs facts from your own documents, a retrieval approach joins the picture. Our article on RAG explains it, and the two methods complement each other.
What is structured output's impact on cost and latency?
The schema approach often lowers total cost, because fewer parsing failures mean fewer repeat calls. Even so, keep two points in mind.
First, the schema and its descriptions are part of the request. A long schema means extra input tokens on every call. Keeping the field count low reduces both cost and delay. Check current prices and limits on the provider's official page.
Second, sending a stable schema in the same form on every call may let you benefit from provider caching mechanisms. This varies by provider, so read the documentation.
Third, limit the output length on purpose. Also, overly long text fields raise cost and increase the risk of truncation. Mention a short length expectation in the description of fields such as summaries.
Finally, measure latency with your real data. A test with one tiny example tells you little about production behavior.
What should you watch for on security and privacy?
A schema protects the format, but it does not provide security on its own. Therefore, if you work with text from outside, you must plan for some risks.
For example, a user can write a sentence into a form that tries to instruct the model. This is called prompt injection, and a schema alone does not stop it. The model may still produce a value that fits the schema but misleads you. We explain this topic in our article on prompt injection.
These measures help in practice:
- Never turn model output directly into a database query or a system command.
- Validate every field with your own rules.
- For forms with personal data, document which data goes to the provider under the privacy laws that apply to you.
- Do not send sensitive fields such as ID numbers unless you need them.
- Decide how long you keep logs of structured records, and who can read them.
Note: This article is not legal advice. Talk to your legal adviser about your personal data processes.
How do you measure whether it works?
After going live, "it seems to work" is not enough. Instead, a few simple metrics keep quality and cost under control.
| Metric | What it shows | If it is low |
|---|---|---|
| Format pass rate | How often a reply passes the schema | Simplify the schema, check provider limits |
| Value accuracy | Field accuracy against a human sample | Improve descriptions and add examples |
| Retry rate | Share of records that fail the first time | Find the failing field in your logs |
| Human review rate | Share of records in the review queue | Revisit your rules and empty value policy |
| Time per record | Latency impact | Keep the schema stable, remove extra fields |
A small weekly sample is enough. For example, you can check twenty random records by hand each week. That number is only an example calculation, so adjust it to your own volume.
What checklist should you run before going live?
Use the short checklist below with your team before you launch a structured output flow.
- Write down the target system and the fields it needs.
- Design the schema with the smallest useful set of fields.
- Add a rule for empty values when the model cannot find the data.
- Check the schema features your provider supports in its official documentation.
- Build a second validation layer in your own application.
- Test your business rules separately.
- Add branches for refusals and truncated replies.
- Limit retries and log every failure.
- Measure the pass rate on a small sample of real data.
- Document the flow of personal data.
A flow that passes this list catches most first-week surprises in advance.
How can our team help with structured output projects?
At Talha Aslan and team, we often meet the same problem when we connect AI to workflows: the model works well, but its output does not enter the next system cleanly. In many cases, the fix is not a bigger model. Instead, a better schema and a validation layer solve it.
We can design the process with you for jobs such as moving form data into a CRM, sorting emails or extracting fields from documents. You can review our approach on the AI automation services page. Our AI document processing page also shows related examples.
If you want a refresher on the models underneath, our article on large language models is a good place to return to.
Where should you start?
Start small. First, pick one data flow, such as your contact form or your support inbox. Then write a simple schema with five to eight fields and try it on real examples.
Then measure three things: how often the format is right, how often the values are right and how much manual work you saved. The first measurement will mostly solve the format problem. The second and third show where to focus your validation rules.
When people ask us what is structured output good for, we answer with one sentence. It fixes the format problem, not the quality problem. Once the format is stable, your team can focus on what matters most: the accuracy of values and your business rules.
Remember that structured output is a bridge between AI and software, and both ends of the bridge need a checkpoint. Models and schema features change often, so check the official documentation regularly. The docs from OpenAI, Anthropic and Google are good starting points.



