What Is Information Architecture? How to Build a User Centered Site Structure

Information architecture decides whether visitors find what they came for or leave confused. This guide is not about drawing a menu. Instead, it covers the research methods behind a user centered site structure: content inventory, card sorting, tree testing, content models and taxonomy. I have used this sequence on client projects since 2012, and here is how it works in practice.
What is information architecture on a website?
Information architecture is the practice of organizing, labeling and connecting website content so that it matches how users think about it. Put simply, its goal is clear: visitors should find the right information quickly and with confidence. Menus, categories, URLs and internal links are the visible results of these decisions.
In practice, many teams confuse information architecture with menu design. However, the menu is only the tip of the iceberg. Below it sit the content inventory, the content types, the label vocabulary and the relationships between pages. If you skip these layers, the menu usually breaks during the first growth spurt.
In practice, information architecture answers three questions:
- What content exists, and which types does it fall into?
- Which words and groups do users use when they think about it?
- How do people move from one page to another, and in how many steps?
Once these answers are clear, design and content production both speed up. Nobody argues any more about where a new page belongs. The decision already exists, and the whole team works from the same map.
Is information architecture the same as a sitemap or a category structure?
No. A sitemap is one of the output documents. A category structure is one layer. Information architecture is the full set of decisions behind both. For example, the label vocabulary and the content model never appear in a sitemap, yet they shape the whole structure.
On sites with thousands of products, category depth, filters and crawl budget become a specialist topic of their own. I covered that in my guide to category structure for large websites. This article focuses on the stage before that one: how you discover the right structure by researching real users.
In short, the category guide answers "how do you scale a structure?". This one answers "how do you find the right structure in the first place?". I recommend reading them in that order.
Why should site structure follow users instead of your org chart?
The most common mistake I see is a navigation that copies the company org chart. Sales wants one menu. The technical team wants another. As a result, visitors must learn your internal structure before they can find anything. Users, however, think about the problem they want to solve, not about your departments.
Take a manufacturer that opens three menus called Products, Solutions and Applications. The visitor only asks one thing: which machine fits my industry? So three menus create three times the hesitation. I discussed the specific needs of these sites in my article on manufacturing website design.
Also, company language and search language rarely match. You say "enterprise solutions" while the customer searches for "bulk orders". Therefore, choose labels with user data, not in the meeting room.
Should you start with a content inventory?
Yes. Every project I run starts with a content inventory. It is a list of every page: URL, title, content type, owner and freshness. If there is no site yet, list the planned content instead. Without this list, any grouping you create is a guess.
In practice, my inventory lives in a spreadsheet. Each row gets a decision column: keep, merge, rewrite or remove. That way, dead pages never move into the new structure. You also spot two pages covering the same topic early.
These are the sources I use to build it:
- The URL list from a crawler and the XML sitemap.
- Pages with impressions in Google Search Console.
- Pages with traffic in analytics that are missing from the menu.
- Documents the sales team sends to prospects again and again.
The inventory also makes the workload visible. You know from day one how many pages need rewriting and how many will merge. Consequently, the schedule and budget stay realistic. My Google Search Console guide shows how I pull the page data.
What is a content model and where does it fit?
A content model is a schema that defines the content types on your site and the fields each type contains. For example, a "service" type might have a title, short summary, process steps, FAQ and a related case study. A "blog post" type has different fields.
The model lets you think in types rather than single pages. So when you add a new service, the template already exists. You also define relationships between types. For instance, every service page links to three related posts, and every post points to one main service.
These relationships are the foundation of your internal linking. I covered that side in my internal linking strategy guide. In other words, the content model is the hidden skeleton, and the menu is the skin on top of it.
Sites without a content model treat every page as a separate design request. Two pages of the same type end up with different sections. Visitors then have to learn a new layout each time. A model prevents this inconsistency from the start.
Which fields belong in a content model?
I define four groups of fields for each type. First come identity fields: title, slug and summary. Next come body fields: main text, steps and tables. Then relationship fields: parent category and related content. Finally, management fields: owner and last review date.
| Content type | Required fields | Relationships |
|---|---|---|
| Service | Title, summary, process, FAQ | Related posts, case study |
| Product | Name, specs, image, price info | Category, accessories, documents |
| Blog post | Title, body, author, date | Main service, topic tag |
| Case study | Problem, approach, outcome | Service, industry |
You fill in this table together with the designer and the developer. As a result, the components in Figma match the fields in the CMS one to one. I described that workflow in my Figma web design process article.
What is taxonomy, and how do categories differ from tags?
A taxonomy is the controlled vocabulary you use to classify content. Categories are a hierarchical kind: each item usually belongs to one parent group. Tags, on the other hand, are flat: one item can carry several of them. A good taxonomy uses both on purpose and defines the job of each in writing.
For example, on a law firm website "Practice Area" is hierarchical: employment law, commercial law. "Client Type" is a flat dimension: individual, small business, enterprise. So the same article can appear along two axes.
Without this distinction, tags spiral out of control. Writers add new tags to every post, and hundreds of thin archive pages appear. That is why every taxonomy needs a term list and a single owner. Only that person approves new terms.
I also expect each tag to hold several items before it goes live. A tag with one post is an empty corridor for the user. Until enough content exists, the term waits on a draft list.
How do you choose taxonomy terms?
I combine three sources. First comes search data: which words do people use for this concept? Second, card sorting shows how do participants name their groups? Finally, there is the language that sales and support hear from customers every day.
On the search side, a keyword map helps a lot. You decide which term belongs to which page with the method from my keyword mapping guide. Still, search volume should not decide alone. Clarity comes first.
For each concept, my term list stores one preferred label plus synonyms. For example, the preferred label is "request a quote" and the synonyms are "get pricing" and "price estimate". This way, on site search can route synonyms to the right page.
Multilingual sites need extra care here. Each language has its own search habits, and literal translation often produces the wrong label. My multilingual website SEO guide goes deeper on this.
What is card sorting and how do you run it?
Card sorting is a research method where users receive content topics on cards and group them in a way that makes sense to them. The aim is to reveal the mental model in their heads. Paper cards work fine, and online tools also do the job.
The Nielsen Norman Group card sorting guide describes three variants. In an open sort, participants create and name the groups. In a closed sort, the groups are predefined. A hybrid sort offers some groups and lets participants add their own.
My steps look like this:
- Pick 30 to 60 representative items from the inventory.
- Write each card as a plain phrase without jargon.
- Recruit participants from the real target audience.
- Ask each participant why they built each group.
Keep sessions short. As the number of cards grows, participants get tired and place the last cards carelessly. If the inventory is large, split it into two sessions.
How many participants does a card sort need?
Nielsen Norman Group recommends at least 15 participants for a qualitative card sort. For quantitative studies that you want to generalize, it suggests 30 to 50. If no clear pattern emerges, it advises recruiting more people.
For most business websites, a qualitative study is enough. So I start with around 15 people and focus on the "why" question at the end of each session. Numbers tell me which cards belong together. Conversations, by contrast, tell me why a label fails.
If recruiting is hard, ask existing customers, dealers or prospects from the sales pipeline. However, do not test only with employees. They already know the internal language, so the result reflects the company's mind, not the customer's.
How do you read card sorting results?
The first thing I check is the similarity matrix. It shows how many participants placed two cards in the same group. Cards that often stay together form a natural group. Cards that keep moving between groups point to ambiguous content.
Those ambiguous cards are a valuable finding. Usually the content is doing two jobs, or its title misleads. For example, if a "Support" card goes to pre sales in some sessions and after sales in others, you probably need two separate pages.
I list the group names as well. Specifically, the words participants use most are strong candidates for menu labels. Even so, I do not ship them right away. The next step tests them with a tree test.
Afterwards, I condense everything into a one page summary: strong groups, ambiguous cards and candidate labels. This keeps the stakeholder meeting short. The discussion now rests on user data rather than personal taste.
What is tree testing, and how does it differ from card sorting?
Tree testing is an evaluation method that shows participants only the text hierarchy of a site and asks them to find specific information. There is no visual design, no search box and no promotional content. You measure the structure itself, independent of the interface.
The Nielsen Norman Group article on tree testing puts the difference clearly. Card sorting is a generative method for discovering possible groupings. Tree testing, meanwhile, evaluates a proposed navigation hierarchy. One asks the question; the other checks the answer.
That is why I use them in sequence. Card sorting produces a draft tree. Then a tree test shows whether that draft actually works. After that, I fix the weak branches and test again.
If you already have a website, test the old tree before the redesign too. This gives you a baseline. Consequently, you can prove that the new structure is better, not just newer.
Which tasks and metrics should a tree test use?
Write tasks from real needs. Avoid leading phrasing like "Find the Services menu". Instead, ask "You want a price quote for your company. Where do you click?". Also avoid repeating the exact words that appear in the menu.
These are the core metrics I track:
- Success rate: did the participant reach the correct node?
- Directness: did they go straight there, or wander between branches?
- First click: was the first top level category correct?
- Time: how long did the task take?
NN/g describes directness as a signal of how clear or ambiguous your labels are. For example, high success with low directness means people find the item eventually, but the label makes them hesitate. So that branch needs a better name.
Is a deep hierarchy or a wide hierarchy better?
There is no single answer, but there is a way to find the balance. A deep structure offers few options per level and requires more clicks. A wide structure needs fewer clicks but shows many options on each screen. Both have a cost.
I let a tree test decide. I test the same content with two different trees and compare success and directness. For example, a wide tree with eight top categories against a deeper one with four. Whichever gives clearer results wins, so the debate rests on data rather than taste.
My general rule is to offer as many top level options as users can clearly tell apart. If options look too similar, I reduce them. On the other hand, when users miss the right branch on the first click, the issue is usually the label, not the number.
Then there is the "three click rule". People repeat it often. However, I measure confidence at each click rather than the click count. In short, a clear four click path beats a hesitant two click path.
How do you make labels and menu names clear?
A menu label should answer "what will I find here?" at a glance. Therefore, I avoid creative but vague names. Labels like "Discover" or "Our World" look interesting yet tell the visitor nothing.
These are the rules I apply:
- Use the customer's word, not the company's word.
- Make sure labels on the same level do not overlap.
- Keep each label to two words where possible.
- Use the same word for the same concept everywhere on the site.
Consistency matters most. If the menu says "Get a Quote", the button says "Request Pricing" and the form says "Application", users assume three different actions. I discuss this further in my UX mistakes that kill sales article.
How do navigation systems carry the architecture?
Navigation is the interface that shows the architecture to users. The main menu is only one part. Local menus, breadcrumbs, footer links, related content blocks and on site search also belong to the system. Each serves a different user need.
Breadcrumbs also show users where they are in the hierarchy. Google explains in its breadcrumb structured data documentation that it uses this markup to categorize page information in search results. I covered the markup in my schema markup guide.
On mobile, navigation becomes even more critical. The screen is narrow, so the menu hides and the structure disappears from view. For that reason, I keep key tasks visible inside the page as well. My article on mobile first design explains the approach.
How does information architecture affect SEO?
Information architecture directly affects how search engines understand your site. In its crawlable links documentation, Google states that it uses links to find pages and can follow a elements that have an href attribute. So the links in your architecture form the crawler's road map.
A strong structure also concentrates topical authority. When a main page and its supporting pages link to each other, Google sees their relationship more easily. By contrast, in a messy structure, pages covering the same topic compete with each other.
Still, I never build the architecture for SEO alone. A structure that is clear for people is usually clear for search engines too. My guide on balancing UX and SEO goes into more detail.
Should your URL structure mirror the architecture?
Usually yes, but not rigidly. Because it is readable, a clean URL tells users and search engines where a page sits. For example, "/services/seo-consulting/" shows at once that the page belongs to the services group. That consistency builds trust.
However, avoid deep URLs. If every category level goes into the address, restructuring later requires hundreds of redirects. Therefore, I keep only stable levels in the URL. Temporary campaign groups and frequently changing tags stay out.
Next, keep slugs short, lowercase and free of special characters. When you change the architecture, redirect old addresses correctly. My guide to protecting SEO during a redesign walks through that process.
How can on site search data improve the structure?
The search box tells you, in the user's own words, what they could not find in the menu. So I treat it as a research source. Open the site search report in your analytics tool and list the most frequent queries.
Three patterns stand out in that list:
- Terms that exist in the menu but people still search for: a label problem.
- Terms with no matching content at all: a content gap.
- The same concept spelled in different ways: new synonyms.
For example, if users keep searching "shipping" while the menu says "Delivery Terms", consider renaming the label. I also track queries that land on the "no results" page. That page is the most honest feedback your architecture gets.
Which metrics show whether the architecture works after launch?
Measurement does not stop at launch. Tree test metrics are for the pre launch phase. After launch, look at real behavior and review it next to your marketing goals.
| Metric | What it shows | Source |
|---|---|---|
| Site search rate | Needs the menu does not meet | Analytics |
| Zero result searches | Content or synonym gaps | Analytics |
| Exits from category pages | Groups that miss expectations | Analytics |
| Pages with impressions | Crawlability and topic coverage | Search Console |
| Form and quote conversions | Contribution to business goals | Analytics, CRM |
Never read this table in isolation. For instance, if exits rise, check the content first and the structure second. My guide on setting website conversion goals helps you choose what to track.
Can AI tools help with information architecture?
AI tools save time during preparation, but they do not replace user research. For example, they can roughly cluster a large inventory or simplify card wording. They are also useful for summarizing notes from card sorting sessions.
On the other hand, AI suggested groups reflect an average of training data, not your users' minds. Therefore, I use them to generate hypotheses, not decisions. Every suggestion still goes through card sorting and tree testing.
AI search experiences add another angle. A clear structure and consistent labels help machines understand your content correctly too. I cover this in my article on technical SEO after AI.
What are the most common information architecture mistakes?
The mistakes I meet in the field look alike. First, teams draw a menu without research. Second, they squeeze everything into the main navigation. They also use internal jargon in labels.
Two more show up often. Some teams run a card sort but never validate it with a tree test. Others build the architecture once and never update it as the site grows. As a result, within two years the structure becomes unrecognizable.
Another mistake is trying to fix navigation problems with decoration. Icons and animations in the menu do not repair wrong groupings. Worse, they slow the page down and make mobile use harder.
How long does an information architecture project take?
Duration depends on site size and research depth. Here is a starting range based on my field experience, not a guarantee: for a business website with a few dozen pages, the inventory, one card sort and one round of tree testing usually take two to four weeks.
In practice, a small core team is enough. You need a decision maker, a content owner, a designer and a developer. I also add someone from sales or support, because they know the real customer questions best.
I treat this work as the first phase of any new website. In my web design and SEO consulting projects, the structure gets tested with users before any visual design begins. In short: skeleton first, surface second.
Not sure about your current structure? Start with a small tree test. Give a handful of customers five tasks and watch where they get stuck. Even this short trial shows you which branch deserves attention first.




