How to Use AI for Data Analysis: ChatGPT, Claude, Gemini and Copilot in Practice

How do you use AI for data analysis?
Using AI for data analysis means handing a spreadsheet or exported report to an assistant such as ChatGPT, Claude, Gemini or Copilot and asking it, in plain language, to clean, summarise, chart or write formulas. The assistant speeds up the work; you still check the result and own the decision.
I have worked with ad accounts, sales sheets and analytics reports since 2012. The biggest change of the last two years is the gap between asking a question and seeing an answer. Building a pivot table, fixing a formula and formatting a chart used to take half an hour; today I often see a first draft in a few minutes. That said, I have also seen plenty of cases where the assistant got it wrong. So this guide covers both sides: where it saves time and where it can mislead you.
I built the article around a real workflow. First we look at which assistant suits which job. Then we move through data cleaning, exploratory analysis, visualisation and formula or code generation. Finally, we cover verification, hallucinations and data privacy under GDPR, because that is where the real risk sits.
What does an AI assistant actually do with your data?
An assistant turns your sentence into an operation. For example, when you ask "which campaign brought conversions at the lowest cost?", it groups rows, divides one column by another and sorts the result. Some assistants do this by running Python code in the background; others use the formula and pivot features inside the spreadsheet itself.
This distinction matters. An assistant that runs code actually performs the calculation and can show you that code. A model that only generates text, on the other hand, may produce numbers that look plausible without computing them. As a result, the same question can get answers of very different reliability from two tools.
- Preparation: fixing column names, unifying date formats, finding duplicate rows.
- Exploration: summary statistics, distributions, outliers and segment differences.
- Communication: chart suggestions, a short executive summary, presentation-ready tables.
- Generation: writing Excel formulas, SQL queries or Python scripts.
In other words, the assistant is not an analyst; it is a very fast helper. Asking the right question, knowing your data and placing the result in business context remain your job. The rest of this guide shows how to split the work sensibly between you and the tool, step by step.
Which assistant fits which job?
Every assistant on the market claims to "analyse data", but they run in different places. I recommend choosing by where your data already lives rather than by brand. If the data sits in Excel, the assistant inside Excel creates the least friction. If it sits in Google Sheets, Gemini is the natural choice. For a messy CSV export, a chat assistant with file upload is usually the fastest route.
| Assistant | Where it runs | Strong at | Watch out for |
|---|---|---|---|
| ChatGPT (data analysis) | Chat with file upload | Analysing CSV and Excel files with Python, charts, interactive tables | Privacy settings of your plan and what you upload |
| Claude | Chat with file upload | Reading long files, writing and explaining code, report text | Whether code execution is enabled on your account |
| Gemini (Sheets, Colab) | Google Sheets side panel, Colab notebooks | Formulas, tables, charts, notebook code | Requires an eligible Workspace or Google AI plan |
| Copilot in Excel | Chat pane inside Excel | Formula columns, PivotTables, charts, text themes | Data needs a clean table structure |
| GA4 built-in features | Google Analytics 4 | Anomaly detection, automated insights | Looks only at GA4 data and does not explain causes |
In short, there is no single "best" assistant. In daily work my team and I usually combine two tools: the assistant that lives where the data is, plus a second tool for cross-checking critical numbers.
How does AI for data analysis work in ChatGPT?
According to OpenAI's official help article, ChatGPT can analyse uploaded files, answer questions about the data and create tables or charts. For some tasks it writes and runs Python code in the background. The same page also notes that this Python environment cannot make external web requests, so you need to upload the data the analysis depends on.
In practice, the flow looks like this:
- Upload the file and first ask, "Which columns does this file contain, and what data type is each one?"
- Compare the structure it describes with the structure you know.
- Then ask one clear question: which metric, which breakdown, which date range.
- After you get the result, ask to see the code it used.
OpenAI also recommends clear column names and one record per row for the best results. That matches my own experience: merged cells, two stacked header rows and subtotal lines confuse assistants more than anything else. Therefore two minutes of tidying before the upload often saves the next half hour.
One more habit helps a lot. When the answer arrives, ask the assistant to restate your question in its own words. If its restatement differs from what you meant, you have found the problem before it reaches your report.
Where do Claude and Gemini make a difference?
I find Claude especially useful for long, text-heavy data. Customer reviews, support tickets and open-ended survey answers are good examples: extracting themes and then turning them into a numeric table is one of its strengths. In addition, it explains the code it writes step by step, which helps anyone who wants to read and audit that code.
Gemini makes sense for data that already lives in the Google ecosystem. According to Google's Sheets help page, Gemini can create tables, write formulas, build charts and generate insights from your data. The same page says the feature requires an eligible Workspace or Google AI plan and works best with native Google Sheets files. So if you work with an .xlsx file, you may need to save it as a Google Sheets file first.
In Colab, Gemini helps you generate code in a notebook and explains errors. If your Python knowledge is limited, this is a good way to run an analysis without writing everything from scratch. Still, this is not a pandas tutorial; the goal is not to produce code but to understand the code you ask for.
Put simply, Claude shines when the data is mostly words, and Gemini shines when the data already sits in Sheets. Both still need the same verification habits described later in this guide.
What can you do with Copilot in Excel?
For teams that live in Excel, Copilot is usually the lowest-friction option. Microsoft's support page says Copilot can analyse your data and return insights as charts, PivotTables, summaries, trends or outliers. It can also create new columns calculated from existing data and summarise themes and sentiment in text sets such as surveys or reviews.
Copilot helps me most when I know the logic of a formula but not the exact syntax. For example, saying "add a column that calculates cost per conversion for each campaign" is faster than typing the formula by hand. Moreover, because the result goes straight into the workbook, you can click the cell and read the formula.
On the other hand, Copilot expects data in a proper table structure. On sheets with vague headers or blank rows splitting the data, the output gets weaker. That is why the first step is always the same: format the range as an Excel table, give columns clear names, and only then ask your question.
I also suggest keeping a copy of the original sheet before Copilot adds columns or reorders data. If a change goes wrong, you can compare against the untouched version instead of guessing what moved.
What do the AI features inside GA4 offer?
Google Analytics 4 has a built-in layer called Analytics Intelligence. According to Google's anomaly detection documentation, it predicts the expected value of a metric from historical data and flags the actual value as an anomaly when it falls outside the credible interval. For daily anomalies, the documented training period is 90 days.
This answers the first half of the question "why did traffic drop today?": is the drop actually unusual or not? However, it does not answer the second half, the cause. Whether a campaign paused, a tag broke or a page went offline is something you have to investigate.
My recommendation is to treat GA4 alerts as an early warning and to ask the detailed follow-up questions on exported data in a chat assistant. To keep traffic sources clean, tag your campaign links consistently with a UTM builder; otherwise the assistant analyses messy source data. I cover the tools we use alongside GA4 in the article on website traffic analysis tools.
Also remember that GA4 insights depend on correct tracking. If key events stop firing, the anomaly model sees a real drop in the data even though customers behaved normally. So check your tracking setup whenever an alert looks surprising.
How can AI help with data cleaning?
Most of the time in any analysis goes into cleaning, and this is where assistants genuinely save hours. However, I suggest you start not with "clean this" but with "list the problems first". That way you see what will change before anything changes.
For a good cleaning pass, you can ask for these checks:
- How many blank cells exist and in which columns they cluster.
- Different spellings of the same value, for example "New York", "new york" and "NYC".
- Mixed date formats and time zone differences in date columns.
- Values that look like numbers but sit in the file as text, especially with comma and dot separators.
- Fully duplicated rows versus rows where only the ID repeats.
Decimal separators are a common trap in international data. "1.250" can mean one thousand two hundred fifty or one point two five, depending on the locale of the export. So tell the assistant explicitly which locale the source system used. After cleaning, compare row counts and totals with the original file; an unexpected difference means something disappeared.
Finally, keep a short log of every cleaning step. If a stakeholder later asks why a number changed, you can point to the exact rule that caused it.
Which questions should you ask during exploratory analysis?
Exploratory analysis answers the question "what is in this data?". Assistants are at their most productive here because they offer fast, multi-angle views. Still, layered questions work better than broad ones.
I usually move through three layers. First, the big picture: totals, averages, medians and distributions. Second, breakdowns: differences by channel, device, region or product group. Third, time: weekly trends, seasonality and break points. At each layer I ask the assistant for both the number and the method.
For example, in an e-commerce sales file, asking for "median order value and the average excluding the top one percent" instead of just "average order value" shows whether a few very large orders distort the picture. Also, assistants answer detailed requests like this more consistently than vague ones.
Exploratory analysis generates hypotheses; it does not prove them. Therefore be careful with any sentence where the assistant claims that X causes Y. To confirm a relationship, you need a controlled test, and an A/B test calculator helps you judge whether a difference is statistically meaningful.
A useful closing question for this stage is: "What else would you look at in this data, and why?" The answer often surfaces a segment or time window you had not considered.
How do you visualise results from AI for data analysis?
Assistants create charts quickly, but they do not always choose the right chart. So state the purpose when you ask: comparison, change over time, distribution or part of a whole. When the purpose is clear, the assistant usually picks a suitable chart type.
These are the simple rules I follow:
- Line charts for time series; horizontal bars for category comparisons.
- Pie charts only when there are two or three parts.
- If the axis does not start at zero, say so in the chart title.
- One message per chart; a second message means a second chart.
In addition, ask the assistant to write a one-sentence reading note under each chart. That note tells your audience what to look at when the chart lands in a presentation. I discuss how charts can drive engagement on a website in the article on data visualisation.
Finally, check colours and labels. An AI-generated chart may miss an axis title or use the wrong unit, and small errors like these damage trust in a presentation very quickly. Ask for currency and percentage formats that match your audience, too.
How do you get formulas, SQL and Python code from an assistant?
Generating code and formulas is the most measurable benefit of assistants, because you can run the result and test it. However, a good result requires that you introduce the table: table name, column names, data types and a few sample rows. Without that, the assistant may invent column names that do not exist.
A good request contains four parts:
- Your environment: Excel, Google Sheets, which database, or Python.
- The table structure: column names and types.
- The desired output: which column, which grouping, which sort order.
- Edge cases: blank values, division by zero, date range boundaries.
For example, asking for "a column in the campaign table that divides cost by conversions and stays blank when conversions are zero" prevents a division by zero error from the start. Run any generated SQL on a small date range first and compare the output with a few rows you checked by hand.
This article is not a SQL or pandas course; the aim is that you understand the generated code well enough to read it. Ask the assistant to explain what each line does. If the explanation contains an illogical step, the code itself is probably wrong too.
How do you verify the assistant's results?
Verification is the step of AI for data analysis that you should never skip. Because the assistant is fast, people rush as well; yet a fast wrong answer costs more than a slow right one. Before any report goes out, my team and I run through the same checklist.
- Total check: does the assistant's total match the total in the source system?
- Row check: how many rows remain after filtering, and does that make sense?
- Manual sample: calculate three random rows by hand and compare.
- Second tool: have another assistant or a plain formula recalculate a critical number.
- Code review: do the filters and groupings match the question you asked?
For instance, recalculating a simple ratio like conversion rate with a conversion rate calculator takes seconds. That small step stops a wrong ratio from reaching a board presentation.
In short, treat the assistant's answer like a colleague's first draft: useful and fast, but something you check before you put your name on it.
How do hallucinations show up in data analysis?
A hallucination happens when a model confidently produces information that is not true. In data analysis it usually appears in three ways. First, it invents a column or value that does not exist. Second, it gives an "estimated" number without running any code. Third, it attaches a wrong interpretation to a correctly calculated number.
The third is the most dangerous, because a correct number makes you trust the explanation too. For example, the assistant may correctly calculate that sales rose in a given month and link it to a new campaign, while that month actually included a holiday season. The model does not know the context outside the data.
A few habits reduce hallucinations. Above all, ask the assistant to calculate with code and show that code. Also instruct it to say where it is unsure. Next, separate calculation from interpretation: ask only for the numbers first, then interpret them yourself or in a separate step.
Finally, watch out when the assistant brings in outside figures. If an industry benchmark arrives without a source link, keep it out of your report. A number without a source undermines trust in the rest of the document.
Which data should you never upload under GDPR?
This section is not legal advice; it is a practical framework from the field. Uploading a file with personal data to an AI service can mean transferring that data to a service provider. Consequently, your GDPR obligations apply here too, including questions about legal basis, processing agreements and where the provider stores data.
These are the rules I apply in practice:
- Remove columns with names, phone numbers, email addresses, ID numbers and full addresses before uploading.
- If the analysis does not need identity, replace customer IDs with meaningless codes.
- Never upload special category data such as health information to general chat assistants.
- On business plans, check the settings and contract to see whether your data trains the provider's models.
Most analysis questions do not need personal data at all. "Which city has the highest average basket?" does not require a name column. Shrinking the data to what you need reduces both the risk and the assistant's confusion. I explain the broader website side in the guide on building a GDPR-compliant website.
In which marketing scenarios do we use assistants?
A large part of our daily work involves marketing data, and the scenarios where assistants save the most time are repetitive and predictable. For example, grouping thousands of rows from a Google Ads search terms report by intent takes hours by hand; with an assistant the first draft appears in minutes, and we then review the grouping.
Likewise, on the SEO side, clustering queries exported from Search Console into topics, finding pages with high impressions but low clicks, and turning them into a priority list all speed up with an assistant. Still, we decide which page to touch first based on business goals, not on the assistant's ranking.
Reporting is the third big area. The assistant can draft a short summary from weekly numbers. However, choosing which metric matters depends on your digital marketing KPIs, and no tool can pick those for you.
If you would rather not run this yourself, my team and I analyse the data in our Google Ads management and SEO consulting work and turn it into clear recommendations.
How do you write a good prompt for AI for data analysis?
Prompt writing is a topic of its own, so here I only cover the part specific to data. A good data prompt usually carries four pieces of information: context, data description, question and output format. If one of them is missing, the assistant fills the gap with its own assumption.
Compare these two prompts. The weak one: "Analyse this file." The strong one: "This file holds the last 90 days of orders from an online store; each row is one order. Give me order count, median basket value and return rate by channel as a table, and show the code you used." The second prompt narrows the calculation and makes verification possible.
Also focus each prompt on one question. Asking five questions in one message leads the assistant to answer some of them superficially. Then it becomes easier to check the answer and to see at which step an error appeared.
One last tip: ask the assistant to list its assumptions. The question "What assumptions did you make in this analysis?" often reveals filters and conversions you would otherwise miss.
What are the limits of AI assistants?
Knowing the limits is the precondition for using assistants with the right expectations. First of all, there are file size and context limits; with very large datasets, the assistant may sample the data or skip parts of it. Therefore it is safer to aggregate large data in the database first and give the assistant the summary.
The second limit is business context. The assistant does not know that you changed prices last month, that a product went out of stock or that your tracking code failed for two days. If you do not supply that information, you get results that are correctly calculated but wrongly interpreted.
The third is consistency. Asking the same question twice can produce slightly different answers. So for recurring reports, instead of trusting the assistant to recalculate every time, store the formula or script you generated and verified once, and reuse it.
Finally, responsibility does not transfer. You stand behind the number in the report. If you want to use AI on the customer-facing side of your website, read the guide on using AI on your website.
What are the most common mistakes?
The mistake I see most often is treating the assistant's first answer as final. The first answer is a starting point; it usually becomes clear after one or two follow-up questions. So get used to asking things like "which filter produced this result?" or "did you exclude returned orders?".
The second mistake is merging data from different sources without checking definitions. A conversion in your ad platform and a conversion in GA4 may not share the same definition. The assistant will happily place both numbers side by side, but it does not know why they differ. Consequently, unless you explain the definitions, you end up with a misleading comparison.
The third mistake is working too long in one session. As a chat grows, the assistant may mix up details from earlier steps. For a large analysis, starting each major step in a fresh session with clean data gives more consistent results.
Finally, mixing file versions is common. Save the cleaned file under a new name, so you always know which result came from which data.
What steps should you follow to get started?
Run your first attempt on small, low-risk data. A report without personal data whose answer you already know is ideal, for example last month's sales summary by channel. That way you immediately see where the assistant is right and where it slips.
- Choose the assistant based on where the data lives: Excel, Sheets or chat.
- Remove personal data columns and bring the file into a clean table structure.
- Ask the assistant to describe the data first, then ask one clear question.
- Look at the code or formula and compare totals with the source.
- Store the verified formula or script in a reusable form.
When you repeat these five steps on a few reports, you build your own sense of where the assistant is reliable and where it needs attention. No tool can hand you that judgement ready-made.
If you want to turn your data into a form you can make decisions with, and clarify priorities across ads and SEO, my team and I can plan the process with you from start to finish.




