One of the most time-consuming parts of data analysis has nothing to do with the analysis itself: getting the data into a shape you can actually visualize or work with. Writing SQL, reading through an API’s documentation, copying a table out of a PDF, combining several Excel sheets into one. None of it counts as analysis, and yet it takes real effort every single time.
AI Data, added in Exploratory Desktop v16, replaces that effort with a conversation. Instead of working out how to write the SQL, call the API, parse the PDF, combine the Excel files, or pull data off a website, you simply describe what data you want in your own words, and AI assembles an R script to get it, runs that script, and builds the data frame from the result.
This note walks through what AI Data can do, how to use it, and a few things worth knowing before you start.
AI Data is a chat-based data acquisition screen built into Exploratory. The screen is split into three areas:

What matters here is that what comes back isn’t “data” — it’s an “R script.” AI isn’t fetching an answer from somewhere and displaying it; it assembles a script that says “here’s how you get this data,” and that script runs on your own machine, in your own R.
This design gives you three things.
Turning the tedious, technical parts of data preparation into a conversation, while keeping all the transparency and reproducibility of an R-based data pipeline — that’s what AI Data is about.
You can import essentially any kind of data except images. If R can reach it, consider it within AI Data’s scope.
Internally, AI Data comes with more than 110 reference documents, one per data source, each describing how to retrieve data from it. AI consults the relevant reference first, before writing anything — so instead of a plausible-looking guess, you get a script that follows the correct way to access that particular source.
| Kind | Extensions | How it’s handled |
|---|---|---|
| Tabular text / binary | csv tsv txt
tab log parquet
json |
Read directly as a table |
| Excel | xlsx xls
xlsm |
Pick a sheet, then read it |
| Statistical software files | sav (SPSS) dta (Stata)
sas7bdat xpt (SAS) rds
rda rdata (R) |
Read directly as a table |
| Documents | pdf docx xml
html htm |
Extract the tables inside |
| Folder | — | Handle the files inside together |
Being able to say “combine these into one dataset” — for several Excel sheets, or every CSV in a folder — is what sets this apart from doing it by hand.
PostgreSQL, MySQL / MariaDB, SQL Server, Oracle, Snowflake, BigQuery, Amazon Redshift, Amazon Athena, Amazon Aurora, Presto / Trino, MongoDB, SQLite, Treasure Data, Vertica, Teradata, MS Access, and more — plus generic ODBC / DBI database connections.
Amazon S3, Google Drive, Google Sheets, Google Cloud Storage, Dropbox, Box, SharePoint, Microsoft Azure Storage, Salesforce, Stripe, Intercom, GitHub, Google Analytics, and more.
This is where AI Data has the biggest impact. You don’t need to read the API spec — just name the source, and it goes and gets it. The list below is a representative sample; many more national statistics agencies and domain-specific open data APIs are supported as well.
You can pull a table off a specific page, or ask AI to search the web and find one for you. When you see progress steps like “Searching the web…” and “Fetching web content…” in the chat — shown as the “Searching the Web” and “Fetching Web Page” step cards — this is what’s happening.
Click the + button above the data frame list, and choose AI Data at the top of the menu.

The chat dialog opens, showing the prompt “What data would you like to get?” In the input field below (placeholder text: “Write what data you’d like to get…”), describe what you want in your own words.

From here on, let’s follow a real example. What we have on hand is two workbooks of daily sales data, one for each year:
sales_2024.xlsx — 12 monthly sheets,
2024-01 through 2024-12sales_2025.xlsx — 12 monthly sheets,
2025-01 through 2025-12Each sheet already carries a full daily date —
2024-01-01, 2024-01-02, and so on — right
alongside sales, units_sold, and
customers, and every one of the 24 sheets across both files
shares exactly the same four columns. What makes this tedious isn’t
messy formatting; it’s the sheer number of tabs. The 731 rows are
scattered across 24 separate sheets in two separate workbooks, and Excel
has no built-in “read every sheet from every file in this folder and
stack them into one table” command. Doing it by hand means clicking into
24 tabs, one at a time, and pasting them together — exactly the kind of
data that’s genuinely tedious to assemble manually.
For local files, drag and drop the file or folder onto the chat area.
You can also pick them from the folder icon at the bottom left of the input field (Select files or folders), or from + → Add files or folders.

Once you attach the files, AI Data suggests how to import them, based on what’s inside.

Once you send the request, AI shows what it’s doing as a series of steps — “Looking up reference…”, “Generating R script…”, “Running R script…”, and so on.

When it finishes, the generated R script appears at the top right, and the preview from running it appears at the bottom right.

All 24 sheets from both workbooks come together into a single table
of 731 rows × 6 columns — the original
date, sales, units_sold, and
customers columns, plus a year column parsed
from each file name and a month column parsed from each
sheet name — and AI explains in the chat what it did to combine
them.
For data sources that require authentication, such as databases and APIs, select a connection before making your request. Click the + button in the input field and choose Data connections to see your saved connections.

Pick the connection you want, then describe the data you need.
Aggregate the last 30 days of sales from the orders table, broken down by day and product category.
Get the number of sessions from Google Analytics for the past 90 days, broken down by date, device, and channel.
Once you select a connection, AI builds the script with that source’s type in mind — whether it’s PostgreSQL, Salesforce, or something else.

You can ask for data from a source you haven’t set up a connection for yet — just go ahead and describe it.
Get this fiscal year's monthly sales from the sales table in MySQL.
When no matching connection is found, AI recognizes this, suggests creating one, and opens the connection dialog.

Enter the connection details and save, and a message —
The connection "<name>" has been created. (for
example,
The connection "Redshift (demo)" has been created.) — is
automatically sent into the chat, and the conversation continues
from there. You don’t need to type your request again.

Because AI Data is chat-based, you can look at the preview it produces and give follow-up instructions to add to it or fix anything that catches your eye. Looking at the 731-row sales table from Step 2, for example, adding one more request —
Also add a 7-day moving average of sales, and a column showing the change from the previous day.
— brings the table to 731 rows × 8 columns: a
sales_7day_ma column (a 7-day trailing average, so the
first six rows come back as NA) and a day-over-day change
column, added without disturbing the rows already there.

When you’re happy with the preview, click Save.

It’s imported as a data frame, and from then on you can treat it just like any other data source.

Even after saving, clicking the instruction box shown on the step reopens the chat with the entire conversation intact.

Because the history is still there, you never have to reconstruct what you did last time. Just add your next instruction and continue.

The data you want isn’t always already in a clean, ready-to-use
shape. Take this CO2 emissions dataset from BP’s Statistical Review of
World Energy 2023, for example — it’s 109 rows and 34 columns, and
instead of a single year column, it spreads 1990 through
2022 out across 33 separate year columns, one per year: a wide layout.
Mixed into those 109 rows are eight regional subtotal rows
(Total North America,
Total S. & Cent. America, and so on), broader aggregate
rows like Total World and of which: OECD, nine
blank rows used as section dividers, and six lines of footnotes at the
very bottom explaining things like what n/a means and how
the former USSR’s figures were compiled.

Import the file in this shape and you can’t chart how emissions have changed over time for a given country — there’s no single column you could even put on the axis.
But AI Data reshapes it into a tidy, workable structure automatically, without any extra instruction, reading the shape of the data and adapting to it.
Asking only “Import this CO2 emissions data so I can chart the trend
by country over time,” with nothing said about pivoting or reshaping, is
enough. AI Data recognizes the wide layout, skips the blank divider rows
on its own, and folds the 33 year columns into a single
year column: 3,267 rows × 3 columns, with
country, year, and
co2_emissions_mt — 99 countries and regions across all 33
years.

AI also notices, on its own, that the regional subtotal rows
(Total World, Total OECD, and the rest) and
the footnote rows are still sitting in the result, and says so — it
doesn’t just leave them in silently. Left in, they’d show up in a
chart’s legend as if they were countries. One more instruction —
“Exclude the regional total rows (Total World, Total OECD, and the other
Total rows) and the footnote rows at the bottom of the file, so only
individual countries remain.” — takes the table down to 2,739
rows (83 countries × 33 years), leaving only individual
countries.
From there, charting the trend takes nothing more than setting
year on the X axis, co2_emissions_mt (SUM) on
the Y axis, and country for color, and every country’s
emissions history from 1990 to 2022 is right there.

When you save an AI Data result, the retrieval procedure is automatically saved as a User Skill, so the knowledge of how to get that data accumulates every time you use it.
A skill is a Markdown file, saved under
~/.exploratory/ai_skills/data_source/. It looks like
this:
---
title: Google Sheets Survey Response Customer Data
connectionId: ...
connectionName: ...
description: Join survey responses with customer data by customer ID
keywords:
- survey
- customer data
- Google Sheets
useCount: 0
createdAt: '2026-08-19T14:31:14.781Z'
---
# Request Summary
(a summary of what was requested at the time)
# Script
(the R script that was actually used)
The title, description, and keywords are written by AI, based on that conversation.
The next time you ask for something similar, AI Data searches your saved skills and builds on the previous procedure.
Attaching the same two sales workbooks in a brand-new chat and asking, “Combine all the sheets in both files into one dataset with year and month columns, sorted by date,” gets a different response than the first time. AI Data replies, “There’s a very relevant saved skill for this exact use case. Let me retrieve it,” runs a step labeled “Retrieving the saved skill for combining monthly sheets from multiple yearly Excel sales files,” and then says, “I have a saved skill that matches this request exactly. Let me run it directly.” It skips straight past the structure investigation and retrieves the data using the R script that worked last time as a reference, landing on the same 731 rows × 6 columns as before.

User Skills give you a few concrete benefits.
This matters most for knowledge that only applies in your own environment — the quirky table layout of an internal database, or the specific way one particular API has to be called. AI Data is designed to fit your environment better the more you use it.
AI Data lets you describe the data you want in plain language and have AI assemble and run the R script that retrieves it. What changes with this feature is the time it takes to get from an idea to actually having the data in front of you. Looking up SQL syntax, reading through an API’s documentation, retyping a table out of a PDF by hand — none of that is analysis, but none of it can be skipped either, and turning it into a conversation collapses the distance between deciding you need something and actually having it.
And because what comes out of that conversation is an R script rather than a black-box answer, you can always see exactly what it’s doing, and once you’ve built something, you can re-run it identically, as many times as you like, without going through AI again. Not losing transparency or reproducibility in exchange for convenience — that’s the real value of AI Data.
For frequently asked questions about AI Data, please see this note.
This FAQ note answers a wide range of questions about AI Data. Below is one example — a question about security.
Are API keys or database connection credentials passed to the AI?
No, they are not. Credential handling is designed as follows:
If you’re using a data source that requires an API key, you’ll be prompted to enter it through a dedicated dialog. The value you enter is stored internally within Exploratory and is never passed to the AI as-is.