TL;DR / The Direct Answer: When a client or vendor sends a raw CSV with no data dictionary, don’t scroll through rows guessing what “TRX_STAT_CD” means. Delete the sensitive columns, then paste the header row and the first 50 rows into any AI chat — the free tier is enough — with the prompt below. In five minutes you get a profile of every column, the data-quality problems, and a list of questions to send back. Free plans of ChatGPT, Claude and Gemini accept file uploads too, with daily limits: if your company allows it, upload the full file instead for exact counts.
Who this is for: Analysts and ops teams handling messy raw CSV or Excel exports with no data dictionary.
Skip this if: Your data already sits in a clean, documented warehouse table.
Note: AI pricing, plan names, and product features can change quickly. Re-check official pages before you pay for a tool or choose a plan.
The Problem: The Mystery Data Dump
If you work in operations or analytics, you know the feeling. An email arrives from a client or an external vendor containing a 50-megabyte CSV file named export_final_v3.csv. There is no data dictionary. The column headers are cryptic acronyms. Dates are formatted in three different ways, and half the cells in the "Revenue" column say "9999" or "N/A" instead of zero.
Your job is to build a report from this, but right now, you don't even know what you're looking at. The instinct is to open it in Excel, freeze the top row, and start filtering column by column to figure out the shape of the data. It's tedious, error-prone, and burns hours of your day before you've even done any real analysis.
The Solution: Exploratory Data Analysis (EDA) on Autopilot
In data science, the first step with any new dataset is called Exploratory Data Analysis (EDA). It's a structured way to ask basic questions: How many rows? What are the data types? How many missing values? Are there obvious outliers?
You don’t need to write Python to do this. A sample of 50 rows is enough for an AI to read the structure: what each column seems to hold, which ones mix text and numbers, and where placeholder values hide. It can’t count what it hasn’t seen, so the prompt below tells it to label every number as “in this sample”. You get a map of the territory before you start walking.
Before You Paste: A 60-Second Clean-Up
Client and vendor files often carry personal data. Take one minute in Excel before anything goes into an AI chat:
- Delete columns with personal data: customer names, phone numbers, email IDs, PAN, Aadhaar, bank account numbers, addresses. The AI doesn’t need them to judge the file’s structure.
- Keep the header row and the first 50 rows. Copy them straight from Excel; they paste as tab-separated text, which AI tools read fine.
- Check your company’s AI policy. Many Indian firms ban client data in personal AI accounts. If yours does, use the approved company tool, or share only the header row and a few rows you have typed in yourself.
Build Your CSV Prompt
CSV Interrogation Prompt Builder
Plain Prompt (Copy As-Is)
Role: You are an expert data analyst performing an Exploratory Data Analysis (EDA) on a new, undocumented dataset. Task: Below are the header row and the first 50 rows of a CSV I received with no data dictionary. Do not analyze the business meaning yet. Profile the structure and data quality. Say "in this sample" for every count. Give me: 1. Row and column count. 2. Every column: its inferred data type and a guess at what it represents. 3. Data quality flags: missing values, mixed data types (e.g. text in a number column), placeholders (like '9999', 'NA' or '-'), duplicate rows, and dates that could be DD/MM or MM/DD. 4. Up to 5 questions I should ask the sender. Example Output Format: - **Column Name**: [Name] | **Type**: [Type] | **Likely Means**: [Guess] | **Warning**: [Any data quality issues] Data: [PASTE HEADER ROW + FIRST 50 ROWS HERE]
Follow-up: turn the flags into an email
Turn your questions and data-quality flags into a short, polite email to the person who sent this file. Ask for the data dictionary and a confirmation on each flag. Number the questions so they can reply inline. No more than 150 words. Do not blame anyone.
A Worked Example
Made-up data, for illustration only. Say a collections agency sends a monthly export. You delete the customer name and phone columns, then paste what is left:
| TXN_ID | BR_CD | DT | AMT | STAT_CD | REM |
|---|---|---|---|---|---|
| T-10231 | BLR02 | 03/07/2026 | 1,25,000 | C | |
| T-10232 | PNQ01 | 2026-07-03 | 48500 | P | follow up |
| T-10233 | BLR02 | 07/03/2026 | 9999 | C | NA |
| T-10234 | 04-07-2026 | ₹ 12,400 | X | dup? | |
| T-10235 | HYD1 | 05/07/2026 | - | P | |
| T-10231 | BLR02 | 03/07/2026 | 1,25,000 | C |
What good output looks like
- Size: 6 rows × 6 columns in this sample.
- DT | text | transaction date (guess) | three formats in six rows. “07/03/2026” could be 7 March or 3 July.
- AMT | text, not a number | amount in rupees (guess) | lakh commas, a ₹ sign, “-” for blank, and 9999 looks like a placeholder.
- STAT_CD | code | status: C, P, X might be Closed, Pending, Cancelled (guess) | confirm with the sender.
- BR_CD | code | branch (guess) | one blank; “HYD1” breaks the pattern of three letters and two digits.
- Duplicates: T-10231 appears twice with identical values.
- Biggest risk: summing AMT as-is gives a wrong total, because Excel treats half the values as text.
- Questions for the sender: What do C, P and X mean? Is 9999 a real amount? Which date format does your system export?
A good answer labels every meaning as a guess and every count as “in this sample”. If yours states totals for the whole file from 50 pasted rows, ask it to redo the profile.
Check Before You Trust It
- Spot-check two columns it flagged against the real sheet. If a flag is wrong, tell the AI and ask it to re-check the rest.
- Counts refer to your sample only, unless you uploaded the full file.
- Every “likely means” is a guess until the sender confirms it.
- Every date that could be read two ways is listed.
- No sensitive column went into the chat.
- If you uploaded the full file, the AI showed how it counted. Re-run one count yourself with a filter or a COUNTIF.
The Common Trap
The biggest trap in data analysis is jumping straight into building charts without understanding the mess underneath. If you don't find the missing values, mismatched date formats, and duplicate records first, your final presentation will be built on a foundation of bad math.