Sai Tarun
Case study

AI Data Copilot

Ask a spreadsheet anything in plain English. The AI writes the analysis code, runs it safely, and repairs it when it breaks.

Role
Solo project
Type
Web app, live
Model
GPT-4.1 and 4.1-mini
Stack
Python, Streamlit, Pandas, Plotly

The problem

Most people with a spreadsheet have questions, not Python or SQL. Paste the data into a chatbot and you get an answer that sounds right, with no way to check where the number came from.

The idea

Let the model write code, not answers.

  1. Question
  2. Router
  3. GPT-4.1 writes pandas
  1. Sandbox
  2. Error? Retry, up to 3
  3. Chart and insights
  • Every number comes from pandas running on the real data.
  • The generated code and every attempt are shown to the user.
  • Descriptive questions skip code entirely.

How it works

Profiling
The file is read, duplicate column names are cleaned up, and a profile of types, missing values and unique counts is built.
Routing
GPT-4.1-mini decides whether a question needs analysis or just a written answer, so simple questions stay fast and cheap.
Code
GPT-4.1 writes pandas code against the dataset profile, returning a result table and, when useful, a chart.
Sandbox
Code runs with a minimal set of Python builtins. Anything touching files, the network, the OS or imports is blocked before it runs.
Self-repair
If the code fails, the error goes back to the model as context for a fix, up to three times.
Explaining
GPT-4.1-mini turns the result into three insights, one recommendation and one caveat.
Dashboards
Schema detection finds date, revenue, profit, cost, discount, quantity and category columns, then builds a sales, classification or general dashboard.
Reports
Any analysis can be downloaded as a Markdown report, and every run is kept in a session history.

The hardest part

Letting an AI run code without letting it run anythingExecuting model-written code is the riskiest thing the app does. The answer was layers: screen every snippet before it runs, execute it with almost no builtins, and treat failures as information instead of crashes. A typo in a column name becomes a quiet retry with the error attached, not a stack trace in front of the user.

What I'd build next

Isolation
Move execution into a short-lived container per request. A blocklist is a good first layer, but not a true sandbox.
Evaluation
A set of questions with known answers, run on every change, so prompt edits cannot quietly break the numbers.
Bigger data
Multi-file and SQL sources, so it can work across related tables.